TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

本月最热209

CLEAR: Character Unlearning in Textual and Visual Modalities

Alexey Dontsov, Dmitrii Korzh, Alexey Zhavoronkin +6 authors

CLEAR benchmark evaluates multimodal unlearning methods across textual and visual data, highlighting challenges and demonstrating the effectiveness of $\ell_1$ regularization on LoRA weights in mitigating catastrophic forgetting.

Machine UnlearningMUmultimodal language modelsMLMMsHF ↗arXiv ↗

50 篇论文 · 按点赞排序

02

Differential Transformer

Tianzhu Ye, Li Dong, Yuqing Xia +4 authors

Diff Transformer improves large language models by selectively focusing attention on relevant context and reducing noise, leading to better performance in scaling, long-context modeling, key information retrieval, and in-context learning.

183TransformerDiff TransformerHF ↗arXiv ↗
04

Aria: An Open Multimodal Native Mixture-of-Experts Model

Dongxu Li, Yudong Liu, Haoning Wu +7 authors

Aria is an open multimodal native AI model with best-in-class performance across various tasks, designed with a mixture-of-experts architecture and pre-trained through a four-stage pipeline.

111mixture-of-expert modelvisual tokenHF ↗arXiv ↗
05

Movie Gen: A Cast of Media Foundation Models

Adam Polyak, Amit Zohar, Andrew Brown +85 authors

Movie Gen, a suite of foundation models, generates high-quality videos with synchronized audio, excelling in various tasks through architectural, training, and technical innovations.

100transformervideo tokensHF ↗arXiv ↗
07

GPT-4o System Card

OpenAI, Aaron Hurst, Adam Lerer +416 authors

GPT-4o is an omnimodal autoregressive model trained to handle text, audio, image, and video inputs, offering high-performance outputs across these modalities, with particular strengths in vision and audio.

88autoregressive modelomnimodalHF ↗arXiv ↗
08

Baichuan-Omni Technical Report

Yadong Li, Haoze Sun, Mingan Lin +24 authors

Baichuan-Omni, a 7B open-source Multimodal Large Language Model, excels in processing image, video, audio, and text, showcasing competitive performance across multimodal benchmarks.

88Multimodal Large Language Modelmultimodal training schemaHF ↗arXiv ↗
13

Personalized Visual Instruction Tuning

Renjie Pi, Jianshu Zhang, Tianyang Han +3 authors

A new framework called Personalized Visual Instruction Tuning (PVIT) enhances multimodal large language models to recognize and engage with specific individuals in images, utilizing a curated dataset and benchmarks for evaluation.

70multimodal large language modelsMLLMsHF ↗arXiv ↗
14

Pixtral 12B

Pravesh Agrawal, Szymon Antoniak, Emma Bou Hanna +34 authors

Pixtral-12B, a 12-billion-parameter multimodal language model, excels in both natural language and image understanding, surpassing larger models and introducing an open-source benchmark for evaluation.

70multimodal language modelvision encoderHF ↗arXiv ↗
29

Baichuan Alignment Technical Report

Mingan Lin, Fan Yang, Yanjun Shen +22 authors

Baichuan Alignment provides comprehensive insights into alignment methodologies used in Baichuan models, detailing improvements through Prompt Augmentation System, Supervised Fine-Tuning, and Preference Alignment across various benchmarks.

51Prompt Augmentation SystemSupervised Fine-TuningHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号