TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Oct 7 – Oct 13, 2024
本周最热183

Differential Transformer

Tianzhu Ye, Li Dong, Yuqing Xia +4 authors

Diff Transformer improves large language models by selectively focusing attention on relevant context and reducing noise, leading to better performance in scaling, long-context modeling, key information retrieval, and in-context learning.

TransformerDiff Transformerdifferential attentionsoftmax attentionHF ↗arXiv ↗

50 篇论文 · 按点赞排序

03

Aria: An Open Multimodal Native Mixture-of-Experts Model

Dongxu Li, Yudong Liu, Haoning Wu +7 authors

Aria is an open multimodal native AI model with best-in-class performance across various tasks, designed with a mixture-of-experts architecture and pre-trained through a four-stage pipeline.

111mixture-of-expert modelvisual tokenHF ↗arXiv ↗
05

Personalized Visual Instruction Tuning

Renjie Pi, Jianshu Zhang, Tianyang Han +3 authors

A new framework called Personalized Visual Instruction Tuning (PVIT) enhances multimodal large language models to recognize and engage with specific individuals in images, utilizing a curated dataset and benchmarks for evaluation.

70multimodal large language modelsMLLMsHF ↗arXiv ↗
06

Pixtral 12B

Pravesh Agrawal, Szymon Antoniak, Emma Bou Hanna +34 authors

Pixtral-12B, a 12-billion-parameter multimodal language model, excels in both natural language and image understanding, surpassing larger models and introducing an open-source benchmark for evaluation.

70multimodal language modelvision encoderHF ↗arXiv ↗
20

MLP-KAN: Unifying Deep Representation and Function Learning

Yunhong He, Yifeng Xie, Zhengqing Yuan +1 authors

MLP-KAN integrates MLPs and KANs in a MoE architecture within a transformer framework to adaptively handle both representation and function learning tasks, achieving competitive results across diverse datasets.

31Multi-Layer PerceptronsMLPsHF ↗arXiv ↗
21

Benchmarking Agentic Workflow Generation

Shuofei Qiao, Runnan Fang, Zhisong Qiu +6 authors

WorFBench and WorFEval create a comprehensive framework for evaluating large language models' workflow generation, uncovering gaps in sequence and graph planning and demonstrating improved performance in downstream tasks.

29Large Language Models (LLMs)WorFBenchHF ↗arXiv ↗
24

FAN: Fourier Analysis Networks

Yihong Dong, Ge Li, Yongding Tao +6 authors

A new Fourier-based network architecture, FAN, efficiently models periodic phenomena with fewer parameters and demonstrates superior performance across various tasks.

28FANFourier AnalysisHF ↗arXiv ↗
27

Selective Attention Improves Transformer

Yaniv Leviathan, Matan Kalman, Yossi Matias

Selective Attention reduces unnecessary context elements in the attention mechanism, improving language modeling performance and decreasing memory and compute requirements.

24Selective Attentionstandard attention mechanismHF ↗arXiv ↗
28

MM-Ego: Towards Building Egocentric Multimodal LLMs

Hanrong Ye, Haotian Zhang, Erik Daxberger +9 authors

A multimodal foundation model for egocentric video understanding is developed with a large QA dataset, a challenging benchmark, and a specialized architecture featuring a "Memory Pointer Prompting" mechanism.

22QA dataegocentric video understandingHF ↗arXiv ↗
29

LongGenBench: Long-context Generation Benchmark

Xiang Liu, Peijie Dong, Xuming Hu +1 authors

A new benchmark, LongGenBench, evaluates the long-context generation capabilities of large language models, revealing varying performance degradation across different models and model series.

22Large Language Models (LLMs)long-context generationHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号