TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

May 20 – May 26, 2024
本周最热157

Your Transformer is Secretly Linear

Anton Razzhigaev, Matvey Mikhalchuk, Elizaveta Goncharova +4 authors

Transformer decoders exhibit near-perfect linear relationships between layers, which can be reduced with cosine-similarity-based regularization, leading to improved performance on benchmarks.

transformer decodersProcrustes similarity scoreresidual componentlinear blocksHF ↗arXiv ↗

36 篇论文 · 按点赞排序

03

MoRA: High-Rank Updating for Parameter-Efficient Fine-Tuning

Ting Jiang, Shaohan Huang, Shengyue Luo +8 authors

MoRA, a high-rank updating method using square matrices, enhances the ability of large language models to learn and memorize new knowledge, especially in memory-intensive tasks, compared to LoRA.

49low-rank adaptationparameter-efficient fine-tuningHF ↗arXiv ↗
06

Not All Language Model Features Are Linear

Joshua Engels, Isaac Liao, Eric J. Michaud +2 authors

Research explores multi-dimensional features in language models, discovering interpretable circular representations in GPT-2, Mistral 7B, and Llama 3 8B, which are used for modular arithmetic tasks.

40linear representation hypothesismulti-dimensional featuresHF ↗arXiv ↗
08

Reducing Transformer Key-Value Cache Size with Cross-Layer Attention

William Brandon, Mayank Mishra, Aniruddha Nrusimha +2 authors

Cross-Layer Attention modifies transformer-based autoregressive large language models to reduce KV cache size while maintaining accuracy, enabling longer sequence lengths and larger batch sizes during inference.

32Key-value cachingtransformer-based autoregressive large language modelsHF ↗arXiv ↗
11

Octo: An Open-Source Generalist Robot Policy

Octo Model Team, Dibya Ghosh, Homer Walke +15 authors

Octo, a large transformer-based policy trained on extensive robotic datasets, demonstrates versatility and efficient fine-tuning for diverse robotic platforms and tasks.

27transformer-based policyOpen X-Embodiment datasetHF ↗arXiv ↗
12

Imp: Highly Capable Large Multimodal Models for Mobile Devices

Zhenwei Shao, Zhou Yu, Jun Yu +5 authors

A systematic study of lightweight large multimodal models (LMMs) led to the development of the Imp family, which achieves superior performance compared to larger models and is capable of high-speed inference on mobile devices.

27large language modelslarge multimodal modelsHF ↗arXiv ↗
13

Dense Connector for MLLMs

Huanjin Yao, Wenhao Wu, Taojiannan Yang +7 authors

The Dense Connector enhances Multimodal Large Language Models by integrating multi-layer visual features, improving performance across image and video benchmarks.

24Dense Connectorvision-language connectorHF ↗arXiv ↗
18

Distributed Speculative Inference of Large Language Models

Nadav Timor, Jonathan Mamou, Daniel Korat +6 authors

Distributed speculative inference (DSI) accelerates large language model inference faster than speculative inference (SI) and traditional inference methods, supporting a wider range of drafters and achieving significant speedups.

17distributed speculative inferencespeculative inferenceHF ↗arXiv ↗
20

Thermodynamic Natural Gradient Descent

Kaelan Donatella, Samuel Duffield, Maxwell Aifer +3 authors

A new hybrid digital-analog algorithm efficiently implements natural gradient descent, achieving superior performance in training neural networks compared to existing digital methods.

15natural gradient descentNGDHF ↗arXiv ↗
26

Grounded 3D-LLM with Referent Tokens

Yilun Chen, Shuai Yang, Haifeng Huang +5 authors

Grounded 3D-LLM integrates 3D vision tasks into a unified generative framework using 3D large multi-modal models and scene referent tokens, demonstrating superior performance across various 3D benchmarks.

123D LMMsgrounded 3D-LLMHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号