TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Sep 18 – Sep 24, 2023
本周最热90

LongLoRA: Efficient Fine-tuning of Long-Context Large Language Models

Yukang Chen, Shengju Qian, Haotian Tang +4 authors

LongLoRA extends LLM context sizes efficiently using sparse local attention and parameter-efficient fine-tuning, demonstrating strong results with minimal computational cost.

LongLoRAdense global attentionsparse local attentionshift short attentionHF ↗arXiv ↗

50 篇论文 · 按点赞排序

03

Language Modeling Is Compression

Grégoire Delétang, Anian Ruoss, Paul-Ambroise Duquenne +9 authors

Large language models demonstrate strong compression capabilities, outperforming domain-specific compressors and offering new insights into scaling laws, tokenization, and in-context learning.

85self-supervised modelslarge language modelsHF ↗arXiv ↗
06

FreeU: Free Lunch in Diffusion U-Net

Chenyang Si, Ziqi Huang, Yuming Jiang +1 authors

A method called FreeU improves diffusion U-Net models' generation quality by re-weighting skip connections and backbone features without additional training.

66diffusion U-NetU-NetHF ↗arXiv ↗
07

DreamLLM: Synergistic Multimodal Comprehension and Creation

Runpei Dong, Chunrui Han, Yuang Peng +11 authors

DreamLLM, a framework for Multimodal Large Language Models, directly samples in the multimodal space to enhance comprehension and creation synergy, enabling free-form interleaved content generation.

60generative modelingmultimodal spaceHF ↗arXiv ↗
08

Kosmos-2.5: A Multimodal Literate Model

Tengchao Lv, Yupan Huang, Jingye Chen +11 authors

Kosmos-2.5, a unified multimodal model, generates spatially-aware and structured text from text-intensive images using a Transformer architecture and task-specific prompts.

56multimodal literate modelmachine readingHF ↗arXiv ↗
14

RMT: Retentive Networks Meet Vision Transformers

Qihang Fan, Huaibo Huang, Mingrui Chen +2 authors

The proposed RMT model, combining RetNet and Transformer architectures, introduces explicit spatial distance priors and coordinate-wise decomposition to achieve exceptional performance in computer vision tasks.

34TransformerRetentive Network (RetNet)HF ↗arXiv ↗
19

Baichuan 2: Open Large-scale Language Models

Aiyuan Yang, Bin Xiao, Bingning Wang +49 authors

Baichuan 2, a series of 7 billion and 13 billion parameter multilingual LLMs, achieves competitive performance on benchmarks and excels in specialized domains, with open-source pre-training checkpoints released.

21large language modelsmultilingual language modelsHF ↗arXiv ↗
23

Replacing softmax with ReLU in Vision Transformers

Mitchell Wortsman, Jaehoon Lee, Justin Gilmer +1 authors

Replacing the attention softmax with ReLU and dividing by sequence length in vision transformers can maintain or improve performance compared to using softmax alone.

18attention softmaxReLU-attentionHF ↗arXiv ↗
26

Scaling Laws for Sparsely-Connected Foundation Models

Elias Frantar, Carlos Riquelme, Neil Houlsby +2 authors

Parameter sparsity in Transformers affects their scaling behavior on large datasets; optimal sparsity increases with data volume and different structures can impact performance.

15parameter sparsityscaling lawHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号