TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Apr 8 – Apr 14, 2024
本周最热111

Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention

Tsendsuren Munkhdalai, Manaal Faruqui, Siddharth Gopal

A novel Infini-attention mechanism allows Transformer-based LLMs to handle infinitely long inputs with limited memory and compute, enabling efficient long-context modeling and fast inference.

Infini-attentionTransformer-based Large Language Models (LLMs)compressive memorymasked local attentionHF ↗arXiv ↗

50 篇论文 · 按点赞排序

02

Rho-1: Not All Tokens Are What You Need

Zhenghao Lin, Zhibin Gou, Yeyun Gong +8 authors

Rho-1, a novel language model using Selective Language Modeling, improves efficiency and performance by selectively training on useful tokens rather than all tokens in the corpus.

95next-token predictiontoken-level training dynamicsHF ↗arXiv ↗
03

Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs

Keen You, Haotian Zhang, Eldon Schoop +5 authors

Ferret-UI, a multimodal large language model tailored for mobile UI screens, enhances understanding and interaction through region annotations and a comprehensive dataset of UI tasks, outperforming existing models including GPT-4V.

83multimodal large language modelsMLLMsHF ↗arXiv ↗
04

OmniFusion Technical Report

Elizaveta Goncharova, Anton Razzhigaev, Matvey Mikhalchuk +6 authors

The OmniFusion model, leveraging pretrained LLMs and visual adapters, demonstrates superior performance across multiple visual-language benchmarks compared to open-source alternatives.

78multimodal architecturesOmniFusion modelHF ↗arXiv ↗
05

LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders

Parishad BehnamGhader, Vaibhav Adlakha, Marius Mosbach +3 authors

LLM2Vec transforms decoder-only LLMs into effective text encoders using bidirectional attention, masked next token prediction, and unsupervised contrastive learning, achieving state-of-the-art performance on text embedding tasks.

67decoder-only language modelsLLM2VecHF ↗arXiv ↗
12

JetMoE: Reaching Llama2 Performance with 0.1M Dollars

Yikang Shen, Zhen Guo, Tianle Cai +1 authors

JetMoE-8B, a cost-effective large language model with 8 billion parameters, achieves impressive performance using a Sparsely-gated Mixture-of-Experts architecture, demonstrating efficient use of resources and computational savings.

38Sparsely-gated Mixture-of-ExpertsSMoEHF ↗arXiv ↗
15

Best Practices and Lessons Learned on Synthetic Data for Language Models

Ruibo Liu, Jerry Wei, Fangyu Liu +8 authors

The success of AI models relies on the availability of large, diverse, and high-quality datasets, which can be challenging to obtain due to data scarcity, privacy concerns, and high costs. Synthetic data has emerged as a promising solution by generating artificial data that mimics real-world patterns. This paper provides an overview of synthetic data research, discussing its applications, challenges, and future directions. We present empirical evidence from prior art to demonstrate its effectiveness and highlight the importance of ensuring its factuality, fidelity, and unbiasedness. We emphasize the need for responsible use of synthetic data to build more powerful, inclusive, and trustworthy language models.

32HF ↗arXiv ↗
17

Stream of Search (SoS): Learning to Search in Language

Kanishk Gandhi, Denise Lee, Gabriel Grand +4 authors

Language models can learn complex problem-solving strategies via search by pretraining on streams of search generated by heuristic solvers and further improving with policy-based methods, solving previously unsolvable problems.

30transformer-based language modelpretrainHF ↗arXiv ↗
20

ByteEdit: Boost, Comply and Accelerate Generative Image Editing

Yuxi Ren, Jie Wu, Yanzuo Lu +11 authors

ByteEdit, a feedback learning framework, enhances generative image editing through image reward models, pixel-level coherence, and adversarial learning, significantly improving quality, consistency, and speed compared to existing products.

25diffusion-based generative image editingimage outpaintingHF ↗arXiv ↗
22

CantTalkAboutThis: Aligning Language Models to Stay on Topic in Dialogues

Makesh Narsimhan Sreedhar, Traian Rebedea, Shaona Ghosh +1 authors

The CantTalkAboutThis dataset improves language models' ability to stay on topic during conversations by including distractor turns, making them more resilient and coherent compared to general-purpose instruction-tuned models.

25instruction-tuning datasetstask-oriented interactionsHF ↗arXiv ↗
24

SpatialTracker: Tracking Any 2D Pixels in 3D Space

Yuxi Xiao, Qianqian Wang, Shangzhan Zhang +4 authors

SpatialTracker estimates 3D point trajectories using monocular depth estimation and transformer updates, achieving superior video tracking performance with ARAP constraints and rigidity embeddings.

24monocular depth estimatorstriplane representationHF ↗arXiv ↗
26

WILBUR: Adaptive In-Context Learning for Robust and Accurate Web Agents

Michael Lutz, Arth Bohra, Manvel Saroyan +2 authors

Wilbur uses a differentiable ranking model and instruction synthesis to fine-tune a large language model for web agent tasks, achieving state-of-the-art results with text-only inputs and superior performance on the WebVoyager benchmark.

22differentiable ranking modelinstruction synthesisHF ↗arXiv ↗
27

LLoCO: Learning Long Contexts Offline

Sijun Tan, Xiuyu Li, Shishir Patil +5 authors

LLoCO method enhances large language models' ability to handle long contexts efficiently by combining context compression, retrieval, and parameter-efficient finetuning.

22self-attentioncontext compressionHF ↗arXiv ↗
30

HGRN2: Gated Linear RNNs with State Expansion

Zhen Qin, Songlin Yang, Weixuan Sun +4 authors

The introduction of an outer-product-based state expansion mechanism in hierarchical gated linear RNN (HGRN) enhances its performance and expressiveness without additional parameters, outperforming architectures like Mamba and LLaMa in various tasks.

21hierarchically gated linear RNNlinear attentionHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号