TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Feb 26 – Mar 3, 2024
本周最热630

The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits

Shuming Ma, Hongyu Wang, Lingxiao Ma +7 authors

A 1-bit LLM variant, BitNet b1.58, achieves comparable performance to full-precision models with reduced computational costs and introduces new scaling laws and hardware design opportunities.

BitNet1-bit LLMternaryTransformerHF ↗arXiv ↗

49 篇论文 · 按点赞排序

03

StarCoder 2 and The Stack v2: The Next Generation

Anton Lozhkov, Raymond Li, Loubna Ben Allal +63 authors

StarCoder2, a large language model for code developed through a collaboration with Software Heritage, outperforms other models of similar size on various benchmarks and matches or outperforms larger models in specific areas.

160Large Language ModelsCode LLMsHF ↗arXiv ↗
06

Genie: Generative Interactive Environments

Jake Bruce, Michael Dennis, Ashley Edwards +22 authors

Genie, a 11B parameter unsupervised generative model, creates action-controllable virtual worlds from unlabelled videos using spatiotemporal tokenization and autoregressive dynamics, enabling agent training from unseen video behaviors.

72spatiotemporal video tokenizerautoregressive dynamics modelHF ↗arXiv ↗
10

Nemotron-4 15B Technical Report

Jupinder Parmar, Shrimai Prabhumoye, Joseph Jennings +24 authors

Nemotron-4 15B, a large multilingual language model, excels in English, multilingual, and coding tasks, demonstrating superior performance in multilingual capabilities compared to larger and specialized models.

46large multilingual language modeldownstream evaluation areasHF ↗arXiv ↗
11

FuseChat: Knowledge Fusion of Chat Models

Fanqi Wan, Ziyi Yang, Longguang Zhong +3 authors

FuseChat extends the FuseLLM framework for knowledge fusion of chat LLMs through lightweight fine-tuning and parameter merging, achieving superior performance across various domains.

38knowledge fusionfine-tuningHF ↗arXiv ↗
14

Multi-LoRA Composition for Image Generation

Ming Zhong, Yelong Shen, Shuohang Wang +6 authors

Two training-free methods, LoRA Switch and LoRA Composite, enhance multi-LoRA composition in text-to-image models, leading to improved image synthesis performance.

30Low-Rank AdaptationLoRAHF ↗arXiv ↗
15

Humanoid Locomotion as Next Token Prediction

Ilija Radosavovic, Bike Zhang, Baifeng Shi +5 authors

A causal transformer for next token prediction of sensorimotor trajectories enables zero-shot humanoid walking and generalizes to unseen commands.

28causal transformerautoregressive predictionHF ↗arXiv ↗
16

Do Large Language Models Latently Perform Multi-Hop Reasoning?

Sohee Yang, Elena Gribovskaya, Nora Kassner +2 authors

LLMs exhibit latent multi-hop reasoning for certain complex prompts, with strong evidence of the first reasoning hop and moderate evidence of the second hop, showing scaling trends with model size.

28Large Language Modelsmulti-hop reasoningHF ↗arXiv ↗
21

MOSAIC: A Modular System for Assistive and Interactive Cooking

Huaxiaoyue Wang, Kushal Kedia, Juntao Ren +14 authors

MOSAIC, a modular home robot architecture, collaborates with humans to perform complex cooking tasks using pre-trained models and specific modules, achieving 68.3% success in 60 collaborative trials with various recipes.

24modular architecturehome robotsHF ↗arXiv ↗
24

Training-Free Long-Context Scaling of Large Language Models

Chenxin An, Fei Huang, Jun Zhang +4 authors

Dual Chunk Attention (DCA) enhances Llama2 70B to handle over 100k tokens without fine-tuning by decomposing attention into chunk-based modules, achieving performance comparable to finetuned models.

23Large Language Models (LLMs)input tokensHF ↗arXiv ↗
25

Watermarking Makes Language Models Radioactive

Tom Sander, Pierre Fernandez, Alain Durmus +2 authors

Watermarked training data can be detected with high confidence in LLMs, making it easier to identify if watermarked outputs were used for fine-tuning compared to conventional methods.

23watermarked training datamembership inferenceHF ↗arXiv ↗
28

Video as the New Language for Real-World Decision Making

Sherry Yang, Jacob Walker, Jack Parker-Holder +5 authors

Video generation can serve as a unified interface for diverse tasks and advance real-world AI applications through techniques like in-context learning and reinforcement learning, with opportunities in robotics, self-driving, and science.

21in-context learningplanningHF ↗arXiv ↗
29

GPTVQ: The Blessing of Dimensionality for LLM Quantization

Mart van Baalen, Andrey Kuzmin, Markus Nagel +5 authors

Increasing quantization dimensionality with the GPTVQ method enhances the size-accuracy trade-off for large language models by interleaving quantization and using efficient compression techniques.

21vector quantizationGPTVQHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号