TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

February 2024
本月最热630

The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits

Shuming Ma, Hongyu Wang, Lingxiao Ma +7 authors

A 1-bit LLM variant, BitNet b1.58, achieves comparable performance to full-precision models with reduced computational costs and introduces new scaling laws and hardware design opportunities.

BitNet1-bit LLMternaryTransformerHF ↗arXiv ↗

50 篇论文 · 按点赞排序

07

Chain-of-Thought Reasoning Without Prompting

Xuezhi Wang, Denny Zhou

LLMs can perform chain-of-thought reasoning through top-k decoding without manual prompt engineering, outperforming greedy decoding and showing higher confidence in answers.

111large language models (LLMs)chain-of-thought (CoT) promptingHF ↗arXiv ↗
08

Neural Network Diffusion

Kai Wang, Zhaopan Xu, Yukun Zhou +4 authors

Diffusion models can generate high-performing neural network parameters using an autoencoder, producing new subsets of network parameters with comparable or improved performance.

101diffusion modelsautoencoderHF ↗arXiv ↗
10

OLMo: Accelerating the Science of Language Models

Dirk Groeneveld, Iz Beltagy, Pete Walsh +40 authors

OLMo, an open-sourced language model, provides comprehensive access to training data, code, and architecture, facilitating research and innovation in language modeling.

86open language modellanguage modelsHF ↗arXiv ↗
14

Genie: Generative Interactive Environments

Jake Bruce, Michael Dennis, Ashley Edwards +22 authors

Genie, a 11B parameter unsupervised generative model, creates action-controllable virtual worlds from unlabelled videos using spatiotemporal tokenization and autoregressive dynamics, enabling agent training from unseen video behaviors.

72spatiotemporal video tokenizerautoregressive dynamics modelHF ↗arXiv ↗
15

Grandmaster-Level Chess Without Search

Anian Ruoss, Grégoire Delétang, Sourabh Medapati +5 authors

A large-scale transformer model trained on a vast dataset of chess games outperforms traditional chess engines and other state-of-the-art models without domain-specific tweaks.

70transformer modelsupervised learningHF ↗arXiv ↗
16

Training-Free Consistent Text-to-Image Generation

Yoad Tewel, Omri Kaduri, Rinon Gal +4 authors

ConsiStory achieves state-of-the-art text-to-image subject consistency without fine-tuning by using shared internal activations, attention blocks, and feature injection.

67text-to-image modelssubject consistencyHF ↗arXiv ↗
19

More Agents Is All You Need

Junyou Li, Qin Zhang, Yangbin Yu +2 authors

A sampling-and-voting method enhances large language models' performance by increasing the number of agents, with effectiveness tied to task difficulty.

59large language modelsLLMSHF ↗arXiv ↗
21

Aya Dataset: An Open-Access Collection for Multilingual Instruction Tuning

Shivalika Singh, Freddie Vargus, Daniel Dsouza +30 authors

The initiative builds a human-curated instruction-following dataset spanning 65 languages and creates the largest multilingual collection of instruction-following instances through templating and translating existing datasets across 114 languages, contributing datasets and platforms for participatory research.

57Instruction fine-tuningIFTHF ↗arXiv ↗
22

Generative Representational Instruction Tuning

Niklas Muennighoff, Hongjin Su, Liang Wang +5 authors

GrIT allows large language models to excel at both generative and embedding tasks through task instruction, leading to new state-of-the-art performance without sacrificing efficiency.

54generative representational instruction tuningGRITHF ↗arXiv ↗
26

RAPTOR: Recursive Abstractive Processing for Tree-Organized Retrieval

Parth Sarthi, Salman Abdullah, Aditi Tuli +3 authors

The RAPTOR model enhances retrieval-augmented language models by recursively embedding, clustering, and summarizing text chunks, leading to better performance on question-answering tasks involving complex reasoning.

49retrieval-augmented language modelsRecursive embeddingHF ↗arXiv ↗
27

Nemotron-4 15B Technical Report

Jupinder Parmar, Shrimai Prabhumoye, Joseph Jennings +24 authors

Nemotron-4 15B, a large multilingual language model, excels in English, multilingual, and coding tasks, demonstrating superior performance in multilingual capabilities compared to larger and specialized models.

48large multilingual language modeldownstream evaluation areasHF ↗arXiv ↗
28

FiT: Flexible Vision Transformer for Diffusion Model

Zeyu Lu, Zidong Wang, Di Huang +4 authors

The Flexible Vision Transformer adapts to varied image resolutions and aspect ratios through dynamic tokenization and extrapolation techniques, outperforming traditional methods.

48diffusion modelsDiffusion TransformersHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号