TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Apr 1 – Apr 7, 2024
本周最热113

Jamba: A Hybrid Transformer-Mamba Language Model

Opher Lieber, Barak Lenz, Hofit Bata +19 authors

Jamba, a hybrid Transformer-Mamba MoE model, achieves state-of-the-art performance with high throughput and minimal memory usage, supporting large context lengths.

TransformerMambamixture-of-experts (MoE)interleaving layersHF ↗arXiv ↗

50 篇论文 · 按点赞排序

03

ReFT: Representation Finetuning for Language Models

Zhengxuan Wu, Aryaman Arora, Zheng Wang +4 authors

Representation Finetuning (ReFT) methods, exemplified by Low-rank Linear Subspace ReFT (LoReFT), achieve high efficiency and performance by adapting representations in frozen base models, outperforming state-of-the-art Parameter-efficient Fine-tuning (PEFT) methods.

101Parameter-efficient fine-tuningRepresentation FinetuningHF ↗arXiv ↗
08

Advancing LLM Reasoning Generalists with Preference Trees

Lifan Yuan, Ganqu Cui, Hanbin Wang +12 authors

Eurus, a suite of reasoning-optimized large language models, achieves state-of-the-art performance on various benchmarks through UltraInteract, a large-scale, high-quality alignment dataset, and a novel reward modeling objective.

45large language models (LLMs)Mistral-7BHF ↗arXiv ↗
09

Long-context LLMs Struggle with Long In-context Learning

Tianle Li, Ge Zhang, Quy Duc Do +2 authors

LIConBench evaluates long-context LLMs on extreme-label classification tasks with sequences up to 50K tokens, highlighting performance dips beyond 20K tokens and favoring of recent labels.

36Large Language Modelslong in-context learningHF ↗arXiv ↗
17

HyperCLOVA X Technical Report

Kang Min Yoo, Jaegeun Han, Sookyo In +374 authors

HyperCLOVA X is a multilingual large language model trained in Korean, English, and code, demonstrating strong reasoning, knowledge, and cross-lingual capabilities.

25large language modelsLLMsHF ↗arXiv ↗
22

ReALM: Reference Resolution As Language Modeling

Joel Ruben Antony Moniz, Soundarya Krishnan, Melis Ozyildirim +5 authors

LLMs are effectively utilized to resolve various types of references, including non-conversational entities on screen, by converting reference resolution into a language modeling problem, showing significant improvements over existing systems.

22LLMsreference resolutionHF ↗arXiv ↗
29

Are large language models superhuman chemists?

Adrian Mirza, Nawaf Alampara, Sreekanth Kunchapu +25 authors

ChemBench, an automated evaluation framework, assesses the chemical reasoning of large language models against human experts, revealing their proficiency and shortcomings in specific tasks.

17large language modelsLLMsHF ↗arXiv ↗
30

CosmicMan: A Text-to-Image Foundation Model for Humans

Shikai Li, Jianglin Fu, Kaiyuan Liu +3 authors

CosmicMan generates high-fidelity human images using a specialized data production paradigm and Decomposed-Attention-Refocusing framework within a text-to-image diffusion model.

17text-to-image foundation modelphoto-realisticHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号