TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

本年最热629

The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits

Shuming Ma, Hongyu Wang, Lingxiao Ma +7 authors

A 1-bit LLM variant, BitNet b1.58, achieves comparable performance to full-precision models with reduced computational costs and introduces new scaling laws and hardware design opportunities.

BitNet1-bit LLMternaryTransformerHF ↗arXiv ↗

598 篇论文 · 按点赞排序

02

Qwen2.5 Technical Report

Qwen, An Yang, Baosong Yang +39 authors

Qwen2.5, an enhanced series of large language models, demonstrates superior performance across various benchmarks and use cases through extensive pre-training and advanced post-training techniques.

380large language modelspre-trainingHF ↗arXiv ↗
04

CLEAR: Character Unlearning in Textual and Visual Modalities

Alexey Dontsov, Dmitrii Korzh, Alexey Zhavoronkin +6 authors

CLEAR benchmark evaluates multimodal unlearning methods across textual and visual data, highlighting challenges and demonstrating the effectiveness of $\ell_1$ regularization on LoRA weights in mitigating catastrophic forgetting.

209Machine UnlearningMUHF ↗arXiv ↗
08

Differential Transformer

Tianzhu Ye, Li Dong, Yuqing Xia +4 authors

Diff Transformer improves large language models by selectively focusing attention on relevant context and reducing noise, leading to better performance in scaling, long-context modeling, key information retrieval, and in-context learning.

183TransformerDiff TransformerHF ↗arXiv ↗
10

Qwen2 Technical Report

An Yang, Baosong Yang, Binyuan Hui +55 authors

The Qwen2 series, comprising 0.5 to 72 billion parameter models, surpasses prior open models across language understanding, generation, multilingualism, coding, math, and reasoning, with exceptional performance in benchmarks like MMLU, GPQA, HumanEval, GSM8K, BBH, MT-Bench, Arena-Hard, and LiveCodeBench.

175Mixture-of-Expertslanguage modelsHF ↗arXiv ↗
12

Mixtral of Experts

Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux +23 authors

Mixtral 8x7B, a Sparse Mixture of Experts language model, achieves superior performance across benchmarks by using a selective architecture that leverages fewer active parameters.

162Sparse Mixture of Experts (SMoE)feedforward blocksHF ↗arXiv ↗
13

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Zhe Chen, Weiyun Wang, Yue Cao +37 authors

InternVL 2.5, an advanced multimodal large language model, showcases competitive performance across various benchmarks, including multimodal reasoning and understanding, and is the first open-source model to surpass 70% on the MMMU benchmark using Chain-of-Thought reasoning.

162multimodal large language modelvision encodersHF ↗arXiv ↗
14

Qwen2.5-Coder Technical Report

Binyuan Hui, Jian Yang, Zeyu Cui +14 authors

Qwen2.5-Coder series demonstrates state-of-the-art code generation, completion, reasoning, and repair capabilities using the Qwen2.5 architecture with over 5.5 trillion tokens of training data.

158Qwen2.5-CoderQwen2.5-Coder-1.5BHF ↗arXiv ↗
15

Your Transformer is Secretly Linear

Anton Razzhigaev, Matvey Mikhalchuk, Elizaveta Goncharova +4 authors

Transformer decoders exhibit near-perfect linear relationships between layers, which can be reduced with cosine-similarity-based regularization, leading to improved performance on benchmarks.

157transformer decodersProcrustes similarity scoreHF ↗arXiv ↗
16

StarCoder 2 and The Stack v2: The Next Generation

Anton Lozhkov, Raymond Li, Loubna Ben Allal +63 authors

StarCoder2, a large language model for code developed through a collaboration with Software Heritage, outperforms other models of similar size on various benchmarks and matches or outperforms larger models in specific areas.

157Large Language ModelsCode LLMsHF ↗arXiv ↗
17

Self-Rewarding Language Models

Weizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho +3 authors

A study on Self-Rewarding Language Models shows that using LLM-as-a-Judge prompting for iterative DPO training enhances both instruction-following and self-reward generation, leading to superior performance compared to existing systems.

156Self-Rewarding Language ModelsLLM-as-a-JudgeHF ↗arXiv ↗
24

PaliGemma 2: A Family of Versatile VLMs for Transfer

Andreas Steiner, André Susano Pinto, Michael Tschannen +15 authors

PaliGemma 2 integrates a SigLIP-So400m vision encoder with Gemma 2 models of varying sizes and resolutions, advancing transfer performance across diverse vision-language tasks, including OCR and captioning.

136SigLIP-So400mvision encoderHF ↗arXiv ↗
25

Chameleon: Mixed-Modal Early-Fusion Foundation Models

Chameleon Team

Chameleon is a mixed-modal early-fusion token-based model that achieves state-of-the-art performance across various tasks, including image captioning, text generation, and long-form mixed-modal generation, using a unified architecture.

135early-fusiontoken-basedHF ↗arXiv ↗
1 / 20

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号