TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

本月最热108

BitNet: Scaling 1-bit Transformers for Large Language Models

Hongyu Wang, Shuming Ma, Li Dong +7 authors

BitNet, a 1-bit Transformer architecture, reduces memory and energy consumption while achieving competitive performance in language modeling.

BitNet1-bit TransformerBitLinearnn.LinearHF ↗arXiv ↗

50 篇论文 · 按点赞排序

07

Llemma: An Open Language Model For Mathematics

Zhangir Azerbayev, Hailey Schoelkopf, Keiran Paster +6 authors

Llemma, a large language model pretrained on mathematical data, outperforms existing models and demonstrates tool use and formal theorem proving capabilities.

57large language modelmathematicsHF ↗arXiv ↗
10

Matryoshka Diffusion Models

Jiatao Gu, Shuangfei Zhai, Yizhe Zhang +2 authors

Matryoshka Diffusion Models use a NestedUNet architecture for joint denoising at multiple resolutions, enabling efficient high-resolution image and video synthesis.

46diffusion modelshigh-resolution image and video synthesisHF ↗arXiv ↗
11

In-Context Learning Creates Task Vectors

Roee Hendel, Mor Geva, Amir Globerson

In-Context Learning in Large Language Models can be understood as compressing a training set into a task vector that modulates a transformer for output generation.

43in-context learninglarge language modelsHF ↗arXiv ↗
12

4K4D: Real-Time 4D View Synthesis at 4K Resolution

Zhen Xu, Sida Peng, Haotong Lin +5 authors

4K4D, a 4D point cloud representation with a hybrid appearance model and differentiable depth peeling algorithm, achieves high-speed and high-quality dynamic view synthesis at 4K resolution.

404D point cloud representationhardware rasterizationHF ↗arXiv ↗
13

Table-GPT: Table-tuned GPT for Diverse Table Tasks

Peng Li, Yeye He, Dror Yashar +6 authors

A new table-tuning paradigm improves language models' understanding and generalization in table-related tasks by fine-tuning them with synthesized table data, resulting in enhanced performance.

40table-tuningtable-understandingHF ↗arXiv ↗
19

JudgeLM: Fine-tuned Large Language Models are Scalable Judges

Lianghui Zhu, Xinggang Wang, Xinlong Wang

Large Language Models fine-tuned as scalable judges (JudgeLM) achieve state-of-the-art performance in evaluating open-ended benchmarks through a comprehensive dataset and benchmark, enhancing judgment efficiency and accuracy.

35Large Language Modelsfine-tuningHF ↗arXiv ↗
21

FP8-LM: Training FP8 Large Language Models

Houwen Peng, Kan Wu, Yixuan Wei +17 authors

A new FP8 automatic mixed-precision framework for training large language models reduces memory usage and increases speed compared to BF16 and Nvidia Transformer Engine.

34FP8low-bit data formatsHF ↗arXiv ↗
25

VeRA: Vector-based Random Matrix Adaptation

Dawid Jan Kopiczko, Tijmen Blankevoort, Yuki Markus Asano

Vector-based Random Matrix Adaptation (VeRA) reduces the number of trainable parameters by 10x compared to LoRA while maintaining performance, and is demonstrated on benchmarks like GLUE and E2E, showing its utility in instruction-following.

31Low-rank adaptationLoRAHF ↗arXiv ↗
30

PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Xi Chen, Xiao Wang, Lucas Beyer +16 authors

PaLI-3, a smaller vision language model, outperforms larger models on multimodal tasks, particularly localization and text understanding, using SigLIP pretraining and achieves state-of-the-art results in multilingual cross-modal retrieval.

29Vision TransformerSigLIPHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号