TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Oct 16 – Oct 22, 2023
本周最热108

BitNet: Scaling 1-bit Transformers for Large Language Models

Hongyu Wang, Shuming Ma, Li Dong +7 authors

BitNet, a 1-bit Transformer architecture, reduces memory and energy consumption while achieving competitive performance in language modeling.

BitNet1-bit TransformerBitLinearnn.LinearHF ↗arXiv ↗

42 篇论文 · 按点赞排序

04

Llemma: An Open Language Model For Mathematics

Zhangir Azerbayev, Hailey Schoelkopf, Keiran Paster +6 authors

Llemma, a large language model pretrained on mathematical data, outperforms existing models and demonstrates tool use and formal theorem proving capabilities.

56large language modelmathematicsHF ↗arXiv ↗
05

Table-GPT: Table-tuned GPT for Diverse Table Tasks

Peng Li, Yeye He, Dror Yashar +6 authors

A new table-tuning paradigm improves language models' understanding and generalization in table-related tasks by fine-tuning them with synthesized table data, resulting in enhanced performance.

40table-tuningtable-understandingHF ↗arXiv ↗
06

4K4D: Real-Time 4D View Synthesis at 4K Resolution

Zhen Xu, Sida Peng, Haotong Lin +5 authors

4K4D, a 4D point cloud representation with a hybrid appearance model and differentiable depth peeling algorithm, achieves high-speed and high-quality dynamic view synthesis at 4K resolution.

394D point cloud representationhardware rasterizationHF ↗arXiv ↗
08

VeRA: Vector-based Random Matrix Adaptation

Dawid Jan Kopiczko, Tijmen Blankevoort, Yuki Markus Asano

Vector-based Random Matrix Adaptation (VeRA) reduces the number of trainable parameters by 10x compared to LoRA while maintaining performance, and is demonstrated on benchmarks like GLUE and E2E, showing its utility in instruction-following.

31Low-rank adaptationLoRAHF ↗arXiv ↗
12

PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Xi Chen, Xiao Wang, Lucas Beyer +16 authors

PaLI-3, a smaller vision language model, outperforms larger models on multimodal tasks, particularly localization and text understanding, using SigLIP pretraining and achieves state-of-the-art results in multilingual cross-modal retrieval.

29Vision TransformerSigLIPHF ↗arXiv ↗
19

Context-Aware Meta-Learning

Christopher Fifty, Dennis Duan, Ronald G. Junkins +4 authors

A meta-learning algorithm learns new visual concepts during inference without fine-tuning, leveraging a pre-trained feature extractor and sequence modeling.

17meta-learningLarge Language ModelsHF ↗arXiv ↗
21

AutoMix: Automatically Mixing Language Models

Aman Madaan, Pranjal Aggarwal, Ankit Anand +10 authors

AutoMix strategically routes queries to larger LLMs based on a noise-refined self-verification mechanism, improving the incremental benefit per cost.

14AutoMixfew-shot self-verificationHF ↗arXiv ↗
27

Interactive Task Planning with Language Models

Boyi Li, Philipp Wu, Pieter Abbeel +1 authors

A framework uses language models for interactive task planning, allowing generalization to new goals with simple task guidelines and precise replanning.

12interactive task planninglanguage modelsHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号