TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

November 2024
本月最热133

LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Guowei Xu, Peng Jin, Ziang Wu +5 authors

LLaVA-CoT is a vision-language model that achieves improved reasoning performance through structured multistage processing and test-time scaling, outperforming larger models with a smaller training dataset.

chain-of-thought promptingvisual question answeringmultistage reasoningstructured reasoning annotationsHF ↗arXiv ↗

50 篇论文 · 按点赞排序

03

ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Kevin Qinghong Lin, Linjie Li, Difei Gao +6 authors

ShowUI is a vision-language-action model that enhances GUI assistants by using UI-guided token selection and interleaved vision-language-action streaming, achieving high accuracy and efficiency in zero-shot screenshot grounding across different environments.

90vision-language-action modelUI-Guided Visual Token SelectionHF ↗arXiv ↗
08

LLaMA-Mesh: Unifying 3D Mesh Generation with Language Models

Zhengyi Wang, Jonathan Lorraine, Yikai Wang +4 authors

The work demonstrates the capability of LLMs to generate 3D meshes from text by introducing a novel approach to tokenize 3D mesh data, allowing the unification of 3D and text modalities without expanding the model's vocabulary.

78large language modelsLLMsHF ↗arXiv ↗
09

Generative World Explorer

Taiming Lu, Tianmin Shu, Alan Yuille +2 authors

Generative World Explorer (Genex) enables agents to mentally explore 3D environments using imagined observations, updating their beliefs without physical exploration to improve decision-making.

77Generative World ExplorerGenexHF ↗arXiv ↗
12

BitNet a4.8: 4-bit Activations for 1-bit LLMs

Hongyu Wang, Shuming Ma, Furu Wei

BitNet a4.8 enhances the efficiency of large language models through 4-bit quantization and sparsification, achieving equivalent performance to BitNet b1.58 with reduced inference costs.

701-bit Large Language Models (LLMs)BitNet b1.58HF ↗arXiv ↗
14

TÜLU 3: Pushing Frontiers in Open Language Model Post-Training

Nathan Lambert, Jacob Morrison, Valentina Pyatkin +20 authors

T\"ULU 3, an open-source family of post-trained language models, introduces transparent training data, recipes, and advanced techniques to match or surpass proprietary models in performance.

68supervised finetuning (SFT)Direct Preference Optimization (DPO)HF ↗arXiv ↗
21

RedPajama: an Open Dataset for Training Large Language Models

Maurice Weber, Daniel Fu, Quentin Anthony +16 authors

The RedPajama datasets are introduced to address core challenges for open-source language models by providing transparent data curation, large volumes of high-quality text, and quality signals for web data analysis.

60decoder-only language modelsHF ↗arXiv ↗
24

Star Attention: Efficient LLM Inference over Long Sequences

Shantanu Acharya, Fei Jia, Boris Ginsburg

Star Attention improves inference efficiency of large language models on long sequences by using block-sparse approximation, reducing memory and time without significant accuracy loss.

53Transformer-based Large Language Models (LLMs)self-attention mechanismHF ↗arXiv ↗
28

Hymba: A Hybrid-head Architecture for Small Language Models

Xin Dong, Yonggan Fu, Shizhe Diao +10 authors

Hymba, a family of small language models with a hybrid-head architecture combining transformer attention and state space models, achieves state-of-the-art performance with improved efficiency and reduced cache size.

50hybrid-head parallel architecturetransformer attention mechanismsHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号