TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Nov 18 – Nov 24, 2024
本周最热133

LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Guowei Xu, Peng Jin, Ziang Wu +5 authors

LLaVA-CoT is a vision-language model that achieves improved reasoning performance through structured multistage processing and test-time scaling, outperforming larger models with a smaller training dataset.

chain-of-thought promptingvisual question answeringmultistage reasoningstructured reasoning annotationsHF ↗arXiv ↗

50 篇论文 · 按点赞排序

03

Generative World Explorer

Taiming Lu, Tianmin Shu, Alan Yuille +2 authors

Generative World Explorer (Genex) enables agents to mentally explore 3D environments using imagined observations, updating their beliefs without physical exploration to improve decision-making.

77Generative World ExplorerGenexHF ↗arXiv ↗
05

RedPajama: an Open Dataset for Training Large Language Models

Maurice Weber, Daniel Fu, Quentin Anthony +16 authors

The RedPajama datasets are introduced to address core challenges for open-source language models by providing transparent data curation, large volumes of high-quality text, and quality signals for web data analysis.

60decoder-only language modelsHF ↗arXiv ↗
07

Hymba: A Hybrid-head Architecture for Small Language Models

Xin Dong, Yonggan Fu, Shizhe Diao +10 authors

Hymba, a family of small language models with a hybrid-head architecture combining transformer attention and state space models, achieves state-of-the-art performance with improved efficiency and reduced cache size.

50hybrid-head parallel architecturetransformer attention mechanismsHF ↗arXiv ↗
11

VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models

Ziqi Huang, Fan Zhang, Xiaojie Xu +14 authors

VBench++ is a comprehensive benchmark suite for video generation that evaluates models across specific, hierarchical dimensions using fine-grained metrics and human preference annotations, offering insights into model strengths, weaknesses, and gaps compared to image generation models.

33video generation qualitysubject identity inconsistencyHF ↗arXiv ↗
12

The Dawn of GUI Agent: A Preliminary Case Study with Claude 3.5 Computer Use

Siyuan Hu, Mingyu Ouyang, Difei Gao +1 authors

The recently released model, Claude 3.5 Computer Use, stands out as the first frontier AI model to offer computer use in public beta as a graphical user interface (GUI) agent. As an early beta, its capability in the real-world complex environment remains unknown. In this case study to explore Claude 3.5 Computer Use, we curate and organize a collection of carefully designed tasks spanning a variety of domains and software. Observations from these cases demonstrate Claude 3.5 Computer Use's unprecedented ability in end-to-end language to desktop actions. Along with this study, we provide an out-of-the-box agent framework for deploying API-based GUI automation models with easy implementation. Our case studies aim to showcase a groundwork of capabilities and limitations of Claude 3.5 Computer Use with detailed analyses and bring to the fore questions about planning, action, and critic, which must be considered for future improvement. We hope this preliminary exploration will inspire future research into the GUI agent community. All the test cases in the paper can be tried through the project: https://github.com/showlab/computer_use_ootb.

33graphical user interface (GUI)API-based GUI automationHF ↗arXiv ↗
14

Natural Language Reinforcement Learning

Xidong Feng, Ziyu Wan, Haotian Fu +6 authors

Natural Language Reinforcement Learning (NLRL) redefines traditional RL concepts in a language-based framework, leveraging large language models to achieve efficient and interpretable policy improvement.

30Reinforcement LearningNLRLHF ↗arXiv ↗
20

Top-nσ: Not All Logits Are You Need

Chenxia Tang, Jianchun Liu, Hongli Xu +1 authors

A novel sampling method called top-$n\sigma$ for large language models improves reasoning task performance by filtering pre-softmax logits and maintaining consistent results across different temperatures.

24greedy decodinglow-temperature samplingHF ↗arXiv ↗
21

Ultra-Sparse Memory Network

Zihao Huang, Qiyang Min, Hongzhi Huang +4 authors

UltraMem introduces a large-scale, ultra-sparse memory layer to Transformer models, reducing inference latency and improving performance.

23Mixture of ExpertsTransformer modelsHF ↗arXiv ↗
22

Stable Flow: Vital Layers for Training-Free Image Editing

Omri Avrahami, Or Patashnik, Ohad Fried +4 authors

The work presents an automatic method to identify crucial layers in Diffusion Transformer models for consistent image editing and demonstrates its effectiveness through qualitative and quantitative evaluations.

22Diffusion Transformer (DiT)flow-matchingHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号