TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

598 篇论文 · 按点赞排序

483

VILA^2: VILA Augmented VILA

Yunhao Fang, Ligeng Zhu, Yao Lu +6 authors

A novel data augmentation approach iteratively improves visual language model data quality and performance using self-augmentation and specialist-augmentation, leading to state-of-the-art results on MMMU tasks.

41visual language modelslarge language modelsHF ↗arXiv ↗
487

Automated Design of Agentic Systems

Shengran Hu, Cong Lu, Jeff Clune

A new research area, Automated Design of Agentic Systems (ADAS), uses a meta agent to automatically create powerful agentic system designs, demonstrating superior performance and robustness across various domains.

40Automated Design of Agentic SystemsADASHF ↗arXiv ↗
488

VMamba: Visual State Space Model

Yue Liu, Yunjie Tian, Yuzhong Zhao +5 authors

VMamba, a novel state space model architecture, combines global receptive fields and dynamic weights from ViTs with linear complexity, outperforming established models as image resolution increases.

40Convolutional Neural NetworksVision TransformersHF ↗arXiv ↗
489

Style-Friendly SNR Sampler for Style-Driven Generation

Jooyoung Choi, Chaehun Shin, Yeongtak Oh +2 authors

The Style-friendly SNR sampler modifies the noise level distribution during fine-tuning to improve style alignment in diffusion models, enabling better capture of unique artistic styles.

40diffusion modelssignal-to-noise ratio (SNR)HF ↗arXiv ↗
492

Not All Language Model Features Are Linear

Joshua Engels, Isaac Liao, Eric J. Michaud +2 authors

Research explores multi-dimensional features in language models, discovering interpretable circular representations in GPT-2, Mistral 7B, and Llama 3 8B, which are used for modular arithmetic tasks.

40linear representation hypothesismulti-dimensional featuresHF ↗arXiv ↗
494

Language Model Can Listen While Speaking

Ziyang Ma, Yakun Song, Chenpeng Du +5 authors

A novel listening-while-speaking language model (LSLM) enhances real-time, full-duplex speech interaction by integrating speech generation and real-time audio input with an optimal fusion strategy.

40full duplex modelinginteractive speech language modelsHF ↗arXiv ↗
499

HARE: HumAn pRiors, a key to small language model Efficiency

Lingyun Zhang, Bin jin, Gaojian Ge +7 authors

A principle for leveraging human priors in data construction is proposed to improve small language models, demonstrating favorable performance on large benchmarks in resource-constrained settings.

40large language modelssmall language modelsHF ↗arXiv ↗
504

Stronger Models are NOT Stronger Teachers for Instruction Tuning

Zhangchen Xu, Fengqing Jiang, Luyao Niu +2 authors

The Larger Models' Paradox reveals that larger models are not always better teachers for fine-tuning smaller models, and a new metric, Compatibility-Adjusted Reward (CAR), is introduced to measure and improve the effectiveness of response generators.

39instruction tuninglarge language models (LLMs)HF ↗arXiv ↗
508

FuseChat: Knowledge Fusion of Chat Models

Fanqi Wan, Ziyi Yang, Longguang Zhong +3 authors

FuseChat extends the FuseLLM framework for knowledge fusion of chat LLMs through lightweight fine-tuning and parameter merging, achieving superior performance across various domains.

39knowledge fusionfine-tuningHF ↗arXiv ↗
17 / 20

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号