TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

598 篇论文 · 按点赞排序

218

LLaVA-OneVision: Easy Visual Task Transfer

Bo Li, Yuanhan Zhang, Dong Guo +7 authors

LLaVA-OneVision is a unified multimodal model that advances performance across single-image, multi-image, and video scenarios with strong transfer learning capabilities.

61multimodal modelsLMMsHF ↗arXiv ↗
219

Multi-Head Mixture-of-Experts

Xun Wu, Shaohan Huang, Wenhui Wang +1 authors

MH-MoE enhances SMoE by splitting tokens into sub-tokens processed by diverse experts in parallel, improving expert activation, context understanding, and overfitting, and demonstrating effectiveness across various modeling tasks.

61Sparse Mixtures of Experts (SMoE)Multi-Head Mixture-of-Experts (MH-MoE)HF ↗arXiv ↗
227

Hermes 3 Technical Report

Ryan Teknium, Jeffrey Quesnelle, Chen Guang

Hermes 3, a neutrally-aligned instruct and tool use model with strong reasoning and creative capabilities, achieves top performance on public benchmarks.

60instruct-tuned modelslarge language modelsHF ↗arXiv ↗
232

MambaByte: Token-free Selective State Space Model

Junxiong Wang, Tushaar Gangavarapu, Jing Nathan Yan +1 authors

MambaByte, a token-free byte-level autoregressive model, demonstrates computational efficiency and competitive performance compared to subword token-based models, with the added benefit of fast inference due to linear length scaling.

59MambaBytetoken-free language modelsHF ↗arXiv ↗
236

RedPajama: an Open Dataset for Training Large Language Models

Maurice Weber, Daniel Fu, Quentin Anthony +16 authors

The RedPajama datasets are introduced to address core challenges for open-source language models by providing transparent data curation, large volumes of high-quality text, and quality signals for web data analysis.

59decoder-only language modelsHF ↗arXiv ↗
240

Evolutionary Optimization of Model Merging Recipes

Takuya Akiba, Makoto Shing, Yujin Tang +2 authors

An evolutionary algorithm automates model merging, discovering effective combinations of open-source models without additional training data or compute, achieving state-of-the-art performance on benchmarks and culturally-aware tasks.

58evolutionary algorithmsfoundation modelsHF ↗arXiv ↗
8 / 20

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号