TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

50 篇论文 · 按点赞排序

33

LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Fuzhao Xue, Yukang Chen, Dacheng Li +15 authors

LongVILA, a full-stack solution for long-context vision-language models, introduces Multi-Modal Sequence Parallelism for efficient training and inference, and a five-stage training pipeline that enhances long video processing and captioning performance.

52Multi-Modal Sequence ParallelismMM-SPHF ↗arXiv ↗
34

Med42-v2: A Suite of Clinical LLMs

Clément Christophe, Praveen K Kanithi, Tathagata Raha +2 authors

Med42-v2 enhances Llama3 with clinical data to improve performance in healthcare settings, outperforming generic models across medical benchmarks.

52large language modelsLLMsHF ↗arXiv ↗
38

VITA: Towards Open-Source Interactive Omni Multimodal LLM

Chaoyou Fu, Haojia Lin, Zuwei Long +12 authors

VITA, an open-source Multimodal Large Language Model, excels in processing Video, Image, Text, and Audio with seamless interaction, showcasing advancements in multimodal understanding and human-computer interaction.

50Multimodal Large Language ModelMixtralHF ↗arXiv ↗
41

Foundation Models for Music: A Survey

Yinghao Ma, Anders Øland, Anton Ragni +40 authors

A review of foundation models in music, including large language models and latent diffusion models, highlights their impact, architectural choices, and the need for ethical considerations in music applications.

44large language modelslatent diffusion modelsHF ↗arXiv ↗
47

Automated Design of Agentic Systems

Shengran Hu, Cong Lu, Jeff Clune

A new research area, Automated Design of Agentic Systems (ADAS), uses a meta agent to automatically create powerful agentic system designs, demonstrating superior performance and robustness across various domains.

40Automated Design of Agentic SystemsADASHF ↗arXiv ↗
49

Language Model Can Listen While Speaking

Ziyang Ma, Yakun Song, Chenpeng Du +5 authors

A novel listening-while-speaking language model (LSLM) enhances real-time, full-duplex speech interaction by integrating speech generation and real-time audio input with an optimal fusion strategy.

40full duplex modelinginteractive speech language modelsHF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号