TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Dec 9 – Dec 15, 2024
本周最热162

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Zhe Chen, Weiyun Wang, Yue Cao +37 authors

InternVL 2.5, an advanced multimodal large language model, showcases competitive performance across various benchmarks, including multimodal reasoning and understanding, and is the first open-source model to surpass 70% on the MMMU benchmark using Chain-of-Thought reasoning.

multimodal large language modelvision encoderslanguage modelsdataset sizesHF ↗arXiv ↗

50 篇论文 · 按点赞排序

02

Phi-4 Technical Report

Marah Abdin, Jyoti Aneja, Harkirat Behl +24 authors

A 14-billion parameter language model surpasses its teacher model in STEM-focused QA through strategic use of synthetic data, improved data quality, and enhanced training techniques.

122training recipedata qualityHF ↗arXiv ↗
05

ProcessBench: Identifying Process Errors in Mathematical Reasoning

Chujie Zheng, Zhenru Zhang, Beichen Zhang +6 authors

ProcessBench evaluates models' ability to identify errors in mathematical reasoning steps, showing that existing process reward models struggle with difficult problems and underperform compared to critic models and a fine-tuned PRM.

86ProcessBenchprocess reward modelsHF ↗arXiv ↗
06

STIV: Scalable Text and Image Conditioned Video Generation

Zongyu Lin, Wei Liu, Chen Chen +14 authors

STIV, a text-image-conditioned video generation method integrating Diffusion Transformer and classifier-free guidance, achieves state-of-the-art performance in text-to-video, text-image-to-video, and image-to-video tasks.

74video generationmodel architecturesHF ↗arXiv ↗
10

Multimodal Latent Language Modeling with Next-Token Diffusion

Yutao Sun, Hangbo Bao, Wenhui Wang +5 authors

LatentLM integrates continuous and discrete data using causal Transformers, VAEs, and next-token diffusion, achieving superior performance in multimodal tasks like image generation, large language model integration, and text-to-speech synthesis.

49LatentLMcausal TransformersHF ↗arXiv ↗
11

Lyra: An Efficient and Speech-Centric Framework for Omni-Cognition

Zhisheng Zhong, Chengyao Wang, Yuqi Liu +12 authors

As Multi-modal Large Language Models (MLLMs) evolve, expanding beyond single-domain capabilities is essential to meet the demands for more versatile and efficient AI. However, previous omni-models have insufficiently explored speech, neglecting its integration with multi-modality. We introduce Lyra, an efficient MLLM that enhances multimodal abilities, including advanced long-speech comprehension, sound understanding, cross-modality efficiency, and seamless speech interaction. To achieve efficiency and speech-centric capabilities, Lyra employs three strategies: (1) leveraging existing open-source large models and a proposed multi-modality LoRA to reduce training costs and data requirements; (2) using a latent multi-modality regularizer and extractor to strengthen the relationship between speech and other modalities, thereby enhancing model performance; and (3) constructing a high-quality, extensive dataset that includes 1.5M multi-modal (language, vision, audio) data samples and 12K long speech samples, enabling Lyra to handle complex long speech inputs and achieve more robust omni-cognition. Compared to other omni-methods, Lyra achieves state-of-the-art performance on various vision-language, vision-speech, and speech-language benchmarks, while also using fewer computational resources and less training data.

48Multi-modal Large Language Models (MLLMs)omni-modelsHF ↗arXiv ↗
13

Evaluating and Aligning CodeLLMs on Human Preference

Jian Yang, Jiaxi Yang, Ke Jin +7 authors

A human-curated benchmark (CodeArena) and a large synthetic instruction corpus (SynCode-Instruct) are introduced to evaluate code LLMs based on human preference alignment, revealing performance differences between open-source and proprietary models.

48code large language modelscode generationHF ↗arXiv ↗
27

Maya: An Instruction Finetuned Multilingual Multimodal Model

Nahid Alam, Karthik Reddy Kanjula, Surya Guthikonda +16 authors

Maya is an open-source multimodal multilingual model that addresses the limitations of Vision-Language Models in handling low-resource languages and cultural nuances by introducing a multilingual pretraining dataset and analyzing/removing toxicity.

27Vision-Language ModelsVLMsHF ↗arXiv ↗
28

The BrowserGym Ecosystem for Web Agent Research

Thibault Le Sellier De Chezelles, Maxime Gasse, Alexandre Drouin +17 authors

BrowserGym provides a standardized environment for evaluating web agents using LLMs, facilitating consistent benchmarking and improving agent development and analysis.

24Large Language Modelsgym-like environmentHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号