TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Mar 24 – Mar 30, 2025
本周最热174

Qwen2.5-Omni Technical Report

Jin Xu, Zhifang Guo, Jinzheng He +11 authors

Qwen2.5-Omni is a multimodal model that processes text, images, audio, and video in a streaming fashion and generates text and speech using a dual-track architecture, achieving state-of-the-art performance on multimodal benchmarks.

block-wise processingTMRoPE (Time-aligned Multimodal RoPE)Thinker-Talker architecturedual-track autoregressive modelHF ↗arXiv ↗

50 篇论文 · 按点赞排序

03

Video-T1: Test-Time Scaling for Video Generation

Fangfu Liu, Hanyang Wang, Yimo Cai +3 authors

Test-Time Scaling (TTS) in video generation improves video quality by adaptively sampling from noise space with feedback mechanisms, particularly demonstrated with the Tree-of-Frames method.

90Test-Time Scaling (TTS)video generationHF ↗arXiv ↗
05

Video-R1: Reinforcing Video Reasoning in MLLMs

Kaituo Feng, Kaixiong Gong, Bohao Li +5 authors

Video-R1, leveraging rule-based reinforcement learning and temporal information, enhances video reasoning in multimodal large language models using a combination of video and image data.

79rule-based reinforcement learningRLHF ↗arXiv ↗
08

Wan: Open and Advanced Large-Scale Video Generative Models

WanTeam, Ang Wang, Baole Ai +59 authors

Wan, a comprehensive suite of video foundation models built on the diffusion transformer paradigm, advannces video generation by introducing a novel VAE, scalable pre-training strategies, and large-scale data curation, offering superior performance and versatility across various applications with both large and efficient models.

71diffusion transformerVAEHF ↗arXiv ↗
10

Gemma 3 Technical Report

Gemma Team, Aishwarya Kamath, Johan Ferret +213 authors

Gemma 3 introduces vision capabilities, broader language coverage, and extended context length, featuring an optimized architecture and post-training enhancements to outperform previous versions.

57multimodal modelsvision understandingHF ↗arXiv ↗
13

A Comprehensive Survey on Long Context Language Modeling

Jiaheng Liu, Dawei Zhu, Zhiqi Bai +34 authors

This survey examines recent advancements in long-context language modeling, focusing on methods for obtaining, training, deploying, and evaluating models capable of processing extensive textual inputs.

49Long Context Language Models (LCLMs)long context processingHF ↗arXiv ↗
20

LEGO-Puzzles: How Good Are MLLMs at Multi-Step Spatial Reasoning?

Kexian Tang, Junyao Gao, Yanhong Zeng +6 authors

LEGO-Puzzles evaluates the spatial understanding and sequential reasoning capabilities of Multimodal Large Language Models (MLLMs) through LEGO-based tasks, revealing significant limitations compared to human performance.

35spatial reasoningsequential reasoningHF ↗arXiv ↗
24

Gemini Robotics: Bringing AI into the Physical World

Gemini Robotics Team, Saminda Abeyruwan, Joshua Ainslie +115 authors

Gemini Robotics, an advanced Vision-Language-Action model based on Gemini 2.0, controls robots with multimodal reasoning, handling complex tasks and adapting to new environments and instructions.

31Vision-Language-Actionmultimodal reasoningHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号