TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

606 篇论文 · 按点赞排序

481

Sekai: A Video Dataset towards World Exploration

Zhen Li, Chuanhao Li, Xiaofeng Mao +17 authors

Sekai, a worldwide video dataset with comprehensive annotations, is introduced to support world exploration applications, enhancing video generation models.

67first-person viewworldwide video datasetHF ↗arXiv ↗
482

π^3: Scalable Permutation-Equivariant Visual Geometry Learning

Yifan Wang, Jianjun Zhou, Haoyi Zhu +7 authors

A permutation-equivariant neural network, $\pi^3$, reconstructs visual geometry without a fixed reference view, achieving state-of-the-art performance in camera pose estimation, depth estimation, and point map reconstruction.

67feed-forward neural networkpermutation-equivariant architectureHF ↗arXiv ↗
484

Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Zhenyi Liao, Qingsong Xie, Yanhao Zhang +4 authors

The study enhances visual-spatial reasoning in multi-modal large language models through GRPO training using the VSI-100k dataset, demonstrating significant performance improvements over base models.

67multi-modal large language modelsvisual-spatial intelligenceHF ↗arXiv ↗
487

π_RL: Online RL Fine-tuning for Flow-based Vision-Language-Action Models

Kang Chen, Zhihao Liu, Tonghe Zhang +10 authors

The framework π<sub>RL</sub> uses reinforcement learning to train flow-based Vision-Language-Action models, addressing challenges with intractable action log-likelihoods and achieving significant performance improvements over supervised fine-tuning.

66reinforcement learningsupervised fine-tuningHF ↗arXiv ↗
489

Describe Anything: Detailed Localized Image and Video Captioning

Long Lian, Yifan Ding, Yunhao Ge +8 authors

The Describe Anything Model (DAM) leverages a focal prompt and localized vision backbone to achieve detailed localized captioning, outperforming existing models on various benchmarks through a semi-supervised data pipeline.

66focal promptlocalized vision backboneHF ↗arXiv ↗
491

Kanana: Compute-efficient Bilingual Language Models

Kanana LLM Team, Yunju Bak, Hojin Lee +26 authors

Kanana, a series of bilingual language models, achieves superior performance in Korean and competitive performance in English with lower computational costs through efficient pre-training and post-training techniques.

66high quality data filteringstaged pre-trainingHF ↗arXiv ↗
495

LFM2 Technical Report

Alexander Amini, Anna Banaszak, Harold Benoit +30 authors

LFM2, a family of compact foundation models, achieves high efficiency and performance on-device through hardware-in-the-loop architecture search and advanced training techniques, supporting various tasks including multimodal applications.

65Liquid Foundation Modelshardware-in-the-loop architecture searchHF ↗arXiv ↗
497

Wan: Open and Advanced Large-Scale Video Generative Models

WanTeam, Ang Wang, Baole Ai +59 authors

Wan, a comprehensive suite of video foundation models built on the diffusion transformer paradigm, advannces video generation by introducing a novel VAE, scalable pre-training strategies, and large-scale data curation, offering superior performance and versatility across various applications with both large and efficient models.

65diffusion transformerVAEHF ↗arXiv ↗
503

BANG: Dividing 3D Assets via Generative Exploded Dynamics

Longwen Zhang, Qixuan Zhang, Haoran Jiang +4 authors

BANG is a generative approach using latent diffusion models and temporal attention to enable intuitive, part-level decomposition of 3D objects with precise control and multimodal interaction.

65latent diffusion modelexploded dynamicsHF ↗arXiv ↗
506

Scaling Test-time Compute for LLM Agents

King Zhu, Hanhao Li, Siwei Wu +12 authors

Systematic exploration of test-time scaling methods in large language agents reveals that computational scaling improves performance, especially through parallel sampling, sequential revision, effective verification, and increased rollout diversity.

64parallel sampling algorithmssequential revision strategiesHF ↗arXiv ↗
508

An Empirical Study of GPT-4o Image Generation Capabilities

Sixiang Chen, Jinbin Bai, Zhuoran Zhao +16 authors

An empirical study of GPT-4o's image generation capabilities across multiple tasks reveals its strengths and limitations compared to other models, highlighting the importance of architectural design and data scaling in unified generative frameworks.

64GANdiffusion modelsHF ↗arXiv ↗
17 / 21

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号