TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Apr 20 – Apr 26, 2026
本周最热254

Tstars-Tryon 1.0: Robust and Realistic Virtual Try-On for Diverse Fashion Items

Mengting Chen, Zhengrui Chen, Yongchao Du +16 authors

A commercial-scale virtual try-on system achieves high success rates, photorealistic results, and real-time performance through integrated system design and multi-stage training.

virtual try-onimage generationimage editingphotorealistic resultsHF ↗arXiv ↗

50 篇论文 · 按点赞排序

02

LLaDA2.0-Uni: Unifying Multimodal Understanding and Generation with Diffusion Large Language Model

Inclusion AI, Tiwei Bie, Haoxing Chen +15 authors

LLaDA2.0-Uni is a unified discrete diffusion language model that integrates multimodal understanding and generation through a semantic discrete tokenizer, MoE-based backbone, and diffusion decoder, achieving performance comparable to specialized vision-language models while enabling efficient inference and high-fidelity image generation.

244discrete diffusionlarge language modelHF ↗arXiv ↗
03

AgentSPEX: An Agent SPecification and EXecution Language

Pengcheng Wang, Jerry Huang, Jiarui Yao +7 authors

AgentSPEX is a domain-specific language and framework for creating structured, modular, and interpretable large language model agent workflows with explicit control flow and state management.

168agent specification and execution languagecontrol flowHF ↗arXiv ↗
08

OpenGame: Open Agentic Coding for Games

Yilei Jiang, Jinyuan Hu, Qianyin Xiao +8 authors

OpenGame is an open-source agentic framework for end-to-end web game creation that uses specialized code models and evaluation benchmarks to overcome challenges in interactive application development.

87Large Language Modelscode agentsHF ↗arXiv ↗
09

Near-Future Policy Optimization

Chuanyu Qin, Chenxu Yang, Qingyi Si +6 authors

Mixed-policy reinforcement learning approach using near-future policy optimization to accelerate convergence and improve performance by balancing trajectory quality and variance.

77reinforcement learningverifiable rewardsHF ↗arXiv ↗
10

Elucidating the SNR-t Bias of Diffusion Probabilistic Models

Meng Yu, Lei Sun, Jianhao Zeng +2 authors

Diffusion probabilistic models suffer from SNR-timestep bias during inference, which is addressed through a differential correction method that processes frequency components separately, improving generation quality across multiple models with minimal computational cost.

72diffusion probabilistic modelsSignal-to-Noise Ratio-timestep biasHF ↗arXiv ↗
12

Qwen3.5-Omni Technical Report

Qwen Team

Qwen3.5-Omni is a large-scale multimodal model with hundreds of billions of parameters that excels in audio-visual understanding and generation, featuring advanced architectures and novel capabilities like Audio-Visual Vibe Coding.

61Hybrid Attention Mixture-of-ExpertsMoEHF ↗arXiv ↗
14

DR-Venus: Towards Frontier Edge-Scale Deep Research Agents with Only 10K Open Data

Venus Team, Sunhao Dai, Yong Deng +10 authors

DR-Venus-4B is a 4-billion-parameter deep research agent trained entirely on open data using agentic supervised fine-tuning and reinforcement learning with turn-level rewards to achieve superior performance on research benchmarks while maintaining edge-scale deployment advantages.

54agentic supervised fine-tuningagentic reinforcement learningHF ↗arXiv ↗
16

MultiWorld: Scalable Multi-Agent Multi-View Video World Models

Haoyu Wu, Jiwen Yu, Yingtian Zou +1 authors

MultiWorld is a unified framework for multi-agent multi-view world modeling that achieves accurate multi-agent control while maintaining multi-view consistency through specialized modules for condition handling and global state encoding.

49video world modelsaction-conditioned video generationHF ↗arXiv ↗
17

PersonaVLM: Long-Term Personalized Multimodal LLMs

Chang Nie, Chaoyou Fu, Yifan Zhang +2 authors

A novel personalized multimodal language model framework called PersonaVLM is introduced that enables long-term personalization through memory retention, multi-turn reasoning, and response alignment capabilities.

46Multimodal Large Language Modelspersonalized multimodal agent frameworkHF ↗arXiv ↗
18

EasyVideoR1: Easier RL for Video Understanding

Chuanyu Qin, Chenxu Yang, Qingyi Si +6 authors

EasyVideoR1 presents an efficient reinforcement learning framework for video understanding that improves training throughput, supports diverse video tasks, and enables joint image-video training with comprehensive evaluation across multiple benchmarks.

42reinforcement learning from verifiable rewardslarge vision-language modelsHF ↗arXiv ↗
25

ShadowPEFT: Shadow Network for Parameter-Efficient Fine-Tuning

Xianming Li, Zongxi Li, Tsz-fung Andrew Lee +3 authors

ShadowPEFT is a parameter-efficient fine-tuning framework that performs layer-level refinement through depth-shared shadow modules, offering competitive performance with reduced computational overhead compared to traditional low-rank adaptation methods.

30parameter-efficient fine-tuninglow-rank adaptationHF ↗arXiv ↗
28

Image Generators are Generalist Vision Learners

Valentin Gabeur, Shangbang Long, Songyou Peng +22 authors

Image generation pretraining enables vision models to develop strong visual understanding capabilities, achieving state-of-the-art performance on diverse vision tasks through lightweight instruction-tuning while maintaining generation abilities.

27generative pretrainingvision modelsHF ↗arXiv ↗
29

PlayCoder: Making LLM-Generated GUI Code Playable

Zhiyuan Peng, Wei Tao, Xin Yin +3 authors

Large language models struggle to generate logically correct GUI applications, prompting the development of PlayEval benchmark and PlayCoder framework that uses multi-agent approaches to improve functional correctness through iterative repair.

26large language modelscode generationHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号