TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

May 25 – May 31, 2026

50 篇论文 · 按点赞排序

32

How LoRA Remembers? A Parametric Memory Law for LLM Finetuning

Ziwen Xu, Haiwen Hong, Linsong Yu +4 authors

Research investigates the quantitative limits of parametric memory in large language models using LoRA as a probe, establishing a power law relationship and developing a threshold-guided optimization method for improved memory performance.

45Large Language ModelsLow-Rank AdaptationHF ↗arXiv ↗
33

Toward Native Multimodal Modeling: A Roadmap

Siyu An, Junru Lu, Junnan Dong +18 authors

Native multimodal modeling advances beyond traditional fusion approaches by integrating modalities inherently within a unified transformer framework, enabling seamless understanding and generation across diverse input-output configurations.

44multimodal modelingnative multimodal modelingHF ↗arXiv ↗
34

GenClaw: Code-Driven Agentic Image Generation

Junyan Ye, Jun He, Zilong Huang +4 authors

GenClaw presents a code-driven agentic image generation framework that enables precise visual construction through conceptualization, sketching, and coloring stages, integrating programmatic logic with generative models.

42image generation modelsmultimodal agentsHF ↗arXiv ↗
39

ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence

Rui Meng, Bhavana Dalvi Mishra, Jiefeng Chen +10 authors

Autonomous research agents exhibit verifiability issues like fabricated citations and unreproducible results, which are addressed through a framework ensuring evidence traceability and an end-to-end system maintaining integrity throughout research processes.

38Chain-of-EvidenceScientistOneHF ↗arXiv ↗
41

Native Audio-Visual Alignment for Generation

Longbin Ji, Guan Wang, Xuan Wei +6 authors

NAVA enables joint audio-video generation with improved synchronization and controllability through native audio-visual alignment and context-conditioned denoising.

36joint audio-video generationdual-tower designsHF ↗arXiv ↗
45

Soap2Soap: Long Cinematic Video Remaking via Multi-Agent Collaboration

Yiren Song, Huilin Zhong, Kevin Qinghong Lin +2 authors

A multi-agent framework called Soap2Soap is presented for long-horizon video-to-video generation that maintains narrative structure and character identity across extended sequences through consistent semantic backbone and visual reference anchors.

35video-to-video generationcinematic remakingHF ↗arXiv ↗
47

GEM: Generative Supervision Helps Embodied Intelligence

Ruowen Zhao, Bangguo Li, Zuyan Liu +9 authors

GEM is a vision-language model that integrates depth map generation during pre-training to improve embodied intelligence and physical operation capabilities in robotics.

33Embodied Vision-Language ModelsVision-Language-Action frameworksHF ↗arXiv ↗
49

JLT: Clean-Latent Prediction in Latent Diffusion Transformers

Funing Fu, Tenghui Wang, Junyong Cen +2 authors

Latent diffusion models using clean-data prediction outperform velocity prediction in compressed representations, demonstrating that prediction targets are geometrically dependent rather than algebraically interchangeable.

33flow matchingclean-data predictionHF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号