TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Aug 31 – Sep 6, 2026

50 篇论文 · 按点赞排序

32

Normalized Low-Rank Adaptation

Jiale Kang, Ziyin Yue, Zheng Zhan +2 authors

Normalized Low-Rank Adaptation stabilizes LoRA training by normalizing down-projection matrices, accelerating convergence and improving performance without extra parameters or inference cost.

52low-rank adaptationLoRAHF ↗arXiv ↗
34

H3-World: Turning Language Understanding into World Control

Danze Chen, Zeqing Wang, Ziyue Lin +2 authors

We present H3-World, an efficient framework that turns the 33B MiniMax-H3 video generator into an interactive world model. Our key finding is that, as large video generators become more capable, language is emerging as a natural interface for control. MiniMax-H3, for example, already supports zero-shot control of character behavior and camera motion through natural-language instructions. Building on this, H3-World turns this coarse language interface into precise, temporally grounded world control, without introducing dedicated action modules. Specifically, we represent each action as a structured combination of character and camera instructions, and align them with the corresponding temporal video latents. To make the control temporally precise, we further introduce temporal attention routing, which restricts each instruction to its intended time interval and reduces control leakage across actions. Importantly, H3-World directly reuses the semantic representations learned during large-scale video pretraining and requires only lightweight adaptation. With only 8,000 gameplay samples, 10,000 LoRA optimization steps, and 0.199% trainable parameters, H3-World achieves effective character and camera control while preserving strong generation quality. It also generalizes to unseen scenarios. These results show that the control capabilities emerging in large video generators can be efficiently transformed into interactive world control.

50HF ↗arXiv ↗
44

Revisiting Local Context for Long-Horizon Streaming 3D Reconstruction

Jiarong Han, Jincheng Xiong, Yuzhou Liu +6 authors

ABot-Recon achieves stable long-horizon streaming 3D reconstruction by using only local temporal context and frame-independent predictions composed sequentially, reducing drift via a lightweight temporal refiner and composition-aware pose loss.

33streaming 3D reconstructionKV featuresHF ↗arXiv ↗
45

On the Design Fundamentals of Pixel Text Representation Learning

Chaohao Yuan, Ruifeng Yuan, Zhuoxu Huang +4 authors

Pixel Linguist II improves visual text encoding through variable resolution training, natural image-text grounding, layout-aware rendering, and multilingual curricula, achieving state-of-the-art results and strong compression robustness.

32visual text representation learningvariable image resolutionsHF ↗arXiv ↗
48

Fast Weight Attention for Continual Learning

Yifan Zhang, Steve Ta, Jasper Zhang +8 authors

Recurrent fast-weight memories and selective state-space models are analyzed as online learning rules under autoregressive semantics, yielding normalized update families with stable renormalization that improve length extrapolation.

32fast-weight memoriesselective state-space modelsHF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号