TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

50 篇论文 · 按点赞排序

31

PixelSmile: Toward Fine-Grained Facial Expression Editing

Jiabin Hua, Hengyuan Xu, Aojie Li +4 authors

A diffusion framework called PixelSmile is introduced that disentangles facial expression semantics through symmetric joint training and contrastive learning to enable precise, controllable, and fine-grained expression editing with robust identity preservation.

118diffusion frameworkfacial expression editingHF ↗arXiv ↗
39

SkillNet: Create, Evaluate, and Connect AI Skills

Yuan Liang, Ruobin Zhong, Haoming Xu +46 authors

SkillNet introduces an open infrastructure for systematically accumulating and transferring AI skills through a unified ontology, significantly improving agent performance across multiple domains.

95AI agentsskill consolidationHF ↗arXiv ↗
40

Towards a Medical AI Scientist

Hongtao Wu, Boyun Zheng, Dingjie Song +5 authors

Medical AI Scientist represents the first autonomous research framework designed for clinical applications, enabling evidence-based hypothesis generation and manuscript drafting through clinician-engineer collaboration across three research modes.

94autonomous research frameworkclinical autonomous researchHF ↗arXiv ↗
43

SWE-rebench V2: Language-Agnostic SWE Task Collection at Scale

Ibragim Badertdinov, Maksim Nekrashevich, Anton Shevtsov +1 authors

A large-scale dataset of software engineering tasks spanning multiple programming languages and repositories was created using an automated pipeline that generates executable environments and filters unreliable instances through LLM validation.

92reinforcement learningsoftware engineering agentsHF ↗arXiv ↗
49

UniG2U-Bench: Do Unified Models Advance Multimodal Understanding?

Zimo Wen, Boxiu Li, Wanbo Zhang +11 authors

Unified multimodal models show mixed performance in generation-to-understanding tasks, with specific subtasks benefiting from enhanced spatial and reasoning capabilities while overall performance lags behind specialized vision-language models.

88Unified multimodal modelsVision-Language ModelsHF ↗arXiv ↗
50

Mixture-of-Depths Attention

Lianghui Zhu, Yuxin Fang, Bencheng Liao +10 authors

Scaling depth is a key driver for large language models (LLMs). Yet, as LLMs become deeper, they often suffer from signal degradation: informative features formed in shallow layers are gradually diluted by repeated residual updates, making them harder to recover in deeper layers. We introduce mixture-of-depths attention (MoDA), a mechanism that allows each attention head to attend to sequence KV pairs at the current layer and depth KV pairs from preceding layers. We further describe a hardware-efficient algorithm for MoDA that resolves non-contiguous memory-access patterns, achieving 97.3% of FlashAttention-2's efficiency at a sequence length of 64K. Experiments on 1.5B-parameter models demonstrate that MoDA consistently outperforms strong baselines. Notably, it improves average perplexity by 0.2 across 10 validation benchmarks and increases average performance by 2.11% on 10 downstream tasks, with a negligible 3.7% FLOPs computational overhead. We also find that combining MoDA with post-norm yields better performance than using it with pre-norm. These results suggest that MoDA is a promising primitive for depth scaling. Code is released at https://github.com/hustvl/MoDA .

83HF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号