Memory in the Age of AI Agents
Yuyang Hu, Shichun Liu, Yanwei Yue +44 authors
This survey provides an updated overview of agent memory research, distinguishing its forms, functions, and dynamics, and highlights emerging research directions.
Explore · 每周精选
发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。
Kling Team, Jialu Chen, Yuanzheng Ci +65 authors
Kling-Omni is an end-to-end generative framework that unifies video generation, editing, and reasoning from multimodal inputs to produce high-fidelity video content.
50 篇论文 · 按点赞排序
Yuyang Hu, Shichun Liu, Yanwei Yue +44 authors
This survey provides an updated overview of agent memory research, distinguishing its forms, functions, and dynamics, and highlights emerging research directions.
Haolong Yan, Jia Wang, Xin Huang +94 authors
A self-evolving training pipeline with the Calibrated Step Reward System and GUI-MCP protocol improve GUI automation efficiency, accuracy, and privacy in real-world scenarios.
Taewoong Kang, Kinam Kim, Dohyeon Kim +3 authors
EgoX framework generates egocentric videos from exocentric inputs using video diffusion models with LoRA adaptation, unified conditioning, and geometry-guided self-attention for coherence and visual fidelity.
Zefan Cai, Haoyi Qiu, Tianyi Ma +9 authors
MMGR evaluates video and image models across reasoning abilities in multiple domains, revealing performance gaps and highlighting limitations in perceptual data and causal correctness.
Weizhou Shen, Ziyi Yang, Chenliang Li +11 authors
QwenLong-L1.5 enhances long-context reasoning through data synthesis, stabilized reinforcement learning, and memory-augmented architecture, achieving superior performance on benchmarks and general domains.
Pengcheng Jiang, Jiacheng Lin, Zhiyi Shi +31 authors
This paper presents a framework for agent and tool adaptation in agentic AI systems, clarifying design strategies and identifying open challenges for improving AI capabilities.
Jingfeng Yao, Yuda Song, Yucong Zhou +1 authors
A unified visual tokenizer pre-training framework (VTP) improves generative performance by optimizing image-text contrastive, self-supervised, and reconstruction losses, leading to better scaling properties and higher zero-shot accuracy and faster convergence.
Jia-Nan Li, Jian Guan, Wei Wu +1 authors
ReFusion, a novel masked diffusion model, improves performance and efficiency by using slot-based parallel decoding, achieving superior results compared to autoregressive models and traditional masked diffusion models.
Sihan Xu, Ziqiao Ma, Wenhao Chai +5 authors
Generative pretraining using next embedding prediction outperforms traditional self-supervised methods in visual learning tasks, achieving high accuracy on ImageNet and effective transfer to semantic segmentation.
Tiwei Bie, Maosong Cao, Kun Chen +28 authors
LLaDA2.0 converts auto-regressive models into discrete diffusion large language models with a novel training scheme, achieving superior performance and efficiency at scale.
Jianxiong Gao, Zhaoxi Chen, Xian Liu +7 authors
LongVie 2, an end-to-end autoregressive framework, enhances controllability, visual quality, and temporal consistency in video world models through three progressive training stages.
Wenqiang Sun, Haiyu Zhang, Haoyuan Wang +7 authors
WorldPlay is a streaming video diffusion model that achieves real-time, interactive world modeling with long-term geometric consistency by using a Dual Action Representation, Reconstituted Context Memory, and Context Forcing.
Shengming Yin, Zekai Zhang, Zecheng Tang +11 authors
A diffusion model for decomposing RGB images into semantically disentangled RGBA layers enables consistent image editing through independent manipulation of individual layers.
Jiaqi Wang, Weijia Wu, Yi Zhan +6 authors
The Video Reality Test benchmark evaluates the realism and detection of AI-generated ASMR videos with audio, revealing that even the best models can deceive VLMs and humans, highlighting limitations in perceptual fidelity and audio-visual consistency.
Haoyu Dong, Pengkun Zhang, Yan Gao +5 authors
A finance and accounting benchmark named Finch is presented, featuring real-world enterprise workflows with complex, multi-modal tasks derived from authentic corporate environments.
Jingzhe Ding, Shengda Long, Changxin Pu +45 authors
NL2Repo Bench evaluates long-horizon software development capabilities of coding agents by assessing their ability to generate complete Python libraries from natural-language requirements.
Mengzhang Cai, Xin Gao, Yu Li +13 authors
OpenDataArena (ODA) is an open platform that benchmarks post-training datasets for Large Language Models (LLMs) using a unified pipeline, multi-dimensional scoring, and data lineage exploration to enhance reproducibility and understanding of data impacts on model behavior.
Zhenyang Cai, Jiaming Zhang, Junjie Zhao +21 authors
DentalGPT, a specialized multimodal large language model for dentistry, achieves superior performance in disease classification and dental visual question answering through high-quality domain data and staged adaptation techniques.
Zicong Cheng, Guo-Wei Yang, Jia Li +3 authors
Diffusion large language model drafters address limitations of autoregressive decoding in speculative frameworks through parallel processing and improved probabilistic modeling, achieving significant speedup in code generation tasks.
Zitian Gao, Lynx Chen, Yihao Xiao +5 authors
The Universal Reasoning Model enhances Universal Transformers with short convolution and truncated backpropagation to improve reasoning performance on ARC-AGI tasks.
Lanxiang Hu, Siqi Kou, Yichao Fu +5 authors
Parallel decoding for transformers is improved through a progressive distillation method that maintains causal inference properties while achieving significant speedup, combined with multi-block decoding that further reduces inference latency.
Kling Team, Jialu Chen, Yikang Ding +25 authors
KlingAvatar 2.0 addresses inefficiencies in generating long-duration, high-resolution videos by using a spatio-temporal cascade framework with a Co-Reasoning Director and Negative Director for improved multimodal instruction alignment.
Jingdi Lei, Di Zhang, Soujanya Poria
Error-Free Linear Attention (EFLA) is a stable, parallelizable, and theoretically sound linear-time attention mechanism that outperforms DeltaNet in language modeling and downstream tasks.
HyperAI Team, Yuchen Liu, Kaiyang Han +26 authors
HyperVL, an efficient multimodal large language model for on-device inference, uses image tiling, Visual Resolution Compressor, and Dual Consistency Learning to reduce memory usage, latency, and power consumption while maintaining performance.
Heyi Chen, Siyan Chen, Xin Chen +193 authors
Seedance 1.5 pro, a dual-branch Diffusion Transformer model, achieves high-quality audio-visual synchronization and generation through cross-modal integration, post-training optimizations, and an acceleration framework.
Yuran Wang, Bohan Zeng, Chengzhuo Tong +6 authors
Scone is a unified method for subject distinction in multi-subject image generation that uses semantic bridging and two-stage training to improve both composition and subject identification.
Zhiyuan Li, Chi-Man Pun, Chen Fang +2 authors
PersonaLive is a diffusion-based portrait animation framework that improves real-time performance through hybrid implicit signals, appearance distillation, and autoregressive streaming generation.
Boxin Wang, Chankyu Lee, Nayeon Lee +9 authors
Cascaded domain-wise reinforcement learning (Cascade RL) is proposed to enhance general-purpose reasoning models, achieving state-of-the-art performance across benchmarks and outperforming the teacher model in coding competitions.
Chun-Wei Tuan Mu, Jia-Bin Huang, Yu-Lun Liu
Generative Refocusing uses semi-supervised learning with DeblurNet and BokehNet to achieve high-quality single-image refocusing with controllable bokeh and text-guided adjustments.
北京市昌平区探索星信息技术及软件开发工作室
京ICP备2026059466号