TensorX
返回文献探索

Paper · arXiv 2505.02156

Think on your Feet: Adaptive Thinking via Reinforcement Learning for Social Agents

Minzheng Wang, Yongbin Li, Haobo Wang, Xinghua Zhang, Nan Xu, Bingli Wu, Fei Huang, Haiyang Yu, Wenji Mao

18 upvotesMay 4, 2025arXiv 预印本
AI 摘要

A novel adaptive mode learning framework significantly enhances social intelligence simulation by dynamically selecting reasoning modes based on context, reducing token usage and improving performance compared to existing methods.

adaptive mode learningthinking modesdeep contemplationadaptive mode policy optimizationmulti-granular designcontext-aware switchingtoken-efficient reasoningdepth-adaptive processingsocial intelligence tasks

Abstract

Effective social intelligence simulation requires language agents to dynamically adjust reasoning depth, a capability notably absent in current approaches. While existing methods either lack this kind of reasoning capability or enforce uniform long chain-of-thought reasoning across all scenarios, resulting in excessive token usage and inappropriate social simulation. In this paper, we propose Adaptive Mode Learning (AML) that strategically selects from four thinking modes (intuitive reaction rightarrow deep contemplation) based on real-time context. Our framework's core innovation, the Adaptive Mode Policy Optimization (AMPO) algorithm, introduces three key advancements over existing methods: (1) Multi-granular thinking mode design, (2) Context-aware mode switching across social interaction, and (3) Token-efficient reasoning via depth-adaptive processing. Extensive experiments on social intelligence tasks confirm that AML achieves 15.6% higher task performance than state-of-the-art methods. Notably, our method outperforms GRPO by 7.0% with 32.8% shorter reasoning chains. These results demonstrate that context-sensitive thinking mode selection, as implemented in AMPO, enables more human-like adaptive reasoning than GRPO's fixed-depth approach

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Think on your Feet: Adaptive Thinking via Reinforcement Learning for Social Agents | TensorX