TensorX
返回文献探索

Paper · arXiv 2401.02072

ICE-GRT: Instruction Context Enhancement by Generative Reinforcement based Transformers

Chen Zheng, Ke Sun, Da Tang, Yukun Ma, Yuyu Zhang, Chenguang Xi, Xun Zhou

10 upvotesJanuary 4, 2024arXiv 预印本
AI 摘要

ICE-GRT, a Reinforcement Learning from Human Feedback approach using Proximal Policy Optimization, enhances domain-specific performance while maintaining general task capabilities in Large Language Models.

Large Language ModelsLLMsICE-GRTReinforcement Learning from Human FeedbackRLHFProximal Policy OptimizationPPOin-domain scenariosSupervised Fine-Tuningdomain-specific tasksgeneral Language tasksAppropriate DataReward Size ScalingKL-ControlAdvantage Normalization

Abstract

The emergence of Large Language Models (LLMs) such as ChatGPT and LLaMA encounter limitations in domain-specific tasks, with these models often lacking depth and accuracy in specialized areas, and exhibiting a decrease in general capabilities when fine-tuned, particularly analysis ability in small sized models. To address these gaps, we introduce ICE-GRT, utilizing Reinforcement Learning from Human Feedback (RLHF) grounded in Proximal Policy Optimization (PPO), demonstrating remarkable ability in in-domain scenarios without compromising general task performance. Our exploration of ICE-GRT highlights its understanding and reasoning ability to not only generate robust answers but also to provide detailed analyses of the reasons behind the answer. This capability marks a significant progression beyond the scope of Supervised Fine-Tuning models. The success of ICE-GRT is dependent on several crucial factors, including Appropriate Data, Reward Size Scaling, KL-Control, Advantage Normalization, etc. The ICE-GRT model exhibits state-of-the-art performance in domain-specific tasks and across 12 general Language tasks against equivalent size and even larger size LLMs, highlighting the effectiveness of our approach. We provide a comprehensive analysis of the ICE-GRT, underscoring the significant advancements it brings to the field of LLM.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
ICE-GRT: Instruction Context Enhancement by Generative Reinforcement based Transformers | TensorX