TensorX
返回文献探索

Paper · arXiv 2405.01481

NeMo-Aligner: Scalable Toolkit for Efficient Model Alignment

Gerald Shen, Zhilin Wang, Olivier Delalleau, Jiaqi Zeng, Yi Dong, Daniel Egert, Shengyang Sun, Jimmy Zhang, Sahil Jain, Ali Taghibakhshi, Markel Sanz Ausin, Ashwath Aithal, Oleksii Kuchaiev

30 upvotesMay 2, 2024arXiv 预印本
AI 摘要

NeMo-Aligner is a toolkit for aligning large language models with human values using scalable techniques like RLHF, DPO, SteerLM, and SPIN, optimized for use with hundreds of GPUs and parameter-efficient fine-tuning.

NeMo-AlignerReinforcement Learning from Human Feedback (RLHF)Direct Preference Optimization (DPO)SteerLMSelf-Play Fine-Tuning (SPIN)Parameter Efficient Fine-Tuning (PEFT)

Abstract

Aligning Large Language Models (LLMs) with human values and preferences is essential for making them helpful and safe. However, building efficient tools to perform alignment can be challenging, especially for the largest and most competent LLMs which often contain tens or hundreds of billions of parameters. We create NeMo-Aligner, a toolkit for model alignment that can efficiently scale to using hundreds of GPUs for training. NeMo-Aligner comes with highly optimized and scalable implementations for major paradigms of model alignment such as: Reinforcement Learning from Human Feedback (RLHF), Direct Preference Optimization (DPO), SteerLM, and Self-Play Fine-Tuning (SPIN). Additionally, our toolkit supports running most of the alignment techniques in a Parameter Efficient Fine-Tuning (PEFT) setting. NeMo-Aligner is designed for extensibility, allowing support for other alignment techniques with minimal effort. It is open-sourced with Apache 2.0 License and we invite community contributions at https://github.com/NVIDIA/NeMo-Aligner

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号