TensorX
返回文献探索

Paper · arXiv 2402.16822

Rainbow Teaming: Open-Ended Generation of Diverse Adversarial Prompts

Mikayel Samvelyan, Sharath Chandra Raparthy, Andrei Lupu, Eric Hambro, Aram H. Markosyan, Manish Bhatt, Yuning Mao, Minqi Jiang, Jack Parker-Holder, Jakob Foerster, Tim Rocktäschel, Roberta Raileanu

16 upvotesFebruary 26, 2024arXiv 预印本
AI 摘要

Rainbow Teaming, an approach to generating diverse adversarial prompts, enhances the robustness of large language models across multiple domains by improving their safety without compromising general performance.

large language models (LLMs)adversarial promptsquality-diversity problemopen-ended searchsafetyquestion answeringcybersecurityfine-tuningsynthetic dataopen-ended self-improvement

Abstract

As large language models (LLMs) become increasingly prevalent across many real-world applications, understanding and enhancing their robustness to user inputs is of paramount importance. Existing methods for identifying adversarial prompts tend to focus on specific domains, lack diversity, or require extensive human annotations. To address these limitations, we present Rainbow Teaming, a novel approach for producing a diverse collection of adversarial prompts. Rainbow Teaming casts adversarial prompt generation as a quality-diversity problem, and uses open-ended search to generate prompts that are both effective and diverse. It can uncover a model's vulnerabilities across a broad range of domains including, in this paper, safety, question answering, and cybersecurity. We also demonstrate that fine-tuning on synthetic data generated by Rainbow Teaming improves the safety of state-of-the-art LLMs without hurting their general capabilities and helpfulness, paving the path to open-ended self-improvement.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Rainbow Teaming: Open-Ended Generation of Diverse Adversarial Prompts | TensorX