TensorX
返回文献探索

Paper · arXiv 2407.21772

ShieldGemma: Generative AI Content Moderation Based on Gemma

Wenjun Zeng, Yuchi Liu, Ryan Mullins, Ludovic Peran, Joe Fernandez, Hamza Harkous, Karthik Narasimhan, Drew Proud, Piyush Kumar, Bhaktipriya Radharapu, Olivia Sturman, Oscar Wahltinez

15 upvotesJuly 31, 2024arXiv 预印本
AI 摘要

ShieldGemma, a suite of LLM-based safety content moderation models, outperforms existing solutions across various harm types and demonstrates strong generalization even when trained on synthetic data.

LLM-based safety content moderation modelsGemma2safety riskssexually explicitdangerous contentharassmenthate speechLlama GuardWildCardAU-PRCLLM-based data curation pipelinesynthetic data

Abstract

We present ShieldGemma, a comprehensive suite of LLM-based safety content moderation models built upon Gemma2. These models provide robust, state-of-the-art predictions of safety risks across key harm types (sexually explicit, dangerous content, harassment, hate speech) in both user input and LLM-generated output. By evaluating on both public and internal benchmarks, we demonstrate superior performance compared to existing models, such as Llama Guard (+10.8\% AU-PRC on public benchmarks) and WildCard (+4.3\%). Additionally, we present a novel LLM-based data curation pipeline, adaptable to a variety of safety-related tasks and beyond. We have shown strong generalization performance for model trained mainly on synthetic data. By releasing ShieldGemma, we provide a valuable resource to the research community, advancing LLM safety and enabling the creation of more effective content moderation solutions for developers.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
ShieldGemma: Generative AI Content Moderation Based on Gemma | TensorX