TensorX
返回文献探索

Paper · arXiv 2406.20094

Scaling Synthetic Data Creation with 1,000,000,000 Personas

Xin Chan, Xiaoyang Wang, Dian Yu, Haitao Mi, Dong Yu

107 upvotesJune 28, 2024arXiv 预印本
AI 摘要

Persona-driven synthetic data creation using a large language model and 1 billion diverse personas for various applications demonstrates versatility, scalability, and flexibility.

large language model (LLM)persona-driven data synthesisPersona Hubsynthetic datamathematical reasoninglogical reasoninguser promptsknowledge-rich textsgame NPCstools (functions)

Abstract

We propose a novel persona-driven data synthesis methodology that leverages various perspectives within a large language model (LLM) to create diverse synthetic data. To fully exploit this methodology at scale, we introduce Persona Hub -- a collection of 1 billion diverse personas automatically curated from web data. These 1 billion personas (~13% of the world's total population), acting as distributed carriers of world knowledge, can tap into almost every perspective encapsulated within the LLM, thereby facilitating the creation of diverse synthetic data at scale for various scenarios. By showcasing Persona Hub's use cases in synthesizing high-quality mathematical and logical reasoning problems, instructions (i.e., user prompts), knowledge-rich texts, game NPCs and tools (functions) at scale, we demonstrate persona-driven data synthesis is versatile, scalable, flexible, and easy to use, potentially driving a paradigm shift in synthetic data creation and applications in practice, which may have a profound impact on LLM research and development.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号