TensorX
返回文献探索

Paper · arXiv 2501.18511

WILDCHAT-50M: A Deep Dive Into the Role of Synthetic Data in Post-Training

Benjamin Feuer, Chinmay Hegde

20 upvotesJanuary 30, 2025arXiv 预印本
AI 摘要

The WILDCHAT-50M dataset, comprising responses from over 50 open-weight models, enables large-scale comparative analysis and demonstrates improved performance in SFT mixtures with fewer samples.

post-trainingDPOdistillationsynthetic data generating modelsLLM judgesWildChat datasetparameter-efficientRE-WILDSFT mixTulu-3 SFT mixture

Abstract

Language model (LLM) post-training, from DPO to distillation, can refine behaviors and unlock new skills, but the open science supporting these post-training techniques is still in its infancy. One limiting factor has been the difficulty of conducting large-scale comparative analyses of synthetic data generating models and LLM judges. To close this gap, we introduce WILDCHAT-50M, the largest public chat dataset to date. We extend the existing WildChat dataset to include responses not only from GPT, but from over 50 different open-weight models, ranging in size from 0.5B to 104B parameters. We conduct an extensive comparative analysis and demonstrate the potential of this dataset by creating RE-WILD, our own public SFT mix, which outperforms the recent Tulu-3 SFT mixture from Allen AI with only 40% as many samples. Our dataset, samples and code are available at https://github.com/penfever/wildchat-50m.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号