TensorX
返回文献探索

Paper · arXiv 2409.19020

DiaSynth -- Synthetic Dialogue Generation Framework

Sathya Krishnan Suresh, Wu Mengjun, Tushar Pranav, Eng Siong Chng

20 upvotesSeptember 25, 2024arXiv 预印本
AI 摘要

DiaSynth generates high-quality, contextually rich, and domain-specific synthetic dialogues using a Large Language Model with Chain of Thought reasoning, demonstrating competitive performance in comparison to in-domain datasets.

synthetic dialogue generationLarge Language Model (LLM)Chain of Thought (CoT)simulated personassubtopicsconversational characteristicsin-domain datafew-shot examplesDialogSumSAMSum

Abstract

The scarcity of domain specific dialogue datasets across various domains, from academic topics to everyday conversations, limits the development of dialogue systems for various applications. Existing research is often constrained either by dialogue datasets that are too general or by niche domain dialogue datasets whose scale does not match the required scale for training dialogue systems. To address this gap, we introduce DiaSynth - a synthetic dialogue generation framework capable of generating high quality, contextually rich dialogues across a wide range of domains. Our approach differs from existing frameworks by dynamically generating dialogues that incorporate simulated personas, subtopics, and diverse conversational characteristics, using a Large Language Model (LLM) with Chain of Thought (CoT) reasoning to create contextually rich, domain-specific dialogues that closely mimic natural human interactions. DiaSynth produces tailored dialogues that emulate realistic conversations. We perform our experiments by generating synthetic data using different LLMs and few-shot examples from DialogSum and SAMSum. The pretrained language models fine-tuned on the synthetic data outperform the base models by 16.47%, while the comparison between models fine-tuned on in-domain data and synthetic data shows that the synthetic data is able to capture 90.48% of the distribution of the in-domain data. The quality of the data generated also scales with the size of LLMs. These results validate DiaSynth's potential as a robust alternative to traditional data collection methods.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
DiaSynth -- Synthetic Dialogue Generation Framework | TensorX