TensorX
返回文献探索

Paper · arXiv 2409.08239

Source2Synth: Synthetic Data Generation and Curation Grounded in Real Data Sources

Alisia Lupidi, Carlos Gemmell, Nicola Cancedda, Jane Dwivedi-Yu, Jason Weston, Jakob Foerster, Roberta Raileanu, Maria Lomeli

20 upvotesSeptember 12, 2024arXiv 预印本
AI 摘要

Source2Synth improves LLM performance in structured reasoning and tool usage scenarios by generating high-quality synthetic data points and filtering out low-quality ones.

Large Language Modelsmulti-hop question answering (MHQA)tabular question answering (TQA)synthetic dataanswerability filteringSource2SynthWikiSQLHotPotQA

Abstract

Large Language Models still struggle in challenging scenarios that leverage structured data, complex reasoning, or tool usage. In this paper, we propose Source2Synth: a new method that can be used for teaching LLMs new skills without relying on costly human annotations. Source2Synth takes as input a custom data source and produces synthetic data points with intermediate reasoning steps grounded in real-world sources. Source2Synth improves the dataset quality by discarding low-quality generations based on their answerability. We demonstrate the generality of this approach by applying it to two challenging domains: we test reasoning abilities in multi-hop question answering (MHQA), and tool usage in tabular question answering (TQA). Our method improves performance by 25.51% for TQA on WikiSQL and 22.57% for MHQA on HotPotQA compared to the fine-tuned baselines.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Source2Synth: Synthetic Data Generation and Curation Grounded in Real Data Sources | TensorX