TensorX
返回文献探索

Paper · arXiv 2401.10225

ChatQA: Building GPT-4 Level Conversational QA Models

Zihan Liu, Wei Ping, Rajarshi Roy, Peng Xu, Mohammad Shoeybi, Bryan Catanzaro

36 upvotesJanuary 18, 2024arXiv 预印本
AI 摘要

ChatQA uses a two-stage instruction tuning method with a dense retriever to improve zero-shot conversational QA performance, outperforming GPT-4 on multiple datasets.

ChatQAconversational QAtwo-stage instruction tuningdense retrievermulti-turn QA datasetzero-shot conversational QAlarge language models (LLMs)synthetic data

Abstract

In this work, we introduce ChatQA, a family of conversational question answering (QA) models, that obtain GPT-4 level accuracies. Specifically, we propose a two-stage instruction tuning method that can significantly improve the zero-shot conversational QA results from large language models (LLMs). To handle retrieval in conversational QA, we fine-tune a dense retriever on a multi-turn QA dataset, which provides comparable results to using the state-of-the-art query rewriting model while largely reducing deployment cost. Notably, our ChatQA-70B can outperform GPT-4 in terms of average score on 10 conversational QA datasets (54.14 vs. 53.90), without relying on any synthetic data from OpenAI GPT models.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号