TensorX
返回文献探索

Paper · arXiv 2309.09530

Adapting Large Language Models via Reading Comprehension

Daixuan Cheng, Shaohan Huang, Furu Wei

82 upvotesSeptember 18, 2023arXiv 预印本
AI 摘要

Transforming raw corpora into reading comprehension texts enhances large language models' question-answering abilities across various domains and improves performance on general benchmarks.

large language modelspre-trainingdomain-specific corporadomain knowledgeprompting abilityreading comprehensionscalablegeneral benchmarks

Abstract

We explore how continued pre-training on domain-specific corpora influences large language models, revealing that training on the raw corpora endows the model with domain knowledge, but drastically hurts its prompting ability for question answering. Taken inspiration from human learning via reading comprehension--practice after reading improves the ability to answer questions based on the learned knowledge--we propose a simple method for transforming raw corpora into reading comprehension texts. Each raw text is enriched with a series of tasks related to its content. Our method, highly scalable and applicable to any pre-training corpora, consistently enhances performance across various tasks in three different domains: biomedicine, finance, and law. Notably, our 7B language model achieves competitive performance with domain-specific models of much larger scales, such as BloombergGPT-50B. Furthermore, we demonstrate that domain-specific reading comprehension texts can improve the model's performance even on general benchmarks, showing the potential to develop a general model across even more domains. Our model, code, and data will be available at https://github.com/microsoft/LMOps.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号