TensorX
返回文献探索

Paper · arXiv 2402.07625

AutoMathText: Autonomous Data Selection with Language Models for Mathematical Texts

Yifan Zhang, Yifan Luo, Yang Yuan, Andrew Chi-Chih Yao

18 upvotesFebruary 12, 2024arXiv 预印本
AI 摘要

A novel strategy using meta-prompted language models autonomously selects high-quality mathematical data for continual pretraining, significantly improving a language model's mathematical reasoning and reducing token usage.

meta-prompted language modelszero-shot verifierscontinual pretrainingAutoMathText datasetMistral language modeltoken efficiencyMATH dataset

Abstract

To improve language models' proficiency in mathematical reasoning via continual pretraining, we introduce a novel strategy that leverages base language models for autonomous data selection. Departing from conventional supervised fine-tuning or trained classifiers with human-annotated data, our approach utilizes meta-prompted language models as zero-shot verifiers to autonomously evaluate and select high-quality mathematical content, and we release the curated open-source AutoMathText dataset encompassing over 200GB of data. To demonstrate the efficacy of our method, we continuously pretrained a 7B-parameter Mistral language model on the AutoMathText dataset, achieving substantial improvements in downstream performance on the MATH dataset with a token amount reduced by orders of magnitude compared to previous continuous pretraining works. Our method showcases a 2 times increase in pretraining token efficiency compared to baselines, underscoring the potential of our approach in enhancing models' mathematical reasoning capabilities. The AutoMathText dataset is available at https://huggingface.co/datasets/math-ai/AutoMathText. The code is available at https://github.com/yifanzhang-pro/AutoMathText.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
AutoMathText: Autonomous Data Selection with Language Models for Mathematical Texts | TensorX