TensorX
返回文献探索

Paper · arXiv 2310.19019

TeacherLM: Teaching to Fish Rather Than Giving the Fish, Language Modeling Likewise

Nan He, Hanyu Lai, Chenyang Zhao, Zirui Cheng, Junting Pan, Ruoyu Qin, Ruofan Lu, Rui Lu, Yunchen Zhang, Gangming Zhao, Zhaohui Hou, Zhiyuan Huang, Shaoqing Lu, Ding Liang, Mingjie Zhan

9 upvotesOctober 29, 2023arXiv 预印本
AI 摘要

TeacherLM-7.1B significantly enhances NLP tasks through zero-shot performance and data augmentation, surpassing larger models and improving various student models.

Large Language ModelsNLP tasksTeacherLM-7.1Bchain of thoughtcommon mistakesMMLUdata augmentationOPT seriesBLOOM seriesmulti-task settingzero-shot score

Abstract

Large Language Models (LLMs) exhibit impressive reasoning and data augmentation capabilities in various NLP tasks. However, what about small models? In this work, we propose TeacherLM-7.1B, capable of annotating relevant fundamentals, chain of thought, and common mistakes for most NLP samples, which makes annotation more than just an answer, thus allowing other models to learn "why" instead of just "what". The TeacherLM-7.1B model achieved a zero-shot score of 52.3 on MMLU, surpassing most models with over 100B parameters. Even more remarkable is its data augmentation ability. Based on TeacherLM-7.1B, we augmented 58 NLP datasets and taught various student models with different parameters from OPT and BLOOM series in a multi-task setting. The experimental results indicate that the data augmentation provided by TeacherLM has brought significant benefits. We will release the TeacherLM series of models and augmented datasets as open-source.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
TeacherLM: Teaching to Fish Rather Than Giving the Fish, Language Modeling Likewise | TensorX