TensorX
返回文献探索

Paper · arXiv 2309.16583

GPT-Fathom: Benchmarking Large Language Models to Decipher the Evolutionary Path towards GPT-4 and Beyond

Shen Zheng, Yuyu Zhang, Yijie Zhu, Chenguang Xi, Pengyang Gao, Xun Zhou, Kevin Chen-Chuan Chang

13 upvotesSeptember 28, 2023arXiv 预印本
AI 摘要

GPT-Fathom evaluates leading LLMs using a standardized suite to provide insights into the evolution from GPT-3 to GPT-4, addressing aspects like reasoning capabilities, SFT, RLHF, and alignment.

large language modelsLLMGPT-FathomOpenAI Evalsbenchmarkscapability categoriessystematic evaluationOpenAI's legacy modelsGPT-3GPT-4cherry-pickingSFTRLHFalignment tax

Abstract

With the rapid advancement of large language models (LLMs), there is a pressing need for a comprehensive evaluation suite to assess their capabilities and limitations. Existing LLM leaderboards often reference scores reported in other papers without consistent settings and prompts, which may inadvertently encourage cherry-picking favored settings and prompts for better results. In this work, we introduce GPT-Fathom, an open-source and reproducible LLM evaluation suite built on top of OpenAI Evals. We systematically evaluate 10+ leading LLMs as well as OpenAI's legacy models on 20+ curated benchmarks across 7 capability categories, all under aligned settings. Our retrospective study on OpenAI's earlier models offers valuable insights into the evolutionary path from GPT-3 to GPT-4. Currently, the community is eager to know how GPT-3 progressively improves to GPT-4, including technical details like whether adding code data improves LLM's reasoning capability, which aspects of LLM capability can be improved by SFT and RLHF, how much is the alignment tax, etc. Our analysis sheds light on many of these questions, aiming to improve the transparency of advanced LLMs.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
GPT-Fathom: Benchmarking Large Language Models to Decipher the Evolutionary Path towards GPT-4 and Beyond | TensorX