TensorX
返回文献探索

Paper · arXiv 2306.16793

Benchmarking Large Language Model Capabilities for Conditional Generation

Joshua Maynez, Priyanka Agrawal, Sebastian Gehrmann

7 upvotesJune 29, 2023arXiv 预印本
AI 摘要

Researchers adapt application-specific generation benchmarks to pre-trained large language models (PLMs) to evaluate their limitations and capabilities across different dimensions like scale, architecture, and language.

pre-trained large language modelsPLMsnatural language processingautoregressive PLMsfew-shot learninggenerationclassificationregressionbenchmarkingnatural language generation

Abstract

Pre-trained large language models (PLMs) underlie most new developments in natural language processing. They have shifted the field from application-specific model pipelines to a single model that is adapted to a wide range of tasks. Autoregressive PLMs like GPT-3 or PaLM, alongside techniques like few-shot learning, have additionally shifted the output modality to generation instead of classification or regression. Despite their ubiquitous use, the generation quality of language models is rarely evaluated when these models are introduced. Additionally, it is unclear how existing generation tasks--while they can be used to compare systems at a high level--relate to the real world use cases for which people have been adopting them. In this work, we discuss how to adapt existing application-specific generation benchmarks to PLMs and provide an in-depth, empirical study of the limitations and capabilities of PLMs in natural language generation tasks along dimensions such as scale, architecture, input and output language. Our results show that PLMs differ in their applicability to different data regimes and their generalization to multiple languages and inform which PLMs to use for a given generation task setup. We share best practices to be taken into consideration when benchmarking generation capabilities during the development of upcoming PLMs.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Benchmarking Large Language Model Capabilities for Conditional Generation | TensorX