TensorX
返回文献探索

Paper · arXiv 2312.10523

Paloma: A Benchmark for Evaluating Language Model Fit

Ian Magnusson, Akshita Bhagia, Valentin Hofmann, Luca Soldaini, Ananya Harsh Jha, Oyvind Tafjord, Dustin Schwenk, Evan Pete Walsh, Yanai Elazar, Kyle Lo, Dirk Groeneveld, Iz Beltagy, Hannaneh Hajishirzi, Noah A. Smith, Kyle Richardson, Jesse Dodge

12 upvotesDecember 16, 2023arXiv 预印本
AI 摘要

Paloma analyzes the performance of language models across 585 diverse text domains to provide a comprehensive and comparable assessment of their domain fit and cost-effectiveness.

perplexityPerplexity Analysis for Language Model Assessment (Paloma)text domainsnytimes.comr/depressionbenchmark contaminationPareto efficiencyparameter counttraining token countCommon Crawldomain fit

Abstract

Language models (LMs) commonly report perplexity on monolithic data held out from training. Implicitly or explicitly, this data is composed of domainsx2013varying distributions of language. Rather than assuming perplexity on one distribution extrapolates to others, Perplexity Analysis for Language Model Assessment (Paloma), measures LM fit to 585 text domains, ranging from nytimes.com to r/depression on Reddit. We invite submissions to our benchmark and organize results by comparability based on compliance with guidelines such as removal of benchmark contamination from pretraining. Submissions can also record parameter and training token count to make comparisons of Pareto efficiency for performance as a function of these measures of cost. We populate our benchmark with results from 6 baselines pretrained on popular corpora. In case studies, we demonstrate analyses that are possible with Paloma, such as finding that pretraining without data beyond Common Crawl leads to inconsistent fit to many domains.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Paloma: A Benchmark for Evaluating Language Model Fit | TensorX