TensorX
返回文献探索

Paper · arXiv 2412.17758

In Case You Missed It: ARC 'Challenge' Is Not That Challenging

Łukasz Borchmann

17 upvotesDecember 23, 2024arXiv 预印本
AI 摘要

The evaluation setup in ARC Challenge affects perceived model difficulty, leading to false implications of reasoning deficits, whereas fairer methods reduce performance gaps and reveal superhuman capabilities in tasks like OpenBookQA.

evaluation setupARC ChallengeARC Easymodern LLMsanswer choicesreasoning deficitsSIQAOpenBookQAmultiple-choice evaluationsmodel capabilities

Abstract

ARC Challenge appears more difficult than ARC Easy for modern LLMs primarily due to an evaluation setup that prevents direct comparison of answer choices rather than inherent complexity. Although some researchers have quietly shifted to a more appropriate scheme over the last year, the implications of this change have yet to be widely acknowledged. We highlight this overlooked shift, show how similar evaluation practices falsely imply reasoning deficits in other benchmarks, and demonstrate that fairer methods dramatically reduce performance gaps (e.g. on SIQA) and even yield superhuman results (OpenBookQA). In doing so, we reveal how evaluation shapes perceived difficulty and offer guidelines to ensure that multiple-choice evaluations accurately reflect actual model capabilities.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
In Case You Missed It: ARC 'Challenge' Is Not That Challenging | TensorX