TensorX
返回文献探索

Paper · arXiv 2410.07985

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Bofei Gao, Feifan Song, Zhe Yang, Zefan Cai, Yibo Miao, Qingxiu Dong, Lei Li, Chenghao Ma, Liang Chen, Runxin Xu, Zhengyang Tang, Benyou Wang, Daoguang Zan, Shanghaoran Quan, Ge Zhang, Lei Sha, Yichang Zhang, Xuancheng Ren, Tianyu Liu, Baobao Chang

32 upvotesOctober 10, 2024arXiv 预印本
AI 摘要

A new benchmark evaluates LLMs on Olympiad-level mathematics, revealing challenges even for advanced models.

large language modelsLLMsmathematical reasoningGSM8KMATHOpenAI o1benchmarkcompetition-level problemshuman annotationmathematical reasoning at Olympiad levelsub-domainsdifficulty levelsOpenAI o1-miniOpenAI o1-previewOlympiad-level mathematical reasoning

Abstract

Recent advancements in large language models (LLMs) have led to significant breakthroughs in mathematical reasoning capabilities. However, existing benchmarks like GSM8K or MATH are now being solved with high accuracy (e.g., OpenAI o1 achieves 94.8% on MATH dataset), indicating their inadequacy for truly challenging these models. To bridge this gap, we propose a comprehensive and challenging benchmark specifically designed to assess LLMs' mathematical reasoning at the Olympiad level. Unlike existing Olympiad-related benchmarks, our dataset focuses exclusively on mathematics and comprises a vast collection of 4428 competition-level problems with rigorous human annotation. These problems are meticulously categorized into over 33 sub-domains and span more than 10 distinct difficulty levels, enabling a holistic assessment of model performance in Olympiad-mathematical reasoning. Furthermore, we conducted an in-depth analysis based on this benchmark. Our experimental results show that even the most advanced models, OpenAI o1-mini and OpenAI o1-preview, struggle with highly challenging Olympiad-level problems, with 60.54% and 52.55% accuracy, highlighting significant challenges in Olympiad-level mathematical reasoning.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models | TensorX