TensorX
返回文献探索

Paper · arXiv 2306.08997

Exploring the MIT Mathematics and EECS Curriculum Using Large Language Models

Sarah J. Zhang, Samuel Florin, Ariel N. Lee, Eamon Niknafs, Andrei Marginean, Annie Wang, Keith Tyser, Zad Chin, Yann Hicke, Nikhil Singh, Madeleine Udell, Yoon Kim, Tonio Buonassisi, Armando Solar-Lezama, Iddo Drori

10 upvotesJune 15, 2023arXiv 预印本
AI 摘要

Large language models, including GPT-3.5 and GPT-4, can effectively solve a significant portion of MIT Mathematics and EECS curriculum questions, with GPT-4 achieving near-perfect performance through prompt engineering and fine-tuning.

large language modelsGPT-3.5GPT-4prompt engineeringfine-tuningfew-shot learning

Abstract

We curate a comprehensive dataset of 4,550 questions and solutions from problem sets, midterm exams, and final exams across all MIT Mathematics and Electrical Engineering and Computer Science (EECS) courses required for obtaining a degree. We evaluate the ability of large language models to fulfill the graduation requirements for any MIT major in Mathematics and EECS. Our results demonstrate that GPT-3.5 successfully solves a third of the entire MIT curriculum, while GPT-4, with prompt engineering, achieves a perfect solve rate on a test set excluding questions based on images. We fine-tune an open-source large language model on this dataset. We employ GPT-4 to automatically grade model responses, providing a detailed performance breakdown by course, question, and answer type. By embedding questions in a low-dimensional space, we explore the relationships between questions, topics, and classes and discover which questions and classes are required for solving other questions and classes through few-shot learning. Our analysis offers valuable insights into course prerequisites and curriculum design, highlighting language models' potential for learning and improving Mathematics and EECS education.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号