TensorX
返回文献探索

Paper · arXiv 2502.16069

Curie: Toward Rigorous and Automated Scientific Experimentation with AI Agents

Patrick Tser Jern Kon, Jiachen Liu, Qiuyi Ding, Yiming Qiu, Zhenning Yang, Yibo Huang, Jayanth Srinivasa, Myungjin Lee, Mosharaf Chowdhury, Ang Chen

20 upvotesFebruary 22, 2025arXiv 预印本
AI 摘要

Curie is an AI agent framework that improves rigor in scientific experimentation through reliability, control, and interpretability, demonstrating a significant improvement over baselines in answering experimental questions.

large language models (LLMs)AI agent frameworkintra-agent rigor moduleinter-agent rigor moduleexperiment knowledge moduleexperimental benchmarkcomputer science domainsopen-source projects

Abstract

Scientific experimentation, a cornerstone of human progress, demands rigor in reliability, methodical control, and interpretability to yield meaningful results. Despite the growing capabilities of large language models (LLMs) in automating different aspects of the scientific process, automating rigorous experimentation remains a significant challenge. To address this gap, we propose Curie, an AI agent framework designed to embed rigor into the experimentation process through three key components: an intra-agent rigor module to enhance reliability, an inter-agent rigor module to maintain methodical control, and an experiment knowledge module to enhance interpretability. To evaluate Curie, we design a novel experimental benchmark composed of 46 questions across four computer science domains, derived from influential research papers, and widely adopted open-source projects. Compared to the strongest baseline tested, we achieve a 3.4times improvement in correctly answering experimental questions.Curie is open-sourced at https://github.com/Just-Curieous/Curie.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号