TensorX
返回文献探索

Paper · arXiv 2409.06185

Can Large Language Models Unlock Novel Scientific Research Ideas?

Sandeep Kumar, Tirthankar Ghosal, Vinayak Goyal, Asif Ekbal

15 upvotesSeptember 10, 2024arXiv 预印本
AI 摘要

LLMs, particularly Claude-2, generate more novel and diverse research ideas compared to other models, as assessed through human evaluation across multiple domains.

Large Language ModelsChatGPTArtificial Intelligencefuture research ideasClaude-2GPT-4GPT-3.5Geminihuman evaluationnoveltyrelevancefeasibility

Abstract

"An idea is nothing more nor less than a new combination of old elements" (Young, J.W.). The widespread adoption of Large Language Models (LLMs) and publicly available ChatGPT have marked a significant turning point in the integration of Artificial Intelligence (AI) into people's everyday lives. This study explores the capability of LLMs in generating novel research ideas based on information from research papers. We conduct a thorough examination of 4 LLMs in five domains (e.g., Chemistry, Computer, Economics, Medical, and Physics). We found that the future research ideas generated by Claude-2 and GPT-4 are more aligned with the author's perspective than GPT-3.5 and Gemini. We also found that Claude-2 generates more diverse future research ideas than GPT-4, GPT-3.5, and Gemini 1.0. We further performed a human evaluation of the novelty, relevancy, and feasibility of the generated future research ideas. This investigation offers insights into the evolving role of LLMs in idea generation, highlighting both its capability and limitations. Our work contributes to the ongoing efforts in evaluating and utilizing language models for generating future research ideas. We make our datasets and codes publicly available.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Can Large Language Models Unlock Novel Scientific Research Ideas? | TensorX