TensorX
返回文献探索

Paper · arXiv 2309.06126

AstroLLaMA: Towards Specialized Foundation Models in Astronomy

Tuan Dung Nguyen, Yuan-Sen Ting, Ioana Ciucă, Charlie O'Neill, Ze-Chang Sun, Maja Jabłońska, Sandor Kruk, Ernest Perkowski, Jack Miller, Jason Li, Josh Peek, Kartheik Iyer, Tomasz Różański, Pranav Khetarpal, Sharaf Zaman, David Brodrick, Sergio J. Rodríguez Méndez, Thang Bui, Alyssa Goodman, Alberto Accomazzi, Jill Naiman, Jesse Cranney, Kevin Schawinski, UniverseTBD

18 upvotesSeptember 12, 2023arXiv 预印本
AI 摘要

AstroLLaMA, a fine-tuned large language model using astronomy abstracts, significantly reduces perplexity and improves text completions and embeddings in specialized astronomical domains.

AstroLLaMALLaMA-2causal language modelingperplexityembedding extractionautomatic paper summarizationconversational agent development

Abstract

Large language models excel in many human-language tasks but often falter in highly specialized domains like scholarly astronomy. To bridge this gap, we introduce AstroLLaMA, a 7-billion-parameter model fine-tuned from LLaMA-2 using over 300,000 astronomy abstracts from arXiv. Optimized for traditional causal language modeling, AstroLLaMA achieves a 30% lower perplexity than Llama-2, showing marked domain adaptation. Our model generates more insightful and scientifically relevant text completions and embedding extraction than state-of-the-arts foundation models despite having significantly fewer parameters. AstroLLaMA serves as a robust, domain-specific model with broad fine-tuning potential. Its public release aims to spur astronomy-focused research, including automatic paper summarization and conversational agent development.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
AstroLLaMA: Towards Specialized Foundation Models in Astronomy | TensorX