Search NASA⌕ Search

DOE OSTI · 2997504

Achieving GPT-4o level performance in astronomy with a specialized 8B-parameter large language model

Abstract

AstroSage-Llama-3.1-8B is a domain-specialized natural-language AI assistant tailored for research in astronomy, astrophysics, cosmology, and astronomical instrumentation. Trained on the complete collection of astronomy-related arXiv papers from 2007 to 2024 along with millions of synthetically-generated question-answer pairs and other astronomical literature, AstroSage-Llama-3.1-8B demonstrates remarkable proficiency on a wide range of questions. AstroSage-Llama-3.1-8B scores 80.9% on the AstroMLab-1 benchmark, greatly outperforming all models—proprietary and open-weight—in the 8-billion parameter class, and performing on par with GPT-4o. This achievement demonstrates the potential of domain specialization in AI, suggesting that focused training can yield capabilities exceeding those of much larger, general-purpose models. AstroSage-Llama-3.1-8B is freely available, enabling widespread access to advanced AI capabilities for astronomical education and research.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

de Haan, Tijmen [High Energy Accelerator Research Organization (KEK), Tsukuba (Japan)], Ting, Yuan-Sen [The Ohio State University, Columbus, OH (United States)], Ghosal, Tirthankar [Oak Ridge National Laboratory (ORNL), Oak Ridge, TN (United States)], Nguyen, Tuan Dung [University of Pennsylvania, Philadelphia, PA (United States)], Accomazzi, Alberto [Harvard-Smithsonian Center for Astrophysics, Cambridge, MA (United States)], Wells, Azton [Argonne National Laboratory (ANL), Argonne, IL (United States)], Ramachandra, Nesar [Argonne National Laboratory (ANL), Argonne, IL (United States)], Pan, Rui [Hong Kong University of Science and Technology (HKUST) (Hong Kong)], Sun, Zechang [Tsinghua University, Beijing (China)]. 2025-04-21. Achieving GPT-4o level performance in astronomy with a specialized 8B-parameter large language model. https://doi.org/10.1038/s41598-025-97131-y

Cite the original work for its findings. Save a collection to share your selection of sources.