DOE OSTI · 2999076
Evaluation of LLM-Generated Kokkos Code Using Compile-Time and Run-Time Testing
Abstract
Due to the growing use of large language models (LLMs) by developers and researchers, it has become essential to reliably evaluate their ability to generate code that uses specialized libraries. We explore the use of compile-time and run-time evaluation of LLM-generated Kokkos code through extending the methods used by OpenAI with the HumanEval dataset. Our evaluation framework is based on the first 40 prompts from the Kokkos138 dataset. We start by discussing two different forms of LLM prompting, using entirely plain English or providing pseudocode for added context. These two methods are used to generate Kokkos code with the Llama-3.1-8B-Instruct and CodeQwen1.5-7B-Chat models. We found that both forms of prompting led to high failure rates and difficulties with reliably parsing LLM-generated code, while prompts with pseudocode for context generally led to improved results on more complicated tests.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Ellingwood, Nathan David [Sandia National Laboratories (SNL-NM), Albuquerque, NM (United States)] (ORCID:0000000276229667), Siefert, Christopher [Sandia National Laboratories (SNL-NM), Albuquerque, NM (United States)] (ORCID:000900032116125X), Jermann, Emil Joseph [Sandia National Laboratories (SNL-NM), Albuquerque, NM (United States)]. 2025-10-01. Evaluation of LLM-Generated Kokkos Code Using Compile-Time and Run-Time Testing. https://doi.org/10.2172/2999076
Cite the original work for its findings. Save a collection to share your selection of sources.