Search NASA⌕ Search

SEARCH · Search NASA

Results for “AI Evaluations”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Opportunities in AI/ML for the Rubin LSST Dark Energy Science Collaboration

The Vera C. Rubin Observatory's Legacy Survey of Space and Time (LSST) will produce unprecedented volumes of heterogeneous astronomical data (images, catalogs, and alerts) that challenge traditional analysis pipelines. The LSST Dark Energy Science Collaboration (DESC) aims to derive robust constraints on dark energy and dark matter from these data, requiring methods that are statistically powerful, scalable, and operationally reliable. Artificial intelligence and machine learning (AI/ML) are already embedded across DESC science workflows, from photometric redshifts and transient classification to weak lensing inference and cosmological simulations. Yet their utility for precision cosmology hinges on trustworthy uncertainty quantification, robustness to covariate shift and model misspecification, and reproducible integration within scientific pipelines. This white paper surveys the current landscape of AI/ML across DESC's primary cosmological probes and cross-cutting analyses, revealing that the same core methodologies and fundamental challenges recur across disparate science cases. Since progress on these cross-cutting challenges would benefit multiple probes simultaneously, we identify key methodological research priorities, including Bayesian inference at scale, physics-informed methods, validation frameworks, and active learning for discovery. With an eye on emerging techniques, we also explore the potential of the latest foundation model methodologies and LLM-driven agentic AI systems to reshape DESC workflows, provided their deployment is coupled with rigorous evaluation and governance. Finally, we discuss critical software, computing, data infrastructure, and human capital requirements for the successful deployment of these new methodologies, and consider associated risks and opportunities for broader coordination with external actors.

Aubourg, Eric [APC, Paris] (ORCID:000000025592023X↗

An integrated environment for tactical guidance research and evaluation

NASA-Langley's Tactical Guidance Research and Evaluation System (TGRES) constitutes an integrated environment for the development of tactical guidance algorithms and evaluating the effects of novel technologies; the modularity of the system allows easy modification or replacement of system elements in order to conduct evaluations of alternative technologies. TGRES differs from existing systems in its capitalization on AI programming techniques for guidance-logic implementation. Its ability to encompass high-fidelity, six-DOF simulation models will facilitate the analysis of complete aircraft dynamics.

Goodrich, Kenneth H.↗

Advancing Sustainability in Data Centers: Evaluation of Hybrid Air/Liquid Cooling Schemes for IT Payload Using Sea Water

Abstract-The growth in cloud computing, Big Data, AI and high-performance computing (HPC) necessitate the deployment of additional data centers (DC's) with high energy demands. The unprecedented increase in the Thermal Design Power (TDP) of the computing chips will require innovative cooling techniques. Furthermore, DC's are increasingly limited in their ability to add powerful GPU servers by power capacity constraints. As cooling energy use accounts for up to 40% of DC energy consumption, creative cooling solutions are urgently needed to allow deployment of additional servers, enhance sustainability and increase energy efficiency of DC's. The information in this study is provided from Start Campus' Sines facility supported by Alfa Laval for the heat exchanger and CO 2 emission calculations. The study evaluates the performance and sustainability impact of various data center cooling strategies including an air-only deployment and a subsequent hybrid air/water cooling solution all utilizing sea water as the cooling source. Here we evaluate scenarios from 3 MW to 15+1 MW of IT load in 3 MW increments which correspond to the size of heat exchangers used in the Start Campus' modular system design. This study also evaluates the CO 2 emissions compared to a conventional chiller system for all the presented scenarios. Results indicate that the effective use of the sea water cooled system combined with liquid cooled systems improve the efficiency of the DC, plays a role in decreasing the CO 2 emissions and supports in achieving sustainability goals.

97 MATHEMATICS AND COMPUTING↗

Measurement of the 252 Cf ⁢(sf) prompt fission neutron spectrum utilizing 12 C ⁡(𝑛, 𝑛) and 9 Be ⁢(𝑛, 𝑛) neutron scattering reference measurements

The 252 Cf spontaneous fission (sf), prompt fission neutron spectrum (PFNS) is a fundamental quantity for nuclear physics measurements of neutron-emitting reactions. This energy distribution of neutrons emitted from fission has been considered a neutron data standard for decades and has been utilized as a reference for neutron detection efficiency, validation of Monte Carlo simulations, benchmarking of dosimetry standards, and more. A significant portion of the global collection of nuclear data on neutron-induced reactions is correlated with the 252 Cf ⁢(sf) PFNS. Despite the reliance on this quantity by the nuclear physics community, the historical collection of 252 Cf PFNS measurements display systematic disagreements that are not understood or easily explained. These experimental discrepancies could potentially bias the 252 Cf PFNS Standard evaluation. On top of this, these past experiments frequently employed correlated experimental measurement or analysis methods. The artificial intelligence (AI)/machine learning (ML)-informed californium chi-nuclear data experiment (AIACHNE) project was formed to (a) investigate these discrepancies utilizing AI/ML methods to identify outlying regions of literature data, assign these regions to features of the experiment itself, and perform an improved evaluation of the 252 Cf PFNS and (b) perform a new experimental measurement of this quantity designed to improve upon the existing literature database. Here, in this work, we report on the AIACHNE 252 Cf PFNS experiment utilizing a new analysis method uncorrelated with all previous measurements: neutron efficiency determinations based on elastic neutron scattering on 12 C and 9 Be . This new method provides an independent test of the existing literature data and evaluation of the 252 Cf ⁢(sf) PFNS. The method is described with detailed covariance quantification procedures, as well as a direct discussion of the sources of uncertainty described as requirements in the “Templates” series of papers. The 252 Cf ⁢(sf) PFNS reported in this work agrees well with the overall shape of the existing standard PFNS evaluation as well as many literature measurements, thus verifying the current evaluation utilizing new techniques. However, the results suggest that there are deficiencies in the angle-differential 12 C and 9 Be ⁢(𝑛, 𝑛) evaluated nuclear data, which produce unphysical structures in the reported result. While these structures are relatively minor, they become obvious because of the high statistical precision of the data and the expected smooth continuity of the 252 Cf ⁢(sf) PFNS.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

An evaluation of Ada for Al applications

Expert system technology seems to be the most promising type of Artificial Intelligence (AI) application for Ada. An expert system implemented with an expert system shell provides a highly structured approach that fits well with the structured approach found in Ada systems. The current commercial expert system shells use Lisp. In this highly structured situation a shell could be built that used Ada just as well. On the other hand, if it is necessary to deal with some AI problems that are not suited to expert systems, the use of Ada becomes more problematical. Ada was not designed as an AI development language, and is not suited to that. It is possible that an application developed in say, Common Lisp could be translated to Ada for actual use in a particular application, but this could be difficult. Some standard Ada packages could be developed to make such a translation easier. If the most general AI programs need to be dealt with, a Common Lisp system integrated with the Ada Environment is probably necessary. Aside from problems with language features, Ada, by itself, is not well suited to the prototyping and incremental development that is well supported by Lisp.

Wallace, David R.↗

Ten questions concerning Large Language Models (LLMs) for building applications

Large Language Models (LLMs) are emerging as powerful AI tools capable of transforming how building information is collected, processed, analyzed, and applied across diverse research areas. Their capabilities can help building operators, facility managers and other stakeholders such as designers, architects and engineers by providing actionable insights for decision-making across planning, construction, operations, and maintenance of buildings and facilities. This paper explores ten key questions concerning the role of LLMs in shaping sustainable, intelligent, and human-centric buildings. From fundamental definitions to advanced applications, we examine how LLMs facilitate decision-making across the life cycle of buildings and energy systems. LLMs can enhance life cycle assessments (LCA), building energy simulations, and real-time data integration, empowering more efficient and adaptive human-AI environments. They can also contribute to streamlining regulatory compliance, improving post-occupancy evaluations, and fostering more inclusive and participatory design processes. Additionally, this paper addresses the ethical challenges posed by LLMs, such as bias, data privacy, and environmental impacts, and explores their potentials in advancing intelligent digital twins (DT) for ongoing building operations and maintenance. Built upon our applied research using LLMs and the review of tools, datasets, and research gaps, we provide a forward-looking perspective on how LLMs can drive innovation, collaboration, and productivity in the built environment while supporting ethical and effective implementation.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Hardware-in-the-Loop Evaluation for Potential High Limit Estimation-Based PV Plant Active Control: Preprint

This paper validates the effectiveness of an Artificial Intelligence (AI)-driven PV plant control and optimization approach, namely, the Automated Learner for Intermittency Control by Extrapolation (ALICE), in empowering PV plant as a dependable grid reliability service provider. The validation is performed in a realistic laboratory controller-hardware-in-the-loop (CHIL) environment, leveraging accurate PV plant modeling and standard industrial communication protocol. Simulation results, considering both varying weather conditions and active control scenarios, demonstrate the superior performance of ALICE in improving the grid service delivery precision and reducing the over-curtailment compared to a state-of-the-art approach, i.e., reference-control grouping based approach. Such a work could help mitigate risks and provide practical guidance during the field deployment of ALICE, while establishing a standardized testing framework for evaluating various PV active control strategies.

hardware-in-the-loop↗

A space systems perspective of graphics simulation integration

Creation of an interactive display environment can expose issues in system design and operation not apparent from nongraphics development approaches. Large amounts of information can be presented in a short period of time. Processes can be simulated and observed before committing resources. In addition, changes in the economics of computing have enabled broader graphics usage beyond traditional engineering and design into integrated telerobotics and Artificial Intelligence (AI) applications. The highly integrated nature of space operations often tend to rely upon visually intensive man-machine communication to ensure success. Graphics simulation activities at the Mission Planning and Analysis Division (MPAD) of NASA's Johnson Space Center are focusing on the evaluation of a wide variety of graphical analysis within the context of present and future space operations. Several telerobotics and AI applications studies utilizing graphical simulation are described. The presentation includes portions of videotape illustrating technology developments involving: (1) coordinated manned maneuvering unit and remote manipulator system operations, (2) a helmet mounted display system, and (3) an automated rendezous application utilizing expert system and voice input/output technology.

Brown, R.↗

chatHPC: Empowering HPC users with large language models

The ever-growing number of pre-trained large language models (LLMs) across scientific domains presents a challenge for application developers. While these models offer vast potential, fine-tuning them with custom data, aligning them for specific tasks, and evaluating their performance remain crucial steps for effective utilization. However, applying these techniques to models with tens of billions of parameters can take days or even weeks on modern workstations, making the cumulative cost of model comparison and evaluation a significant barrier to LLM-based application development. To address this challenge, we introduce an end-to-end pipeline specifically designed for building conversational and programmable AI agents on high performance computing (HPC) platforms. Our comprehensive pipeline encompasses: model pre-training, fine-tuning, web and API service deployment, along with crucial evaluations for lexical coherence, semantic accuracy, hallucination detection, and privacy considerations. Here, we demonstrate our pipeline through the development of chatHPC, a chatbot for HPC question answering and script generation. Leveraging our scalable pipeline, we achieve end-to-end LLM alignment in under an hour on the Frontier supercomputer. We propose a novel self-improved, self-instruction method for instruction set generation, investigate scaling and fine-tuning strategies, and conduct a systematic evaluation of model performance. The established practices within chatHPC will serve as a valuable guidance for future LLM-based application development on HPC platforms.

97 MATHEMATICS AND COMPUTING↗

Hubble Space Telescope Fine Guidance Sensor Post-Flight Bearing Inspection

Aerospace mechanism engineering success stories often, if not always, consist of overcoming developmental, test and flight anomalies. Many times it is these anomalies that stimulate technology growth and more reliable future systems. However, one must learn from these to achieve an ultimately successful mission. It is not often that a spacecraft engineer is able to inspect hardware that has flown in orbit for several years. However, in February 1997, the Fine Guidance Sensor-I (FGS-1) was removed from the Hubble Space Telescope (HST) and returned to NASA Goddard Space Flight Center (GSFC) during the second Servicing Mission (SM2). At the time of removal, FGS-1 had nearly 7 years of service and the bearings in the Star Selector Servos (SSS) had accumulated approximately 25 million Coarse Track (CT) cycles. The main reason for its replacement was due to a bearing torque anomaly leading to stalling of the B Star Selector Servo (SSS-B) when reversing direction during a vehicle offset maneuver, referred to herein as a Reversal Bump (RB). The returned HST FGS SSS bearings were disassembled for post-service condition assessment to better understand the actual cause of the torque spikes, identify potential process/design improvements, and provide information for remedial on-orbit operation modifications. The methods and technology utilized for this inspection are not unique to this system and can be adapted to most investigation ai varying stages of the mechanism life from development, through testing, io post night evaluation. The systematic methods used for the HST Fine Guidance Sensor (FGS) SSS and specific findings are the subjects presented in this paper. The lessons learned include the importance of cleanliness and handling for precision instrument bearings and the potential effects from contamination. The paper describes in detail, the analytical techniques used for the SSS and their importance in this investigation. Inspection analytical data and photographs are included throughout the paper.

Pellicciotti, Joseph↗

A Comparison of Seasonal and Interannual Variability of Soil Dust Aerosols Over the Atlantic Ocean as Inferred by the Toms AI and AVHRR AOT Retrievals

The seasonal cycle and interannual variability of two estimates of soil (or 'mineral') dust aerosols are compared: Advanced Very High Resolution Radiometer (AVHRR) aerosol optical thickness (AOT) and Total Ozone Mapping Spectrometer (TOMS) aerosol index (AI), Both data sets, comprising more than a decade of global, daily images, are commonly used to evaluate aerosol transport models. The present comparison is based upon monthly averages, constructed from daily images of each data set for the period between 1984 and 1990, a period that excludes contamination from volcanic eruptions. The comparison focuses upon the Northern Hemisphere subtropical Atlantic Ocean, where soil dust aerosols make the largest contribution to the aerosol load, and are assumed to dominate the variability of each data set. While each retrieval is sensitive to a different aerosol radiative property - absorption for the TOMS AI versus reflectance for the AVHRR AOT - the seasonal cycles of dust loading implied by each retrieval are consistent, if seasonal variations in the height of the aerosol layer are taken into account when interpreting the TOMS AI. On interannual time scales, the correlation is low at most locations. It is suggested that the poor interannual correlation is at least partly a consequence of data availability. When the monthly averages are constructed using only days common to both data sets, the correlation is substantially increased: this consistency suggests that both TOMS and AVHRR accurately measure the aerosol load in any given scene. However, the two retrievals have only a few days in common per month so that these restricted monthly averages have a large uncertainty. Calculations suggest that at least 7 to 10 daily images are needed to estimate reliably the average dust load during any particular month, a threshold that is rarely satisfied by the AVHRR AOT due to the presence of clouds in the domain. By rebinning each data set onto a coarser grid, the availability of the AVHRR AOT is increased during any particular month, along with its interannual correlation with the TOMS AI The latter easily exceeds the sampling threshold due to its greater ability to infer the aerosol load in the presence of clouds. Whether the TOMS AI should be regarded as a more reliable indicator of interannual variability depends upon the extent of contamination by sub-pixel clouds.

Cakmur, R. V.↗

PowerModel-AI: A First On-the-Fly Machine-Learning Predictor for AC Power Flow Solutions

The real-time creation of machine-learning models via active or on-the-fly learning has attracted considerable interest across various scientific and engineering disciplines. These algorithms enable machines to build models autonomously while remaining operational. Through a series of query strategies, the machine can evaluate whether newly encountered data fall outside the scope of the existing training set. In this study, we introduce PowerModel-AI, an end-to-end machine learning software designed to accurately predict AC power flow solutions. We present detailed justifications for our model design choices and demonstrate that selecting the right input features effectively captures load flow decoupling inherent in power flow equations. Our approach incorporates on-the-fly learning, where power flow calculations are initiated only when the machine detects a need to improve the dataset in regions where the model’s suboptimal performance is based on specific criteria. Otherwise, the existing model is used for power flow predictions. This study includes analyses of five Texas A&M synthetic power grid cases, encompassing the 14-, 30-, 37-, 200-, and 500-bus systems. The training and test datasets were generated using PowerModels.jl, an open-source power flow solver/optimizer developed at Los Alamos National Laboratory, NM, USA.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Artificial Intelligence in Nuclear Safeguards; Evaluating Safeguards and Security Risks and Benefits for Advanced and Small Modular Reactor Deployments

Rapidly growing interest in advanced and small modular reactor (A/SMR) technologies presents challenges as well as opportunities for implementing international safeguards and security. A/SMR deployments are expected to be more numerous, more geographically dispersed, and more varied in their designs, placing new demands on the data systems and analytical tools used to support oversight (Alberti et al., 2023; Canadian Nuclear Safety Commission et al., 2024). Because of this variability, the importance and reliance on data systems for A/SMR deployments is expected to be higher than for previous reactor generations. Artificial Intelligence and Machine Learning (AI/ML) offer potential capabilities to address the high variability inherent in A/SMR technology. The beneficiaries of AI-assisted tools include facility operators, government regulators, IAEA inspectors, and A/SMR vendors. This report analyzes how AI/ML-assisted technologies can strengthen the implementation of IAEA safeguards and security measures. It also identifies AI-assisted tools to strengthen operator, facility, and regulator knowledge management practices and examines the potential risks AI/ML-based tools may introduce to IAEA safeguards and security efforts. It concludes with a set of hypothetical, standards-style requirements for AI/ML systems used in safeguards contexts, grounded in an inspector-centric view of system verification. Despite the potential benefits of AI/ML systems, understanding potential intentional and unintentional failure modes is critical for ensuring adequate protection of nuclear materials and facilities. Unique features of A/SMRs including sealed cores, remote and novel paradigms of operation, off-site reactor fabrication, novel fuel forms, and varied refueling requirements, introduce challenges for traditional safeguards technological approaches (Pensado et al., 2024; Federation of American Scientists, 2025). AI/ML systems deployed to address these challenges may introduce new risks requiring systematic evaluation rooted in both AI-specific risk frameworks, such as the NIST AI Risk Management Framework (NIST AI RMF), and established cyber risk management standards such as NIST SP 800-30 (National Institute of Standards and Technology [NIST], 2023; NIST, 2012).

97 MATHEMATICS AND COMPUTING↗

Hardware-in-the-Loop Evaluation for Potential High Limit Estimation-Based PV Plant Active Control

This paper validates the efficacy of an artificial intelligence (AI)-based photovoltaic (PV) plant control and optimization approach in enabling PV plants as accountable grid reliability service providers. The validation is performed in a realistic laboratory controller-hardware-in-the-loop environment, leveraging accurate PV plant modeling and standard industrial communication protocols. Through simulations that account for diverse weather conditions and active control scenarios, the results highlight the superior performance of the AI-based solution in comparison to a state-of-the-art reference-control grouping-based approach. Such a finding contributes to mitigating the risk of overcurtailment and uninstructed deviations of active PV plant controls, and offers practical guidance for its field deployment. Furthermore, it establishes a standardized testing framework for comparing various PV active control strategies.

hardware-in-the-loop↗

System Diagnostic Builder - A rule generation tool for expert systems that do intelligent data evaluation

Consideration is given to the System Diagnostic Builder (SDB), an automated knowledge acquisition tool using state-of-the-art AI technologies. The SDB employs an inductive machine learning technique to generate rules from data sets that are classified by a subject matter expert. Thus, data are captured from the subject system, classified, and used to drive the rule generation process. These rule bases are used to represent the observable behavior of the subject system, and to represent knowledge about this system. The knowledge bases captured from the Shuttle Mission Simulator can be used as black box simulations by the Intelligent Computer Aided Training devices. The SDB can also be used to construct knowledge bases for the process control industry, such as chemical production or oil and gas production.

Nieten, Joseph↗

ChatBLAS: The First AI-Generated and Portable BLAS Library

We present ChatBLAS, the first AI-generated and portable Basic Linear Algebra Subprograms (BLAS) library on different CPU/GPU configurations. The purpose of this study is (i) to evaluate the capabilities of current large language models (LLMs) to generate a portable and HPC library for BLAS operations and (ii) to define the fundamental practices and criteria to interact with LLMs for HPC targets to elevate the trustworthiness and performance levels of the AI-generated HPC codes. The generated C/C++ codes must be highly optimized using device-specific solutions to reach high levels of performance. Additionally, these codes are very algorithm-dependent, thereby adding an extra dimension of complexity to this study. We used OpenAI’s LLM ChatGPT and focused on vector-vector BLAS level-1 operations. ChatBLAS can generate functional and correct codes, achieving high-trustworthiness levels, and can compete or even provide better performance against vendor libraries.

Valero Lara, Pedro↗

Integration of Solid Oxide Fuel Cell Systems Into Artificial Intelligence Data Centers

This report presents the results of a techno-economic analysis (TEA) that evaluates the economic benefits of integrating solid oxide fuel cell (SOFC) systems with artificial intelligence (AI) data centers. The analysis was completed in two phases: a scoping-level analysis was performed to identify impactful integration opportunities, followed by a more detailed TEA. Results show that, due to their modularity, SOFC can meet the 99.999% availability requirement of data centers with minimal additional costs. Heat integration via absorption chillers decreases data center electricity consumption at the tradeoff of increased water consumption. Higher SOFC exhaust temperatures are important for achieving larger electricity savings. Finally, power electronics integration with SOFC direct current electricity can reduce electricity consumption by 9 percent and reduce water consumption by 6.4 percent.

20 FOSSIL-FUELED POWER PLANTS↗

Preliminary assessment of airborne imaging spectrometer and airborne thematic mapper data acquired for forest decline areas in the Federal Republic of Germany

This study evaluated the utility of data collected by the high-spectral resolution airborne imaging spectrometer (AIS-2, tree mode, spectral range 0.8-2.2 microns) and the broad-band Daedalus airborne thematic mapper (ATM, spectral range 0.42-13.0 micron) in assessing forest decline damage at a predominantly Scotch pine forest in the FRG. Analysis of spectral radiance values from the ATM and raw digital number values from AIS-2 showed that higher reflectance in the near infrared was characteristic of high damage (heavy chlorosis, limited needle loss) in Scotch pine canopies. A classification image of a portion of the AIS-2 flight line agreed very well with a damage assessment map produced by standard aerial photointerpretation techniques.

Herrmann, Karin↗