Search NASA⌕ Search

SEARCH · Search NASA

Results for “AI Evaluations”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Data Readiness for AI: A 360-Degree Survey

Artificial Intelligence (AI) applications critically depend on data. Poor-quality data produces inaccurate and ineffective AI models that may lead to incorrect or unsafe use. Evaluation of data readiness is a crucial step in improving the quality and appropriateness of data usage for AI. R&D efforts have been spent on improving data quality. However, standardized metrics for evaluating data readiness for use in AI training are still evolving. In this study, we perform a comprehensive survey of metrics used to verify data readiness for AI training. This survey examines more than 140 papers published by ACM Digital Library, IEEE Xplore, journals such as Nature, Springer, and Science Direct, and online articles published by prominent AI experts. This survey aims to propose a taxonomy of data readiness for AI (DRAI) metrics for structured and unstructured datasets. We anticipate that this taxonomy will lead to new standards for DRAI metrics that would be used for enhancing the quality, accuracy, and fairness of AI training and inference.

97 MATHEMATICS AND COMPUTING↗

Dual Context: Leveraging Structured Application Context for Code Generation and Runtime Feature Activation via Chat Interfaces

Integrating artificial intelligence (AI) capabilities into software applications typically involves two common paths. For developers, AI assists in generating and documenting source code and other related software engineering efforts. For users, AI assists them through question-and-answer exchanges via chatbots. Both approaches have their value, but neither effectively leverages the modularity of component-based architectures that modern web application frameworks offer. We implement a proof of concept within a centralized suite of applications used for the Atmospheric Radiation Measurement (ARM) Data Center Operational Tools, where we introduce a third integration path through the ARM Context Engine (ACE). ACE is a context driven system that uses structured contextual specifications to enable Large Language Models (LLMs) to render interactive and feature-rich user interface (UI) components directly within chat responses, alongside or in place of conventional text outputs. These specifications serve two important purposes across what we call code context and UI context. Code context provides AI-assisted development tools with structured application knowledge beyond raw code, including component relationships, architectural patterns and schematic information, enabling the generation of consistent, well-structured code. UI context defines the rules for enabling and rendering component features at runtime based on the user's natural language input, allowing end users to activate capabilities such as data export, filtering, and pagination within chat responses, without requiring code changes or redeployment. We demonstrate, through a comparative evaluation against general-purpose AI chatbots, that context-driven component rendering provides interactive capabilities that text-based responses cannot replicate, including deterministic component behavior, application-consistent design language, and on-demand feature activation. A development effort comparison further shows that features that traditionally require multi-step development cycles can be activated with a single naturallanguage request. In this ongoing work, we present ACE as an emerging approach to AI integration that positions modular, well-documented software architecture as the foundation for AI-ready applications. ACE treats context as a shared resource across both development and user-facing AI, bringing cohesion to conventionally disconnected efforts, bridging developer tooling and end-user capabilities within a single framework.

Tadimeti, Vijay [ORNL]↗

Calibrating AIS images using the surface as a reference

A method of evaluating the initial assumptions and uncertainties of the physical connection between Airborne Imaging Spectrometer (AIS) image data and laboratory/field spectrometer data was tested. The Tuscon AIS-2 image connects to lab reference spectra by an alignment to the image spectral endmembers through a system gain and offset for each band. Images were calibrated to reflectance so as to transform the image into a measure that is independent of the solar radiant flux. This transformation also makes the image spectra directly comparable to data from lab and field spectrometers. A method was tested for calibrating AIS images using the surface as a reference. The surface heterogeneity is defined by lab/field spectral measurements. It was found that the Tuscon AIS-2 image is consistent with each of the initial hypotheses: (1) that the AIS-2 instrument calibration is nearly linear; (2) the spectral variance is caused by sub-pixel mixtures of spectrally distinct materials and shade, and (3) that sub-pixel mixtures can be treated as linear mixtures of pure endmembers. It was also found that the image can be characterized by relatively few endmembers using the AIS-2 spectra.

Smith, M. O.↗

Development and performance evaluation of active insulation systems using solid-state thermal switches

Traditional building envelopes have passive insulation systems that cannot respond to dynamic changes in the environment. An Active Insulation System (AIS) consists of Active Insulation Materials (AIMs) that dynamically vary the thermal conductivity of the insulation system. Several researchers have evaluated the impact of AIS on building thermal and energy performance by using simulation tools. Up to 70% savings in annual heating and cooling energy and significant reductions in peak demand have been predicted for some climates with wall systems employing AIS. However, materials and assembly development still need a cost-effective product that achieves the required performance. Here, in this study, we present the process of developing an AIS that we will install in a test hut for its performance evaluation. Minimum performance criteria of the AIS system are developed based on R-low/R-high ratio, required time and efficiency to switch states, and cost estimates. The following steps during this study are creating the concept to meet the requirements, predicting the performance via simulations, developing the experimental setup for bench-scale testing, and finally, constructing a full-scale wall assembly and monitoring the performance when exposed to environmental chamber tests. The selected approach uses off-the-shelf products to create an AIS that can switch R-value between 0.98 ft 2 ·°F·h/BTU (0.173 m 2 ·K/W) and 5.81 ft 2 ·°F·h/BTU (1.02 m 2 ·K/W) and have a switching time of less than one minute between R-high and R-low.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Intelligent fault diagnosis and failure management of flight control actuation systems

The real-time fault diagnosis and failure management (FDFM) of current operational and experimental dual tandem aircraft flight control system actuators was investigated. Dual tandem actuators were studied because of the active FDFM capability required to manage the redundancy of these actuators. The FDFM methods used on current dual tandem actuators were determined by examining six specific actuators. The FDFM capability on these six actuators was also evaluated. One approach for improving the FDFM capability on dual tandem actuators may be through the application of artificial intelligence (AI) technology. Existing AI approaches and applications of FDFM were examined and evaluated. Based on the general survey of AI FDFM approaches, the potential role of AI technology for real-time actuator FDFM was determined. Finally, FDFM and maintainability improvements for dual tandem actuators were recommended.

Bonnice, William F.↗

An evaluation of imaging spectrometry for estimating forest canopy chemistry

High spectral resolution Airborne Imaging Spectrometer (AIS) data were acquired over 20 well-studied Wisconsin forest sites to evaluate the potential of remote sensing for estimating forest canopy chemistry. Intensive nutrient cycling research in these forests demonstrates that canopy lignin content is strongly related to measured annual nitrogen mineralization at the undisturbed sites and may serve as an accurate index for nitrogen cycling rates. Ground measurements were made of foliar biomass and canopy nitrogen and lignin content, the latter within two weeks of the AIS overflight. The spectral data were transformed using derivative techniques modified from laboratroy spectroscopy. Stepwise regression assisted in determining combinations of wavelengths most highly correlated with canopy chemistry and biomass. Strong correlations between AIS data and total canopy lignin content in deciduous forests and canopy lignin concentration (total lignin/biomass) in both deciduous and coniferous stands indicate that imaging spectrometry can be used to estimate canopy lignin content and, from that, the spatial distribution of annual nitrogen mineralization rates.

Wessman, Carol A.↗

Using Artificial Intelligence to Improve Reliability and Operational Efficiency of Small-Scale Hydroelectric Distributed Generation

Reliability and resilience are critical concerns for distributed generation (DG) at the rural electric level. The integration of renewable energy sources, such as small-scale hydroelectric distributed generators (hydro DGs), introduces operational challenges, particularly regarding aging infrastructure and grid stability. Artificial Intelligence (AI)-driven Machine Learning (ML) models and applications of Large Language Models (LLMs) offer promising solutions for optimizing DG operations and enhancing resilience. This paper explores AI-based models for improving efficiency, fault resolution, and outage mitigation in small-scale hydro DGs. Furthermore, it highlights the development of a centralized, AI-powered information portal for rural electric cooperatives and municipalities. The research evaluates hydro DG plant models and discusses the applicability of AI-powered question-answering tools for real-time operations, focusing on statistical data, load flow, voltage regulation, and generation power. The findings demonstrate AI’s potential to transform DG management to ensure greater stability and resilience in rural electric grids.

Bhattacharyya, Arjun [ORNL] (ORCID:000900060976046↗

Human-centered automation and AI - Ideas, insights, and issues from the Intelligent Cockpit Aids research effort

A development status evaluation is presented for the NASA-Langley Intelligent Cockpit Aids research program, which encompasses AI, human/machine interfaces, and conventional automation. Attention is being given to decision-aiding concepts for human-centered automation, with emphasis on inflight subsystem fault management, inflight mission replanning, and communications management. The cockpit envisioned is for advanced commercial transport aircraft.

Abbott, Kathy H.↗

Reimagining metal-organic framework discovery: Integrating experiment, computation, and artificial intelligence

The traditional development of novel metal–organic frameworks (MOFs) is often hindered by challenges such as synthetic accessibility and time- and resource-intensive experimentation. High-throughput, automated experimental and computational techniques have enabled rapid chemical space exploration and theoretical MOF design. When combined with artificial intelligence (AI), these methods can be used to lead autonomous laboratories to new frontiers for MOF discovery, where these materials can be designed for a specific application, efficiently synthesized, characterized, and evaluated. Here, this perspective highlights the role of AI in advancing automated MOF synthesis and characterization, computational MOF design and screening, and the integration of these approaches within autonomous workflows to ultimately enable the MOF laboratories of the future.

Gaidimas, Madeleine A. [Northwestern University, E↗

AI Benchmark Democratization and Carpentry

Benchmarks are a cornerstone of modern machine learning, enabling reproducibility, comparison, and scientific progress. However, AI benchmarks are increasingly complex, requiring dynamic, AI-focused workflows. Rapid evolution in model architectures, scale, datasets, and deployment contexts makes evaluation a moving target. Large language models often memorize static benchmarks, causing a gap between benchmark results and real-world performance. Beyond traditional static benchmarks, continuous adaptive benchmarking frameworks are needed to align scientific assessment with deployment risks. This calls for skills and education in AI Benchmark Carpentry. From our experience with MLCommons, educational initiatives, and programs like the DOE's Trillion Parameter Consortium, key barriers include high resource demands, limited access to specialized hardware, lack of benchmark design expertise, and uncertainty in relating results to application domains. Current benchmarks often emphasize peak performance on top-tier hardware, offering limited guidance for diverse, real-world scenarios. Benchmarking must become dynamic, incorporating evolving models, updated data, and heterogeneous platforms while maintaining transparency, reproducibility, and interpretability. Democratization requires both technical innovation and systematic education across levels, building sustained expertise in benchmark design and use. Benchmarks should support application-relevant comparisons, enabling informed, context-sensitive decisions. Dynamic, inclusive benchmarking will ensure evaluation keeps pace with AI evolution and supports responsible, reproducible, and accessible AI deployment. Community efforts can provide a foundation for AI Benchmark Carpentry.

von Laszewski, Gregor [Virginia U.]↗

Event Report for The Ethical Artificial Intelligence Quantification Workshop

Artificial Intelligence (AI) is a powerful emerging technology area which requires special attention to using it ethically. AI ethics is still an emerging field, and the partners for this workshop and report seek to move AI ethics discussion ahead by experimenting with ways to measure AI ethics criteria. The following document describes the outcomes and learnings from The Ethical Artificial Intelligence Quantification Workshop held at the National Institute for Aerospace (NIA), Hampton, Virginia on May 12th, 2022. The purpose of the workshop was for participants to evaluate and experiment-with the methodology and process presented by AIEthics.World in cooperation with Intel Corporation. The meeting participants learned about the Ethical AI Certification and Maturity Model™ and applied the methodology to selected notional AI systems. The workshop facilitated the evaluation of the maturity of the AI system according to ethical considerations relevant to NASA, NIA and other participants. The workshop consisted of three main phases. The first phase focused on understanding and summarizing NASA’s ethical approaches, mission and values based on published documentation, discussions and individual insights & opinions of participants. This information was prioritized, weighted, ordered, and quantified in phase two, to formulate an alignment between human values (ethics) and their applicability to AI systems during all lifecycle phases. The first two phases were summarized as a form of ethical genealogy for artificial intelligence, specific to NASA’s ethical approaches. In the third and last phase of the workshop the participants evaluated notional examples of artificial intelligence to qualify and quantify its ability to adhere to the organizational ethics approaches, using the Ethical AI Certification and Maturity Model™. The workshop uses the concept of genealogy, in the traditional sense: the study and traceability of lines of ancestors in the process of evolutionary development from earlier forms. However, as it is applied to an Ethical AI definition, it is providing the insights to the necessary and mandatory traceability of content, data, metrics, telemetry, elements, and structures which are used in the AI’s lifecycle to foster and measure AI ethics in all steps of its lifecycle. The Ethical Artificial Intelligence Quantification Workshop provided NASA with the opportunity to apply the Ethical AI Certification and Maturity Model™, in combination with existing and well-known decision-making and quality control methods to identify the metrics and measurements for an Ethical AI and assess its ethical condition and quality aligned with NASA ethics approaches. The result of the workshop is the capacity for NASA to apply the maturity model assessment to its AI Systems as desired and if necessary, publish the ability of these AI Systems to adhere to the organizational ethical goals. AI ethics frameworks need to be customized for each application domain, for example, individual NASA Mission Directorates. General principles that work in one area such as AI/Machine Learning-based text analysis (the ethics of information-extraction) may need to be adapted for another such as sense-and-avoid decision-making in a flight environment. The workshop was conducted among approximately twenty NASA subject matter experts, so the elements noted above should be considered examples, not definitive NASA ethical AI principles, genealogy, etc. Generating a definitive AI ethics framework for an organization as diverse as NASA would require far more discussion, debate, review, etc. However, the workshop provided valuable insight into mechanisms and processes for quantifying AI ethical qualities.

Artificial Intelligence↗

The Use of AIS Data for Identifying and Mapping Calcareous Soils in Western Nebraska

The identification of calcareous soils, through unique spectral responses of the vegetation to the chemical nature of calcareous soils, can improve the accuracy of delineating the boundaries of soil mapping units over conventional field techniques. The objective of this experiment is to evaluate the use of the Airborne Imaging Spectrometer (AIS) in the identification and delineation of calcareous soils in the western Sandhills of Nebraska. Based upon statistical differences found in separating the spectral curves below 1.3 microns, calcareous and non-calcareous soils may be identified by differences in species of vegetation. Additional work is needed to identify biogeochemical differences between the two soils.

Samson, S. A.↗

Demonstration Trials of AI/ML Edge+Cloud Suite (CRADA Final Report)

PACE AI and LBNL partnered under this CRADA to test and evaluate the PACE5 edge node prototype, an AI/ML edge and cloud-based suite, at FLEXLAB.The objective of the test was to evaluate the PACE5 edge node prototype's ability to perform demand shed and take to dynamic price signals, and to demonstrate advanced fault detection and microgrid monitoring capabilities.

97 MATHEMATICS AND COMPUTING↗