Search NASA⌕ Search

SEARCH · Search NASA

Results for “data coverage assessment”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Data coverage assessment on neural network based digital twins for autonomous control system

We report in a recently developed Nearly Autonomous Management and Control (NAMAC) system, neural networks (NNs) are used to develop digital twins for diagnosis (DT-Ds). However, NNs are not usually considered extrapolation models and may result in large errors if they are applied to unseen data outside the training data (uncovered). In this study, we propose a data coverage assessment (DCA) to determine if the NN-based DT-Ds are extrapolated based on their epistemic uncertainty. The uncertainty quantification algorithms and uncertainty thresholds are selected based on the confusion matrix of classifying evaluation data into covered or uncovered data. To demonstrate the adaptability of the proposed framework, we applied it to a basic feedforward neural network and a more advanced recurrent neural network based on a more nonlinear database. Case studies show that the proposed framework can distinguish unseen data for both basic and advanced applications with proper uncertainty quantification algorithms and thresholds.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Trustworthiness modeling and evaluation for a nearly autonomous management and control system

The Nearly Autonomous Management and Control (NAMAC) system supports the advanced reactor operation by recommending control actions to operators based on real-time measurements and digital twins (DTs) learning from the knowledge base. To enable the safe and reliable use of autonomous technologies, NAMAC and its recommendations should be trustworthy to operators and regulators at both the design and operation stages. This study proposes a NAMAC trustworthiness modeling and evaluation framework supported by trustworthiness ontologies and evidence-based approaches. The development-time and run-time ontologies are separately constructed and then converted to Bayesian networks to quantitatively evaluate the NAMAC trustworthiness. This evaluation is demonstrated by collecting and characterizing evidence from NAMAC practices, such as the development and assessment of the NAMAC system, data coverage assessment, and the training and optimizations of neural-network-based DTs. Our proposed approach can aggregate various trustworthiness attributes of complex artificial-intelligence-supported systems for safety-critical applications. It also considers the interaction between different DTs and extends beyond the trustworthiness evaluation of a single DT. In conclusion, the evidence-based method enhances the transparency of the trustworthiness modeling and evaluation processes and helps identify uncertainties and subjectivity involved in the processes.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Active Learning Meets Foundation Models: Fast Remote Sensing Data Annotation for Object Detection

Object detection in remote sensing demands extensive, high-quality annotations—a process that is both labor-intensive and time-consuming. In this work, we introduce a real-time active learning and semi-automated labeling framework that leverages foundation models to streamline dataset annotation for object detection in remote sensing imagery. For example, by integrating a Segment Anything Model (SAM), our approach generates mask-based bounding boxes that serve as the basis for dual sampling: (a) uncertainty estimation to pinpoint challenging samples, and (b) diversity assessment to ensure broad data coverage. Furthermore, our Dynamic Box Switching Module (DBS) addresses the well-known cold start problem for object detection models by replacing its suboptimal initial predictions with SAM-derived masks, thereby enhancing early-stage localization accuracy. Extensive evaluations on multiple remote sensing datasets plus a real-world user study, demonstrate that our framework not only reduces annotation effort, but also significantly boosts detection performance compared to traditional active learning sampling methods. The code for training and the user interface will be made available.

Burges, Marvin [ORNL] (ORCID:0000000312690769)↗

DER Cybersecurity Standards: Assessment and Gap Analysis

The purpose of this report is to share the comprehensive gap analysis of existing cybersecurity standards applicable to Distributed Energy Resources (DERs) within the electric power sector. This analysis aims to identify critical deficiencies in current standards, assess their alignment with industry needs, and provide actionable recommendations for enhancing cybersecurity measures. The scope encompasses various DER technologies, including solar, wind, energy storage, and hydrogen fuel cells, and emphasizes the significance of establishing robust cybersecurity frameworks and standards to safeguard these increasingly integrated systems. The report provides valuable insights for stakeholders in the DER ecosystem, including manufacturers, utilities, and regulators. It underscores the importance of continued development and refinement of cybersecurity standards to keep up with the technical advances in DERs and associated cybersecurity challenges. The analysis evaluated IEC, IEEE, ISA, ISO, and UL standards relevant to DER cybersecurity. Standards were assessed on their coverage of key requirements including data availability, integrity, confidentiality, access control, authentication, encryption, and system hardening. For each standard, the analysis assessed its alignment with current industry practices, regulatory compliance, effectiveness in addressing known risks, coverage of emerging risks, and how it promotes interoperability. The evaluation also considered potential integration challenges and barriers to adoption.

97 MATHEMATICS AND COMPUTING↗

United States CMM Insights Dataset

Geo-data science applications for critical mineral analysis, development acceleration, economic impact assessment, and project efficiency. Coverage: 10 years (2014-2023), 3 geographic levels (county, state, tract), 45 features Data Categories: - census: 30 features (e.g., asian population percentage, black population percentage, citizen voting age population percentage) - ejscreen: 3 features (e.g., P_DSLPM, P_PM25, P_PWDIS) - energyexpenditure: 12 features (e.g., Housing adjustment factor, Income adjustment factor, Monthly electricity cost)

AS↗

Prediction and Experimental Verification of Electrolyte Solvation Structure from an OMol25-Trained Interatomic Potential

A molecular-level understanding of electrolyte solvation structure and ion–ion correlations is critical to developing next-generation battery chemistries. Atomistic simulation capabilities with sufficient accuracy, speed, and transferability to deliver reliable structural insights while avoiding arduous system-specific reparameterization are thus highly desirable. Machine learning interatomic potentials (MLIPs) trained on large, chemically diverse data sets are revolutionizing computational chemistry, enabling molecular dynamics simulations of battery electrolytes with near-DFT accuracy over 10,000× faster than DFT. While previous MLIP training data sets with suitable elemental coverage for electrolytes have been based on inorganic materials, the Open Molecules 2025 (OMol25) data set provides large-scale molecular DFT MLIP training data with broad elemental coverage and specifically samples tens of millions of electrolyte configurations. Here, we integrate computational modeling with experimental validation to systematically assess the ability of large-scale MLIPs pretrained on materials data or on OMol25 to accurately resolve nanoscale structural organization and ion-solvation characteristics in Na-ion battery electrolytes across diverse physicochemical conditions and compositional regimes. We find that the OMol25-trained Universal Model of Atoms (UMA-OMol) predicts experimentally measured densities and X-ray structure factors in substantially better agreement compared to state-of-the-art models trained only on inorganic materials data. Using UMA-OMol, we further analyze systematic trends in solvation structure as a function of cation identity, anion chemistry, salt concentration, and solvent topology. We observe that increasing system temperature amplifies the heterogeneity within the solvation environment, perturbing cation–solvent interactions and promoting the formation of contact ion pairs (CIPs). Moreover, subtle variations in the solvent topology of glyme-based electrolytes cause pronounced changes in ion correlations and solvation structure. The experimental agreement and microscopic insights shown here position OMol25-trained MLIPs as a practical route to predictive, high-throughput electrolyte simulations beyond the limits of classical force fields and direct DFT molecular dynamics, serving as a powerful tool for accelerating the design of next-generation Na-ion battery electrolytes and beyond.

MLIPs↗

Analysis of Selected Publicly Available Geothermal Exploration Data Gaps

As part of a United States Department of Energy (DOE) supported retrospective analysis of DOE's Play Fairway Analysis (PFA) projects, the National Renewable Energy Laboratory (NREL) compiled and analyzed publicly available geothermal exploration datasets to identify and highlight data gaps in areas prospective for hosting geothermal resources. The analysis was intended to understand the existing geographic coverage of selected datasets commonly utilized both by the PFA projects and geothermal developers during resource assessments including geologic mapping, temperature gradient drilling, and aeromagnetic, gravimetric, and lidar surveys. Results indicate that broad areas of the western United States estimated to have geothermal potential lack sufficient geologic and geophysical coverage necessary for even regional resource exploration. The study directly informed the recent Geoscience Data Acquisition for Western Nevada, or GeoDAWN - which united DOE's Geothermal Technologies Office (GTO) with the U.S. Geological Survey (USGS) of the U.S. Department of the Interior to assist U.S. needs for energy and critical minerals. The study also has the potential to inform public investment in further data acquisition for characterization of the Earth both for geothermal and other natural resource assessments.

data↗

High-Resolution South American Wind Resource Data Downscaled with Generative Machine Learning Conditioned on Near-Surface Observations

High-resolution historical wind data was developed for the entirety of South America using the innovative Super-Resolution for Renewable Resource Data (sup3r) machine learning framework. The publicly available Sup3rWind South America dataset represents a significant advancement in wind resource data generation, leveraging generative machine learning conditioned on near-surface observations from the Meteorological Assimilation Data Ingest System (MADIS) to efficiently and accurately downscale coarse reanalysis data from the European Centre for Medium-Range Weather Forecasts (ERA5). This approach produces fine-scale, spatially and temporally coherent wind and meteorological fields hundreds of times more computationally efficient than traditional numerical weather modeling methods, enabling access to high-fidelity wind information across both continental and offshore regions. Sup3rWind South America builds on the earlier Sup3rWind Ukraine dataset through improvements in model architecture and outputs conditioned on near-surface observation inputs. As with the Ukraine data release, this dataset includes wind speed, wind direction, temperature, relative humidity, and pressure at a horizontal resolution of ~2 km, representing a 15x spatial enhancement relative to the 31 km ERA5 grid. Wind speed and direction are provided at 5-minute resolution, a 12x temporal refinement compared to the hourly ERA5 data, while temperature, relative humidity, and pressure remain at hourly resolution. The data covers all years from 2005 to 2024. Before downscaling, ERA5 inputs were bias-corrected using long-term monthly means and a limited number of quality-controlled observations to align large-scale statistics with regional conditions. The resulting dataset is the first publicly available high-resolution timeseries wind record that provides full spatial coverage of South America. Model validation demonstrates strong agreement with observations across several statistical metrics, consistent with other state-of-the-art high-resolution wind resource datasets. The potential applications of Sup3rWind South America span renewable energy resource assessment, energy system modeling, and grid resilience analysis. The 20-year record and high spatial and temporal resolution support accurate estimation of long-term energy yield and the economic feasibility of potential wind development sites. Continuous coverage across both continental and offshore regions enables comprehensive site prospecting within exclusive economic zones. The 2 km, 5-minute resolution data provide the spatial and temporal variability required for power system simulation, operational planning, and regional risk assessments.

17 WIND ENERGY↗

Quantitative Performance Assessment of Proxy Apps and Parents

The ECP Proxy Application Project has an annual milestone to assess the state of ECP proxy applications. Our FY21 milestone (ADCD-504-11) proposed to: Assess the performance and fidelity of proxy applications, including those in the ECP Proxy App Suite, relative to the ECP Application workload on heterogeneous platforms. Use proxy applications and selected ECP applications to assess the utility of critical elements of the Exascale toolchain, especially tools used to collect performance data. Identify gaps in coverage and/or common situations in which proxies may fail to adequately represent ECP applications.

97 MATHEMATICS AND COMPUTING↗

High redshift LBGs from deep broadband imaging for future spectroscopic surveys

Lyman break galaxies (LBGs) are promising probes for clustering measurements at high redshift, z > 2, a region only covered so far by Lyman-α forest measurements. Here, in this paper, we investigate the feasibility of selecting LBGs by exploiting the existence of a strong deficit of flux shortward of the Lyman limit, due to various absorption processes along the line of sight. The target selection relies on deep imaging data from the HSC and CLAUDS surveys in the g, r, z and u bands, respectively, with median depths reaching 27 AB in all bands. The selections were validated by several dedicated spectroscopic observation campaigns with DESI. Visual inspection of spectra has enabled us to develop an automated spectroscopic typing and redshift estimation algorithm specific to LBGs. Based on these data and tools, we assess the efficiency and purity of target selections optimised for different purposes. Selections providing a wide redshift coverage retain 57% of the observed targets after spectroscopic confirmation with DESI, and provide an efficiency for LBGs of 83 ± 3%, for a purity of the selected LBG sample of 90 ± 2%. This would deliver a confirmed LBG density of ~ 620 deg$^{-2}$ in the range 2.3 < z < 3.5 for a r-band limiting magnitude r < 24.2. Selections optimised for high redshift efficiency retain 73% of the observed targets after spectroscopic confirmation, with 89 ± 4% efficiency for 97 ± 2% purity. This would provide a confirmed LBG density of ~ 470 deg$^{-2}$ in the range 2.8 < z < 3.5 for a r-band limiting magnitude r < 24.5.A preliminary study of the LBG sample 3d-clustering properties is also presented and used to estimate the LBG linear bias. A value of b$_{LBG}$ = 3.3 ± 0.2 (stat.) is obtained for a mean redshift of 2.9 and a limiting magnitude in r of 24.2, in agreement with results reported in the literature.

79 ASTRONOMY AND ASTROPHYSICS↗

Remote Sensing of Live Fuel Moisture for Wildfires Using SMAP Satellite Observations

Live Fuel Moisture (LFM) is a critical parameter for wildfire risk assessment, traditionally measured by labor-intensive field sampling. However, sampled LFM data are influenced by site-specific factors, such as local vegetation types and plant traits, and are often collected retrospectively after wildfire events, making it difficult to obtain pre-fire data for predictive applications. Here, we evaluate the relationship between LFM and Vegetation Water Content (VWC) and Soil Moisture (SM) retrieved from SMAP L-band brightness temperature using the Maximum Entropy Production (MEP) approach. The MEP-retrieved VWC exhibited strong correlation with in situ measurements of LFM ( r > 0.6) in the Western U.S. The integration of high-resolution vegetation coverage data enhances the detection of sub-grid vegetation heterogeneity. This study demonstrates the operational potential of remote sensing derived VWC as a scalable proxy of LFM, supporting its application in regional assessment of wildfire risk.

Cho, Kyeungwoo [Georgia Institute of Technology, A↗

Exploring Sustainability in Scientific Software through Code Quality & Test Coverage Metrics

Context: Scientific open-source software (SciOSS) plays a foundational role in research and engineering, yet its long-term sustainability has often been overlooked and remains a significant concern. Objective: This study investigates the long-term sustainability of SciOSS through code and test quality metrics. Method: We analyze CASS Software Portfolio projects, classifying them by sustainability and comparing their code structure, test coverage, and links between code quality and testing across the dataset. Results: Sustainable projects show higher, more consistent test coverage and clearer code-test correlations, while unsustainable ones show weaker patterns. Overall, test coverage is low in scientific software, and high complexity and coupling reduce testability. Conclusion: In this study, we present a practical, data-driven approach for assessing sustainability in scientific software, offering a foundation for evaluating long-term software health and supporting future efforts in quality assurance and sustainability monitoring.

Md mushfiqur rahman, Sheikh [University of Tenness↗

Weather radar utility in hazard detection and response

Publicly accessible weather radar data have significant capabilities for meteorological measurements and predictions and, further, have the potential to measure nonmeteorological events that include smoke, ash, and debris plumes as well as explosions. The ability to identify and track nonmeteorological events can be of assistance in emergency response, hazard mitigation, and related activities in locations where radar coverage both exists and is recorded and accessible to the user. Here, in this study, events from multiple locations in the United States that are reported in news outlets are assessed using a manual inspection process of Level 2 weather radar data to identify anthropogenic and nonbiological returns. Explosive events are also identified, and a large high-altitude debris cloud from the intentional destruction of the SpaceX Starship is tracked across a wide area. Finally, future efforts using a machine learning model are discussed as a means of automating the process and potentially enabling near-real-time nonmeteorological event identification in the same areas where the data are accessible. Using weather radar data can be a valuable new tool for Department of Defense systems to aid in military awareness, and for interagency emergency response and forensic mission experts to consider national weather service data in their mission profiles. Radar data can be effective in detecting several common types of emergencies and inform and aid response personnel.

54 ENVIRONMENTAL SCIENCES↗

A Risk-Informed Approach to Trustworthiness Assessment in Digital Twins-Based Autonomous Control

In autonomous control systems, digital twins (DTs) are used to perform diagnostic and prognostic functions. The trustworthiness of these DTs is dependent on quality and coverage of the training data, model accuracy and integrity of sensor data. This work introduces a methodology to determine the trustworthiness of a DT system given faulty sensor data using a risk informed approach. Bayesian Belief Networks (BBNs) are used to propagate uncertainties and determine the probability of trustable recommendations. The decision to trust the control action provided by the DT is based on the DT output, expert opinion, and severity of problems. The performance of DTs is reliant on the data they are trained on. When they encounter out of distribution data, the trustworthiness of the recommendations decreases. To address this issue, we include an expert component that provides input on sensor degradation. For this, we utilize a generative artificial intelligence (AI) model, such as Generative Pretrained Transformer (GPT). The GPT functions as an expert with broad knowledge. The GPT is fine-tuned to understand and discriminate sensor degradation scenarios using manufactured data. This methodology is demonstrated through a case study on a Nearly Autonomous Management and Control System (NAMAC) during a steady state scenario. Various sensor degradation types with different severity levels are considered. Degraded sensor data is processed by the DT system and the fine-tuned GPT. Finally, using the BBN, we combine the GPT information and the DT output with its sources of uncertainty. This provides an output regarding the trustworthiness of the DT recommendation.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

A multi-scale temperature-based strategy to map hydrologic exchange flows in highly dynamic systems

Mapping and quantifying hydrologic exchange flows (HEFs) is critical to environmental monitoring and remediation at contaminated sites; however, these objectives are challenging in highly dynamic systems, e.g., along dam-regulated rivers, where HEFs vary rapidly. Direct seepage measurements are labor-intensive and difficult to automate, whereas indirect (e.g., thermal) and remote sensing methods have potential to allow continuous monitoring with limited field effort. We present a preliminary assessment of a multi-scale temperature-based strategy for monitoring HEFs along the Hanford Reach of the Columbia River, in eastern WA, United States. Five thermal methods were assessed. First, a vertical temperature profile (VTP) was installed into the streambed. The VTP data were analyzed using a data assimilation algorithm designed for automated real-time estimation in dynamic systems. Second, a thermal infrared (TIR) camera was used in roving surveys to identify seeps. Third, a TIR camera was stationed at the VTP site to collect images at 1-h intervals. Together, the two TIR datasets provided a basis to assess the potential for drone-based TIR. Fourth, temperature was measured at the sediment/water interface to assess fiber-optic distributed temperature sensing. Fifth, imagery from the ECOSTRESS satellite mission was acquired to assess the potential of spaceborne thermal monitoring. Based on our preliminary assessment, VTP, TIR, and bed temperature measurements provide complementary spatial coverage, temporal sampling, and resolution; these methods have potential for long-term, automated monitoring of HEFs. The publicly available spaceborne imagery, however, proved inadequate because of insufficient spatial resolution and data gaps resulting from cloud cover and revisit frequency.

54 ENVIRONMENTAL SCIENCES↗

Evaluation of Differential Peptide Loading on Tandem Mass Tag-Based Proteomic and Phosphoproteomic Data Quality

Global and phosphoproteome profiling has demonstrated great utility for the analysis of clinical specimens. One major barrier to the broad clinical application of proteomic profiling is the large amount of biological material required, particularly for phosphoproteomics—currently on the order of 25 mg wet tissue weight, depending on tissue type. For hematopoietic cancers such as acute myeloid leukemia (AML), the sample requirement is in excess of 10 million (1E7) peripheral blood mononuclear cells (PBMCs). Throughout the course of a prospective study, this requirement will certainly exceed what is obtainable from many of the individual patients/timepoints. For this reason, we were interested in examining the impact of differential peptide loading across multiplex channels on proteomic data quality. Methods: To achieve this, we tested a range of channel loading amounts (20, 40, 100, 200, and 400 µg of tryptic peptides, or approximately the material obtainable from 5E5, 1E6, 2.5E6, 5E6, and 1E7 AML patient cells) to assess proteome coverage, quantification reproducibility and accuracy in experiments utilizing isobaric tandem mass tag (TMT) labeling. As expected, we found that fewer missing values are observed in TMT channels with higher peptide loading amounts compared to those with lower loading. Moreover, channels with lower loading amounts have greater quantitative variability than channels with higher loading amounts. Statistical analysis of the differences in means among the five loading groups showed that the 20 µg loading group was significantly different from the 400 µg loading group. However, no significant differences were detected among the 40, 100, 200 and 400 µg loading groups. Conclusions: These assessment data demonstrate the practical limits of loading differential quantities of peptides across channels in TMT multiplexes, and provide a basis for designing the optimal clinical proteomics study when specimen quantities are limited.

59 BASIC BIOLOGICAL SCIENCES↗

Developing new pathways for energy and environmental decision-making in India: a review

Abstract India faces a dual challenge of economic development and responding to climate change. Although India’s per capita emissions are well below global average, the country is one of the world’s largest greenhouse gas emitters. Indian policymakers and stakeholders require high-quality data and research to assess low-emissions, sustainable development strategies. Peer-reviewed literature is a key source of this information and also a key venue for conversation amongst research leaders. This paper examines the recent peer-reviewed literature on India’s 2030 and 2050 pathways. We conducted a systematic literature review to identify key quantitative national modeling studies. From the 34 studies identified, we synthesized scenario data to draw common conclusions and identify critical research gaps. The main focus was on examining the coverage and the state of information available on low-carbon pathways. Overall, we find a few scenarios that are potentially consistent with a 2070 net-zero goal, but more limited assessment of pathways to reach net-zero emissions before this date. Mitigation pathways with greater ambition are required across all energy sectors to ensure a smooth transition to net-zero emissions by or before 2070. The scenarios confirm that reducing emissions to below 2 GtCO 2 yr −1 by mid-century would necessitate significant transformations of the Indian energy sector, such as, a decrease in unabated coal power capacity, transportation modal shift, and industrial process switching. The assessment also finds substantial differences in final energy estimates reported across studies, particularly in transportation. The lack of consistency in, and transparency about underlying drivers, assumptions, and even outputs across studies points to the critical need for the sorts of coordinated, multi-model studies that have proven exceptionally valuable for decision makers in other major emitting countries.

54 ENVIRONMENTAL SCIENCES↗