Search NASA⌕ Search

SEARCH · Search NASA

Results for “data gap”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Expediting field-effect transistor chemical sensor design with neuromorphic spiking graph neural networks

Improving the sensitive and selective detection of analytes in a variety of applications requires accelerating the rational design of field-effect transistor (FET) chemical sensors. Achieving high-performance detection relies on identifying optimal probe materials that can effectively interact with target analytes, a process traditionally driven by chemical intuition and time-consuming trial-and-error methods. To address the difficulties in probe screening for FET sensor development, this work presents a methodology that combines neuromorphic machine learning (ML) architectures, specifically a hybrid spiking graph neural network (SGNN), with an enriched dataset of physicochemical properties through semi-automated data extraction using large language models. Achieving a classification accuracy of 0.89 in predicting sensor sensitivity categories, the SGNN model outperformed traditional ML techniques by leveraging its ability to capture both global physicochemical properties and sparse topological features through a hybrid modeling framework. Next-generation sensor design was informed by the actionable insights into the connections between material properties and sensing performance offered by the SGNN framework. Through virtual screening for the detection of per- and polyfluoroalkyl substances (PFAS) as a use case, the effectiveness of the SGNN model was further validated. Density functional theory simulations confirmed graphene as a promising active material for PFAS detection as suggested by the SGNN framework. By bridging gaps in predictive modeling and data availability, this integrated approach provides a strong foundation for accelerating advancements in FET sensor design and innovation.

Ferreira, Rodrigo Pires [Univ. of Chicago, IL (Uni↗

Verification, Validation, and Calibration Through a Causal Lens

While typical validation and verification approaches focus on identifying the associations between data elements using statistical and machine learning methods, the novel methods in this paper focus instead on identifying causal relationships between data elements. Statistical and machine-learning-based approaches are strictly data-driven, meaning that they provide quantitative comparison measures between data sets without explicitly considering the hypotheses behind them. This can lead to the erroneous conclusion that, if two data sets are close enough, the models that generated them are similar. In addition, when experimental and simulated data differ to an extent that fails to meet the acceptance criteria, calibration techniques are used to tweak simulation model parameters to reduce the gap between the two types of data. This produces the false expectation that a simulation model will match reality. The methods presented in this paper move away from these strictly data-driven methods for validation and calibration toward more robust, model-driven methods based on causal inference. Causal inference aims to identify the possible mechanisms that might have generated data. Thus, this analysis targets the prediction of the effects when one (or more) of the identified mechanisms are altered. There are many approaches to identify, quantify, and illustrate causal relationships. For the scope of this paper, directed graphs are employed as causal models. If the directed graph lacks cycles, it is known as a directed acyclic graph. A node in such a graph represents an observed data element while a directed edge connecting two nodes represents a causal relationship between two variables. The developed causal methods are designed to extract causal models from simulation models and experimental data. Causal models capture the causal relationships between data elements (e.g., simulated and experimental data). In this context, validation and verification are performed by comparing causal models. The proposed approach does not only inform system analysts on how a simulation model matches real-world data, but also identifies elements of the simulation model that should be revised when discrepancies between simulation and experimental data are observed. Through these causal methods, analysts can identify the portion of the model equation(s) that are behind an edge connecting two variables. Hence, once the structural differences between causal models have been determined, model calibration can occur by changing only those model parameters that impact the identified causal relationships.

97 MATHEMATICS AND COMPUTING↗

LandScan mosaic enables high-resolution gridded population estimates with explicit uncertainty

Gridded population datasets represent high-resolution distributions of human occupancy, enabling informed decision-making across a broad range of fields. These data products are valuable for assessing environmental risk, urban development, disaster preparedness and resource allocation—areas where accurate population estimates directly enhance policy effectiveness and optimize resource distribution. Despite the importance of gridded population datasets, traditional population modeling approaches often overlook inherent uncertainties in the estimation process. This limitation can create a false sense of certainty in population estimates, potentially leading to flawed decisions by those who rely on the data. To address this methodological gap, we introduce a probabilistic machine learning modeling framework, LandScan Mosaic, that explicitly incorporates uncertainty into the population modeling process. Our approach systematically quantifies uncertainty in three key modeling parameters of the LandScan HD gridded population dataset: building use types, floor counts, and occupancy rates. By employing Monte Carlo simulations, we propagate these uncertainties through the modeling process, yielding probability distributions of population counts in place of deterministic point estimates. We demonstrate the practical application of this framework in Iloilo City, Philippines, using structured decision-making techniques and our probabilistic estimates to identify and prioritize areas most affected by projected flooding, supporting targeted interventions that address both economic and social risks. In doing so, we propose a population-specific approach for incorporating confidence into structured decision making processes. Through a comparative analysis with conventional deterministic approaches and point estimate approaches, including LandScan HD and WorldPop, we evaluate how the incorporation of machine learning and uncertainty influences decision rankings. This research advances population distribution modeling by offering a robust, quantitative approach that explicitly accounts for uncertainty in the underlying data, along with guidance for how users can apply uncertainty in their decision-making.

Environmental sciences↗

Magnon gap tuning in lithium-doped MnTe

Data in this DOI includes: 1) Powder Diffraction at HFIR POWDER (HB-2A) of pure MnTe and 5%-lithium doped MnTe at various temperatures ranging from T = 4 - 360 K for pure MnTe and T = 4 - 290 K for 5% lithium doped MnTe. 2) Inelastic neutron scattering at ARCS of pure MnTe and 5% lithium doped MnTe at T = 10 K measured with incident energy Ei = 30 meV and 150 meV, and corresponding background files (i.e. empty can measurements)

diffraction↗

Roadmap for transforming heterogeneous catalysis with artificial intelligence

Artificial intelligence (AI) is poised to transform heterogeneous catalysis, opening avenues for catalytic materials discovery. By uncovering intricate patterns in high-dimensional data, AI has been reshaping our pursuit of sustainable catalytic processes across the energy, environmental and chemical sectors. This promise, however, hinges on overcoming fundamental barriers, including limitations in data availability and quality, challenges in the generalizability and interpretability of data-augmented decisions, and the persistent gap between in silico predictions and experiments. Furthermore, we outline a forward-looking roadmap for deeply integrating AI into heterogeneous catalysis with an AI-ready data ecosystem, multimodal foundation models, and ultimately autonomous laboratories to accelerate the development of next-generation catalytic technologies via AI-empowered human–machine collaboration.

Computational methods↗

Carbon cycling across ecosystem succession in a north temperate forest: Controls and management implications

Despite decades of progress, much remains unknown about successional trajectories of carbon (C) cycling in north temperate forests. Drivers and mechanisms of these changes, including the role of different types of disturbances, are particularly elusive. To address this gap, we synthesized decades of data from experimental chronosequences and long-term monitoring at a well-studied, regionally representative field site in northern Michigan, USA. Our study provides a comprehensive assessment of changes in above- and belowground ecosystem components over two centuries of succession, links temporal dynamics in C pools and fluxes with underlying drivers, and offers several conceptual insights to the field of forest ecology. Our first advance shows how temporal dynamics in some ecosystem components are consistent across severe disturbances that reset succession and partial disturbances that slightly modify it: both of these disturbance types increase soil N availability, alter fungal community composition, and alter growth and competitive interactions between short-lived pioneer and longer-lived tree taxa. Further, these changes in turn affect soil C stocks, respiratory emissions, and other belowground processes. Second, we show that some other ecosystem components have effects on C cycling that are not consistent over the course of succession. For example, canopy structure does not influence C uptake early in succession but becomes important as stands develop, and the importance of individual structural properties changes over the course of two centuries of stand development. Third, we show that in recent decades, climate change is masking or overriding the influence of community composition on C uptake, while respiratory emissions are sensitive to both climatic and compositional change. In synthesis, we emphasize that time is not a driver of C cycling; it is a dimension within which ecosystem drivers such as canopy structure, tree and microbial community composition change. Changes in those drivers, not in forest age, are what control forest C trajectories, and those changes can happen quickly or slowly, through natural processes or deliberate intervention. Stemming from this view and a whole-ecosystem perspective on forest succession, we offer management applications from this work and assess its broader relevance to understanding long-term change in other north temperate forest ecosystems.

54 ENVIRONMENTAL SCIENCES↗

Winds of fortune? Understanding the geographic, sociodemographic, and temporal distribution of benefit mechanisms from land-based wind projects in the United States

Despite decades of research on factors shaping local responses to wind development, there is relatively little known about benefit mechanisms (e.g., agreements, funds, donations) used by developers in the U.S. land-based wind sector. To address this gap, we collected benefit mechanism data across all current utility-scale land-based wind projects installed between 1982 and 2024 (n = 1047), finding that just under one-third of projects had a benefit mechanism attached to them. We find the use of benefit mechanisms has become more common over time, is associated with larger projects, and varies by region. In terms of host community characteristics, the use of benefit mechanisms is associated with characteristics like higher education level, higher percent white, higher percent Republican, higher decision-making capacity, lower unemployment rate, and higher poverty rate. Building on a theoretical framework of purposes, we discuss what these findings could suggest about the motivations driving developers' use of these mechanisms, such as increasing local acceptance of a wind project or supporting distributive fairness. This first-of-its-kind study builds a comprehensive understanding of how benefit mechanisms have been used in the U.S. wind industry throughout its history, which can inform future approaches to benefit-sharing across sectors.

17 WIND ENERGY↗

Scanning tunneling spectroscopy of surface-oxidized Gallium-stabilized δ -phase plutonium

Here, the first contiguous experimental scanning tunneling spectroscopy (STS) measurements of the electronic structure of a plutonium (Pu) surface are presented and compared to current theoretical results. The experiments took place under UHV conditions, for which Pu is known to retain a native surface oxide. The STS data shows that no band gap is observable at room temperature. Theoretical calculations predict that bulk Pu oxides are semiconductors, but the spectral results for the cleaned Pu surface native oxide under UHV conditions could be better characterized as that of a semimetal.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Linear Discriminant Analysis-Based Machine Learning and All-Atom Molecular Dynamics Simulations for Probing Electro-Osmotic Transport in Cationic-Polyelectrolyte-Brush-Grafted Nanochannels

Deciphering the correct mechanisms governing certain phenomena in polyelectrolyte (PE) brush grafted systems, revealed through atomistic simulations, is an extremely challenging problem. In a recent study, our all-atom molecular dynamics (MD) simulations revealed a non-linearly large electroosmotic (EOS) flow (in the presence of an applied electric field) in nanochannels grafted with PMETAC [Poly(2-(methacryloyloxy)ethyl trimethylammonium chloride] brushes. Given the lack of any formal procedure that would have directed us to identify the correct factors responsible for such an occurrence, we needed to spend several months and devote significant analyses to unravel the involved mechanisms. In this paper, we propose a Linear Discriminant Analysis (LDA) based Machine Learning (ML) approach to address this gap. At first, we obtain data on certain basic features from the all-atom MD data. These basic features represent the number of atoms of certain species around one atom of another (or same) species. Here, we obtain such data on basic features for a reference case (case of an EOS flow in PMETAC-brush-grafted nanochannels with a smaller electric field) and a perturbed case (case of an EOS flow in PMETAC-brush-grafted nanochannels with a larger electric field) in bins in which the nanochannel half height has been divided into. These datasets are high-dimensional dataset, to which the LDA is applied. This leads to the projection of the data (between the reference and the perturbed states) in a highly separated form on a 1D line. From such LDA calculations, we are able to identify the relative importance of the different basic features in ensuring this separation of the data (between the reference and the perturbed states) on the 1D line. This relative importance of the different basic features is quantified as “importance scores” for the different features, which in turn tell us what to study and where to study. Such knowledge enables us to rapidly identify the key factors responsible for the non-linearly large EOS transport in PMETAC-brush-grafted nanochannels.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Interplay between Mixed and Pure Exciton States Controls Singlet Fission in Rubrene Single Crystals

Singlet fission (SF) is a multielectron process in which one singlet exciton S converts into a pair of separated triplet excitons T. SF is widely studied as it may help overcome the Shockley−Queisser efficiency limit for semiconductor photovoltaic cells. To elucidate and control the SF mechanism, great attention has been given to the identification of intermediate states in SF materials, which often appear elusive due to the complexity and fast time scales of the SF process. Here, we apply 14 fs-1 ms transient absorption techniques to high-purity rubrene single crystals to disentangle the intrinsic fission dynamics from the effects of defects and grain boundaries and to identify reliably the fission intermediates. Our data demonstrates that above-gap excitation directly generates a hybrid vibronically assisted mixture of singlet state and triplet-pair multiexciton [S/TT], which rapidly (<100 fs) and coherently branches into pure singlet or triplet excitations. The relaxation of [S/TT] to S is followed by a relatively slow and temperature-activated (48 meV activation energy) incoherent fission process. The SF competing pathways and intermediates revealed here unify the observations and models presented in previous studies of SF in rubrene and offer alternative strategies for the development of SF-enhanced photovoltaic materials.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Re-examining urban rainfall enhancement over North America

Abstract Quantifying intensification/suppression of precipitation over urban areas relative to their rural surroundings can inform efforts to reduce urban flooding. Few studies have systematically addressed whether urban areas exhibit a higher/lower probability of precipitation and/or higher/lower annual total precipitation and/or intensification/weakening of intense precipitation events relative to nearby rural areas across a range of hydroclimatic conditions and urban contexts. Here we address this literature gap using the IMERG V07 data set and analyses of rural and urban samples drawn from 47 conurbations across North America. Specifically, we quantify whether/how precipitation regimes over the urban grid cells differ from those in rural grid cells located 100–250 km from the city center and at a similar elevation. As in previous research, there is evidence that both the probability of precipitation and annual total precipitation are typically higher in the urban grid cells. However, most conurbations have lower upper percentile precipitation rates in the urban sample and lower median precipitation rates above the 95th percentile than are present in samples drawn from rural grid cells. Thus, these conurbations are not, on average, intensifying high-magnitude precipitation events over urban grid cells. Further, the total volume of water accumulated at the surface during events of equivalent duration is not systematically higher over the urban areas, and 20 year return period values of 30 min and wettest pentad precipitation are also not systematically higher over the urban areas. The nature of urban modification of precipitation is a strong function of the prevailing hydroclimate. For example, the heaviest rainfall periods are enhanced over urban grid cells within regional hydroclimates where the overall probability of precipitation and annual total precipitation are low. Conversely, there is evidence for urban suppression of the highest percentile precipitation rates in wetter hydroclimates.

Pryor, Sara C. (ORCID:0000000348473440)↗

EPWgen (EPW generator) [SWR-26-017]

EPWgen fetches hourly station observations (NOAA/Meteostat), fills gaps with MERRA2 reanalysis, merges data, and writes EPW files with computed headers (HDD/CDD, ground temperatures). It also runs QC checks, supports CSV-driven batch and metered-variable exports, and provides a PyQt5 GUI with mapping and progress tracking for single/multi-year workflows.

Bianchi, Carlo [National Laboratory of the Rockies↗

An open-access simulated earthquake ground-motion database for an M7 Hayward Fault earthquake in the San Francisco Bay Region

Comprehensive understanding of earthquake ground motions, particularly in the near-fault region of large-magnitude events, is limited by gaps in strong-motion data. This challenge is prominent in areas with high seismic hazard but infrequent large earthquakes where data is sparse and difficult to interpret. These data limitations lead to uncertainties in the development of site-specific ground motions, which are crucial for engineering risk assessments. To address these challenges, physics-based regional-scale ground-motion simulations have been developed. With the emergence of exaflop-scale computing ecosystems, it is now possible to simulate regional earthquake processes at unprecedented fidelity and generate the large number of fault rupture realizations necessary to characterize both intra- and inter-event ground-motion variability. This article introduces a new database of simulated earthquake ground motions, created for applications in earthquake engineering, earthquake planning, and emergency response. The inaugural version of the database features simulated ground motions for a magnitude 7 Hayward Fault earthquake in the San Francisco Bay Region (SFBR), using the EarthQuake SIMulation (EQSIM) simulation framework and the Graves–Pitarka kinematic rupture model. The aim is to provide high-fidelity, spatially dense, three-component motions generated on the Department of Energy’s (DOE) newest generation of graphics processing unit (GPU)-accelerated supercomputers. These motions are being made openly available to the engineering, scientific, and disaster planning communities. In addition, this work develops protocols for the efficient dissemination of these large data sets and emphasizes community engagement to build confidence in their application. This article discusses the methodology behind the data, underlying software verification and validation, scalable data management, and a user interface for data access. The goal is to facilitate widespread use and elicit expert feedback to maximize the utility and exploitation of simulated motions. While the initial focus is on the San Francisco Region, simulations for additional regions will be added as the DOE program progresses.

Simulated ground-motion database↗

Oklo Sponsored Testing using PELICAN: System, Subchannel, and High-Fidelity Software Validation (Final CRADA Report)

Argonne National Laboratory (the Contractor), located in Lemont IL, and Oklo, Inc., (the Participant), headquartered in Santa Clara, CA, propose to enter into a Cooperative Research and Development Agreement (CRADA) to perform a gap analysis of thermal hydraulic data, perform prototypical fuel assembly pressure drop and cavitation model validation, generate the experimental data as well as the corresponding uncertainties for this matrix and, finally, develop the validation models with the Argonne system level code SAS4A/SASSYS-1, the Argonne subchannel analysis code DASSH, the Argonne high fidelity code Nek5000, and/or the Idaho National Laboratory code Pronghorn Subchannel. The work outlined below will significantly improve the experimental and validation database currently available for liquid metal fast reactors, thus making it a viable part of a comprehensive reactor design and licensing suite to be used by the participant.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Elucidating Processes Controlling Arctic Atmospheric Aerosol Sources, Aging, and Mixing States (Final Report)

Atmospheric aerosols play critical roles in the Earth’s energy budget, directly by scattering or absorbing solar and terrestrial radiation and indirectly by serving as seeds (nuclei) for cloud droplet and ice crystal formation and by depositing on snow and ice surface, thereby changing the surface albedo. These effects are dependent on aerosol particle size and chemical composition and impact the hydrological cycle as well. This project provided single-particle size and chemical composition measurements across the entire annual cycle in the high Arctic and in the Alaskan Arctic during fall – winter, addressing the most significant gaps in Arctic aerosol observational data. These needs were based on recent rapid sea ice loss across the entire Arctic, as well as the major annual delays in sea ice freeze-up during fall in the Chukchi Sea and increased wintertime sea ice fracturing in the Beaufort Sea, both off the North Slope of Alaska. Two DOE Atmospheric Radiation Measurement (ARM) field campaigns were conducted for atmospheric aerosol sampling. The Aerosols during the Polar Utqiagvik Night (APUN – ‘snow on ground’ in Iñupiaq) ARM field campaign at Utqagivik, Alaska was conducted from Oct. 28 – Dec. 22, 2018. Aerosol sizing instrumentation and a single-particle mass spectrometer were successfully deployed for size-resolved number concentration measurements and measurements of individual particle size and chemical composition, respectively. These results show the influence of locally-produced sea spray aerosol, with high cloud-forming potential, due to delayed sea ice freeze-up in the fall. During the 2019‐2020 international Multidisciplinary drifting Observatory for the Study of Arctic Climate (MOSAiC) expedition, daily atmospheric aerosol particles were collected aboard the German icebreaker Polarstern in the Central Arctic from Nov. 2019 – Oct. 2020. Sea salt aerosol and marine organics were observed year-round during MOSAiC with varying morphologies and sources. These findings are important because most Arctic models do not include a sea spray aerosol source, despite this source increasing with declining sea ice extent. In addition to collecting new samples and data, this project also conducted further analysis of previously collected single-particle chemical composition measurements within the North Slope of Alaska oil fields and at Utqiaġvik, AK, during Aug. – Sep. 2015 and 2016 field campaigns. This work resulted in the discovery of chemical reactions of oil field combustion emissions occurring within fog droplets across the North Slope of Alaska and forming secondary aerosol, showing the impact of Arctic oil field emissions beyond black carbon aerosol and greenhouse gases. In addition, the distribution of chemical species across the aerosol population within the oil fields was quantified, using these data and a previously development framework. We also presented the first ambient evidence of the collision of two atmospheric particles resulting in formation of an organic-coated ammonium sulfate particle of marine origin, which has implications for cloud formation with declining sea ice extent. Overall, this project has elucidated connections between seawater biogeochemistry, resource extraction activities, atmospheric composition, clouds, and the energy budget of the Arctic region. The results of this project are expected to improve weather and sea ice forecasting for security and development in the Arctic and beyond.

54 ENVIRONMENTAL SCIENCES↗

Grid Operator Analytics and Assessment Tools for Inverter- Based Resources Dominated Grid (GOAAT-IBR) Project Update

This presentation provides an update on the OPTIMA GOAAT project, with emphasis on the cloud-native data platform developed in-house to ingest, manage, and operationalize high-resolution power system data. Since our last NASPI presentation, accessible via OSTI ID #2671437, the project team advanced the design and deployment of a scalable architecture capable of handling both synchronized and non-synchronized streams, including PMU, point-on-wave (POW), COMTRADE, and SCADA data. These materials review the project status, recent progress, and key lessons learned. The core of the presentation examines the architecture and engineering of our cloud-native ingestion and data management platform. We then explain how pipelines were designed to collect, normalize, time-align, store, and serve heterogeneous data at scale. We will discuss design choices such as data models, streaming versus batch ingestion, storage tiers, and interoperability with analytics applications. Practical experiences with cloud-native technologies were shared during the event, including benefits, limitations, and integration challenges in a utility environment, along with methods used to improve performance, reduce latency, and optimize resource usage. The presentation also showcases user interface designs and visualization tools that convert raw measurements and analytics results into intuitive, actionable insights for operators and engineers. During the presentation examples were provided demonstrating how visualization, event views, and summarized analytics enhance situational awareness and support operational decision-making. These use cases illustrate how a well-designed data infrastructure can bridge the gap between high-volume measurements and practical grid operations.

Aminifar, Farrokh↗

Myriad World Baseline: Global Geodemographic Estimates

The LandScan Myriad World Baseline (MWB) method produces global, residential (nighttime/home-location) gridded geodemographic estimates based on 5-year age/gender cohorts—at 30-arcsecond (≈1 km) resolution. MWB is designed to fill gaps where detailed, georeferenced survey data (e.g., Demographic and Health Surveys (DHS)) are missing or outdated, and to provide a baseline that can support human security analysis, including consequence assessment, “patterns of life” modeling, and scenario-based population futures. MWB’s workflow spatializes household-level age/gender characteristics from the GLOPOP-S dataset by conflating household and gridded expected relative wealth adapted from Global Gridded Relative Deprivation Index (GRDI), then adjusts them to a target year of interest. Age/gender estimates are then applied to harmonize lowest-administrative-level statistics with LandScan residential counts, yielding final geodemographic estimates. Two validation case studies are presented: Ghana (2021) and Tokyo/Kanagawa, Japan (2020), illustrating spatial variability in demographic cohorts and comparing MWB outputs to official gridded statistics. Results show close overall alignment relative to validation criteria including population pyramids and age-dependency ratios.

Tuccillo, Joe [ORNL] (ORCID:0000000259300943)↗

Manufacturing Facility Inventory National Dataset (M-FIND)

This asset provides a high-fidelity, validated inventory of manufacturing facilities across the United States, filling a critical gap in publicly available industrial data. By integrating and cross-referencing thirteen distinct data sources, this dataset moves beyond the limitations of single-source registries to provide a harmonized list that includes precise geographic coordinates, industrial subsector designations, and—crucially—parcel-level spatial boundaries.

Billings, Blake [ORNL] (ORCID:0000000186021600)↗