Search NASA⌕ Search

SEARCH · Search NASA

Results for “data gap”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Ion correlations explain kinetic selectivity in diffusion-limited solid-state synthesis reactions

Establishing viable solid-state synthesis pathways for novel inorganic materials remains a major challenge in materials science. Previous pathway design methods using pairwise reaction approaches have navigated the thermodynamic landscape with first-principles data but lack kinetic information, limiting their effectiveness. This gap leads to suboptimal precursor selection and predictions, especially for reactions forming competing phases with similar formation energies, where ion diffusion is a critical influence. Here we demonstrate an inorganic synthesis framework by incorporating machine learning-derived transport properties through ‘liquid-like’ product layers into a thermodynamic cellular reaction model. In the Ba–Ti–O system, known for its competitive polymorphism, we obtain accurate predictions of phase formation with varying BaO:TiO2 ratios as a function of time and temperature. We find that diffusion–thermodynamics interplay governs phase compositions, with cross-ion transport coefficients critical for predicting diffusion-limited selectivity. This work bridges length scales and timescales by integrating solid-state reaction kinetics with first-principles thermodynamics and spatial reactivity.

Atomistic models↗

PAH101: A GW+BSE Dataset of 101 Polycyclic Aromatic Hydrocarbon (PAH) Molecular Crystals

Abstract The excited-state properties of molecular crystals are important for applications in organic electronic devices. TheGWapproximation and Bethe-Salpeter equation (GW+BSE) is the state-of-the-art method for calculating the excited-state properties of crystalline solids with periodic boundary conditions. We present the PAH101 dataset ofGW+BSE calculations for 101 molecular crystals of polycyclic aromatic hydrocarbons (PAHs) with up to ~500 atoms in the unit cell. To the best of our knowledge, this is the firstGW+BSE dataset for molecular crystals. The data records include theGWquasiparticle band structure, the fundamental band gap, the static dielectric constant, the first singlet exciton energy (optical gap), the first triplet exciton energy, the dielectric function, and optical absorption spectra for light polarized along the three lattice vectors. The dataset can be used to (i) discover materials with desired electronic/optical properties, (ii) identify correlations between DFT andGW+BSE quantities, and (iii) train machine learned models to help in materials discovery efforts.

Science & Technology - Other Topics↗

Forty-year hydropower generation reanalysis for Conterminous United States

First published in 2022, the RectifHyd dataset provides hydrologically consistent estimates of monthly net generation for approximately 1,500 hydropower plants in the United States, addressing a gap in industrial surveys that have collected monthly generation data from only ~10% of plants post-2003. Here we present RectifHydPlus—an extended and enhanced dataset that improves on both the proxy information and temporal downscaling methodology adopted in RectifHyd. In addition to providing updated estimates of historical monthly generation for 590 plants with >10 MW nameplate capacity from 1980 through 2019, RectifHydPlus adds a hydrological control dataset that isolates the influence of historical water availability on generation. The new hydrological control dataset is suited to applications seeking to represent the capabilities of the contemporary fleet subject to historical interannual variability in climate. RectifHydPlus also includes a forty-year, daily-resolution, spill-adjusted water release time series for each dam, allowing users to aggregate generation estimates to the desired temporal resolution.

Hydroelectricity↗

Carbon source–driven metabolic and regulatory remodeling defines phenomic states in Lipomyces starkeyi

Lipomyces is a genus of oleaginous yeasts with potential for contributing to reliable biomanufacturing supply chains. However, progress in advanced strain designs and engineering efforts are still constrained by a lack of understanding of the underlying molecular drivers of Lipomyces phenotypes. To address this gap, we collected a suite of multi-omic data to dissect how carbon source availability reshapes the metabolic network, lipid allocation, and regulatory architecture of Lipomyces starkeyi. We observed that glucose promotes biosynthetic and proliferative processes supported by abundant energy and carbon intermediates, xylose enhances redox-balancing mechanisms centered on the pentose phosphate pathway, and glycerol activates respiratory metabolism, ß-oxidation, and the glyoxylate cycle. Lipid species distributions remained consistent in both nitrogen replete and depleted conditions across the carbon sources, indicating robust production mechanisms. Regulatory protein identification and network analysis revealed glycerol-driven respiratory growth favors regulatory programs integrating stress tolerance, redox balance, and lipid-associated metabolism, whereas xylose growth activates compensatory transcriptional responses aimed at maintaining mitochondrial function. Nitrogen limitation modulates the strength of these responses but does not fundamentally alter their direction, reinforcing carbon source as the dominant driver of regulatory architecture. Taken together, this data enhances the understanding of Lipomyces molecular rearrangements and provides a foundation for further development of predictive phenotypic tools in this genus.

Biotechnology↗

Machine learning and deep learning tools for the automated capture of cancer surveillance data

The National Cancer Institute and the Department of Energy strategic partnership applies advanced computing and predictive machine learning and deep learning models to automate the capture of information from unstructured clinical text for inclusion in cancer registries. Applications include extraction of key data elements from pathology reports, determination of whether a pathology or radiology report is related to cancer, extraction of relevant biomarker information, and identification of recurrence. With the growing complexity of cancer diagnosis and treatment, capturing essential information with purely manual methods is increasingly difficult. These new methods for applying advanced computational capabilities to automate data extraction represent an opportunity to close critical information gaps and create a nimble, flexible platform on which new information sources, such as genomics, can be added. This will ultimately provide a deeper understanding of the drivers of cancer and outcomes in the population and increase the timeliness of reporting. These advances will enable better understanding of how real-world patients are treated and the outcomes associated with those treatments in the context of our complex medical and social environment.

60 APPLIED LIFE SCIENCES↗

ORNL/mind_the_gap

Mind the Gap is an algorithm for detecting areas of missing data in building footprints.

Gonzales, Jack Joseph↗

Quality-Controlled Meteorological Data from the Flood Control District of Maricopa County (FCDMC) Network, Phoenix, Arizona (1987-2024)

This dataset contains 15- or 30-minute interval meteorological data from the Flood Control District of Maricopa County (FCDMC), Arizona, USA, covering eight key variables across multiple sensor stations between 1987 and 2024. Each variable is stored as a separate CSV file, containing time-series data that have undergone rigorous quality control (QC) procedures and, where appropriate, short-gap interpolation for consistency. The quality control (QC) pipeline consisted of four sequential tests: (1) a range test to ensure all values fall within physically realistic limits, (2) a step test to identify abrupt and implausible changes between consecutive records, (3) a proximity test that validates flagged values from step test using data from nearby stations and exceedance probability thresholds, and (4) a persistence test to detect and remove periods of unrealistically constant readings. These thresholds were calibrated to Arizona’s environmental conditions and sensor specifications. After QC, short gaps (≤2 hours) were linearly interpolated to ensure consistent temporal resolution, except for wind variables. Due to a major upgrade in FCDMC’s data transmission system, only ALERT-2 protocol data (2016–2024) for wind variables are included; earlier ALERT-1 data were excluded because of irregular sampling and high missing rates. This dataset supports regional climate and infrastructure resilience studies by providing standardized, high-resolution meteorological data for the greater Phoenix metropolitan area.

54 ENVIRONMENTAL SCIENCES↗

HP-MDR: High-performance and Portable Data Refactoring and Progressive Retrieval with Advanced GPUs

Scientific applications produce vast amounts of data, posing grand challenges in the underlying data management and analytic tasks. Progressive compression is a promising way to address this problem, as it allows for on-demand data retrieval with significantly reduced data movement cost. However, most existing progressive methods are designed for CPUs, leaving a gap for them to unleash the power of today’s heterogeneous computing systems with GPUs.In this work, we propose HP-MDR, a high-performance and portable data refactoring and progressive retrieval framework for GPUs. Our contributions are four-fold: (1) We carefully optimize the bitplane encoding and lossless encoding, two key stages in progressive methods, to achieve high performance on GPUs; (2) We propose pipeline optimization and incorporate it with data refactoring and progressive retrieval workflows to further enhance the performance for large data process; (3) We leverage our framework to enable high-performance data retrieval with guaranteed error control for common Quantities of Interest; (4) We evaluate HP-MDR and compare it with state of the arts using five real-world datasets. Experimental results demonstrate that HP-MDR delivers an average 13.68 × and 6.31 × throughput in data refactoring and progressive retrieval tasks, respectively. It also leads to 11.22 × throughput for recomposing required data representations under Quantity-of-Interest error control and 6.04 × performance for the corresponding end-to-end data retrieval, when compared with state-of-the-art solutions.

Li, Yanliang [University of Oregon]↗

Facilitating Data Collection of Maintenance Events to Populate the Hydrogen Component Reliability Database (HyCReD)

The Hydrogen Component Reliability Database (HyCReD) is a collaborative project between the National Renewable Energy Laboratory, the University of Maryland, and hydrogen stakeholders to improve safety and reliability for hydrogen facilities by implementing component reliability data taxonomies that support hydrogen infrastructure failure rate analysis. The project aims to quantify failure rates of hydrogen components through high-quality data collection and analysis on root causes and maintenance needed. HyCReD provides a common database for cataloging hydrogen component failures which exists for reliability research in many other mature industries [2]. The database fills a gap for the hydrogen community by providing a scientifically rigorous approach to quantitative risk assessment (QRA), prognostic health management (PHM), and reliability-centered maintenance (RCM) analysis. High level results will be aggregated and anonymized to protect company sensitive information; detailed results will be used to help address issues of hydrogen components. These advanced analytics will support accelerated deployment of hydrogen infrastructure by enabling better: design and safety of projects (safety codes and standards development), infrastructure reliability and cost (component failure rates, maintenance protocols), and component R&D needs (robust supply chain). A key to a successful HyCReD implementation is facilitating the ease of reporting and data quality in the database that can be used for analysis. Maintenance data was a previously identified gap in initial efforts to populate and validate the database taxonomies [3]. Collection of maintenance data will be instrumental in identifying failure modes and rates, identifying incipient component failures or reduced performance, cataloging best practices for maintenance routines and methods for prognostic health management, and quantifying the risk and effect of different failure modes. Several key priorities are identified for streamlined data collection to achieve quality and detailed failure data: Applicability, Ease of Use, Accessibility, and Information Security. The HyCReD team has now begun deployment of the database to several companies and groups that have signed non-disclosure agreements to facilitate the data collection of failures in industry hydrogen refueling station infrastructure. This paper will provide an update into the process of HyCReD deployment including the development of a coding guide for facility personnel to reference and ensure data quality and consistency from one station to another as well as implementation of contextually dependent data fields of system taxonomy and formatted entries to provide ease of use. The goal is to communicate the lessons learned from the roll-out to technicians and engineers in the field, and the addition of need for high level of security to protect all stakeholders.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

From Bricks to Clicks: Mapping the White Space in Building Innovation

It is a critical national imperative to transform the buildings sector, yet innovation is impeded by deployment failures that leave promising technologies stranded. Conventional market reports and techno-economic analysis provide an insufficient understanding of markets and resource allocation for emerging building technologies. They omit crucial commercialization factors such as ecosystem maturity and adoption friction, where the coordinated participation of a network of suppliers, contractors, financiers, regulators, and integrators is required to scale solutions. This study addresses these gaps by introducing an evaluation framework grounded in front-line data from six years of the DOE's IMPEL incubator, comprising experience from 300 building-sector innovators and the adjacent, complex ecosystem. Our methodology synthesizes top-down market analysis with bottom-up, practitioner-level data across five megatrends: (M1) Affordable materials and industrialized construction; (M2) Healthy and efficient mechanical systems; (M3) Intelligent building operations; (M4) Buildings as grid assets; and (M5) High-density power and cooling for data centers and therein identify twelve "white space" technology opportunities. Next, we develop a multi-criteria scoring rubric to rank these opportunities based on parameters, i.e., Affordability, Quality of Life, Reliability, and Security, yielding composite ‘Demand’ and ‘Maturity’ indices. Our results indicate that the most significant white spaces may not be incremental products but a new class of ‘Ecosystem Enablers’, such as logistics platforms, orchestration layers, and automated compliance software that solve structural deployment gaps. This paper summarizes this transparent, evidence-based, practitioner-informed evaluation framework for policymakers and investors to re-evaluate policy and resource allocation and unlock scalable market transformation.

Singh, Reshma↗

A ModEx Framework for Watershed Subsurface Investigation With Limited Geophysical Data Using Machine Learning and Hydrologic Modeling

Abstract Subsurface heterogeneity influences watershed hydrology strongly but remains difficult to characterize at catchment scales with sparse and costly field data. Geophysical surveys such as electromagnetic induction (EMI) provide local spatial subsurface images yet scaling them to watershed scales and converting EMI‐derived resistivity into hydraulic properties remains a challenge. We present a Model–Experiment (ModEx) framework that integrates limited EMI data with machine learning (ML) and hydrologic modeling to improve process representation and guide field investigations. Sparse EMI surveys were scaled to the catchment scale using a Random Forest model, and the resulting resistivity fields were combined with nearby borehole constraints to parameterize a hydrologic model. The EMI‐informed hydrological simulations improved predictions of streamflow sustained by subsurface flow and shallow saturation patterns. By combining EMI data and ML with hydrologic modeling, the ModEx framework guides future subsurface surveys, providing a transferable and efficient strategy for data–model integration across diverse watersheds. Plain Language Summary Mapping the underground network of soil and rock that controls water is essential for predicting floods and droughts, but seeing underground is difficult and expensive. We cannot drill everywhere, so scientists use geophysical tools to scan broad areas. There are two key challenges: these geophysical scans are often sparse across the whole watershed, and the geophysical data is hard to translate into water‐related properties. We used artificial intelligence to solve these problems. We taught a computer to find patterns linking the limited geophysical data to the land surface properties. This allowed it to fill in the gaps and create a complete, useful subsurface map for the entire watershed. This new map improves hydrologic simulations, leading to more accurate predictions of water movement in the watershed. It also helps scientists build better models with less data and generates a priority map showing where to measure next, making future investigations more efficient. Key Points Limited EMI scaled with ML improves catchment‐scale subsurface parameterization for hydrologic models The framework integrates hydrologic modeling with limited geophysical data to support subsurface investigation design ModEx framework offers a transferable data–model integration strategy that quantifies and reduces uncertainty guiding watershed studies

Chen, Hang↗

Development of Predictive Model for Accurate Rupture Time from Multi-Axial Creep in Alloy 709 with Physics-Based Simulations

A physics-based model is developed to predict multiaxial creep behavior in Alloy 709 (A709), an advanced austenitic stainless steel intended for high-temperature applications such as Sodium Fast Reactors (SFRs). Compared to conventional stainless steels like 316H, A709 offers superior high-temperature performance; however, comprehensive data on its multiaxial creep response remain limited. To address this gap, a crystal plasticity finite element (CPFE) framework is used to simulate the deformation and failure mechanisms of A709 under multiaxial loading conditions. The model incorporates an extended Hu-Cocks dislocation creep formulation that accounts for precipitation effects, along with the Sham–Needleman model to capture grain boundary cavitation-driven failure. These advanced constitutive models enable a detailed understanding of the interplay between microstructural evolution and macroscopic creep response. Furthermore, the study evaluates the predictive accuracy of various effective stress measures in estimating creep rupture life, leveraging simulated multiaxial creep data. The findings provide critical insights into the applicability of different stress measures for engineering design and life prediction of A709 components operating under complex loading conditions. This work contributes to improving the reliability of high-temperature structural components by advancing predictive modeling capabilities for advanced austenitic steels.

Alloy 709↗

Efficient distributed continual learning for steering experiments in real-time

Deep learning has emerged as a powerful method for extracting valuable information from large volumes of data. However, when new training data arrives continuously (i.e., is not fully available from the beginning), incremental training suffers from catastrophic forgetting (i.e., new patterns are reinforced at the expense of previously acquired knowledge). Training from scratch each time new training data becomes available would result in extremely long training times and massive data accumulation. Rehearsal-based continual learning has shown promise for addressing the catastrophic forgetting challenge, but research to date has not addressed performance and scalability. To fill this gap, we propose an approach based on a distributed rehearsal buffer that efficiently complements data-parallel training on multiple GPUs to achieve high accuracy, short runtime, and scalability. It leverages a set of buffers (local to each GPU) and uses several asynchronous techniques for updating these local buffers in an embarrassingly parallel fashion, all while handling the communication overheads necessary to augment input minibatches using unbiased, global sampling. We further propose a generalization of rehearsal buffers to support both classification and generative learning tasks, as well as more advanced rehearsal strategies (notably Dark Experience Replay, leveraging knowledge distillation). We illustrate this approach with a real-life HPC streaming application from the domain of ptychographic image reconstruction. Furthermore, we run extensive experiments on up to 128 GPUs of the ThetaGPU supercomputer to compare our approach with baselines representative of training-from-scratch (the upper bound in terms of accuracy) and incremental training (the lower bound). Results show that rehearsal-based continual learning achieves a top-5 validation accuracy close to the upper bound, while simultaneously exhibiting a runtime close to the lower bound.

Asynchronous data management↗

Leveraging generative artificial intelligence to bridge domain gaps in wind turbine research

A central challenge in wind turbine health monitoring is the scarcity of real-world data due to limited instrumentation, leading researchers to rely on simulation models that often suffer from reduced fidelity. However, even within simulation environments, discrepancies arise because of modeling assumptions, and configuration fidelities, creating domain gaps that limit the transferability of learned representations. Here, to investigate domain translation under controlled conditions, this project explores the use of generative artificial intelligence, specifically cycle-consistent generative adversarial networks (CGANs), to bridge the gap between OpenFAST simulation models representing 1.5 MW and 5 MW wind turbines. A physics-informed CGAN architecture is introduced, where a simplified turbine tower dynamics model is incorporated into the training loss to ensure physically consistent outputs. Quantitative results showed moderate to high agreement in frequency-domain features. Incorporating the physics-informed loss function improved the R 2 values by 30%, reduced the RMSE from 1.39 to 1.1 m/s 2 , and reduced training time by 82%. Furthermore, under increased turbulence intensity (IEC Category A), the RMSE remained stable at approximately 1.1 m/s 2 . While the present study is entirely simulation-based, it establishes a pipeline for evaluating physics-informed generative domain translation, which may serve as a foundation for future simulation-to-reality validation studies.

17 WIND ENERGY↗

Adaptive anomaly detection for identifying attacks in cyber-physical systems: A systematic literature review

Modern cyberattacks in cyber-physical systems (CPS) rapidly evolve and cannot be deterred effectively with most current methods, which focus on characterizing past threats. Adaptive anomaly detection (AAD) is among the most promising techniques to detect evolving cyberattacks, with an emphasis on fast data processing and model adaptation. AAD has been researched extensively; however, to the best of our knowledge, our work is the first systematic literature review (SLR) on current research in this field. We present a comprehensive SLR, gathering 397 relevant papers and systematically analyzing 65 of them (47 research and 18 survey papers) on AAD in CPS from 2013 to November 2023. We introduce a novel taxonomy considering attack types, CPS application, learning paradigm, data management, and algorithms. Our findings show that most studies addressed either model adaptation or data processing, but rarely both simultaneously. This indicates a research gap in fully adaptive solutions. We also categorize algorithms, datasets, and attack characteristics, and summarize strengths and weaknesses across the literature. Our review provides a structured and accessible reference for researchers and practitioners, offering insights into key trends and highlighting limitations in current approaches. Finally, we outline several future research directions, including the need for integrated real-time processing and adaptive learning, explainability, and uncertainty quantification in AAD for CPS.

Adaptation↗

Rising infrastructure inequalities accompany urbanization and economic development

Impending global urban population growth is expected to occur with considerable infrastructure expansion. However, our understanding of attendant infrastructure inequalities is limited, highlighting a critical knowledge gap in the sustainable development implications of urbanization. Using satellite data from 2000 to 2019, we examine country-level population-adjusted biases in infrastructure distribution within and between regions of varying urbanization levels and derive four key findings. First, we find long-run positive associations between infrastructure inequalities and both urbanization and economic development. Second, our estimates highlight increasing infrastructure inequalities across most of the countries examined. Third, we find greater future infrastructure inequality increases in the global south, where inequalities will rise more in countries with substantial urban primacy. Fourth, we find that infrastructure inequality may evolve differently than economic inequalities. Overall, advancing sustainable development vis-à-vis urbanization and economic development will require intentional infrastructure planning for spatial equity.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Roadmap for the future of extreme wildfire events

Background Extreme wildfire events (EWEs) represent a growing threat globally, posing substantial risks to ecosystems, human communities, and infrastructure. Despite increased recognition of their ecological, social, and economic significance, current definitions of EWEs vary widely, reflecting disciplinary biases and regional contexts. This article emerges from an interdisciplinary workshop convened to reassess and refine the definition of EWEs, examine their impacts across ecological and social dimensions, and identify critical knowledge gaps impeding our understanding of these infrequent but important events. Results Our synthesis highlights significant limitations with existing definitions, particularly their reliance on subjective thresholds and their emphasis on extreme fire behavior alone. EWEs encompass a spectrum of complex, multi-dimensional phenomena that extend beyond immediate biophysical characteristics to include cumulative social, economic, and ecological impacts. These impacts often manifest over extended timeframes and include hazardous environmental contamination, severe geomorphic disturbances, ecosystem transformations, and unintended consequences of post-fire management actions. Current wildfire modeling frameworks inadequately capture these compounding factors, particularly the interactions among social systems, ecological conditions, and extreme fire behavior. To overcome these issues, we advocate for an interdisciplinary and context-sensitive approach to defining and studying EWEs. This revised definition emphasizes wildfires exhibiting anomalies in fire behavior, ecological outcomes, or social impacts relative to historically observed baselines, accommodating variability across different geographic regions and ecological settings. Conclusions Adopting an interdisciplinary framework that integrates biophysical and social sciences will enhance the predictive capability of wildfire models and improve resilience planning and response strategies. Filling identified knowledge gaps—such as limited high-quality empirical fire behavior data and insufficient integration of social dynamics into modeling—will better prepare communities and ecosystems to cope with and adapt to EWEs. This inclusive approach underscores the necessity for collaboration across disciplines and sectors, essential to managing extreme wildfires in an era of increasing climatic and ecological uncertainty.

54 ENVIRONMENTAL SCIENCES↗

Cyber Resilience and Social Equity: Twin Pillars of a Sustainable Energy Future

This paper examines the intersection of security and accessibility within energy systems amidst the rise of grid modernization and digitization, especially considering the regulatory changes and the imperatives of inclusive energy strategies. It addresses the dual need for secure, resilient infrastructure and a commitment to mitigate energy poverty while maintaining equitable access to energy. Amid escalating cybersecurity and physical threats, the paper advocates for sustainable energy delivery systems that ensure robust defenses without compromising the goals of reducing energy poverty and ensuring energy security. This paper identifies the pressing need for Cyber-Informed Engineering (CIE) and Secure-by-Design (SbD) principles, highlighting how these strategies can protect critical infrastructure and democratize access to secure energy, particularly for disadvantaged communities. The analysis underscores the challenges presented by the expansion of attack surfaces, interoperability requirements, and grid-edge analytics, offering innovative solutions that leverage advanced technologies and data-driven insights. Furthermore, this paper addresses the workforce development gap, emphasizing the necessity for public-private partnerships and vendor engagement in creating a skilled cybersecurity workforce. This paper has a dual focus on both the technological aspect of cybersecurity and the social dimension of equity within the context of sustainable energy development. It suggests a comprehensive examination of how these two critical elements interact and support the overarching goal of a sustainable energy future.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗