Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data Integration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21

Using Machine Learning to Understand Electric and Hybrid Vehicles Ownership in Burdened and Nonburdened Communities

Transitioning to electric and hybrid vehicles (EHVs) for all communities is a pivotal step toward sustainable transportation and environmental conservation. This paper aims to understand the adoption of EHVs, focusing on burdened communities (BCs) in the United States. The EHV ownership-based analysis combines two datasets—behavioral data from the Puget Sound Regional Travel Survey integrated with BCs (Justice40) data covering transportation insecurity, environmental burden, social vulnerability, health vulnerability, and climate and disaster risk burden. After creating this unique database, descriptive analysis and modeling are used to analyze the data and predict EHV ownership in the future. Specifically, we use a new method that combines particle swarm optimization (PSO) with a stacking model named PSO-Stacking, which incorporates heterogeneous base learners of machine learning and deep learning. PSO applies a customized objective function to select the optimal hyperparameters for heterogeneous learners within the stacking model, effectively addressing challenges such as multicollinearity, data imbalance, nonlinearity, and overfitting. The proposed solution covers more accurate results than standard benchmark models for EHV ownership in BCs and non-BCs. In addition, the results of the PSO-Stacking method are explained using the local interpretable model-agnostic explanations technique. Results show a negative correlation between the BCs indicators, that is, higher transportation insecurity associated with lower EHV ownership. Furthermore, BCs have higher future climate risk scores, diesel particulate matter levels, and PM2.5 in the air than non-BCs because of higher conventional vehicle ownership. These communities are at higher risk and can benefit from electrification, EV infrastructure, and EV policies to address environmental challenges.

Aslam, Zeeshan [ORNL]↗

Enabling kilometer-scale E3SM land model simulation over North America: A new integrated framework solution

This study introduces a novel framework designed to enhance the performance, scalability, and portability of the kilometer-scale E3SM Land Model (km-ELM) within the E3SM modeling infrastructure. By seamlessly integrating cutting-edge data tools, we address existing challenges such as slow performance, limited scalability, and difficulties in software integration in current data-driven ELM simulation over large geographic areas. Our innovative approach leverages the KiloCraft data toolkit to generate unified inputs for simulations ranging from a single-cite case, to a 72,083-cell regional case to a continental configuration encompassing 21.6 million land grid cells at a 1 km × 1 km resolution. We conduct extensive strong- and weak-scaling experiments on three state-of-the-art supercomputers, utilizing up to 100,800 CPU cores across 2400 compute nodes to evaluate end-to-end metrics including wall-clock time, simulation-years-per-day (SYPD), initialization costs, and I/O throughput. Our results reveal the land (LND) component’s efficient scaling, demonstrating near-ideal weak scaling and strong-scaling parallel efficiencies reaching up to 87% at 50,400 cores. We confirm portability and reproducibility through bitwise-equivalent outputs across different machines using identical inputs over supported machines. Notably, at extreme scales, we identify I/O as a critical bottleneck and that leads to effective solution with the SCORPIO/ADIOS stack. Collectively, these findings validate the deployment of km-ELM at a continental scale with high parallel efficiency and provide essential guidance on configuration, decomposition, and I/O settings for optimized kilometer-scale land simulations in E3SM. This work emphasizes the innovative design and practical solutions that enhance the operational capabilities of km-ELM, focusing on software performance and scalability while leaving detailed scientific evaluations of simulated land processes for future investigations.

E3SM land model (ELM), km-ELM, scalability, perfor↗

An ultra-fast method for generating synthetic down-scattered neutron data for inertial confinement fusion implosions

In inertial confinement fusion experiments at the National Ignition Facility, asymmetries are probed by a variety of neutron diagnostics, including neutron imaging systems, real-time neutron activation diagnostics (RTNADs), and neutron spectrometers. It is often useful to generate synthetic data based on these diagnostics to validate and tune models. However, current methods of doing so using Monte Carlo particle tracing are time-consuming. In this paper, an ultra-fast method is presented for generating synthetic neutron images, RTNAD data, and spectrometry data using line integrals and 3D convolutions. While it does not contain as much physics as particle tracing codes, it is thousands of times faster and produces nearly identical data. This enables analysis techniques that depend on generating large amounts of synthetic data, which will prove very useful for the study of asymmetries going forward.

Deuterium↗

PIPES (Pipeline for Integrated Projects in Energy Systems) [SWR-24-89]

The Pipeline for Integrated Projects in Energy Systems (PIPES) is a comprehensive project, data, and workflow management tool designed for integrated modeling teams. PIPES facilitates the management of data requirements, tasks, and progress tracking, serving as a higher-level integration layer that works across various data and modeling software. This tool integrates models, data, and tools to perform large-scale, integrated analysis work at scale. PIPES is designed to streamline integrated modeling projects, enhance collaboration, and ensure the quality and efficiency of data management and workflow processes. https://github.com/nrel-pipes/pipes-api https://github.com/nrel-pipes/pipes-web https://github.com/nrel-pipes/nrel-pipes

Gu, Jianli↗

From 2D to 4D: a containerized workflow and browser to explore dynamic chromatin architecture

Background Characterizing the physical organization of the genome is essential for understanding long-range gene regulation, chromatin compartmentalization, and epigenetic accessibility. Hi-C experiments generate two-dimensional (2D) genome-wide contact maps of chromatin interactions by capturing the spatial proximity between genomic loci, which reveal interaction frequencies but lack the spatial resolution needed to interpret the three-dimensional (3D) genome structure(s). Emerging evidence suggests that epigenetic regulation is closely linked to 3D genome architecture, and that structural changes over time (4D) drive key biological processes in development, disease, and environmental response. Thus, integrating 3D structure with functional data is critical for a more complete understanding of genome regulation. Previous work, most notably the 4DHiC chromosome modeling framework, has shown that physical multi-dimensional modeling approaches rooted in polymer physics and molecular dynamics can resolve these structures at biologically meaningful resolutions by integrating temporal Hi-C data with physical constraints to uncover dynamic chromosome reorganization. Thus, molecular dynamics simulations, constrained by Hi-C contact matrices, can resolve fine-scale structural changes and reveal functionally significant transitions in chromatin conformation. Results Herein, we present the 4D Genome Browser Workflow (4DGBWorkflow) and the 4D Genome Browser (4DGB). The algorithm is based on the 4DHiC method, and the containerized tool is an end-to-end workflow that can transform, filter, and view 4D epigenomics and chromatin datasets, allowing non-specialists to apply three-dimensional modeling principles to diverse datasets and experimental conditions. The software executes on a laptop running macOS, Linux or Windows. From input Hi-C files (.hic), the 4DGBWorkflow produces 3D reconstructions of chromosomes, integrates the reconstruction with track data (e.g., epigenetic marks, transcriptome profiles), and provides comparative visualization of the results in a single workflow. Conclusions The 4DGBWorkflow and 4D Genome Browser are open-source tools for comparative analysis and visualization of 4D chromosome datasets, including chromatin architecture and epigenomic signals. Automatic integration of Hi-C data with molecular dynamics democratizes the construction of time resolved 3D genome structures, simplifying complex simulations and data integration schemes.

3D Genome Browser↗

Bayesian Conavigation: Dynamic Designing of the Material Digital Twins via Active Learning

Scientific advancement is universally based on the dynamic interplay between theoretical insights, modeling, and experimental discoveries. However, this feedback loop is often slow, including delayed community interactions and the gradual integration of experimental data into theoretical frameworks. This challenge is particularly exacerbated in domains dealing with high-dimensional object spaces, such as molecules and complex microstructures. Hence, the integration of theory within automated and autonomous experimental setups, or theory in the loop-automated experiment, is emerging as a crucial objective for accelerating scientific research. The critical aspect is to use not only theory but also on-the-fly theory updates during the experiment. Furthermore, we introduce a method for integrating theory into the loop through Bayesian conavigation of theoretical model space and experimentation. Our approach leverages the concurrent development of surrogate models for both simulation and experimental domains at the rates determined by latencies and costs of experiments and computation, alongside the adjustment of control parameters within theoretical models to minimize epistemic uncertainty over the experimental object spaces. This methodology facilitates the creation of digital twins of material structures, encompassing both the surrogate model of behavior that includes the correlative part and the theoretical model itself. While being demonstrated here within the context of functional responses in ferroelectric materials, our approach holds promise for broader applications, such as the exploration of optical properties in nanoclusters, microstructure-dependent properties in complex materials, and properties of molecular systems.

Microscopy↗

Data Generation for Machine Learning Interatomic Potentials and Beyond

The field of data-driven chemistry is undergoing an evolution, driven by innovations in machine learning models for predicting molecular properties and behavior. Recent strides in ML-based interatomic potentials have paved the way for accurate modeling of diverse chemical and structural properties at the atomic level. The key determinant defining MLIP reliability remains the quality of the training data. A paramount challenge lies in constructing training sets that capture specific domains in the vast chemical and structural space. This Review navigates the intricate landscape of essential components and integrity of training data that ensure the extensibility and transferability of the resulting models. We delve into the details of active learning, discussing its various facets and implementations. We outline different types of uncertainty quantification applied to atomistic data acquisition and the correlations between estimated uncertainty and true error. The role of atomistic data samplers in generating diverse and informative structures is highlighted. Furthermore, we discuss data acquisition via modified and surrogate potential energy surfaces as an innovative approach to diversify training data. The Review also provides a list of publicly available data sets that cover essential domains of chemical space.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Measurement of the time-integrated 𝐶⁢𝑃 asymmetry in 𝐷 0 → 𝐾$^{0}_{S}$𝐾$^{0}_{S}$ decays using opposite-side flavor tagging at Belle and Belle II

We measure the time-integrated 𝐶⁢𝑃 asymmetry in 𝐷 0 → 𝐾$^{0}_{S}$𝐾$^{0}_{S}$ decays reconstructed in 𝑒 + ⁢𝑒 − → $c\bar{c}$ events collected by the Belle and Belle II experiments. The corresponding data samples have integrated luminosities of 980 and 428 fb −1 , respectively. To infer the flavor of the 𝐷 0 meson, we exploit the correlation between the flavor of the reconstructed decay and the electric charges of particles reconstructed in the rest of the 𝑒 + ⁢𝑒 − → $c\bar{c}$ event. This results in a sample which is independent from any other previously used at Belle or Belle II. The result, 𝐴 𝐶⁢𝑃 ⁡(𝐷 0 → 𝐾$^{0}_{S}$𝐾$^{0}_{S}$)=(1.3±2.0±0.2)%, where the first uncertainty is statistical and the second systematic, is consistent with previous determinations and with 𝐶⁢𝑃 symmetry.

CP violation↗

Myna: Connecting powder bed fusion build data to simulation tools for digital twin applications

Additive manufacturing (AM), as a digital process, can generate a detailed digital thread linking a part’s design and manufacturing to its operational performance. As AM systems advance, an increasing amount of process data is stored in manufacturing databases. In principle, this data can be utilized by simulation-based digital twin approaches, such as real-time process control and asynchronous post-processing guidance. However, few tools currently exist for systematically integrating digital thread data with computational tools. Here, in this study, we propose a software package, called Myna, for connecting data from powder bed fusion processes to simulation tools. The utility of such a platform is demonstrated using build data from the Oak Ridge National Laboratory Manufacturing Demonstration Facility “Peregrine v2023-10” public dataset to automatically configure and run 54 semi-analytical 3DThesis melt pool simulations, 78 numerical Additive FOAM melt pool simulations, and 3 ExaCA microstructure simulations. The simulated, spatially registered microstructures are then compared directly with electron backscatter diffraction characterization of the corresponding as-built part locations. The resulting simulated microstructure showed variation as a function of process parameters, particularly stripe width; however, the experimental data had little variation between the microstructure texture and grain size resulting from different processing conditions. Analysis of the discrepancies suggest that it is possible a two-phase ferritic-austenitic solidification model is needed to accurately predict grain size and texture for certain stainless steel 316L feedstock compositions under powder bed fusion conditions, providing direction for future research. As illustrated here, due to the number and complexity of the simulations involved in AM process-structure–property predictions, automated methods to connect process data and simulations will remain necessary tools for testing hypotheses and implementing digital twin applications.

Knapp, Gerald L. [Oak Ridge National Laboratory (O↗

Hybrid data-driven cement-stabilized soil design: An integration of machine learning, multi-objective optimization, and life cycle assessment

Soil stabilization is crucial in geotechnical engineering, yet conventional methods are often time-consuming, resource-intensive, and environmentally unsustainable. Despite growing interest in Machine Learning (ML) and optimization tools for mix design, few studies integrate these methods with decision-making techniques and environmental assessment to support practical implementation. This study proposes a hybrid data-driven framework for predicting strength, optimizing mix compositions, and evaluating environmental impacts via life cycle assessment of cement-stabilized soft soils. Six ML models were evaluated, and the top-performing eXtreme Gradient Boosting (XGB) model was further improved using the Grey Wolf Optimizer (GWO). The optimized XGB-GWO model, integrated with a polynomial cost function, served as the objective function in a multi-objective optimization problem solved via the Non-Dominated Sorting Genetic Algorithm II (NSGA-II), with final mix selection guided by the entropy-weighted TOPSIS method. Validation through a case study produced mix designs offering superior strength-cost trade-offs, with the optimal mix achieving 2243.2 kPa unconfined compressive strength and a 16.07 % reduction in carbon emissions compared to the highest-cost design. In conclusion, this study offers a sustainable, scalable approach to soil stabilization and supports informed decision-making in construction.

Life cycle assessment↗

Temporal sequence transformer to advance long-term streamflow prediction

Accurate streamflow prediction is crucial for understanding climate change impacts on water resources and for effective management of extreme hydrological events. While Long Short-Term Memory (LSTM) networks have been the dominant data-driven approach for streamflow forecasting, recent advancements in transformer architectures for time series tasks have shown promise in outperforming traditional LSTM models. This study introduces a transformer-based model that integrates historical streamflow data with climatic variables to enhance streamflow prediction accuracy. We evaluated our transformer model against a benchmark LSTM across five diverse basins in the United States. Results demonstrate that the transformer architecture consistently outperforms the LSTM model across all evaluation metrics, highlighting its potential as a more effective tool for hydrological forecasting. This research contributes to the ongoing development of advanced AI techniques for improved water resource management and climate change adaptation strategies.

Singh, Ruhaan [Farragut High School]↗

Stratigraphy‐Induced Localization of Microseismicity During CO 2 Injection in Illinois Basin

Abstract Subsurface fluid injection stimulates complex hydromechanical interaction, necessitating the integration of geomechanical data across spatial and temporal scales to consider the sophisticated behavior. Induced seismic response is usually associated with the complex reservoir architecture and pre‐existing features that are three‐dimensional, such as local stratigraphy, fractures, faults, and other discontinuities. This study encompasses laboratory characterization of the coupled hydromechanical response of cores extracted from rock formations in Illinois Basin: reservoir ‐ Mt. Simon sandstone, basal seal ‐ Argenta sandstone, and crystalline basement ‐ Precambrian rhyolite. High‐resolution numerical modeling allows considering the three‐dimensional complexity of the Illinois Basin Decatur Project with spatial resolution comparable to one of the active seismic surveys. A detailed reconstruction of the evolving state of stress in formations lacking direct stress measurements is achieved by numerical modeling that integrated laboratory‐derived hydromechanical properties, a porosity‐permeability relationship, active seismic data, and an inverted three‐dimensional porosity distribution. It appears that the microseismic clusters, mainly observed in the crystalline basement during the injection, are linked to zones experiencing more critically stressed conditions prior to injection. These zones have a potential for reactivation during the injection and are attributed to the specific local stratigraphy of the injection site, as well as transfer of triggering perturbations during the injection.

Bondarenko, N. [University of Illinois Urbana‐Cham↗

Roadmap for the future of extreme wildfire events

Background Extreme wildfire events (EWEs) represent a growing threat globally, posing substantial risks to ecosystems, human communities, and infrastructure. Despite increased recognition of their ecological, social, and economic significance, current definitions of EWEs vary widely, reflecting disciplinary biases and regional contexts. This article emerges from an interdisciplinary workshop convened to reassess and refine the definition of EWEs, examine their impacts across ecological and social dimensions, and identify critical knowledge gaps impeding our understanding of these infrequent but important events. Results Our synthesis highlights significant limitations with existing definitions, particularly their reliance on subjective thresholds and their emphasis on extreme fire behavior alone. EWEs encompass a spectrum of complex, multi-dimensional phenomena that extend beyond immediate biophysical characteristics to include cumulative social, economic, and ecological impacts. These impacts often manifest over extended timeframes and include hazardous environmental contamination, severe geomorphic disturbances, ecosystem transformations, and unintended consequences of post-fire management actions. Current wildfire modeling frameworks inadequately capture these compounding factors, particularly the interactions among social systems, ecological conditions, and extreme fire behavior. To overcome these issues, we advocate for an interdisciplinary and context-sensitive approach to defining and studying EWEs. This revised definition emphasizes wildfires exhibiting anomalies in fire behavior, ecological outcomes, or social impacts relative to historically observed baselines, accommodating variability across different geographic regions and ecological settings. Conclusions Adopting an interdisciplinary framework that integrates biophysical and social sciences will enhance the predictive capability of wildfire models and improve resilience planning and response strategies. Filling identified knowledge gaps—such as limited high-quality empirical fire behavior data and insufficient integration of social dynamics into modeling—will better prepare communities and ecosystems to cope with and adapt to EWEs. This inclusive approach underscores the necessity for collaboration across disciplines and sectors, essential to managing extreme wildfires in an era of increasing climatic and ecological uncertainty.

54 ENVIRONMENTAL SCIENCES↗

Taylor-Expansion-Based Robust Power Flow in Unbalanced Distribution Systems: A Hybrid Data-Aided Method

Traditional power flow methods often adopt certain assumptions designed for passive balanced distribution systems, thus lacking practicality for unbalanced operation. moreover, their computation accuracy and efficiency are heavily subject to unknown errors and bad data in measurements or prediction data of distributed energy resources (ders). to address these issues, this paper proposes a hybrid data-aided robust power flow algorithm in unbalanced distribution systems, which combines taylor series expansion knowledge with a data-driven regression technique. the proposed method initiates a linearization power flow model to derive an explicitly analytical solution by modified taylor expansion. to mitigate the approximation loss that surges due to the der integration and bad data, we further develop a data-aided robust support vector regression approach to estimate the errors efficiently. comparative analysis in the 13-bus and 123-bus ieee unbalanced feeders shows that the proposed hybrid algorithm achieves superior computational efficiency, with guaranteed accuracy and robustness against outliers.

data-driven↗

A Refined Method to Translate Solar Data Quality Assessment Flags to Estimated Measurement Uncertainty

Integrating solar resource uncertainties due to radiometer measurement performance and operational data quality assessment can provide improved estimates of economic bankability, system design performance, and compliance of solar energy conversion systems. Estimating radiometer measurement uncertainty is an established procedure consistent with recognized best practices and international guidelines. SERI QC is a robust solar data quality assessment software tool that has been in continuous use for more than three decades. This report, the fourth of six for the Data Quality and Uncertainty Integration Project, presents a refined algorithm description for software to translate solar resource data quality assessment results into estimated uncertainty values in a Solar Resource Operational Uncertainty Integrator (SROUI) application. This algorithm requires three-component solar irradiance measurements - global horizontal irradiance, direct normal irradiance, and diffuse horizontal irradiance - collected at 1- to 60-minute intervals, as described in the previous deliverables. The development of this report as Deliverable 6.4 was an iterative process that included reviews and feedback from the project team on initial drafts designed to refine how the new software could best support determining solar resource data uncertainty. The results of this effort will contribute to the final software system development by National Renewable Energy Laboratory staff.

14 SOLAR ENERGY↗

NeuDiff Agent: a governed AI workflow for single-crystal neutron crystallography

Large-scale facilities increasingly face analysis and reporting latency as a limiting step in scientific throughput, particularly for structural studies that require iterative reduction, integration, refinement and validation. To improve the time to result and analysis efficiency, NeuDiff Agent is introduced as a governed, tool-using AI workflow for TOPAZ at the Spallation Neutron Source. NeuDiff Agent takes instrument data through reduction, integration, refinement and validation to a validated crystal structure and a publication-ready CIF. NeuDiff Agent coordinates established crystallographic tools under explicit governance by restricting actions to allowlisted tools, enforcing fail-closed verification gates at key workflow boundaries, and capturing complete provenance for inspection, auditing and controlled replay. The present benchmark is limited to structural crystallography for periodic structures; magnetic structure analysis and incommensurate or superspace refinement are outside the scope of the current workflow. Performance is assessed using a fixed prompt protocol and repeated end-to-end runs with two large language model backends, with user and machine time partitioned and intervention burden and recovery behaviors quantified under gating. In a reference-case benchmark, NeuDiff Agent reduces wall time from 435 min (manual) to 86.5 ± 4.7 to 94.4 ± 3.5 min (4.6–5.0× faster) while producing a validated CIF with no checkCIF level A or B alerts. These results establish a practical route to deploy agentic AI in facility crystallography while preserving traceability and publication-facing validation requirements.

Xiao, Zhongcan [ORNL] (ORCID:0000000220761961)↗

Genesis Data Card Schema, Template and Supporting Tools

Genesis Data Cards provide a standardized template and schema for documenting scientific datasets in support of discovery, access, interoperability, reusability, governed use, and AI usability. This release of the Genesis Data Card repository includes a versioned Markdown template, a LinkML schema with generated Pydantic and JSON artifacts, schema documentation, and example completed data cards. Validation tooling is provided to ensure that completed data cards conform to the schema prior to submission. Accompanying documentation for the structured metadata is provided as a Field Reference Guide. The schema and accompanying template provided in this repository address the call for actionable context that enables humans and AI systems to find, access, interpret, cite, and reuse data, and, when appropriate, integrate it into AI and machine learning workflows. The data card is intended to serve as a common metadata artifact intended to support standardized, cross-program dataset documentation across Department of Energy (DOE)-aligned efforts, including but not limited to Genesis Mission-related implementations, the Office of Science, National Nuclear Security Administration (NNSA), and Advanced Simulation and Computing (ASC) data governance and stewardship initiatives.

data card↗

Event generators for high-energy physics experiments

We provide an overview of the status of Monte-Carlo event generators for high-energy particle physics. Guided by the experimental needs and requirements, we highlight areas of active development, and opportunities for future improvements. Particular emphasis is given to physics models and algorithms that are employed across a variety of experiments. These common themes in event generator development lead to a more comprehensive understanding of physics at the highest energies and intensities, and allow models to be tested against a wealth of data that have been accumulated over the past decades. A cohesive approach to event generator development will allow these models to be further improved and systematic uncertainties to be reduced, directly contributing to future experimental success. Event generators are part of a much larger ecosystem of computational tools. They typically involve a number of unknown model parameters that must be tuned to experimental data, while maintaining the integrity of the underlying physics models. Making both these data, and the analyses with which they have been obtained accessible to future users is an essential aspect of open science and data preservation. It ensures the consistency of physics models across a variety of experiments.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗