Search NASA⌕ Search

SEARCH · Search NASA

Results for “data integration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22

Myna: Connecting powder bed fusion build data to simulation tools for digital twin applications

Additive manufacturing (AM), as a digital process, can generate a detailed digital thread linking a part’s design and manufacturing to its operational performance. As AM systems advance, an increasing amount of process data is stored in manufacturing databases. In principle, this data can be utilized by simulation-based digital twin approaches, such as real-time process control and asynchronous post-processing guidance. However, few tools currently exist for systematically integrating digital thread data with computational tools. Here, in this study, we propose a software package, called Myna, for connecting data from powder bed fusion processes to simulation tools. The utility of such a platform is demonstrated using build data from the Oak Ridge National Laboratory Manufacturing Demonstration Facility “Peregrine v2023-10” public dataset to automatically configure and run 54 semi-analytical 3DThesis melt pool simulations, 78 numerical Additive FOAM melt pool simulations, and 3 ExaCA microstructure simulations. The simulated, spatially registered microstructures are then compared directly with electron backscatter diffraction characterization of the corresponding as-built part locations. The resulting simulated microstructure showed variation as a function of process parameters, particularly stripe width; however, the experimental data had little variation between the microstructure texture and grain size resulting from different processing conditions. Analysis of the discrepancies suggest that it is possible a two-phase ferritic-austenitic solidification model is needed to accurately predict grain size and texture for certain stainless steel 316L feedstock compositions under powder bed fusion conditions, providing direction for future research. As illustrated here, due to the number and complexity of the simulations involved in AM process-structure–property predictions, automated methods to connect process data and simulations will remain necessary tools for testing hypotheses and implementing digital twin applications.

Knapp, Gerald L. [Oak Ridge National Laboratory (O↗

Hybrid data-driven cement-stabilized soil design: An integration of machine learning, multi-objective optimization, and life cycle assessment

Soil stabilization is crucial in geotechnical engineering, yet conventional methods are often time-consuming, resource-intensive, and environmentally unsustainable. Despite growing interest in Machine Learning (ML) and optimization tools for mix design, few studies integrate these methods with decision-making techniques and environmental assessment to support practical implementation. This study proposes a hybrid data-driven framework for predicting strength, optimizing mix compositions, and evaluating environmental impacts via life cycle assessment of cement-stabilized soft soils. Six ML models were evaluated, and the top-performing eXtreme Gradient Boosting (XGB) model was further improved using the Grey Wolf Optimizer (GWO). The optimized XGB-GWO model, integrated with a polynomial cost function, served as the objective function in a multi-objective optimization problem solved via the Non-Dominated Sorting Genetic Algorithm II (NSGA-II), with final mix selection guided by the entropy-weighted TOPSIS method. Validation through a case study produced mix designs offering superior strength-cost trade-offs, with the optimal mix achieving 2243.2 kPa unconfined compressive strength and a 16.07 % reduction in carbon emissions compared to the highest-cost design. In conclusion, this study offers a sustainable, scalable approach to soil stabilization and supports informed decision-making in construction.

Life cycle assessment↗

Temporal sequence transformer to advance long-term streamflow prediction

Accurate streamflow prediction is crucial for understanding climate change impacts on water resources and for effective management of extreme hydrological events. While Long Short-Term Memory (LSTM) networks have been the dominant data-driven approach for streamflow forecasting, recent advancements in transformer architectures for time series tasks have shown promise in outperforming traditional LSTM models. This study introduces a transformer-based model that integrates historical streamflow data with climatic variables to enhance streamflow prediction accuracy. We evaluated our transformer model against a benchmark LSTM across five diverse basins in the United States. Results demonstrate that the transformer architecture consistently outperforms the LSTM model across all evaluation metrics, highlighting its potential as a more effective tool for hydrological forecasting. This research contributes to the ongoing development of advanced AI techniques for improved water resource management and climate change adaptation strategies.

Singh, Ruhaan [Farragut High School]↗

Stratigraphy‐Induced Localization of Microseismicity During CO 2 Injection in Illinois Basin

Abstract Subsurface fluid injection stimulates complex hydromechanical interaction, necessitating the integration of geomechanical data across spatial and temporal scales to consider the sophisticated behavior. Induced seismic response is usually associated with the complex reservoir architecture and pre‐existing features that are three‐dimensional, such as local stratigraphy, fractures, faults, and other discontinuities. This study encompasses laboratory characterization of the coupled hydromechanical response of cores extracted from rock formations in Illinois Basin: reservoir ‐ Mt. Simon sandstone, basal seal ‐ Argenta sandstone, and crystalline basement ‐ Precambrian rhyolite. High‐resolution numerical modeling allows considering the three‐dimensional complexity of the Illinois Basin Decatur Project with spatial resolution comparable to one of the active seismic surveys. A detailed reconstruction of the evolving state of stress in formations lacking direct stress measurements is achieved by numerical modeling that integrated laboratory‐derived hydromechanical properties, a porosity‐permeability relationship, active seismic data, and an inverted three‐dimensional porosity distribution. It appears that the microseismic clusters, mainly observed in the crystalline basement during the injection, are linked to zones experiencing more critically stressed conditions prior to injection. These zones have a potential for reactivation during the injection and are attributed to the specific local stratigraphy of the injection site, as well as transfer of triggering perturbations during the injection.

Bondarenko, N. [University of Illinois Urbana‐Cham↗

Roadmap for the future of extreme wildfire events

Background Extreme wildfire events (EWEs) represent a growing threat globally, posing substantial risks to ecosystems, human communities, and infrastructure. Despite increased recognition of their ecological, social, and economic significance, current definitions of EWEs vary widely, reflecting disciplinary biases and regional contexts. This article emerges from an interdisciplinary workshop convened to reassess and refine the definition of EWEs, examine their impacts across ecological and social dimensions, and identify critical knowledge gaps impeding our understanding of these infrequent but important events. Results Our synthesis highlights significant limitations with existing definitions, particularly their reliance on subjective thresholds and their emphasis on extreme fire behavior alone. EWEs encompass a spectrum of complex, multi-dimensional phenomena that extend beyond immediate biophysical characteristics to include cumulative social, economic, and ecological impacts. These impacts often manifest over extended timeframes and include hazardous environmental contamination, severe geomorphic disturbances, ecosystem transformations, and unintended consequences of post-fire management actions. Current wildfire modeling frameworks inadequately capture these compounding factors, particularly the interactions among social systems, ecological conditions, and extreme fire behavior. To overcome these issues, we advocate for an interdisciplinary and context-sensitive approach to defining and studying EWEs. This revised definition emphasizes wildfires exhibiting anomalies in fire behavior, ecological outcomes, or social impacts relative to historically observed baselines, accommodating variability across different geographic regions and ecological settings. Conclusions Adopting an interdisciplinary framework that integrates biophysical and social sciences will enhance the predictive capability of wildfire models and improve resilience planning and response strategies. Filling identified knowledge gaps—such as limited high-quality empirical fire behavior data and insufficient integration of social dynamics into modeling—will better prepare communities and ecosystems to cope with and adapt to EWEs. This inclusive approach underscores the necessity for collaboration across disciplines and sectors, essential to managing extreme wildfires in an era of increasing climatic and ecological uncertainty.

54 ENVIRONMENTAL SCIENCES↗

Taylor-Expansion-Based Robust Power Flow in Unbalanced Distribution Systems: A Hybrid Data-Aided Method

Traditional power flow methods often adopt certain assumptions designed for passive balanced distribution systems, thus lacking practicality for unbalanced operation. moreover, their computation accuracy and efficiency are heavily subject to unknown errors and bad data in measurements or prediction data of distributed energy resources (ders). to address these issues, this paper proposes a hybrid data-aided robust power flow algorithm in unbalanced distribution systems, which combines taylor series expansion knowledge with a data-driven regression technique. the proposed method initiates a linearization power flow model to derive an explicitly analytical solution by modified taylor expansion. to mitigate the approximation loss that surges due to the der integration and bad data, we further develop a data-aided robust support vector regression approach to estimate the errors efficiently. comparative analysis in the 13-bus and 123-bus ieee unbalanced feeders shows that the proposed hybrid algorithm achieves superior computational efficiency, with guaranteed accuracy and robustness against outliers.

data-driven↗

A Refined Method to Translate Solar Data Quality Assessment Flags to Estimated Measurement Uncertainty

Integrating solar resource uncertainties due to radiometer measurement performance and operational data quality assessment can provide improved estimates of economic bankability, system design performance, and compliance of solar energy conversion systems. Estimating radiometer measurement uncertainty is an established procedure consistent with recognized best practices and international guidelines. SERI QC is a robust solar data quality assessment software tool that has been in continuous use for more than three decades. This report, the fourth of six for the Data Quality and Uncertainty Integration Project, presents a refined algorithm description for software to translate solar resource data quality assessment results into estimated uncertainty values in a Solar Resource Operational Uncertainty Integrator (SROUI) application. This algorithm requires three-component solar irradiance measurements - global horizontal irradiance, direct normal irradiance, and diffuse horizontal irradiance - collected at 1- to 60-minute intervals, as described in the previous deliverables. The development of this report as Deliverable 6.4 was an iterative process that included reviews and feedback from the project team on initial drafts designed to refine how the new software could best support determining solar resource data uncertainty. The results of this effort will contribute to the final software system development by National Renewable Energy Laboratory staff.

14 SOLAR ENERGY↗

NeuDiff Agent: a governed AI workflow for single-crystal neutron crystallography

Large-scale facilities increasingly face analysis and reporting latency as a limiting step in scientific throughput, particularly for structural studies that require iterative reduction, integration, refinement and validation. To improve the time to result and analysis efficiency, NeuDiff Agent is introduced as a governed, tool-using AI workflow for TOPAZ at the Spallation Neutron Source. NeuDiff Agent takes instrument data through reduction, integration, refinement and validation to a validated crystal structure and a publication-ready CIF. NeuDiff Agent coordinates established crystallographic tools under explicit governance by restricting actions to allowlisted tools, enforcing fail-closed verification gates at key workflow boundaries, and capturing complete provenance for inspection, auditing and controlled replay. The present benchmark is limited to structural crystallography for periodic structures; magnetic structure analysis and incommensurate or superspace refinement are outside the scope of the current workflow. Performance is assessed using a fixed prompt protocol and repeated end-to-end runs with two large language model backends, with user and machine time partitioned and intervention burden and recovery behaviors quantified under gating. In a reference-case benchmark, NeuDiff Agent reduces wall time from 435 min (manual) to 86.5 ± 4.7 to 94.4 ± 3.5 min (4.6–5.0× faster) while producing a validated CIF with no checkCIF level A or B alerts. These results establish a practical route to deploy agentic AI in facility crystallography while preserving traceability and publication-facing validation requirements.

Xiao, Zhongcan [ORNL] (ORCID:0000000220761961)↗

Genesis Data Card Schema, Template and Supporting Tools

Genesis Data Cards provide a standardized template and schema for documenting scientific datasets in support of discovery, access, interoperability, reusability, governed use, and AI usability. This release of the Genesis Data Card repository includes a versioned Markdown template, a LinkML schema with generated Pydantic and JSON artifacts, schema documentation, and example completed data cards. Validation tooling is provided to ensure that completed data cards conform to the schema prior to submission. Accompanying documentation for the structured metadata is provided as a Field Reference Guide. The schema and accompanying template provided in this repository address the call for actionable context that enables humans and AI systems to find, access, interpret, cite, and reuse data, and, when appropriate, integrate it into AI and machine learning workflows. The data card is intended to serve as a common metadata artifact intended to support standardized, cross-program dataset documentation across Department of Energy (DOE)-aligned efforts, including but not limited to Genesis Mission-related implementations, the Office of Science, National Nuclear Security Administration (NNSA), and Advanced Simulation and Computing (ASC) data governance and stewardship initiatives.

data card↗

Event generators for high-energy physics experiments

We provide an overview of the status of Monte-Carlo event generators for high-energy particle physics. Guided by the experimental needs and requirements, we highlight areas of active development, and opportunities for future improvements. Particular emphasis is given to physics models and algorithms that are employed across a variety of experiments. These common themes in event generator development lead to a more comprehensive understanding of physics at the highest energies and intensities, and allow models to be tested against a wealth of data that have been accumulated over the past decades. A cohesive approach to event generator development will allow these models to be further improved and systematic uncertainties to be reduced, directly contributing to future experimental success. Event generators are part of a much larger ecosystem of computational tools. They typically involve a number of unknown model parameters that must be tuned to experimental data, while maintaining the integrity of the underlying physics models. Making both these data, and the analyses with which they have been obtained accessible to future users is an essential aspect of open science and data preservation. It ensures the consistency of physics models across a variety of experiments.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

WELLBASE - An Interactive Platform for Wellbore Material Assessment

This project seeks to build an open-source wellbore material data repository with adequate material performance and contextual data to support Geological Carbon Storage (GCS). By appropriately evaluating the data types as mentioned earlier made available by the WELLBASE tool, stakeholders can make more informed decisions regarding well selections, risk assessment, and economic analysis for geologic carbon storage projects. Advanced Natural Language Processing models and other custom python scripts will be deployed in an automated process to extract unstructured data from documents, reports, and web applications and subsequently parse to more usable formats. The processed data will then be integrated into a robust and comprehensive database architecture, optimizing data accessibility, and usability for analytical purposes. The final data products will be accessible through a user-friendly visualization platform that will allow users to query and visualize the data, as well as download data in usable formats.

Tetteh, Daniel A.↗

Mechanical separations of corn stover anatomical fractions in an integrated feedstock preprocessing system: An experimental and data-driven modeling study

High variabilities of material attributes in lignocellulosic biomass present risks for biofuel and biochemical productions and must be mitigated via preprocessing. Since almost no mechanical device is originally designed for processing biomass, how to operate existing apparatuses with efficient performance has not been investigated extensively. This work presents a study on an integrated screening and air classification to separate cobs and stalks from husks and leaves in corn stover. Prototype machine learning models were developed to assess the feasibility of predicting the process outcome based on the measurable parameters. The models trained upon limited experimental data rendered decent predictive accuracy of yield and purity. The experimental data and modeling results collectively suggest decreasing throughput leads to a higher purity. To the contrary, if throughput increases, a lower purity is likely. A possible trade-off between yield and purity of the separated streams indicates the need for optimal combinations of feedstock size, moisture, and throughput to achieve optimized separations. The results of this study also suggest the need to further improve model predictability by developing more accurate formulations for physics governing the integrated unit operations. To accomplish this, additional experimental data needs to be generated for model training.

09 - BIOMASS FUELS↗

Downscaled Earth System Model Data for Resilient Energy System Planning

The second-generation Sup3rCC dataset provides high-resolution meteorological data generated through the downscaling of multiple earth system models (ESMs) from the Coupled Model Intercomparison Project Phase 6 (CMIP6). This downscaling is performed through application of a generative machine learning approach called Super-Resolution for Renewable Resource Data (sup3r). This dataset builds on the first-generation Sup3rCC data by applying improved bias correction methods and adding downscaled precipitation to the output variables. In this presentation, we explore the output characteristics of the dataset and various validation analyses. We also present and discuss plans for the integration of this data into power system planning models using a decision-making under deep uncertainty (DMDU) methodology.

97 MATHEMATICS AND COMPUTING↗

Path-Integrated X-Ray Digital Image Correlation using Synthetic Reference Images

X-rays can provide images when an object is visibly obstructed, allowing for motion measurements via x-ray digital image correlation (DIC). However, x-ray images are path-integrated and contain data for all objects between the source and detector. If multiple objects are present in the x-ray path, conventional DIC algorithms may fail to correlate the x-ray images. A new DIC algorithm called path-integrated (PI)-DIC addresses this issue by reformulating the matching criterion for DIC to account for multiple, independently-moving objects. PI-DIC requires a set of reference x-ray images of each independent object. However, due to experimental constraints, such reference images might not be obtainable from the experiment. Here, this work focuses on the reliability of synthetically-generated reference images, in such cases. A simplified exemplar is used for demonstration purposes, consisting of two aluminum plates with tantalum x-ray DIC patterns undergoing independent rigid translations. Synthetic reference images based on the “as-designed” DIC patterns were generated. However, PI-DIC with the synthetic images suffered some biases due to manufacturing defects of the patterns. A systematic study of seven identified defect types found that an incorrect feature diameter was the most influential defect. Synthetic images were re-generated with the corrected feature diameter, and PI-DIC errors were improved by a factor of 3-4. Final biases ranged from 0.00-0.04 px, and standard uncertainties ranged from 0.06-0.11 px. In conclusion, PI-DIC accurately measured the independent displacement of two plates from a single series of path-integrated x-ray images using synthetically-generated reference images, and the methods and conclusions derived here can be extended to more generalized cases involving stereo PI-DIC for arbitrary specimen geometry and motion. This work thus extends the application space of x-ray imaging for full-field DIC measurements of multiple surfaces or objects in extreme environments where optical DIC is not possible.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Leveraging Large Language Models for Real-World Data Evidence: A Framework for Automated Treatment Extraction and Data Harmonization

Background: The ability to comprehensively collect treatment information from cancer patient medical records would enable studies to evaluate real-world benefits and risks tied to specific treatments. Currently, it is difficult to system- atically collect high-quality treatment information because it is often stored in unstructured text. Manually extracting and standardizing drug and regimen data is time-intensive. Recent advances in large language models (LLMs) offer a potential solution for automated extraction of structured treatment information from clinical text. Objective: This study systematically evaluates the utility of four LLMs from the Llama family for automated extraction of oncology treatment information from clinical text. This information can guide researchers using cancer registry data to provide insights into cancer care and outcomes beyond clinical trials. Methods: Four instruction-tuned Llama models with varying parameter counts (1B, 3B, 8B, and 70B) were evaluated for their ability to extract treatment information from clinical documents. A unified oncology knowledge base integrating seven major public data sources was developed to standardize and normalize extracted entities—a critical step for harmonizing data from diverse sources. Extracted treatment data were compared against expert-annotated ground truth. Model performance was assessed using accuracy metrics (Precision, Recall, F1-Score) and opera- tional feasibility metrics, including processing speed and structural compliance of the output. Results: A strong positive correlation was observed between model size and extraction accuracy. F1-score improved from 0.609 for the 1B model to 0.710 (3B), 0.807 (8B), and 0.828 (70B). While larger models demonstrated superior accuracy and compliance, they incurred higher computational costs. The modest performance difference between 8B and 70B suggests diminishing returns with increasing model size. Conclusions: LLMs represent a viable technology for automating oncology treatment extraction. The 8B-parameter model emerged as a highly effective option, balancing high accuracy and computational efficiency. Selecting an appropriate LLM for deployment in cancer registries involves a trade-off between desired accuracy and available operational resources. Harmonizing extracted entities with the oncology knowledge base facilitates standardized integration into common data models, enhancing data quality for real-world evidence analyses.

artificial intelligence↗

GPS-supported smartphone app-based integrated travel diary and time-use data collection: challenges and lessons learned

Travel behaviour and time-use data are two vital data sources for travel demand modelling. Travel behaviour is traditionally collected through household travel surveys, enhanced by using GPS-supported smartphone apps for passive location data collection. However, recruiting individuals willing to install these apps with sustained motivation to continue participation has been a critical challenge. This paper shares insights from a travel and time-use data collection procedure in Chicago and Sydney using the Fourstep app. Social media platforms were utilised as a solution to recruit participants in Chicago, where an international market research company failed to accomplish the task. This paper also discusses the challenges we faced and suggests ways to overcome them, offering valuable guidance to researchers in recruiting participants for smartphone application-based data collection. It also offers an analysis of travel, time-use, and travel-based multitasking behaviours based on the data collected from the Chicago and Sydney samples.

GPS-supported smartphone apps↗

Time and Frequency Analysis of Load Profile Data

Technology advancements and integration of modern advanced metering systems can monitor, forecast, inform, control, and operate the building's mechanical, electrical, and plumbing (MEP) systems. They offer a higher level of information, which can contribute to making smart buildings more energy efficient and to making them closer to becoming grid-interactive energy efficient buildings (GEB). This paper builds on the ongoing research on variability analysis of a case study building with a 1-minute load profile and examines the Discrete Wavelet Transform (DWT) process in the frequency domain to quantify the signal's energy in each bandwidth, with respect to each end-use category. Moreover, the amount of variability in the total variability is not similar among the end-use categories. This information is needed to understand the behavior of the variability in the frequency domain for future applications, such as generating synthetic load profiles with a similar frequency spectrum as the measured signal.

decomposition↗

Laue-DIALS: Open-source software for polychromatic x-ray diffraction data

Most x-ray sources are inherently polychromatic. Polychromatic (“pink”) x-rays provide an efficient way to conduct diffraction experiments as many more photons can be used and large regions of reciprocal space can be probed without sample rotation during exposure—ideal conditions for time-resolved applications. Analysis of such data is complicated, however, causing most x-ray facilities to discard >99% of x-ray photons to obtain monochromatic data. Key challenges in analyzing polychromatic diffraction data include lattice searching, indexing and wavelength assignment, correction of measured intensities for wavelength-dependent effects, and deconvolution of harmonics. We recently described an algorithm, Careless, that can perform harmonic deconvolution and correct measured intensities for variation in wavelength when presented with integrated diffraction intensities and assigned wavelengths. Here, we present Laue-DIALS, an open-source software pipeline that indexes and integrates polychromatic diffraction data. Laue-DIALS is based on the dxtbx toolbox, which supports the DIALS software commonly used to process monochromatic data. As such, Laue-DIALS provides many of the same advantages: an open-source, modular, and extensible architecture, providing a robust basis for future development. We present benchmark results showing that Laue-DIALS, together with Careless, provides a suitable approach to the analysis of polychromatic diffraction data, including for time-resolved applications.

97 MATHEMATICS AND COMPUTING↗