Search NASA⌕ Search

SEARCH · Search NASA

Results for “Research data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

VA EDH Data Curation Documentation FY25-Q1

This data source documentation report provides researchers with valuable insights into the structure, contents, and data sources used to compile the datasets. It specifically covers the Fiscal Year 2024, Fourth Quarter (FY25-Q1) dataset curation documentation for the Environmental Determinants of Health (EDH) project.

97 MATHEMATICS AND COMPUTING↗

Performance and Reliability Assessment of the U.S. Department of Energy Atmospheric Radiation Measurement (ARM) Data Advisor (ADA)

The Atmospheric Radiation Measurement (ARM) User Facility provides one of the world's largest openly accessible repositories of atmospheric observations through the ARM Data Discovery platform. Although the repository contains more than three decades of measurements collected from permanent observatories, mobile facilities, aircraft campaigns, and field experiments, identifying appropriate datasets can be challenging, particularly for new users unfamiliar with ARM instrumentation and datastream organization. To improve data accessibility, the ARM Data Center developed the ARM Data Advisor (ADA), an artificial intelligence-powered assistant designed to facilitate scientific data discovery, dataset interpretation, and user guidance. This report evaluates ADA's performance as a domain-specific scientific assistant using realistic atmospheric science workflows. The evaluation examines five key capabilities: data retrieval and curation efficiency, hallucination resistance, scientific reasoning, response to ambiguous queries, and content retention and session continuity. Representative prompts were developed to simulate typical interactions between researchers and the ARM Data Discovery platform, and ADA's responses were assessed for retrieval completeness, scientific accuracy, consistency, and practical usefulness. In these representative tests, ADA reduced the complexity of discovering and accessing ARM datasets by recommending appropriate datastreams, explaining instrumentation, interpreting metadata, and assisting with data processing workflows. ADA also exhibits strong domain knowledge of atmospheric science terminology and generally resists hallucination by acknowledging unavailable datasets and requesting clarification when appropriate. Overall, the results indicate that ADA represents a promising advancement in scientific data discovery within the ARM User Facility and has considerable potential to improve researcher productivity, particularly for new users and interdisciplinary scientists seeking efficient access to ARM observations.

Salvador, Christian [ORNL] (ORCID:0000000283287777↗

A Comment on “Deep Proteogenomics of a Photosynthetic Cyanobacterium”

Proteomic researchers strive to achieve complete annotation of protein-coding DNA sequences to provide a foundational context for their relevant biological data. A recent deep proteogenomic study using a photosynthetic cyanobacterium Synechocystis sp. PCC 6803 by Spät et al. proposed 64 refined open reading frames (ORFs). By searching LC-MS/MS data from affinity chromatography-isolated protein complexes, our laboratory identified that six of these high-abundance ORFs possess Nterminal initiation start sites that differ than those proposed in the alternative models. Our findings are supported by highly confident MS2 data, phylogenetic analysis, chemical labeling, and established data from two independent research groups. Based on these highquality experimental identifications, we subsequently propose a standardized strategy and set of criteria for future deep proteogenomic efforts to ensure accurate and stringent proteogenomic annotation.

cyanobacteria↗

FAIRLinked: Data FAIRification Tools for Materials Data Science

FAIRLinked is a software package created to support the FAIRification of materials science data, ensuring proper alignment with FAIR principles: Findable, Accessible, Interoperable, and Reusable. It is built to be compatible with MDS-Onto, an ontology designed to capture the semantics of various types of materials data, enabling integration and sharing across different research workflows. The package is subdivided into three subpackages: InterfaceMDS, RDFTableConversion, and QBWorkflow. The first subpackage, InterfaceMDS allows users to search for terms using either string search or various filters, explore different domains and subdomains, and add terms to MDS-Onto. RDFTableConversion is used for serialization and deserialization of data from CSV into JSONLDs and vice versa in a way that captures the semantics of the data using MDS-Onto. Lastly, QBWorkflow is a serialization and deserialization workflow that incorporates RDF Data Cube vocabulary, useful for working with multidimensional datasets. By offering these packages, FAIRLinked lowers the barrier of creating FAIR, machine-actionable data for researchers in the materials science community.

FAIR↗

SAXS Assistant: Automated SAXS analysis for structural discovery in biologics and polymeric nanoparticles

Small-angle x-ray scattering (SAXS) is a powerful technique for assessing macromolecular structure. High-throughput SAXS is limited by the time-consuming and, at times, subjective nature of SAXS data interpretation. Here, we present SAXS Assistant, a Python-based script that streamlines SAXS data analysis to extract features for machine learning (ML) and key structural parameters, including the Guinier radius of gyration (R g ), pair distance distribution function (PDDF)-derived R g , maximum particle dimension (D max ), and Kratky plots. The script builds upon BioXTAS RAW and validates reliability via Guinier/PDDF R g agreement, an important indicator of well-measured data sets. For assistance in D max estimation, a multilayer perceptron regressor was trained with 1940 data files from the Small Angle Scattering Biological Data Bank. The model achieved a test set performance R 2 = 0.90 and mean absolute error = 11.7 Å. Training exclusively with experimental data translates analyses from researchers, including experts in the field, to the ML model, which helps assess D max estimations from PDDF. Gaussian mixture model clustering was implemented to classify profiles into structural classes based on entries in the Small Angle Scattering Biological Data Bank. Users may therefore assess the similarity between experimental samples and known biomolecular shapes within the mapped repository entries. This probabilistic clustering aids in quantifying information from Kratky and generating shape-descriptive features. SAXS Assistant accelerates SAXS data analysis through enforced quality control, ML-ready outputs, and flags for low-confidence results. In addition to providing the ability to analyze large data sets at high throughput, this tool is versatile and may serve researchers in both biological and synthetic polymer research fields.

36 MATERIALS SCIENCE↗

Landscape analysis of environmental data sources for linkage with SEER cancer patients database

Abstract One of the challenges associated with understanding environmental impacts on cancer risk and outcomes is estimating potential exposures of individuals diagnosed with cancer to adverse environmental conditions over the life course. Historically, this has been partly due to the lack of reliable measures of cancer patients’ potential environmental exposures before a cancer diagnosis. The emerging sources of cancer-related spatiotemporal environmental data and residential history information, coupled with novel technologies for data extraction and linkage, present an opportunity to integrate these data into the existing cancer surveillance data infrastructure, thereby facilitating more comprehensive assessment of cancer risk and outcomes. In this paper, we performed a landscape analysis of the available environmental data sources that could be linked to historical residential address information of cancer patients’ records collected by the National Cancer Institute’s Surveillance, Epidemiology, and End Results Program. The objective is to enable researchers to use these data to assess potential exposures at the time of cancer initiation through the time of diagnosis and even after diagnosis. The paper addresses the challenges associated with data collection and completeness at various spatial and temporal scales, as well as opportunities and directions for future research.

60 APPLIED LIFE SCIENCES↗

Deliverable 6.7-Final Technical Report: Development Summary and Evaluation of the Solar Uncertainty Integrator (SUNI) Software

The Data Quality and Uncertainty Integration Project was a three-year effort to address stakeholder needs for assessing solar radiation resource data quality based on existing tools for estimating radiometer measurement uncertainties and assessing post-measurement data quality. The annual research objectives for the project addressed a logical progression of effort needed to achieve the ultimate project goal of developing the Solar Uncertainty Integrator (SUNI) software. This final technical report summarizes the development process for achieving these key research objectives and addresses the outreach and code development efforts in the final year of the project to develop a new solar irradiance data uncertainty integration software package.

14 SOLAR ENERGY↗

Characteristics and trends of Atlantic tropical cyclones that do and do not develop from African easterly waves

Abstract Atlantic tropical cyclones (TCs) are known to develop from African easterly waves (AEWs) that propagate across North Africa and out over the Atlantic Ocean. The relationship between AEWs and TCs has been the subject of numerous previous studies. There are, however, many Atlantic TCs that do not have AEW origins. In this study, we provide a novel analysis of the characteristics and trends of Atlantic TCs both with and without AEW origins using 43 years of observational and reanalysis data. To conduct this research, we identified TCs with and without AEW origins from the observational record between 1980 and 2022, and ran objective tracking algorithms on reanalysis data to identify the AEWs and TCs during this time period. We found statistically significant differences in the characteristics and environments of TCs with and without AEW origins. TCs with AEW origins are stronger and costlier, experience more favorable environmental conditions, and are more likely to make landfall in the Gulf of Mexico and the Caribbean when compared to TCs without AEW origins. Additionally, the 43‐year increasing trend in Atlantic TC activity is primarily driven by an increase in TCs with AEW origins that is associated with increasing AEW frequency and strength, with anthropogenic aerosols potentially driving this trend. In contrast, we found no trend in TCs without AEW origins.

Bercos‐Hickey, Emily↗

Multi-Modal Bayesian Neural Network Surrogates with Conjugate Last-Layer Estimation

As data collection and simulation capabilities advance, multi-modal learning, the task of learning from multiple modalities and sources of data, is becoming an increasingly important area of research. Surrogate models that learn from data of multiple auxiliary modalities to support the modeling of a highly expensive quantity of interest have the potential to aid outer loop applications such as optimization, inverse problems, or sensitivity analyses when multi-modal data are available. We develop two multi-modal Bayesian neural network surrogate models and leverage conditionally conjugate distributions in the last layer to estimate model parameters using stochastic variational inference (SVI). We provide a method to perform this conjugate SVI estimation in the presence of partially missing observations. Here, we demonstrate improved prediction accuracy and uncertainty quantification compared to unimodal surrogate models for both scalar and time series data.

97 MATHEMATICS AND COMPUTING↗

UNNT: A novel Utility for comparing Neural Net and Tree-based models

The use of deep learning (DL) is steadily gaining traction in scientific challenges such as cancer research. Advances in enhanced data generation, machine learning algorithms, and compute infrastructure have led to an acceleration in the use of deep learning in various domains of cancer research such as drug response problems. In our study, we explored tree-based models to improve the accuracy of a single drug response model and demonstrate that tree-based models such as XGBoost (eXtreme Gradient Boosting) have advantages over deep learning models, such as a convolutional neural network (CNN), for single drug response problems. However, comparing models is not a trivial task. To make training and comparing CNNs and XGBoost more accessible to users, we developed an open-source library called UNNT (A novel Utility for comparing Neural Net and Tree-based models). The case studies, in this manuscript, focus on cancer drug response datasets however the application can be used on datasets from other domains, such as chemistry.

59 BASIC BIOLOGICAL SCIENCES↗

Hybrid learning techniques for scientific data reduction with performance guarantees

The research initiatives supported by the U.S. Department of Energy (DOE) Grant DE-SC0022265 are fundamentally aimed at pioneering advanced machine learning (ML) techniques for scientific data compression within high-performance computing (HPC) environments. This comprehensive body of work addresses the critical challenge posed by the exponential growth of data generated by scientific simulations in domains such as fusion energy, climate modeling, and computational fluid dynamics (CFD). A core objective is to develop compression algorithms that achieve substantial data reduction—often by orders of magnitude—while rigorously ensuring the fidelity of both the primary data (PD) and scientifically crucial derived quantities of interest (QoI). The methodologies deployed under this grant integrate sophisticated deep learning architectures, prominently featuring autoencoders, advanced generative models like conditional diffusion, and hybrid learning techniques. Key innovations include the development of Guaranteed Autoencoders (GAE) and the Guaranteed Conditional Diffusion with Tensor Correction (GCDTC) framework, which provide explicit, instance-level error bounds on reconstructed data. Furthermore, specialized strategies such as nonlinear constraint satisfaction are employed to preserve the integrity of QoI, a vital requirement for the trustworthiness of downstream scientific analyses. This research also focuses on the design and implementation of scalable, GPU-accelerated software pipelines that seamlessly integrate into existing HPC workflows, ensuring both computational efficiency and practical applicability. The CAESAR framework, for example, unifies foundation and generative models to create an adaptive and efficient compression solution for spatio-temporal scientific data. Collectively, these efforts represent a significant advancement in mitigating the scientific data deluge, enabling more effective data management, accelerated scientific discovery, and optimized utilization of HPC resources.

97 MATHEMATICS AND COMPUTING↗

Frameworks, Algorithms, and Scalable Technologies for Mathematics (FASTMath) SciDAC Institute

As computational models scale to larger computers, the rate at which they produce data has far outstripped the same computers ability to write that data and further the file systems ability to store that data. Almost all of the SciDAC applications, but especially those related to fusion solve very large scale PDEs whose scientific output his impacted by this problem. To gain access to dynamics in an exascale simulation that are not identifiable a priori and to make that dynamical data available to machine learning requires fundamental research in the area of in situ data data analytics. Here data analytics includes compression, visualization, uncertainty quantification, and machine learning. This in situ data analytics will enable on-the-fly spatial and temporal compression of solution dynamics, expose that space-time compressed field to machine learning algorithms that have been specialized to work with dynamically evolving data (existing machine learning algorithms treat data sets as static), greatly improving the opportunity for machine learning to provide feedback to the compression, all within an ongoing simulation, without the need to write data to files. The same concepts are also being applied to uncertainty quantification and multi-fidelity modeling which have similar needs for spatial and temporal compression of the ongoing exascale simulation to perform either without the typical, unacceptable writing of data to files.

97 MATHEMATICS AND COMPUTING↗

Machine Learning Approaches to Predicting Induced Seismicity and Imaging Geothermal Reservoir Properties

This project developed machine learning (ML) methods, lab data sets, and field data to advance geothermal exploration and geothermal energy production. The work had three focus areas. One involved the development of ML methods to use microearthquakes (MEQs) for imaging geothermal reservoir properties and improving subsurface characterization – most importantly the evolution of permeability within the evolving reservoir. This part of the work included development of ML approaches for automated MEQ location, focal mechanism determination and identification of earthquake precursors. The second area focused on using MEQ signals generated by geothermal exploration and production to predict the relationship between fluid injection and seismicity. Here, we extended to reservoir scale our success in using ML to predict laboratory earthquakes and fault zone stress state. The third focus area was on lab experiments. Here, we developed new ML models for lab earthquake prediction and identification of precursors to failure to improve earthquake forecasting and early warning in geothermal settings. Major outcomes of our work include ML models that learn from MEQ signals during geothermal exploration and production to predict induced seismicity. MEQs occur naturally in connection with drilling and energy production. We developed ML methods to use the seismic waves from these events to characterize the elastic, hydraulic and poromechanical properties of reservoirs. Our work illuminated fracture geometry and the evolution of fracture permeability by incorporating seismic coda wave analysis and ML methods to relate fluid injection and seismicity. We significantly expanded laboratory earthquake prediction to include methods that use both passive measurements of microearthquakes within the lab fault zones and also active source acoustic measurements of fault zone elastic properties. These methods can now predict fault zone stress state, time to failure and the magnitude of lab earthquakes. Our work showed that repetitive stick- slip failure events during frictional sliding (the lab equivalent of earthquakes) are preceded by a cascade of micro-failure events that radiate energy in a manner that foretells unstable failure – manifest as laboratory MEQs. We documented a mapping between fracture properties and statistical attributes of elastic radiation. We extended existing works to geothermal reservoir scale and developed ML methods to determine reservoir permeability, fracture properties, and their evolution during geothermal energy production. An attractive feature of ML algorithms is their ability to handle big datasets and reveal patterns and correlations that may remain invisible to conventional analyses. Our work connected data from field, laboratory and intermediate scales to study permeability, stress, strength, fracture stiffness and geometry. At the field scale we used data from the Newberry Volcano field site, UtahFORGE, EGS Collab, and also the Bedretto underground research lab in Switzerland. These data sets are bridging the gap between the lab scale, theory, and reservoir scale. Our work produced plain language summaries to improve public understanding of DOE research. We also developed openly distributed ML and seismicity datasets for use by all researchers and we published connections between induced seismicity in geothermal areas and reservoir properties including permeability, fracture properties, and stress state. Our models are designed for the large data sets of induced seismicity typically associated with geothermal sites. We produced labeled event catalogs and used them on geothermal data to assess how ML can facilitate geothermal production and exploration. All datasets are available on the GDR Productivity: The project produced 32 publications in peer reviewed journals (two are in review). It supported the work of 6 PhD students, 40 conference presentations, 6 keynote talks at national meetings, and mentoring and professional development for 4 postdoctoral fellows.

15 GEOTHERMAL ENERGY↗

How initial conditions-, structural-, and parameter-based model uncertainty interact and influence predictions in permafrost ecosystems: Modeling Archive

This dataset contains model output and input data, as well as source code examples for the Terrestrial Ecosystem Model with the Dynamic Vegetation Model and Dynamic Organic Soil (DVM-DOS-TEM) for the field sites Imnavait creek and the Bonanza creek Long Term Ecological Research Network (LTER). The data covers simulations from the last glacial maximum (LGM) until 2100 for a selection of paleo scenarios, setting the mean temperature of the LGM up to 10°C lower than pre-industrial conditions. The model structure was modulated to represent various model versions, and this dataset contains the relevant changes in the source code. The raw output data, the processed statistical data, the setup and processing scripts as well as parameter value distribution files from a parameter sensitivity analysis are included as well. Model outputs include active layer depth, organic soil carbon, soil layer depths, gross primary productivity (GPP) with and without nitrogen limitation, net primary productivity (NPP), soil liquid water content, heterotrophic, maintenance, and growth respiration, soil temperature, and vegetation carbon (*.nc files). The Next-Generation Ecosystem Experiments in the Arctic (NGEE Arctic) project is a research effort to reduce uncertainty in the Department of Energy’s Energy Exascale Earth System Model (E3SM) by developing a predictive understanding of Arctic tundra ecosystems underlain by permafrost and to quantify feedbacks from the Arctic tundra to the Earth system. NGEE Arctic is supported by the Department of Energy's Office of Biological and Environmental Research.Over Phases 1–3, observations made by the NGEE Arctic team across a gradient of permafrost landscapes in Arctic Alaska improved the representation of tundra processes in the land surface component of E3SM (the E3SM Land Model, ELM). Model improvements emphasized unique aspects of permafrost environments and explored reductions in model complexity while retaining predictive power. The Arctic-informed ELM developed by NGEE Arctic has been used to make novel predictions on processes ranging from permafrost thaw to soil biogeochemical cycling to Earth system feedbacks associated with the unique characteristics of tundra plants. In Phase 4, the NGEE Arctic team is evaluating our new predictive understanding under novel conditions across the Arctic domain. In collaboration with partners at long-term pan-Arctic research sites we are examining whether an Arctic-informed ELM can faithfully simulate interactions among surface and subsurface processes at site, regional, and pan-Arctic scales. In turn, we are using variety of tools to dynamically extend and evaluate ELM inference, with an emphasis on data synthesis and pan-Arctic model evaluation, reintegration of code with an evolving E3SM, scaling across heterogeneous Arctic landscapes, and the appropriate representation of the impacts of increasingly frequent Arctic disturbances.

54 ENVIRONMENTAL SCIENCES↗

Long-term measurements of ice nucleating particles at Atmospheric Radiation Measurement (ARM) sites worldwide

Ice nucleating particles (INPs) play a critical role in cloud microphysics and precipitation formation, yet long-term, spatially extensive observational datasets remain limited. Here, we present one of the most comprehensive publicly available datasets of immersion-mode INP concentrations using a single analytical method, generated through the U.S. Department of Energy's (DOE) Atmospheric Radiation Measurement (ARM) user facility. INP filter samples have been collected across a broad range of environments – including agricultural plains, Arctic coastlines, high-elevation mountain sites, marine regions, and urban areas – via fixed observatories, mobile facility deployments, and vertically-resolved tethered balloon system operations. We describe the standardized processing and quality assurance pipeline, from filter collection and processing using the Ice Nucleation Spectrometer to final data products archived on the ARM Data Discovery portal. The dataset includes both total INP concentrations and selectively treated samples, allowing for classification of biological, organic, and inorganic INP types. It features a continuous 5-year record of INP measurements from a central U.S. site, with data collection still ongoing. Seasonal and site-specific differences in INP concentrations are illustrated through intercomparisons at −10 and −20 °C, revealing distinct regional sources and atmospheric drivers. We also outline mechanisms for researchers to access existing data, request additional sample analyses, and propose future field campaigns involving ARM INP measurements. This dataset supports a wide range of scientific applications, from observational and mechanistic studies to model development, and provides critical constraints on aerosol-cloud interactions across diverse atmospheric regimes (Creamean et al., 2024, 2020b; https://doi.org/10.5439/1770816).

Creamean, Jessie M. [Colorado State Univ., Fort Co↗

Time Distribution Analysis for Task Primitives to Support Dynamic Human Reliability Analysis

To support data collection for dynamic human reliability analysis (HRA), this study investigates time distributions for task primitives defined in the Goals, Operators, Methods, and Selection rules (GOMS)–Human Reliability Analysis (HRA) method and Human Reliability data EXtraction (HuREX). GOMS-HRA was developed to provide cognition-based time and human error probability (HEP) information for dynamic HRA calculations within the Human Unimodel for Nuclear Technology to Enhance Reliability (HUNTER) framework, while HuREX is a comprehensive HRA data collection method developed by the Korea Atomic Energy Research Institute (KAERI). In this paper, we examine time distributions by using experimental data collected from the Simplified Human Error Experimental Program (SHEEP) study, which proposes an HRA data collection framework to complement full-scope simulator research and gather input data for dynamic HRA by using simplified simulators such as the Rancor Microworld simulator. This paper investigates whether the time required for GOMS-HRA and HuREX task primitives fits 13 statistical distributions. Additionally, we compare and discuss the time distributions obtained from both student operators and professional operators. The result was that this study identified several time distributions for five GOMS-HRA and four HuREX task primitives. In the future, the results of this study are expected to provide objective reference data on the elapsed time for task primitives and aid in realistically simulating scenarios within dynamic HRA.

Dynamic Human Reliability Analysis↗

Probabilistic Seismic Hazard Assessment for Azerbaijan

Probabilistic Seismic Hazard Assessments (PSHA) underpin the determination of seismic loads in most contemporary seismic provisions of building codes around the world. Modern building codes are migrating towards using the entire uniform hazard spectrum at a range of vibration periods (typically up to 4 s or 10 s) rather than a single peak value such as peak ground acceleration (PGA), thus requiring a larger range of PSHA outputs. The hazard maps in the current version of the building code of Azerbaijan (2011) are in terms of intensity and peak ground acceleration. Recognizing the need for an up-to-date seismic hazard assessment in the country, Seismic Cooperation Program (SCP) under the Lawrence Livermore National Laboratory (LLNL) undertook, in coordination with the Republican Seismic Survey Center of the Azerbaijan National Academy of Sciences and the Azerbaijan Scientific Research Institute of Construction and Architecture, a PSHA study that reflects new seismic data recorded locally and recent research conducted nationally and regionally since the last update to the building code. The PSHA framework for this project was designed to help develop a new earthquake catalogue, to incorporate a novel characterization of ground motions that specifically reflects the attenuation characteristics in the eastern Caucasus, and to generate hazard information in a form that is useful for an update of the current building code or the development of a new building code for Azerbaijan. The project also aimed to provide training and support for the local seismologists and engineers related to the seismic hazard models, probabilistic seismic hazard results and their use towards changes in the building code.

58 GEOSCIENCES↗

Developing ML/AI Methods for High-Throughput Characterization of Multiple-Sensor Streams of Tokamak Dynamics for High-Speed Control (Final Report)

This project evaluated and developed new mathematical and algorithmic techniques capable of handling (in real-time) the growing amounts of data generated by modern fusion research. While existing numerical linear algebra (NLA) methods provide the backbone to classical data analysis and algorithms, these methods fundamentally do not port to distributed architectures nor do they allow low-latency data reduction for control. Motivated by the needs for modern fusion reactors, this project explored and implemented new numerical methods to characterize plasma dynamics, respond in real-time to discharge evolution, and to process massive-scale data accurately and rapidly more fully. This project links expertise in multiple-sensor diagnostics of tokamak plasma dynamics from Columbia University’s Plasma Physics Laboratory with expertise in massive-scale data reduction and extreme data control algorithms at Columbia University’s Data Science Institute. This interdisciplinary project (i) applied machine learning methods, (ii) implemented a properly-trained neural-network for very fast processing of high-speed plasma videography, and (ii) developed the applied mathematical methods, based on randomized-NLA (rNLA) routines, for data analysis, reduction, and real-time control. The Columbia University High Beta Tokamak-Extended Pulse (HBT-EP) facility provided data to test new algorithms and partnership with Columbia University's Data Sciences Institute evaluated the broader use of new algorithms for many challenging control applications.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗