Search NASA⌕ Search

SEARCH · Search NASA

Results for “Science Data Processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Cosmological Hydrodynamics at Exascale: A Trillion-Particle Leap in Capability

Resolving the most fundamental questions in cosmology requires simulations that match the scale, fidelity, and physical complexity demanded by next-generation sky surveys. To achieve the realism needed for this critical scientific partnership, detailed gas dynamics must be treated self-consistently with gravity for end-to-end modeling of structure formation. Exascale computing enables simulations that span survey-scale volumes while incorporating key astrophysical processes that shape complex cosmic structures. We present results from CRK-HACC, a cosmological hydrodynamics code built for extreme scalability. Using separation-of-scale techniques, GPU-resident tree solvers, in situ analysis pipelines, and multi-tiered I/O, CRK-HACCexecuted Frontier-E: a four trillion particle full-sky simulation, over an order of magnitude larger than previous efforts. The run achieved 513.1 PFLOPs peak performance, processing 46.6 billion particles per second and writing more than 100 PB of data in just over one week of runtime. Frontier-E marks a significant advance in predictive modeling for next-generation cosmological science.

Frontiere, Nicholas [Argonne National Laboratory (↗

Deployment and Evaluation of SciStream on OLCF's Advanced Computing Ecosystem (ACE)

The growing demand for real-time analysis, experimental steering, and decision-making in scientific workflows has created a need for tightly coupled integrations between experimental facilities and high-performance computing (HPC) systems. The Department of Energy’s Integrated Research Infrastructure (IRI) initiative highlights data streaming as a key capability for enabling memory-to-memory data transfers, bypassing the limitations of traditional store-and-forward models. SciStream is a toolkit developed by researchers at Argonne National Laboratory (ANL) to support such streaming by addressing cross-domain security, delegated authentication, and application transparency. We deployed and evaluated SciStream on the Oak Ridge Leadership Computing Facility’s (OLCF) Advanced Computing Ecosystem (ACE) infrastructure, leveraging the Olivine OpenShift cluster and its high-bandwidth Data Streaming Nodes (DSNs) as gateway nodes. Our evaluation included synthetic streaming workloads derived from IRI science workflows, a streaming simulator, and integration with RabbitMQ to handle low-level messaging. This report documents the deployment process, performance evaluation, and challenges encountered, along with opportunities for future improvements.

97 MATHEMATICS AND COMPUTING↗

Lessons Learned from Ecosystem-Scale Experimental Field Studies (Workshop Report)

Efforts to understand and predict ecosystem responses to environmental change require long-term, large-scale, spatially representative experiments and observations that capture natural variability, test predictive models, and generate transferable knowledge. Such studies are indispensable for unraveling the complexities of terrestrial ecosystems and their responses to disturbances and evolving environmental conditions, while generating the data necessary for developing mechanistic models and predictive tools that inform decision-making processes. Having a rich history of designing and executing large-scale ecosystem experiments, the U.S. Department of Energy’s Environmental System Science program convened a workshop in January 2025 that brought together leaders in the field to distill critical lessons from decades of experience in large-scale experiments. The workshop aimed to (1) provide an ecosystem experiment primer for best practices, thus ensuring a high scientific return on investment for funding agencies, and (2) offer a robust framework for the design and management of future research initiatives. This report synthesizes insights and experiences from workshop participants and is structured to capture the entire research life cycle, from goal setting and design to operations, adaptive management, team dynamics, collaborations, and the often overlooked aspect of decommissioning. By synthesizing decision-making and lessons learned across diverse research approaches, the report aims to provide a template of essential factors to consider when designing successful long-term, large-scale ecosystem experiments.

54 ENVIRONMENTAL SCIENCES↗

Generic and ML Workloads in an HPC Datacenter: Node Energy, Job Failures, and Node-Job Analysis

HPC datacenters offer a backbone to the modern digital society. Increasingly, they run Machine Learning (ML) jobs next to generic, compute-intensive workloads, supporting science, business, and other decision-making processes. However, understanding how ML jobs impact the operation of HPC datacenters, relative to generic jobs, remains desirable but understudied. In this work, we leverage long-term operational data, collected from a national-scale production HPC datacenter, and statistically compare how ML and generic jobs can impact the performance, failures, resource utilization, and energy consumption of HPC datacenters. Our study provides key insights, e.g., ML-related power usage causes GPU nodes to run into temperature limitations, median/mean runtime and failure rates are higher for ML jobs than for generic jobs, both ML and generic jobs exhibit highly variable arrival processes and resource demands, significant amounts of energy are spent on unsuccessfully terminating jobs, and concurrent jobs tend to terminate in the same state. We open-source our cleaned-up data traces on Zenodo (https://doi. org/10.5281/zenodo.13685426), and provide our analysis toolkit as software hosted on GitHub (https://github.com/atlarge-research/2024-icpads-hpc-workload-characterization). This study offers multiple benefits for data center administrators, who can improve operational efficiency, and for researchers, who can further improve system designs, scheduling techniques, etc.

crossanalysis↗

EPCAPE Radar b1 Data Processing: Corrections, Calibrations, and Processing Report

The U.S. Department of Energy (DOE)’s Atmospheric Radiation Measurement (ARM) user facility recently deployed its First ARM Mobile Facility (AMF1) to La Jolla, California as part of the Eastern Pacific Cloud Aerosol Precipitation Experiment (EPCAPE) campaign. Some of the goals behind EPCAPE were to characterize the diurnal and seasonal cycles of stratocumulus clouds and to investigate the cloud-aerosol-radiation interactions and feedbacks in the area. The deployment of the AMF1 for a full year from 15 February 2023 to 14 February 2024 aided in addressing these scientific questions. While AMF1 collected data year-round, enhanced measurements were taken during two intensive operational periods (IOPs). The first IOP occurred from April to June and focused on the chemistry of low clouds (EPCAPE_Chem), while the second IOP occurred from July to September and was focused on the radiation of high clouds (EPCAPE_Radiation). Several cloud radars were deployed with AMF1 to collect valuable data on cloud properties that will help users address key science objectives. As in past ARM campaigns, a1-level radar data is extensively analyzed and calibration techniques are performed to generate b1-level data (Matthews et al. 2023, Feng et al. 2024). Radar data at the b1-level are of the highest quality and thus can be used to examine scientific questions. The status of the a1-level data and the a1-to-b1 process for the EPCAPE radars is subsequently detailed in this document.

54 ENVIRONMENTAL SCIENCES↗

Model Data Archive for Manuscript Titled "Evaluation of a Coupled Surface–Subsurface Hydrologic Model Using Dense Water‑Level Sensors in a Mixed Urban–Rural Watershed"

This archive provides scripts, input files, and datasets used for the implementation and evaluation of a fully coupled surface–subsurface hydrologic model in the Neches River Basin, southeast Texas. The study uses the Advanced Terrestrial Simulator (ATS) to simulate coupled surface–subsurface hydrologic processes over a mixed urban–rural watershed and evaluates model performance using a dense network of 136 in situ water-level sensors, nine U.S. Geological Survey (USGS) stream gauges, and SSEBop-derived evapotranspiration estimates during the period October 2014–June 2024. The workflow is implemented primarily in Python 3 using the Watershed Workflow package. The Jupyter notebooks can be executed using open-source software such as Anaconda JupyterLab or Visual Studio Code. Other data files include TXT, CSV, XML, SHP, TIF, NetCDF, HDF5, and ExodusII files, which can be processed using the provided Python scripts. ATS input files are provided in XML format and can be edited using any commonly used text editor. This archive contains: *Scripts and input files used to generate the ATS model setup, including watershed discretization, mesh generation, parameter mapping, and model configuration. *Jupyter notebooks used for preprocessing observational data, evaluating streamflow, water levels, and evapotranspiration, computing performance metrics, and generating the figures presented in the manuscript. *ATS simulation outputs and processed observational datasets, including OneRain and DD6 water-level sensors, USGS streamflow observations, GIS data, and supporting spatial datasets used throughout the study.

Dense water-level sensor network↗

U.S. Agrivoltaics Irradiance Database

This is a foundational data set for research and deployment of agrivoltaics, which is the co-location of agriculture and solar power plants on the same land. This irradiance and shading dataset can be utilized to determine the suitability of agrivoltaics configurations for a given region and crop-type. The data is hourly, 4x4 km resolution across the contiguous United States and Hawaii. It is calculated from the National Solar Radiation Database sites, using the System Advisor Model (SAM) to simulate the shading patterns for 10 common agrivoltaics configurations. Sunlight availability data is reported for 10 locations on the ground between adjacent rows of solar panels, as well as averaged across areas of interest such as the average irradiance in the edge-to-edge open area or across 3-6 planting beds. Other available metrics include the input meteorological data from the NSRDB (e.g. global horizontal irradiance, wind speed, etc.) and estimates for comparing energy and agricultural characteristic across the 10 configurations, including power output per acre or per kW installed capacity and farmable land area per acre.

14 SOLAR ENERGY↗

A Brief Survey of Data Streaming Technologies

Streaming data is data that is emitted at variable volumes in a continuous, incremental manner with the goal of low-latency processing often at a different physical location. Network infrastructure is used to facilitate the connection between data sources and sinks, and must be robust to handle the requirements of the workflow. The U.S. Department of Energy Office of Science (DOE SC) a federal agency supporting fundamental scientific research for energy and the Nation’s largest supporter of basic research in the physical sciences. DOE SC has the responsibility for operating $\mathbf{1 0}$ National Laboratories, and 28 scientific user facilities supporting advanced supercomputers, particle accelerators, large x-ray light sources, neutron scattering sources, and other specialized facilities for nanoscience and genomics. This paper investigates the state of streaming data workfows, and details some of the approaches to this challenging problem.

Kissel, Ezra↗

High fidelity actuator line data from 9 turbine wind farm simulations using ExaWind

This data was generated with the ExaWind code suite (https://github.com/Exawind) to investigate the performance of different Active Wake Mixing turbine control in a wind farm situated in a stable atmospheric boundary layer. All cases correspond to a 3x3 wind farm in a 10km x 10km domain using a total mesh size that varied between 1.6 X 10^9 to 1.85 X 10^9 grid cells. The simulations were run across 1800-2000 GPUs on Frontier. The case description and data generation process is fully documented in Yalla, G. R., Brown, K., Cheung, L., Houck, D., deVelder, N., and Balaji, J. (2025). "Estimating annual energy production of wake mixing control strategies including comparisons to wake steering." Wind Energy Sciences (https://doi.org/10.5194/wes-2025-250).

17 WIND ENERGY↗

High-density Lipoprotein (HDL) Structure and Function Proteomics (JM-DP1)

The purpose of this experiment was to investigate how the interactions between APOA1 and APOA2 on the surface of high-density lipoproteins (HDL) impact particle function by studying the effect of exogenous APOA2 on HDL structure through limited proteolysis. Interactions were investigated on HDL isolated from human blood plasma using structural proteomics tools such as chemical cross-linking and limited proteolysis (LiP). The structural proteomics data was acquired using a Q-Exactive HF-X mass spectrometer and processed using MaxQuant software (v.1.6.17.0).

59 BASIC BIOLOGICAL SCIENCES↗

Metal Scrap Upcycling with Shear Assisted Processing and Extrusion (ShAPE)

The overarching objective of this project is to convert metal scraps, such as aluminum, titanium, and other alloys provided by the industry, into extruded tubing, wires, and rods. Upcycling of scrap will be accomplished using Shear Assisted Processing and Extrusion (ShAPE). This approach is a new solution for recycling. The specific aims of this project are as follows: 1. Receive metal scrap under a Material Transfer Agreement (MTA), in the form of billets, from select industry partners that meet the following requirements: outer diameter of 1.245 inches (+/-0.003 inches), inner diameter drilled with a 0.404-inch drill bit, and a length of 4.0 inches (+/-0.01 inches). The billet must be cast or compacted to greater than 98% density. If the industry partner does not have the capabilities to prepare the billets, PNNL can make introductions to third-party entities as needed. Industry partners will also provide a composition analysis as weight percent. 2. Extrude metal scrap via the ShAPE process at PNNL. 3. Evaluate and benchmark the extrudate material properties per the ASTM B557-15 – Testing Tubulars Standard or similar wire and rod standards. 4. Characterize the extrudate microstructure for any of the following: grain size, second phase composition, and texture. Additional testing may include corrosion per the ASTM B117 standard and electro-potential. 5. Deliver specimens to industry partners under an MTA for additional third-party evaluations. Industry partners will, in return, provide a non-proprietary report on any testing, including testing per the ASTM B557-15, ASTM B117, and other industry standards. 6. PNNL will develop data and insights to support intellectual property capture, a published non-proprietary technical report, research collaborations, and commercialization opportunities.

36 MATERIALS SCIENCE↗

CMIP7 data request: Earth system priorities and opportunities

This paper presents a comprehensive overview of the Coupled Model Intercomparison Project Phase 7 (CMIP7) request for data pertaining to Earth systems science, and provides justification for the resources needed to produce this data. Topics within the CMIP7 Earth System (CMIP7-ES) theme centre around tracking of flows of energy, carbon, water and other fluxes across domains, and constraining feedbacks between these cycles and the climate system. These topics are summarized in this paper as scientific “opportunities” describing specific model intercomparison experiments and use cases for next-generation Earth System Model (ESM) output. These opportunities were submitted by modelling groups and scientific consortia following an extended public consultation process. Contained within each opportunity are requests for groups of Climate & Forecasting (CF) variables, which are bundled into variable groups representing all data required to address the opportunities' needs. Novel opportunities in CMIP7 compared with previous phases will include running `emissions-driven' simulations that integrate carbon emissions and removal scenarios with updated representations of the global carbon cycle, expanded variable groups needed to model marine trophic interactions and biogeochemistry, and data needed to understand the risk of global tipping points, among others. The production of these variables will close key gaps and uncertainties identified during previous rounds of CMIP, and support the 7th Intergovernmental Panel on Climate Change Assessment Report (AR7). We argue that CMIP7-ES data will be broadly used by scientific, policy, governmental, industry, and other communities that rely on climate model projections for research and decision making. As an author group we also reflect on the evolution of the CMIP7-ES data request as a part of a deliberative process in support of the global CMIP program.

54 ENVIRONMENTAL SCIENCES↗

2024 OES-Environmental 2024 State of the Science Report, Chapter 6: Strategies to Aid Consenting Processes for Marine Renewable Energy

While the marine renewable energy (MRE) industry has made positive strides in the past decade, challenges remain that stall forward progress, scaling up, and commercialization. For MRE to provide a viable solution to address the effects of climate change and achieve sustainable development and renewable energy goals, identifying and understanding barriers and opportunities to deployment is key. Barriers to date have included long consenting timelines, costly in-depth baseline data collection and monitoring requirements, and hesitancy in some countries to approve device and array deployments (Copping & Hemery 2020; Kramer et al. 2020). Some of the key drivers behind these barriers are 1) uncertainty about potential effects of MRE on marine animals, habitats, and the environment; 2) lack of familiarity with MRE technologies; or 3) challenges accessing available scientific information (Copping et al. 2020a).

16 TIDAL AND WAVE POWER↗

Advancing electrochemical impedance analysis through innovations in the distribution of relaxation times method

Electrochemical impedance spectroscopy (EIS) is a key tool across various scientific disciplines, including energy sciences, chemistry, and biology, enabling the analysis of electrochemical systems. However, conventional methods for interpreting EIS data are often complex and model dependent. The distribution of relaxation times (DRT) offers a non-parametric approach that simplifies the interpretation process by providing a timescale interpretation of EIS data. This article provides a comprehensive review of current methods for DRT inversion. Additionally, a survey of practitioners highlights key challenges in the field. Here, the findings underscore the need for standardized DRT analysis and benchmarks, as well as the development of automated analysis tools. These advancements would improve the usability and interpretability of EIS data. Ultimately, implementing these improvements could not only propel the field forward but also expand the application of DRT in scientific research by making it accessible to a broader range of researchers, including those without specialized expertise in programming or statistics.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

ARM FY2026 Radar Plan

The U.S. Department of Energy (DOE) Atmospheric Radiation Measurement (ARM) User Facility maintains a suite of advanced atmospheric radar systems that serve as critical tools in ARM’s mission to provide continuous, high-quality observations for advancing the understanding and modeling of atmospheric processes. These radar systems enable detailed characterization of clouds, precipitation, and dynamic structures in the atmosphere, supporting a broad range of scientific applications. The number of deployed systems exceeds what current staffing levels can fully support for continuous 24/7/365 operation. As such, it is essential to have a clearly defined and community-informed plan that prioritizes radar operations and communicates ARM’s strategy for sustaining and evolving these observational assets. This FY2026 Radar Plan outlines ARM’s approach to managing its radar portfolio—balancing scientific impact, operational feasibility, and long-term sustainability. It reflects ARM’s continued commitment to delivering calibrated, well-documented radar data products that enable process-level studies and support the development and evaluation of weather and climate models. Through this plan, ARM aims to ensure transparency in decision-making, alignment with user needs, and support for innovative science across the facility’s fixed and mobile observatories. Given uncertainties around the Fiscal Year (FY) 2026 budget, this plan was developed to assume business as usual and will be updated as budgets and plans may change. It should be noted that, given the limited timeframe involved, this plan will be more succinct than previous plans.

47 OTHER INSTRUMENTATION↗

Data and scripts from: “Denoising autoencoder for reconstructing sensor observation data and predicting evapotranspiration: noisy and missing values repair and uncertainty quantification”

This data package includes data and scripts from the manuscript “Denoising autoencoder for reconstructing sensor observation data and predicting evapotranspiration: noisy and missing values repair and uncertainty quantification”.The study addressed common challenges faced in environmental sensing and modeling, including uncertain input data, missing sensor observations, and high-dimensional datasets with interrelated but redundant variables. Point-scaled meteorological and soil sensor observations were perturbed with noises and missing values, and denoising autoencoder (DAE) neural networks were developed to reconstruct the perturbed data and further predict evapotranspiration. This study concluded that (1) the reconstruction quality of each variable depends on its cross-correlation and alignment to the underlying data structure, (2) uncertainties from the models were overall stronger than those from the data corruption, and (3) there was a tradeoff between reducing bias and reducing variance when evaluating the uncertainty of the machine learning models.This package includes:(1) Four ipython scripts (.ipynb): “DAE_train.ipynb” trains and evaluates DAE neural networks, “DAE_predict.ipynb” makes predictions from the trained DAE models, “ET_train.ipynb” trains and evaluates ET prediction neural networks, and “ET_predict.ipynb” makes predictions from trained ET models.(2) One python file (.py): “methods.py” includes all user-defined functions and python codes used in the ipython scripts.(3) A “sub_models” folder that includes five trained DAE neural networks (in pytorch format, .pt), which could be used to ingest input data before being fed to the downstream ET models in ‘ET_train.ipynb” or ‘ET_predict.ipynb’.(4) Two data files (.csv). Daily meteorological, vegetation, and soil data is in “df_data.csv”, where “df_meta.csv” contains the location and time information of “df_data.csv”. Each row (index) in “df_meta.csv” corresponds to each row in “df_data.csv”. These data files are formatted to follow the data structure requirements and be directly used in the ipython scripts, and they have been shuffled chronologically to train machine learning models. The meteorological and soil data was collected using point sensors between 2019-2023 at(4.a) Three shrub-dominated field sites in East River, Colorado (named “ph1”, “ph2” and “sg5” in “df_meta.csv”, where “ph1” and “ph2” were located at PumpHouse Hillslopes, and “sg5” was at Snodgrass Mountain meadow) and(4.b) One outdoor, mesoscale, and herbaceous-dominated experiment in Berkeley, California (named “tb” in “df_meta.csv”, short for Smartsoils Testbed at Lawrence Berkeley National Lab).- See "df_data_dd.csv" and "df_meta_dd.csv" for variable descriptions and the Methods section for additional data processing steps. See "flmd.csv" and "README.txt" for brief file descriptions.- All ipython scripts and python files are written in and require PYTHON language software.

54 ENVIRONMENTAL SCIENCES↗

FAIRmaterials: Ontology Tools with Data FAIRification in Development

The bilingual FAIRmaterials package simplifies the creation and visualization of materials and data science ontologies. FAIRmaterials, available in the Python and R languages, addresses the complexities associated with traditional ontology editors based on manual user input such as Protege with an intuitive workflow and easy-to-use templates, making it accessible to users both experienced and inexperienced with ontologies. The FAIRmaterials package is its ability to programatically convert simple and structured CSV inputs into rich, well-defined ontologies. This capability is designed to support the findability, accessibility, interoperability, and reusability (FAIR) of research data and serve as a tool in the process of data FAIRification. Its additional features, such as automated ontology merging, static visualizations, and comprehensive documentation for outputs extend its utility, making it a valuable tool for any researcher engaged in knowledge management.

Bradley, Alexander Harding [Case Western Reserve U↗

Data for Genetic Variation in Zea mays Influences Microbial Nitrification and DeNitrification in Conventional Agroecosystems

Nitrogenous fertilizers provide a short-lived benefit to crops in agroecosystems, but stimulate nitrification and denitrification, processes that result in nitrate pollution, N2O production, and reduced soil fertility. Recent advances in plant microbiome science suggest that genetic variation in plants can modulate the composition and activity of rhizosphere N-cycling microorganisms. Here we attempted to determine whether genetic variation exists in Zea mays for the ability to influence the rhizosphere nitrifier and denitrifier microbiome under “real-world” conventional agricultural conditions. To capture an extensive amount of genetic diversity within maize we grew and sampled the rhizosphere microbiome of a diversity panel of germplasm that included ex-PVP inbreds ( Z. mays ssp. mays ), ex-PVP hybrids ( Z. mays ssp. may s), and teosinte ( Z. mays ssp. mexicana and Z. mays ssp. parviglumis ). From these samples, we characterized the microbiome, a suite of microbial genes involved in nitrification and denitrification and carried out N-cycling potential assays. Here we are showing that populations/genotypes of a single species can vary in their ecological interaction with denitrifers and nitrifers. Some hybrid and teosinte genotypes supported microbial communities with lower potential nitrification and potential denitrification activity in the rhizosphere, while inbred genotypes stimulated/did not inhibit these N-cycling activities. These potential differences translated to functional differences in N2O fluxes, with teosinte plots producing less GHG than maize plots. Taken together, these results suggest that Zea genetic variation can lead to changes in N-cycling processes that result in N leaching and N2O production, and thereby are selectable targets for crop improvement. Understanding the underlying genetic variation contributing to belowground microbiome N-cycling into our conventional agricultural system could be useful for sustainability.

Nitrogen↗