Search NASA⌕ Search

SEARCH · Search NASA

Results for “data processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27

Deeplynx Dag Repository

The DeepLynx DAG repository will contain several Airflow DAGs (Directed Acyclic Graphs) which will be used in the context of DeepLynx's deployed Apache Airflow instance. These DAGs will be used for multiple data management tasks for DeepLynx data, including but not limited to: - bringing data from various sources and tools into DeepLynx - managing sequential data workflows, such as running Python scripts on data to perform analysis and returning the results to DeepLynx - performing any necessary transformation or pre-processing on data coming into DeepLynx from external sources or out of DeepLynx to go to external applications

Brownlee, JarenM.↗

Patch2Self2: Self-supervised Denoising on Coresets via Matrix Sketching

Diffusion MRI (dMRI) non-invasively maps brain white matter yet necessitates denoising due to low signal-to-noise ratios. Patch2Self (P2S) employing self-supervised techniques and regression on a Casorati matrix effectively denoises dMRI images and has become the new de-facto standard in this field. P2S however is resource intensive both in terms of running time and memory usage as it uses all voxels (n) from all-but-one held-in volumes (d-1) to learn a linear mapping Phi : \mathbb R ^ n x(d-1) \mapsto \mathbb R ^ n for denoising the held-out volume. The increasing size and dimensionality of higher resolution dMRI acquisitions can make P2S infeasible for large-scale analyses. This work exploits the redundancy imposed by P2S to alleviate its performance issues and inspect regions that influence the noise disproportionately. Specifically this study makes a three-fold contribution: (1) We present Patch2Self2 (P2S2) a method that uses matrix sketching to perform self-supervised denoising. By solving a sub-problem on a smaller sub-space so called coreset we show how P2S2 can yield a significant speedup in training time while using less memory. (2) We present a theoretical analysis of P2S2 focusing on determining the optimal sketch size through rank estimation a key step in achieving a balance between denoising accuracy and computational efficiency. (3) We show how the so-called statistical leverage scores can be used to interpret the denoising of dMRI data a process that was traditionally treated as a black-box. Experimental results on both simulated and real data affirm that P2S2 maintains denoising quality while significantly enhancing speed and memory efficiency achieved by training on a reduced data subset.

Fadnavis, Shreyas↗

Laue-DIALS: Open-source software for polychromatic x-ray diffraction data

Most x-ray sources are inherently polychromatic. Polychromatic (“pink”) x-rays provide an efficient way to conduct diffraction experiments as many more photons can be used and large regions of reciprocal space can be probed without sample rotation during exposure—ideal conditions for time-resolved applications. Analysis of such data is complicated, however, causing most x-ray facilities to discard >99% of x-ray photons to obtain monochromatic data. Key challenges in analyzing polychromatic diffraction data include lattice searching, indexing and wavelength assignment, correction of measured intensities for wavelength-dependent effects, and deconvolution of harmonics. We recently described an algorithm, Careless, that can perform harmonic deconvolution and correct measured intensities for variation in wavelength when presented with integrated diffraction intensities and assigned wavelengths. Here, we present Laue-DIALS, an open-source software pipeline that indexes and integrates polychromatic diffraction data. Laue-DIALS is based on the dxtbx toolbox, which supports the DIALS software commonly used to process monochromatic data. As such, Laue-DIALS provides many of the same advantages: an open-source, modular, and extensible architecture, providing a robust basis for future development. We present benchmark results showing that Laue-DIALS, together with Careless, provides a suitable approach to the analysis of polychromatic diffraction data, including for time-resolved applications.

97 MATHEMATICS AND COMPUTING↗

Privacy Preservation from High-Performance Computing to Autonomous Science [Industrial and Governmental Activities]

High-Performance Computing (HPC) and Leadership-Class Supercomputing are driving forces behind scientific advancements, enabling researchers to tackle complex challenges in physics, chemistry, biology, and engineering. These systems power vast simulations and data analyses, fueling discoveries in fields ranging from materials science to climate modeling. However, their use often involves processing sensitive data—such as proprietary industry simulations, biomedical records, and national security computations—posing significant privacy concerns. In conclusion, this issue is amplified in collaborative environments like Department of Energy (DOE) user facilities, where HPC resources are shared across institutions to foster innovation.

Kotevska, Olivera [Oak Ridge National Laboratory (↗

RhizoGrid Indexed Sorghum Rhizosphere Multi-Omics

PerCon SFA project data dentification of spatially resolved biomarkers of drought in Sorghum bicolor rhizosphere molecular-microbe interactions using a novel root cartography "RhizoGrid" system for sampling plants under drought and control conditions across 10 equally sized root zone environments (4 quadrants each). Each quadrant was sampled and processed for 16S amplicon, metabolomics, and X-ray computed tomography (XCT). Data download includes experimental metadata and results files for 16S rRNA sequence analysis of microbial community assembly (processed data files), liquid chromatography mass spectrometry (LC-MS) metabolomics analysis of microbial community root exudates (processed data files), X-ray computed tomography (XCT) spatial gradient analysis (raw and processed data files) of microbial community composition, and related computational modeling outputs.

59 BASIC BIOLOGICAL SCIENCES↗

CHESS 2025: Leaf Area Index (LAI) for meadow, shrub, tree, and understory vegetation

This dataset contains Leaf Area Index (LAI) measurements made as part of the Colorado Headwaters Ecological Spectroscopy Study (CHESS) during June and July of 2025. Data were collected in the Upper Gunnison Basin, Colorado, across three study domains: the Upper East River (CRBU), Almont Triangle (ALMO), and the Upper Taylor Basin (UPTA). Field observations of LAI were collected within 72 hours of airborne data collection by the National Ecological Observatory Network’s Aerial Observation Platform (NEON AOP). The NEON AOP collected waveform LiDAR (Light Detection and Ranging) and imaging spectrometer data in 426 spectral bands from the visible to shortwave infrared. LAI measurements were collected using the LICOR LAI-2200C Plant Canopy Analyzer following protocols outlined in the instrument manual (LI-COR 2019). Sampling targeted four distinct vegetation types: meadows, shrubs, trees, and aspen forest understory. We have archived data separately by site type because different field methods were used for each. At meadow sites, measurements were made at the four corners of 1m x 1m plots, with the instrument moving inward toward the center of the plot. At shrub sites, we measured the canopies of individual shrubs. At tree sites, we made measurements within a 10m x 10m subplot centered around a focal tree, with 30 observations taken on a regular grid. At aspen understory sites, we measured overstory trees following the tree protocol and understory herbaceous vegetation following the meadow protocol. All measurements included above-canopy (A) and below-canopy (B) readings, with specific protocols for scattering correction measurements in direct-sun conditions. Data were processed using the R package `rlai` (Worsham 2025). This package includes functions to calculate LAI, gap fraction, apparent clumping factor (Ω), scattering correction, and other canopy metrics. Package contents: Full file descriptions appear in ‘flmd.csv’. Files named according to the convention ‘lai_*_summary_data_cleaned.csv’ contain summary values of LAI, apparent clumping factor (Ωapp), and scattering correction factors for each site. These are the analysis-ready products that most data users will work with. Files named ‘lai_*_metadata_cleaned.csv’ contain additional site-level observations made during field collection. We have also archived intermediate and supplementary data for users who wish to check our processing approach or apply alternative methods. ‘raw_lai_2200C.zip’ contains the raw files as read from the LI-COR instrument, with no processing applied, in TXT format. The zip archive contains subdirectories by site type, which are further subdivided by sampling area. Filenames correspond to the sampling site number. ‘intermediate_results.zip’ contains detailed output from the processing routines, in JSON format. The zip archive contains subdirectories by site type; filenames correspond to the sampling site number. ‘scattering_correction_logs.zip’ contains logfiles from the implementation of Kobayashi et al.'s (2013) scattering correction algorithm. The logfiles report values of several parameters at each iteration of the algorithm, as the model converges toward a stable solution. They are intended for users who want to verify scattering correction performance. The zip archive contains subdirectories by site type; filenames correspond to the sampling site number. ‘spot_checks.csv’ reports LAI and other values for a small number of files processed with LI-COR FV2200 software (LI-COR 2013) using the same control parameters as in our R-based approach. Additional metadata are provided in a data dictionary describing column names and definitions (dd.csv), and in a file-level metadata file (flmd.csv). All zip files can be expanded with common archive utilities. TXT, CSV, and JSON files can be ingested into R or Python computing environments or read in common text editor utilities. Geospatial information: Geospatial data for mapping measurement site locations are in the files CHESS_polygons_lai_UTM.geojson, CHESS_polygons_shrub_UTM.geojson, and CHESS_polygons_meadow_UTM.geojson in the companion geospatial package for the 2025 CHESS campaign, ‘CHESS 2025: Location data for field observations and sampling’ (Henderson et al., 2026). CHESS Project Description: The Colorado Headwaters Ecological Spectroscopy Study (CHESS) comprised a multi-week airborne remote sensing and field observation campaign in the Upper Gunnison Basin, Colorado, conducted in June and July of 2025. Airborne remote sensing was conducted by the National Ecological Observatory Network Airborne Observation Platform (NEON AOP), concurrent with a field campaign run by the Rocky Mountain Biological Laboratory (RMBL), the Lawrence Berkeley National Laboratory (LBNL) and SLAC National Accelerator Laboratory Watershed Function Science Focus Area (SFA), and NASA-JPL (Jet Propulsion Laboratory) Earth Surface Mineral Dust Source Investigation (EMIT) program. Between June 10 and July 18, 2025, the NEON AOP flight team collected high-resolution aerial imaging spectroscopy and Light Detection and Ranging (LiDAR) data over three domains: the Upper East River (CRBU), Almont Triangle (ALMO), and the Upper Taylor Basin (UPTA). In coordination with the flights, a field campaign acquired ground-truth observations, including observations of vegetation composition, foliar traits, forest demography, and subsurface properties in 18 core sampling areas within the domains. Additional surface water observations were taken at over 380 point locations. All CHESS campaign datasets can be found within the CHESS ESS-DIVE data portal: https://data.ess-dive.lbl.gov/portals/chess. Funding Acknowledgement: Field and remote-sensing data acquisition was performed under a grant from the National Aeronautics and Space Administration (80NSSC24K1005). This work was also supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. * Todorov and Worsham are co–first authors.

2018 NEON and 2025 CHESS Campaigns↗

Adoption of ROOT RNTuple for the next main event data storage technology in the ATLAS production framework Athena

Since the start of LHC in 2008, the ATLAS experiment has relied on ROOT to provide storage technology for all its processed event data. Internally, ROOT files are organized around TTree structures that are capable of storing complex C++ objects. The capabilities of TTrees developed over the years and are now offering support for advanced concepts like polymorphism, schema evolution and user defined collections and ATLAS makes use of these features to handle its EDM. But some original TTrees concepts, like the POSIX file model and sequential writing, remain unchanged since the beginning and could be an obstacle to achieving the performance required for High Luminosity LHC. With the HL-LHC performance goals in mind, the ROOT project developed a new storage format - the RNTuple. RNTuple, with its accompanying user API, is now in the final development stage and is planned to be production-ready at the end of 2024. Soon after that, the TTree will become a legacy format. ATLAS intends to have its main Event processing framework Athena ready to use RNTuple in the production environment as early as possible. The work on adopting RNTuple as another ROOT storage technology in Athena started already in 2021 and is now nearly complete. Although the initial goal was to focus on derived-AOD products (PHYS and PHYSLITE), with a little added effort all ATLAS data products: RDO, HITS, ESD, AOD and DAOD can be now stored in RNTuple format and transparently read back. In this paper we will describe the current state of RNTuple adoption in the Athena framework and explain the ATLAS EDM requirements that had to be met on the ROOT side to successfully integrate both environments. We will demonstrate the ability to run standard ATLAS production workflows, based on RNTuple as the Event data storage technology, and point out key advantages of the new format.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Core Model Proposal #377: Breaking out food processing sector in GCAM

Purpose: This Core Model Proposal (CMP) expands the representation of detailed industry (CMP-326) in GCAM by separating the food processing sector from the aggregate “other industry sector”. Historical energy use is calibrated to IEA data for food processing, with some infilling for regions with limited IEA data. Food processing is linked to the GCAM food demand module, setting the energy demand for food processing in future periods based on food demand. While the direct price feedback is currently muted and the linkage is represented at the aggregated regional level, this CMP establishes the groundwork for a more detailed connection between the agrifood sectors and energy sectors in future work.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

ML-based Data Assimilation and History Matching: Application to the IBDP CCS Project

It is crucial to monitor the CO2 plume effectively throughout the life cycle of a geologic CO2 sequestration project to ensure safety and storage efficiency. However, the computational cost of existing data assimilation methods can be prohibitively expensive due to the complex physics with multi-component non-isothermal simulation and high dimensionality of large-scale reservoir models. We address this challenge by proposing an accelerated deep learning-based workflow for model calibration and prediction of CO2 plume evolution in the reservoir.The power and efficacy of our workflow is demonstrated by application to the Illinois Basin-Decatur Project (IBDP), a large-scale CO2 storage test in saline aquifer. The data assimilation process is implemented rapidly by the proposed workflow with given field measurements including distributed pressure and temperature sensing (DTS) data at an injection and a monitoring well. CO2 plume evolution is predicted by running the simulations of the calibrated reservoir models.

Nagao, Masahiro↗

Convergence of Emerging Technologies - EAGL Test Information

The Emergency Automatic Gunshot Detection and Lockdown (EAGL) system provides automatic, autonomous, and timely gunshot detection in both indoor and outdoor environments. This system uses both wired and wireless devices. Self-contained wireless EAGL sensors passively “listen” for gunshot events. These devices also perform a single, daily supervisory heartbeat (HB) function to include a device self-check with reporting capability. Transmissions are received by an assigned EAGL Gateway, which translates the RF sensor data to a PoE network format solely for use by the EAGL system server. The server then performs additional processes after data receipt, which include but are not limited to: event validation and logging, GUI presentation, notifications, and other independent operations.

47 OTHER INSTRUMENTATION↗

Understanding the Thermal Physics and Metallurgy of Metal Big Area Additive Manufacturing

The research goal of this EPSCoR-DOE partnership is to mitigate defects in parts made using a new type of additive manufacturing (AM) process called metal Big Area Additive Manufacturing (m-BAAM). To realize this goal, the PIs will detect and correct defects in the part as it is being printed by combining fundamental knowledge of the thermal physics and metallurgy of m-BAAM with in-process sensor data. Developed at the DOE-funded Manufacturing Demonstration Facility at Oak Ridge National Laboratory, the m-BAAM process involves one or more robots working together to produce a part by fusing metal wire layer-by-layer using arc welding. The process can print large metal parts such as turbine blades, which is not possible using other AM processes. In addition, m-BAAM production rates are more than ten times faster than other AM processes while requiring one-tenth of the material cost. Despite their potential to become a critical force multiplier in the energy generation industry, m-BAAM parts may fail to print accurately due to retention of heat and uneven cooling. Overheating and anomalous cooling rates in turn can cause inconsistencies in the microstructure, leading to sudden failure when used in safety-critical applications. In other words, flaw formation in m-BAAM parts is governed by the thermal history – intensity and spatial distribution of heat inside the part during printing. The thermal history is a complex function of the part shape and process settings such as welding energy, path taken by the welding torch for deposition (tool path), wire feed rate, among others.

36 MATERIALS SCIENCE↗

Laser interference structuring of Cu for adhesive joining of HVAC equipment

Due to potential energy and cost savings benefits, adhesive joining has been recently considered for heating, ventilation, air conditioning, and refrigeration (HVAC&R) systems. HVAC adhesive bonding requires cost-effective and adequate surface preparations of adherents. Here, this article investigates the use of abrasion, traditional single-beam laser, and a laser-interference technique as surface preparations of copper surfaces. Surface morphology is characterized using scanning electron microscopy (SEM) and atomic force microscopy (AFM). Effective submicrometer peak-to-valley structuring with a periodicity of ∼2.7μ⁢m was demonstrated for laser-interference processing. Single-lap shear tests were conducted for 89 joints made with bondline thicknesses of 0.15 and 0.3 mm. Data on process variables and measured variables included open-time, bond length, maximum load, displacement at failure, shear lap strength, and failure mode. A statistical analysis was conducted on each lot to determine the lower limit with 95% confidence intervals for displacements at failure and shear lap strengths. A comparison is presented between the properties of laser-structured joints with respect to those prepared by abrasion, which is considered the baseline surface preparation technique. Based on this comparison, one single-beam laser technique and two laser-interference techniques were shown to exhibit vastly superior performance over the joints made with abraded specimens.

Sabau, Adrian S. [Oak Ridge National Laboratory (O↗

Multifidelity_Timeseries

SAND2025-03305O Multifidelity Timeseries is a user-friendly tool designed to create advanced models for analyzing time-series data. It offers three modeling options, allowing users to choose the best fit for their specific needs. The software efficiently processes multiple data sources without the need for complex sampling methods. It helps uncover patterns and insights using data. The result is it is easier to make informed decisions for projects. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Katona, Ryan [Sandia National Lab. (SNL-CA), Liver↗

2024 Standard Scenarios: A U.S. Electricity Sector Outlook

This data corresponds to the 2024 Standard Scenarios report, which contains a suite of forward-looking scenarios of the possible evolution of the U.S. electricity sector through 2050. These files contain modeled projections of the future. Although we strive to capture relevant phenomena as comprehensively as possible, the models used to create this data are unavoidably imperfect, and the future is highly uncertain. Consequentially, this data should not be the sole basis for making decisions. In addition to drawing from multiple scenarios within this set, we encourage analysts to also draw on projections from other sources, to benefit from diverse analytical frameworks and perspectives when forming their conclusions about the future of the power sector. For further discussions about the limitations of the models underlying this data, see section 1.4 of the "ReEDS Documentation" linked below. For scenario descriptions, input assumptions, and metric definitions for the data in these files, see the "2024 Standard Scenarios Report" linked below.

2050↗

Alaska Meteorology, Energy, and Transmission (MET) Toolkit

The Alaska MET (Meteorology, Energy, and Transmission) Toolkit is the National Laboratory of the Rockies' (NLR) new flagship atmospheric dataset, designed to support comprehensive long-term planning and operations across the entire power sector. Serving as the regional counterpart to CONUS-wide HRRR MET Toolkit, this dataset provides a comprehensive, high-fidelity meteorological record covering Alaska.The Alaska MET Toolkit is delivered at an hourly resolution on a standardized 2-km horizontal grid. This dataset is repackaged from the National Oceanic and Atmospheric Administration's (NOAA) operational High-Resolution Rapid Refresh for Alaska (HRRR-AK) forecasts. Spanning from 2019 to 2025, it overcomes the technical barriers of native weather models by providing spatial regridding from the native 3-km HRRR-AK horizontal resolution to a 2-km grid, temporal gap-filling, and vertical interpolation at key energy-relevant heights. By delivering highly accurate, validation-backed data across a comprehensive suite of atmospheric variables - including temperature, pressure, humidity, and wind characteristics - the Alaska MET Toolkit provides a highly accessible and strictly standardized foundation for modern power system modeling.

17 WIND ENERGY↗

Integrated edge-to-exascale workflow for real-time steering in neutron scattering experiments

We introduce a computational framework that integrates artificial intelligence (AI), machine learning, and high-performance computing to enable real-time steering of neutron scattering experiments using an edge-to-exascale workflow. Focusing on time-of-flight neutron event data at the Spallation Neutron Source, our approach combines temporal processing of four-dimensional neutron event data with predictive modeling for multidimensional crystallography. At the core of this workflow is the Temporal Fusion Transformer model, which provides voxel-level precision in predicting 3D neutron scattering patterns. The system incorporates edge computing for rapid data preprocessing and exascale computing via the Frontier supercomputer for large-scale AI model training, enabling adaptive, data-driven decisions during experiments. This framework optimizes neutron beam time, improves experimental accuracy, and lays the foundation for automation in neutron scattering. Although real-time experiment steering is still in the proof-of-concept stage, the demonstrated potential of this system offers a substantial reduction in data processing time from hours to minutes via distributed training, and significant improvements in model accuracy, setting the stage for widespread adoption across neutron scattering facilities and more efficient exploration of complex material systems.

97 MATHEMATICS AND COMPUTING↗

LandScan Mosaic

The LandScan program at Oak Ridge National Laboratory (ORNL), in collaboration with the National Geospatial-Intelligence Agency (NGA), continues to deliver the most accurate and up to date global, high resolution gridded population data. Additionally, the latest advancements in the LandScan HD methodology led to reduced latency in development of rapid updates for geopolitical events. With momentum towards reporting more up to date population estimates, feedback from the user community expressed interest in reporting population estimates in ranges - whether to express a level of uncertainty or confirm to leadership and stakeholders the modeled data are estimates. Building upon the need to understand uncertainty or confidence in the modeled data and report ranges at the global scale, LandScan Mosaic was developed. LandScan Mosaic represents the next generation of high-resolution population modeling, building upon the established success of previous LandScan HD iterations. While LandScan HD employed a deterministic big data fusion approach, LandScan Mosaic enhances this methodology by integrating advanced machine learning techniques to impute missing, yet crucial, population model parameters. This advancement allows for probabilistic modeling of building occupancy and population distribution, incorporating uncertainty quantification through Monte Carlo sampling methods. By combining big data fusion with machine learning-driven imputation and stochastic modeling, LandScan Mosaic provides a more comprehensive and robust representation of population dynamics. LandScan Mosaic will be following the in the footsteps of its longstanding counterpart LandScan Global and releasing a global gridded population raster, at the 3-arcsecond resolution. This technical report documents the current stage of development of LandScan Mosaic, detailing the methodologies and data sources behind the modeling. Stakeholders are encouraged to use this document as an authoritative reference for insight into Mosaic’s data development processes. However, readers should note that LandScan Mosaic remains in a late-stage research and development phase, and methodologies and data presented here are subject to refinements ahead of the anticipated global release in Summer 2025. Feedback and inquiries from users and stakeholders are welcomed as we continue to refine and enhance this important population resource.

97 MATHEMATICS AND COMPUTING↗

WELLBASE - An Interactive Platform for Wellbore Material Assessment

This project seeks to build an open-source wellbore material data repository with adequate material performance and contextual data to support Geological Carbon Storage (GCS). By appropriately evaluating the data types as mentioned earlier made available by the WELLBASE tool, stakeholders can make more informed decisions regarding well selections, risk assessment, and economic analysis for geologic carbon storage projects. Advanced Natural Language Processing models and other custom python scripts will be deployed in an automated process to extract unstructured data from documents, reports, and web applications and subsequently parse to more usable formats. The processed data will then be integrated into a robust and comprehensive database architecture, optimizing data accessibility, and usability for analytical purposes. The final data products will be accessible through a user-friendly visualization platform that will allow users to query and visualize the data, as well as download data in usable formats.

Tetteh, Daniel A.↗