Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data Processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25

Synthetic Streamflow Datasets to Support Emulation of Water Allocations via LSTM

This archive is the data companion to the bonney_et-al_2026_erc metarepo which generates synthetic data, trains an LSTM model, and generates performance metrics on the trained model. While the generation of the synthetic data is fully reprodicible, it is a computationally expensive process. This data archive contains the synthetic datasets needed for training and testing an LSTM model and reproduction of figures and tables. In addition, supplemenatary data products generating and visualizing results is also included, such as geospatial data for the basin. Contents There are two high level directories: `WRAP_archive/` and `repo_data/`. The `WRAP_archive` directory contains compressed intermediate dataproducts from the dataset generation workflow (marked as "I_Dataset_Generation" in the metarepo). These data products are not required by any scripts in the metarepo, but they are archived as they are expensive to generate and may have useful information for other analyses. The `repo_data` directory contains the necessary data for reproducing the workflow in the metarepo and should be decompressed and moved into the top level of the metarepo. Additional details are provided in README.md.

drought↗

The critical importance of software for HEP

Particle physics has an ambitious and broad global experimental programme for the coming decades. Large investments in building new facilities are already underway or under consideration. Scaling the present processing power and data storage needs by the foreseen increase in data rates in the next decade for HL-LHC is not sustainable within the current budgets. As a result, a more efficient usage of computing resources is required in order to realise the physics potential of future experiments. Software and computing are an integral part of experimental design, trigger and data acquisition, simulation, reconstruction, and analysis, as well as related theoretical predictions. A significant investment in computing and software is therefore critical. Advances in software and computing, including artificial intelligence (AI) and machine learning (ML), will be key for solving these challenges. Making better use of new processing hardware such as graphical processing units (GPUs) or ARM chips is a growing trend. This forms part of a computing solution that makes efficient use of facilities and contributes to the reduction of the environmental footprint of HEP computing. The HEP community already provided a roadmap for software and computing for the last EPPSU, and this paper updates that, with a focus on the most resource critical parts of our data processing chain.

97 MATHEMATICS AND COMPUTING↗

Selection Algorithm Improvement for MicroBooNE

Data selection is an extremely important part of data analysis for any experiment. Finding a physics result is often the result of sifting through a massive amount of data, keeping data that we believe to be signal and throwing out data we do not. This process is called data selection. Creating a selection algorithm is an intensive process that must balance keeping enough data to have statistics and maximizing the signal purity of that data. In this study, we used three different reconstruction tools, Pandora, WireCell, and LANTERN, for the MicroBooNE experiment in conjunction to improve the selection algorithm for analysis. For the case of this study, we look into the charged current N proton 0 pions (CCNp0$\pi$) interaction channel. This is the dominant channel for the Short Baseline Neutrino (SBN) program and is expected to be a large contributor to the Deep Underground Neutrino Experiment (DUNE). We first investigated each of the three tools to find out more about their strengths and weaknesses as reconstructions. We then put together a direct comparison of the three methods to find which method or combination of methods would return the best result for us. While the study is ongoing, we have learned a lot about data selection for the experiment and the differences between the reconstruction tools.

Dillon, Brayden [Michigan State U.]↗

Data Selection Improvement For MicroBooNE

Data selection is an extremely important part of data analysis for any experiment. Finding a physics result is often the result of sifting through a massive amount of data, keeping data that we believe to be signal, and throwing out data we do not. This process is called data selection. Creating a selection algorithm is an intensive process that must balance keeping enough data to have statistics and maximizing the signal purity of that data. We also need to choose the right reconstruction method, a tool to take raw data from the detector and convert it into physics results. In this study, we used three different reconstruction tools, Pandora, WireCell, and LANTERN, for the MicroBooNE experiment in conjunction to improve the selection algorithm for analysis. For the case of this study, we look into the charged current N proton 0 pions (CCNp0$\pi$) interaction channel. This is the dominant channel for the Short Baseline Neutrino (SBN) program and is expected to be a large contributor to the Deep Underground Neutrino Experiment (DUNE). We first investigated each of the three tools to find out more about their strengths and weaknesses as reconstructions, and compared them to the truth information directly from the MicroBooNE simulation pipeline. We then put together a direct comparison of the three methods to find which method or combination of methods would return the best result for us. While the study is ongoing, we have learned a lot about data selection for the experiment and the differences between the reconstruction tools.

Dillon, Brayden [Fermilab]↗

A robust synthetic data generation framework for machine learning in high-resolution transmission electron microscopy (HRTEM)

Machine learning techniques are attractive options for developing highly-accurate analysis tools for nanomaterials characterization, including high-resolution transmission electron microscopy (HRTEM). However, successfully implementing such machine learning tools can be difficult due to the challenges in procuring sufficiently large, high-quality training datasets from experiments. In this work, we introduce Construction Zone, a Python package for rapid generation of complex nanoscale atomic structures which enables fast, systematic sampling of realistic nanomaterial structures and can be used as a random structure generator for large, diverse synthetic datasets. Using Construction Zone, we develop an end-to-end machine learning workflow for training neural network models to analyze experimental atomic resolution HRTEM images on the task of nanoparticle image segmentation purely with simulated databases. Further, we study the data curation process to understand how various aspects of the curated simulated data—including simulation fidelity, the distribution of atomic structures, and the distribution of imaging conditions—affect model performance across three benchmark experimental HRTEM image datasets. Using our workflow, we are able to achieve state-of-the-art segmentation performance on these experimental benchmarks and, further, we discuss robust strategies for consistently achieving high performance with machine learning in experimental settings using purely synthetic data. Construction Zone and its documentation are available at https://github.com/lerandc/construction_zone.

36 MATERIALS SCIENCE↗

Carbon-13 NMR spectra of lignin isolated from field grown transgenic poplar

Here we present a curated dataset of a series of 13C nuclear magnetic resonance (NMR) spectra of lignin isolated from transgenic monolignol 4-O-methyltransferase (MOMT4) engineered poplar. The transgenic poplar was collected from a 3-year field trial experiment. The poplar was Soxhlet-extracted with toluene/ethanol and the extractives-free poplar was then ball-milled in a Retsch PM100 planetary ball mill using a porcelain jar with ceramic balls at 600 rpm for 2 h. The ball-milled materials were then subjected to enzymatic hydrolysis for 48 h followed by centrifugation and washing with deionized water. The solid residue was extracted twice with 96:4 (v/v) 1,4-dioxane/water mixture at room temperature overnight. The extracts were combined, rotary evaporated, and freeze-dried to recover the lignin. The dry lignin samples were dissolved in deuterated dimethyl sulfoxide for NMR characterization. 13C experiments were performed in a Bruker Avance III HD 500 MHz NMR spectrometer operating at a frequency of 125.12 MHz for the 13C nucleus using a standard Bruker pulse sequence (zgpg) on a Prodigy platform cryoprobe. The NMR spectra were acquired under the following conditions: spectra width 229 ppm, 64k data points, 1s pulse delay, and 6k scans. All the data was processed using the Bruker’s TopSpin 3.6 software. Additional meta data is embedded in the raw spectra files.

13C NMR, lignin, poplar, field trial, MOMT4, CBI↗

Proton NMR spectra of lignin isolated from field grown transgenic poplar

Here we present a curated dataset of a series of 1H nuclear magnetic resonance (NMR) spectra of lignin isolated from transgenic monolignol 4-O-methyltransferase (MOMT4) engineered poplar. The transgenic poplar was collected from a 2-year-old rotation trees within a three-year field trial experiment. Two replicates were collected for each transgenic poplar for the 1H NMR analysis. The poplar samples were Soxhlet-extracted with toluene/ethanol to remove the extractives and the extractives-free poplar was then ball-milled in a Retsch PM100 planetary ball mill using a porcelain jar with ceramic balls at 600 rpm for 2 h. The ball-milled materials were subjected to enzymatic hydrolysis for 48 h followed by centrifugation and washing with deionized water. The solid residue was extracted twice with 96:4 (v/v) 1,4-dioxane/water mixture at room temperature overnight. The extracts were combined, rotary evaporated, and freeze-dried to recover lignin. The dry lignin samples were dissolved in deuterated dimethyl sulfoxide and transferred into a 5 mm NMR tube. 1H NMR experiments were performed in a Bruker Avance III HD 500 MHz NMR spectrometer operating at a frequency of 125.12 MHz for the 13C nucleus using a standard Bruker pulse sequence (zg) on a Prodigy platform cryoprobe. The NMR spectra were acquired with 16 ppm spectra width, 32k data points, 3s pulse delay, and 16 scans. All the data was processed using the Bruker’s TopSpin 3.6 software. Additional meta data is embedded in the raw spectra files.

1H NMR, lignin, poplar, field trial, MOMT4, CBI↗

System and method for wave prediction

A method and system for prediction of wave properties include collecting time series data streams from one or more wave measurement devices and processing the data using a wave-prediction algorithm to identify the frequency components of the data and compute wave parameters. The wave-field is propagated in space and time to predict wave height, speed, and velocity at a target location. A sliding window approach is used to continuously update the prediction in real-time.

Previsic, Mirko↗

System and method for wave prediction

A method and system for prediction of wave properties include collecting time-series data streams from one or more wave measurement devices and processing the data to identify data parameters to establish boundary conditions of a numerical model. The numerical model may be used to compute a predicted wave field of time-series data for a variety of wave properties at a target location.

Previsic, Mirko↗

Advances in the photon avalanche luminescence of inorganic lanthanide-doped nanomaterials

Photon avalanche (PA)—where the absorption of a single photon initiates a ‘chain reaction’ of additional absorption and energy transfer events within a material—is a highly nonlinear optical process that results in upconverted light emission with an exceptionally steep dependence on the illumination intensity. Over 40 years following the first demonstration of photon avalanche emission in lanthanide-doped bulk crystals, PA emission has been achieved in nanometer-scale colloidal particles. The scaling of PA to nanomaterials has resulted in significant and rapid advances, such as luminescence imaging beyond the diffraction limit of light, optical thermometry and force sensing with (sub)micron spatial resolution, and all-optical data storage and processing. In this review, we discuss the fundamental principles underpinning PA and survey the studies leading to the development of nanoscale PA. Finally, we offer a perspective on how this knowledge can be used for the development of next-generation PA nanomaterials optimized for a broad range of applications, including mid-IR imaging, luminescence thermometry, (bio)sensing, optical data processing and nanophotonics.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Enhancing Discoverability and Management of Atmospheric Data at Scale: Solutions from the ARM Data Center

The Atmospheric Radiation Measurement (ARM) is a multi-laboratory and multi-institutional U.S. Department of Energy (DOE) Office of Science National User Facility. The ARM Data Center (ADC), located at Oak Ridge National Laboratory, collects, archives, and shares vast atmospheric data crucial for climate research. The ADC manages over 7 PB of data from 460 instruments worldwide, processing it into more than 11,000 diverse data products using the Network Common Data Form (NetCDF) for machine-independent accessibility. The primary challenge addressed in this paper is the efficient management and distribution of vast and diverse datasets essential for the climate research community, enhancing accessibility through advanced tools like Data Discovery. The ADC has developed advanced infrastructure and software architecture to handle the continuous influx of heterogeneous data to enhance data discoverability, resulting in increased scientific collaboration. In 2023, users from over 34 countries downloaded and utilized ARM data, resulting in 1,455 publications. The ADC’s efforts have significantly improved the discoverability and usability of atmospheric data, fostering extensive scientific research and collaboration. This paper details the solutions implemented by the ADC team for efficient data discovery and distribution, and it demonstrates ARM’s capability of staging processed data for scientific analysis.

Shah, Chirag [ORNL] (ORCID:0000000203145737)↗

L0 Data from the 2018 NGEE Arctic LiDAR and Imagery Unoccupied Aerial System Campaign at the Teller 27 Field Site, Seward Peninsula, Alaska

Airborne remote sensing data collected from Los Alamos National Laboratory's (LANL) heavy-lift unoccupied aerial system (UAS) hexacopter platform operated by NGEE Arctic scientists from the EES-14 group at Los Alamos National Laboratory. These data were collected in July 2018 at a field site near mile marker 27 along the Teller road between Nome, Alaska and Teller, Alaska. A DJI Matrice 600 Pro Airframe and Routescene UAV LiDARSystem was used to collect LiDAR data along 12 flight paths, and DJI Phantom 4 Advanced was used to collect optical red/green/blue (RGB) imagery at regular intervals along 5 flight paths. This data package contains unprocessed data products (processing level 0) including flight paths, raw photos, and raw lidar data files (*.kml, *.jpg, and *.lpd formats). Ancillary aircraft data, flight mission parameters, and general flight conditions are also included (see Supplemental Files, *.rinex, and *.rtcm3 files). NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research.- The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska. Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

54 ENVIRONMENTAL SCIENCES↗

Electric Vehicle Charging Analytics and Reporting Tool (EV-ChART): Data Format and Preparation Guidance (V.5.0)

The Joint Office of Energy and Transportation maintains the Electric Vehicle Charging Analytics and Reporting Tool (EV-ChART), which provides a centralized hub for submitting electric vehicle (EV) charging infrastructure data directed by the Federal Highway Administration (23 CFR 680.112(a)-(c)). EV-ChART provides a streamlined data submission process and an integrated set of analytic tools, connects to other data sources, and empowers data sharing and access across stakeholders, including the public. Any data shared publicly will be aggregated and anonymized to stay in accordance with 23 CFR 680. This EV-ChART Data Format and Preparation Guidance provides a comprehensive overview of the data reporting requirements as authorized under 23 CFR 680.112(a)-(c)). The guidance is intended to be used alongside the EV-ChART Data Input Template, which defines the tabular data structure that these data submissions must follow. Per 23 CFR 680.112(a)-(c), the annual and quarterly data submissions are required of all National Electric Vehicle Infrastructure (NEVI) Formula Program projects, as well as projects for the construction of publicly accessible EV chargers that are funded with funds made available under Title 23, United States Code, including any EV charging infrastructure project funded with federal funds that is treated as a project on a federal-aid highway. One-time data submissions are required of both the NEVI Formula Program projects and grants awarded under 23 U.S.C. 151(f) for projects that are for EV charging stations located along and designed to serve the users of designated Alternative Fuel Corridors (AFCs). Other information and data required in 23 CFR 680, such as 23 CFR 680.112(d), 23 CFR 680.116(c), and 23 CFR 680.106(a), are not discussed in this guidance.

33 ADVANCED PROPULSION SYSTEMS↗

Robust error calibration for serial crystallography

Serial crystallography is an important technique with unique abilities to resolve enzymatic transition states, minimize radiation damage to sensitive metalloenzymes and perform de novo structure determination from micrometre-sized crystals. This technique requires the merging of data from thousands of crystals, making manual identification of errant crystals unfeasible. cctbx.xfel.merge uses filtering to remove problematic data. However, this process is imperfect, and data reduction must be robust to outliers. We add robustness to cctbx.xfel.merge at the step of uncertainty determination for reflection intensities. This step is a critical point for robustness because it is the first step where the data sets are considered as a whole, as opposed to individual lattices. Robustness is conferred by reformulating the error-calibration procedure to have fewer and less stringent statistical assumptions and incorporating the ability to down-weight low-quality lattices. We then apply this method to five macromolecular XFEL data sets and observe the improvements to each. The appropriateness of the intensity uncertainties is demonstrated through internal consistency. This is performed through theoretical CC 1/2 and I /σ relationships and by weighted second moments, which use Wilson's prior to connect intensity uncertainties with their expected distribution. This work presents new mathematical tools to analyze intensity statistics and demonstrates their effectiveness through the often underappreciated process of uncertainty analysis.

Mittan-Moreau, David W.↗

On the predictability of turbulent fluxes from land: PLUMBER2 MIP experimental description and preliminary results

Accurate representation of the turbulent exchange of carbon, water, and heat between the land surface and the atmosphere is critical for modelling global energy, water, and carbon cycles in both future climate projections and weather forecasts. Evaluation of models' ability to do this is performed in a wide range of simulation environments, often without explicit consideration of the degree of observational constraint or uncertainty and typically without quantification of benchmark performance expectations. We describe a Model Intercomparison Project (MIP) that attempts to resolve these shortcomings, comparing the surface turbulent heat flux predictions of around 20 different land models provided with in situ meteorological forcing evaluated with measured surface fluxes using quality-controlled data from 170 eddy-covariance-based flux tower sites. Predictions from seven out-of-sample empirical models are used to quantify the information available to land models in their forcing data and so the potential for land model performance improvement. Sites with unusual behaviour, complicated processes, poor data quality, or uncommon flux magnitude are more difficult to predict for both mechanistic and empirical models, providing a means of fairer assessment of land model performance. When examining observational uncertainty, model performance does not appear to improve in low-turbulence periods or with energy-balance-corrected flux tower data, and indeed some results raise questions about whether the energy balance correction process itself is appropriate. In all cases the results are broadly consistent, with simple out-of-sample empirical models, including linear regression, comfortably outperforming mechanistic land models. In all but two cases, latent heat flux and net ecosystem exchange of CO 2 are better predicted by land models than sensible heat flux, despite it seeming to have fewer physical controlling processes. Land models that are implemented in Earth system models also appear to perform notably better than stand-alone ecosystem (including demographic) models, at least in terms of the fluxes examined here. The approach we outline enables isolation of the locations and conditions under which model developers can know that a land model can improve, allowing information pathways and discrete parameterisations in models to be identified and targeted for future model development.

54 ENVIRONMENTAL SCIENCES↗

Machine Tool Data Analytics for Digital Twin and Machine Predictive Maintenance

The primary objective of this project is to improve machining process performance using in-process machining data from the machine tool controller and external sensors. Advances in the Industrial Internet of Things (IIoT) enable monitoring of machines using controller data. Examples of the data provided by a controller include execution status of the controller, part count, block of code being executed, door status, tool position, the spindle and axis load, etc. MTConnect and OPC-UA are the two common protocols for capturing machine information. In this collaboration, methods for retrieving the machine controller data from selected machine tool controls and making these data accessible in different subsystems (such as digital twins and machine maintenance portals, etc.) will be developed and tested. In addition, analytics to improve machining process performance (by increasing productivity and reducing downtime) will be developed.

42 ENGINEERING↗

MaPSA Quality Control and AI-Enhanced Grading For the CMS Phase-II Tracker Upgrade

The Compact Muon Solenoid (CMS) experiment will undergo changes as part of the Large Hadron Collider upgrade. The CMS tracker will be upgraded to cope with the new radiation environment and to provide tracking at the first level trigger. This upgrade features a new type of silicon module called PS Module, which combines a Pixel sensor and a Strip sensor in the same module. The pixel portion of the PS module has a sensor bump bonded to 16 Macro Pixel ASICs (MPA) to form a Macro Pixel Sub Assembly (MaPSA). At Fermilab, MaPSAs are tested for quality control before being assembled with the strip sensors, readout and service electronics to form a PS Module. All of this test data is stored in a centralized database, and is used to grade the final module to determine if it will be installed in the detector. The Phase II Outer Tracker Analyzer of Test Outputs (POTATO) is the software that processes this data and determines the module grades. Using recent technologies, an AI agent is being im plemented into POTATO in order to allow users to more efficiently sort through the large amounts of analysis data and ensure that only the user specified data is being considered. This poster will display the process of testing a MaPSA, how that test data is relevant to module assembly and grading, and how the POTATO grading tool is being improved with the use of an embedded AI agent.

Gzamouranis, Olivia [Purdue U.]↗

Machine learning models of intermittent operation of RO wellhead water treatment for salinity reduction and nitrate removal

Machine learning models were developed for intermittent multi-mode operation of a wellhead reverse osmosis water purification and desalination system to predict salt passage, nitrate passage, and permeate flux. The models, based on long short-term memory (LSTM) recurrent neural network (RNN) architecture, included an attention mechanism to increase model performance in proximity of the regulatory limit for nitrate. Training and testing of the models for the Startup, Production, Shutdown and Flushing operational modes were based on operational data (consisting of 22 process variables per data sample) acquired every 2–5 s over a six-month period. The significant sets of model input attributes for the different operational modes were assessed via Spearman ranking correlation, Self-Organizing Map (SOM) analysis and feed forward feature selection (FFFS). Although the variability of nitrate passage, salt passage and permeate flux was significant over the four operational modes, prediction performance for the three outcomes were with R2 and Average Absolute Relative Error (AARE) of 0.78–0.95 and 2.96–6.16 %, respectively. Model updates post membrane elements replacement demonstrated similar levels of prediction accuracy. The study results suggest that there is merit in exploring the utility of multi-mode models for sensor fault detection, data imputation, and for potential use in model-predictive control.

Intermittent RO operation↗