Search NASA⌕ Search

SEARCH · Search NASA

Results for “data processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21

WHONDRS 2016 Sediment Organic Matter Characterization Data from Streams across HJ Andrews Experimental Forest, Oregon

This dataset supports a broader synoptic effort to map morphological, hydrological, chemical, and biological conditions across a fifth-order mountain stream network. Samples were generated through a collaborative synoptic sampling effort in 2016. The dataset provides sediment Fourier Transform Ion Cyclotron Resonance Mass Spectrometry (FTICR-MS) from 60 sites across the HJ Andrews Experimental Forest, Oregon (https://andrewsforest.oregonstate.edu). Related data were collected as part of the event and were published separately in collaboration with other team members. The data are available at http://www.hydroshare.org/resource/ea6c0832885a46c3939e7bb22e48e754 and are described within https://doi.org/10.5194/essd-11-1567-2019 (Ward et al., 2019). The hydroshare data package contains processed FTICR-MS data from the samples included in this data package. The data were processed via Formultitude (previously called Formularity; https://github.com/PNNL-Comp-Mass-Spec/Formultitude). However, we have re-processed the data using Core-MS and included it in this data package. Additional related data collected in 2025 from a similar effort can be found at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/3023310 and http://www.hydroshare.org/resource/b274c4a234bf4b12b7cb8a54a696c629. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. In addition to a readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions. This dataset is comprised of (1) a folder of sample data; (2) data dictionary; (3) file-level metadata; (4); (5) coordinates; and (6) readme. The sample data subfolder contains 12 Tesla (12T) FTICR-MS data. This folder contains the processed data and three subfolders, one containing the .xml files, one containing the CoreMS output files, and the other containing instructions and scripts for processing the files in CoreMS (https://github.com/EMSL-Computing/CoreMS). All files are .csv, .pdf, .R, .xml, .Rmd, .py, .cal, or .json.

Biogeochemistry↗

EDXplorer: A Utility for APA Analysis

The automated particle analysis (APA) method of scanning electron microscopy (SEM) energy dispersive X-ray spectroscopy (EDS/EDX) is a useful tool for analyzing the elemental and morphological data of particulate samples. Often, such datasets have many thousands of particles, and it can be difficult to sift through the data to find meaningful trends. EDXplorer is a software utility for processing APA data. They enable the user to easily load data and determine the important components and aspects of the datasets using a powerful and versatile library of data plotting functions, mainly centered around scatter plots and histograms. These programs are intended to fill a void in the data processing of data from certain instruments, where often the user must rely on their own code or other software that is not user-friendly. EDXplorer is intended as a general plotting utility for browsing through data and discovering data trends.

Moseley, Duncan [ORNL] (ORCID:0000000343518347)↗

FY 2025 Multidimensional Data Correlation Platform: Unified Software Architecture for Advanced Materials and Manufacturing Technologies Data Management and Processing

The Advanced Materials and Manufacturing Technologies (AMMT) program continues to advance a data-driven approach to demonstrate the utility of additive manufacturing for fabricating components for nuclear applications. A key scientific goal is to leverage data to better understand manufacturing outcomes and thereby improve the performance, reliability, and lifespan of nuclear components. Ultimately, this effort supports the development of standards for certification and qualification of additively manufactured components, enabling broader industry adoption. In support of this objective, the AMMT program is building and deploying a data management platform to record, index, analyze, and make available the manufacturing data generated across the AMMT program. In FY 2023, the team conceptualized the architecture of the platform and, in FY 2024, deployed the first functional version at the Oak Ridge National Laboratory (ORNL) Manufacturing Demonstration Facility (MDF). In FY 2025, the platform was officially opened to all AMMT members. To enable this expansion, core modifications and enhancements were developed, including improvements to the user interface and workflows for data entry and retrieval. Most notably, robust security and access control mechanisms were implemented to protect data and manage information sharing. This effort featured a logging system, protected views, and controlled access mechanisms. This report documents these enhancements and the transition of the platform into program-wide use.

36 MATERIALS SCIENCE↗

Towards Online Machine Learning in DUNE Data Acquisition

Processing the large volumes of data produced by liquid argon time projection chamber (LArTPC) experiments presents a significant challenge, especially those at the scale of DUNE. This is a particular challenge when aiming to trigger on low-energy neutrinos from core-collapse supernovae, which are typically buried in a high-rate radiological background. To enable real-time event selection suitable for such rare signals, we are developing machine learning based data filtering methods. In order to demonstrate the feasibility of this approach, we implemented such pipeline using the ICEBERG detector at Fermilab as a small-scale LArTPC, with a focus on identifying Michel electrons as a proxy for low-energy neutrino interactions. This poster will present the current status of integrating these machine learning models into the data acquisition (DAQ) system of this detector.

Dalager, Olivia [Fermilab]↗

Analysis of Rig Parameter Data Using Drilling Process Modeling Constraints, Volume 5: Utah FORGE Well 16B(78)-32

Drill rig parameter measurements are routinely used during deep well construction to monitor and guide drilling conditions for improved performance and reduced costs. While insightful into the drilling process, these measurements are of reduced value without a standard to aid in data evaluation and decision making. In the main body of this work (Volume 1), a method is demonstrated whereby rock reduction model constraints are used to interpret drilling response parameters; the method could be applied in real-time to improve decision-making in the field and to further discern technology performance during post-drilling evaluations. Drilling parameters are evaluated using laboratory-validated rock reduction models for predicting the phenomenological response of drag bits (Detournay and Defourny, 1992) in computational algorithms. The method presented has applicability to development of advanced analytics on future geothermal wells using real-time electronic data recording for improved performance and reduced drilling costs. A drilling cost model is also used to show the tradeoff between rate of penetration and bit life and the influence on interval drilling costs. Details of the bit specifications and performance are cataloged in an independent volume, documented under separate cover, for each of the four wells, and include Volume 2: Utah FORGE 16A(78)-32; Volume 3: Utah FORGE 56-32; Volume 4: Utah FORGE 78B-32 and Volume 5: Utah FORGE 16B(78)-32.

15 GEOTHERMAL ENERGY↗

Analysis of Rig Parameter Data Using Drilling Process Modeling Constraints, Volume 4: Utah FORGE Well 78B-32

Drill rig parameter measurements are routinely used during deep well construction to monitor and guide drilling conditions for improved performance and reduced costs. While insightful into the drilling process, these measurements are of reduced value without a standard to aid in data evaluation and decision making. In the main body of this work (Volume 1), a method is demonstrated whereby rock reduction model constraints are used to interpret drilling response parameters; the method could be applied in real-time to improve decision-making in the field and to further discern technology performance during post-drilling evaluations. Drilling parameters are evaluated using laboratory-validated rock reduction models for predicting the phenomenological response of drag bits (Detournay and Defourny, 1992) in computational algorithms. The method presented has applicability to development of advanced analytics on future geothermal wells using real-time electronic data recording for improved performance and reduced drilling costs. A drilling cost model is also used to show the tradeoff between rate of penetration and bit life and the influence on interval drilling costs. Details of the bit specifications and performance are cataloged in an independent volume, documented under separate cover, for each of the four wells, and include Volume 2: Utah FORGE 16A(78)-32; Volume 3: Utah FORGE 56-32; Volume 4: Utah FORGE 78B-32 and Volume 5: Utah FORGE 16B(78)-32.

15 GEOTHERMAL ENERGY↗

Analysis of Rig Parameter Data Using Drilling Process Modeling Constraints, Volume 3: Utah FORGE Well 56-32

Drill rig parameter measurements are routinely used during deep well construction to monitor and guide drilling conditions for improved performance and reduced costs. While insightful into the drilling process, these measurements are of reduced value without a standard to aid in data evaluation and decision making. In the main body of this work (Volume 1), a method is demonstrated whereby rock reduction model constraints are used to interpret drilling response parameters; the method could be applied in real-time to improve decision-making in the field and to further discern technology performance during post-drilling evaluations. Drilling parameters are evaluated using laboratory-validated rock reduction models for predicting the phenomenological response of drag bits (Detournay and Defourny, 1992) in computational algorithms. The method presented has applicability to development of advanced analytics on future geothermal wells using real-time electronic data recording for improved performance and reduced drilling costs. A drilling cost model is also used to show the tradeoff between rate of penetration and bit life and the influence on interval drilling costs. Details of the bit specifications and performance are cataloged in an independent volume, documented under separate cover, for each of the four wells, and include Volume 2: Utah FORGE 16A(78)-32; Volume 3: Utah FORGE 56-32; Volume 4: Utah FORGE 78B-32 and Volume 5: Utah FORGE 16B(78)-32.

15 GEOTHERMAL ENERGY↗

Analysis of Rig Parameter Data Using Drilling Process Modeling Constraints, Volume 1: Summary of Utah FORGE Wells 16A(78)-32, 56-32, 78B-32 and 16B(78)-32

Drill rig parameter measurements are routinely used during deep well construction to monitor and guide drilling conditions for improved performance and reduced costs. While insightful into the drilling process, these measurements are of reduced value without a standard to aid in data evaluation and decision making. In the main body of this work (Volume 1), a method is demonstrated whereby rock reduction model constraints are used to interpret drilling response parameters; the method could be applied in real-time to improve decision-making in the field and to further discern technology performance during post-drilling evaluations. Drilling parameters are evaluated using laboratory-validated rock reduction models for predicting the phenomenological response of drag bits (Detournay and Defourny, 1992) in computational algorithms. The method presented has applicability to development of advanced analytics on future geothermal wells using real-time electronic data recording for improved performance and reduced drilling costs. A drilling cost model is also used to show the tradeoff between rate of penetration and bit life and the influence on interval drilling costs. Details of the bit specifications and performance are cataloged in an independent volume, documented under separate cover, for each of the four wells, and include Volume 2: Utah FORGE 16A(78)-32; Volume 3: Utah FORGE 56-32; Volume 4: Utah FORGE 78B-32 and Volume 5: Utah FORGE 16B(78)-32.

15 GEOTHERMAL ENERGY↗

Codiscovering graphical structure and functional relationships within data: A Gaussian Process framework for connecting the dots

Most problems within and beyond the scientific domain can be framed into one of the following three levels of complexity of function approximation. Type 1: Approximate an unknown function given input/output data. Type 2: Consider a collection of variables and functions, some of which are unknown, indexed by the nodes and hyperedges of a hypergraph (a generalized graph where edges can connect more than two vertices). Given partial observations of the variables of the hypergraph (satisfying the functional dependencies imposed by its structure), approximate all the unobserved variables and unknown functions. Type 3: Expanding on Type 2, if the hypergraph structure itself is unknown, use partial observations of the variables of the hypergraph to discover its structure and approximate its unknown functions. These hypergraphs offer a natural platform for organizing, communicating, and processing computational knowledge. While most scientific problems can be framed as the data-driven discovery of unknown functions in a computational hypergraph whose structure is known (Type 2), many require the data-driven discovery of the structure (connectivity) of the hypergraph itself (Type 3). We introduce an interpretable Gaussian Process (GP) framework for such (Type 3) problems that does not require randomization of the data, access to or control over its sampling, or sparsity of the unknown functions in a known or learned basis. Its polynomial complexity, which contrasts sharply with the super-exponential complexity of causal inference methods, is enabled by the nonlinear ANOVA capabilities of GPs used as a sensing mechanism.

Science & Technology - Other Topics↗

Scalable Hybrid Learning Techniques for Scientific Data Compression

Data compression is becoming critical for storing scientific data because many scientific applications need to store large amounts of data and post process this data for scientific discovery. Unlike image and video compression algorithms that limit errors to primary data (PD), scientists require compression techniques that accurately preserve derived quantities of interest (QoIs). Here, this article presents a physics-informed compression technique implemented as an end-to-end, scalable, GPU-based pipeline for data compression that addresses this requirement. Our hybrid compression technique combines machine learning techniques and standard compression methods. Specifically, we combine an autoencoder, an error-bounded lossy compressor to provide guarantees on raw data error, and a constraint satisfaction post-processing step to preserve the QoIs within a minimal error (generally less than floating point error). The effectiveness of the data compression pipeline is demonstrated by compressing nuclear fusion simulation data generated by a large-scale fusion code, XGC, which produces hundreds of terabytes of data in a single day. Our approach works within the ADIOS framework and results in compression by a factor of more than 150 while requiring only a few percent of the computational resources necessary for generating the data, making the overall approach highly effective for practical scenarios.

ITER↗

Automated analysis of unlabeled PV data with Solar Data Tools software: Overview and feature updates

Distributed rooftop PV systems: ubiquitous, yet commonly have unlabeled data Difficult or impossible to form a performance index We developed Solar Data Tools (SDT), an open-source Python library for analyzing PV power (and irradiance) time-series data SDT enables analysis of unlabeled PV data—no model, no meteorological data, no performance index required Takes a statistical signal processing approach Data processing steps are largely pre-defined and automatic regardless of system type—from utility tracking systems to multi-pitch rooftop systems

Meyers-Im, Bennet E↗

Multitaper Magnitude‐Squared Coherence for Time Series With Missing Data: Understanding Oscillatory Processes Traced by Multiple Observables

To explore the hypothesis of a common source of variability in two time series, observers may estimate the magnitude-squared coherence (MSC), which is a frequency-domain view of the cross correlation. For time series that do not have uniform observing cadence, MSC can be estimated using Welch's overlapping segment averaging. However, multitaper has superior statistical properties to Welch's method in terms of the tradeoff between bias, variance, and bandwidth. The classical multitaper technique has recently been extended to accommodate time series with underlying uniform observing cadence from which some observations are missing. This situation is common for solar and geomagnetic data sets, which may have gaps due to breaks in satellite coverage, instrument downtime, or poor observing conditions. We demonstrate the scientific use of missing-data multitaper magnitude-squared coherence by detecting known solar mid-term oscillations in simultaneous, missing-data time series of solar Lyman α flux and geomagnetic Disturbance Storm Time index. Due to their superior statistical properties, we recommend that multitaper methods be used for all heliospheric time series with underlying uniform observing cadence.

Astro-statistics techniques (1886)↗

O'Hare Airport roadway traffic prediction via data fusion and Gaussian process regression

This study proposes an approach of leveraging information gathered from multiple traffic data sources at different resolutions to obtain approximate inference on the traffic distribution of Chicago's O'Hare Airport area. Specifically, it proposes the ingestion of traffic datasets at different resolutions to build spatiotemporal models for predicting the distribution of traffic volume on the road network. Due to its good adaptability and flexibility for spatiotemporal data, the Gaussian process (GP) regression was employed to provide short-term forecasts using data collected by loop detectors (sensors) and supplemented by telematics data. The GP regression is used to make predictions of the distribution of the proportion of sensor data traffic volume represented by the telematics data for each location of the sensors. Consequently, the fitted GP model can be used to determine the approximate traffic distribution for a testing location outside of the training points. Policymakers in the transportation sector can find the results of this work helpful for making informed decisions relating to current and future transportation conditions in the area.

42 ENGINEERING↗

ggtaxplot v 0.0.1

ggtaxplot is an R package designed to process and visualize taxonomic data through a taxonomic river plot. This package is ideal for researchers and data scientists who need to visualize taxonomic data. ggtaxplot function processes data and generates a taxonomic river plot, allowing users to visualize the distribution of taxa across different samples.

Coclet, Clement [Lawrence Berkeley National Labora↗

On Road Testing Data

This dataset provides the following on road testing data: - Videos - In-vehicle dash camera videos during different testing scenarios. - Signal controller data - NTCIP log data and processed signal timing data from the corresponding signal controllers - Vehicle data - Vehicle data recorded during the testing, including GNSS, communication, CAN signals.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗