Search NASA⌕ Search

SEARCH · Search NASA

Results for “data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23

Deep learning-driven super-resolution in Raman hyperspectral imaging: Efficient high-resolution reconstruction from low-resolution data

Deep learning (DL) has become an indispensable tool in hyperspectral data analysis, automatically extracting valuable features from complex, high-dimensional datasets. Super-resolution reconstruction, an essential aspect of hyperspectral data, involves enhancing spatial resolution, particularly relevant to low-resolution hyperspectral data. Yet, the pursuit of super-resolution in hyperspectral analysis is fraught with challenges, including acquiring ground truth high-resolution data for training, generalization, and scalability. The pressing issue of extended spectral acquisition times, notably for high-resolution scans, is a significant roadblock in hyperspectral imaging. Super-resolution methods offer a promising solution by providing higher spatial resolution data to expedite data collection and yield more efficient outcomes. This paper delves into a practical application of these concepts using Raman imaging, where spectral acquisition times can be prohibitively long. In this context, DL-based super-resolution models demonstrate their efficacy by predicting and reconstructing high-resolution Raman data from low-resolution input, eliminating the need for resource-intensive high-resolution scans. While previous work often relied on substantial high-resolution datasets, this study showcases the ability to achieve similar outcomes even with limited data, presenting a more practical and cost-effective approach. In conclusion, the results offer a glimpse into the transformative potential of this technology to streamline hyperspectral imaging applications by saving valuable time and resources through the successful generation of high-resolution data from low-resolution inputs.

42 ENGINEERING↗

Streaming Large-Scale Microscopy Data to a Supercomputing Facility

Data management is a critical component of modern experimental workflows. As data generation rates increase, transferring data from acquisition servers to processing servers via conventional file-based methods is becoming increasingly impractical. The 4D Camera at the National Center for Electron Microscopy generates data at a nominal rate of 480 Gbit s -1 (87,000 frames s -1 ⁠), producing a 700 GB dataset in 15 s. To address the challenges associated with storing and processing such quantities of data, we developed a streaming workflow that utilizes a high-speed network to connect the 4D Camera’s data acquisition system to supercomputing nodes at the National Energy Research Scientific Computing Center, bypassing intermediate file storage entirely. In this work, we demonstrate the effectiveness of our streaming pipeline in a production setting through an hour-long experiment that generated over 10 TB of raw data, yielding high-quality datasets suitable for advanced analyses. Additionally, we compare the efficacy of this streaming workflow against the conventional file-transfer workflow by conducting a postmortem analysis on historical data from experiments performed by real users. Our findings show that the streaming workflow significantly improves data turnaround time, enables real-time decision-making, and minimizes the potential for human error by eliminating manual user interactions.

4D-STEM↗

Machine Learning to Select Experiments Driven by Fundamental Science and Applications for Targeted Nuclear Data Improvement

This work describes a blueprint for a process that accelerates progress in science by quantitatively answering the following question: What is the optimal combination of fundamental-science and application-driven experiments to maximally reduce pertinent data uncertainties? Answering this question entails solving a high-dimensional and complex optimization problem that is best solved with advanced statistic techniques often classified as machine learning. We apply this process within the framework of nuclear data with the aim to select an experiment combination that will reduce uncertainties in 239 Pu nuclear data for neutron energies between 1 and 600 keV. In this field, fundamental-physics driven data, called differential, look at one nuclear physics observable at a time. They are contrasted to application-driven, integral, data where one or few resulting values inform a broad set of nuclear data across several nuclides and energies. The candidates for integral experiments are criticality measurements that were refined by a genetic algorithm to be maximally sensitive to 239 Pu fission cross sections in the desired energy range. Twenty-three candidate differential experiments were investigated and span multiple nuclear physics observables (e.g., total, capture cross sections) for isotopes appearing in the integral experiments. The optimal combination among these candidate experiments was investigated via generalized least squares fitting, augmented with Gaussian processes to ameliorate statistical irregularities in data, and the D-optimality criterion. The latter evaluates for each pair of candidates the joint reduction in uncertainties of all 12200 nuclear data appearing in the integral experiments compared to the knowledge we have from 168 past experiments, theory, and nuclear data. We chose as differential measurements those that investigate 63 Cu and 239 Pu total cross sections, based on D-optimality rank and feasibility constraints. Two integral (criticality) experiments were selected: An experiment with Al 2 ⁢O 3 and graphite interleaved with Pu and a thick Cu reflector explores 1–30 keV, while we target the 30–600 keV range with an experiment that swaps boron in place of graphite with a different geometry.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

A machine-learning-driven data labeling pipeline for scientific analysis in MLExchange

This study introduces a novel labeling pipeline to accelerate the labeling process of scientific data sets by using artificial intelligence (AI)-guided tagging techniques. This pipeline includes a set of interconnected web-based graphical user interfaces (GUIs), where Data Clinic and MLCoach enable the preparation of machine learning (ML) models for data reduction and classification, respectively, while Label Maker is used for label assignment. Throughout this pipeline, data can be accessed through a direct connection to a file system or through Tiled for access through Hypertext Transfer Protocol (HTTP). Our experimental results present three use cases where this labeling pipeline has been instrumental for the study of large X-ray scattering data sets in the area of pattern recognition, the remote analysis of resonant soft X-ray scattering data and the fine-tuning process of foundation models. These use cases highlight the labeling capabilities of this pipeline, including the ability to label large data sets in a short period of time, to perform remote data analysis while minimizing data movement and to enhance the fine-tuning process of complex ML models with human involvement.

Chavez, Tanny (ORCID:0000000193172896)↗

An Open-Access Repository of Synchrophasor Data Quality Examples: Curation and Example Applications

Synchrophasor measurements are critical in providing wide-area situational awareness to power system operators. However, data artifacts may be introduced due to various issues such as loss of communication, loss of GPS signal, internal clock error, and vendor-specific implementation of phasor estimation algorithms. Tools designed to provide actionable insights from synchrophasor data, hence, must be designed to be robust to these data quality issues. In this work, two years of synchrophasor data sourced from multiple electric utilities in the United States were analyzed to identify examples of data quality problems. These examples were then labeled and published in the Grid Event Signature Library, a publicly available repository of power system measurements hosted by the Oak Ridge National Laboratory. This paper describes the data curation process, and illustrates two application use cases where the dataset can be valuable to the research community. In the first use case, a random forest classifier is trained to distinguish power system disturbance signatures from data anomalies introduced in synchrophasor measurements due to clock errors. The second use case studies the impact of data quality issues on an example synchrophasor application (specifically, event start time determination). The choice of data quality problems investigated is informed by the examples in the repository curated in this work.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Pulse: An Outlier Sensitive Downsampling Algorithm For Timeseries Data

Pulse is a downsampling algorithm for timeseries data. Frequently datasets become so large that visualization tools and web browsers cannot effectively render graphics due to memory constraints. Downsampling algorithms are commonly applied to minimize the quantity of data required to visualize important features or trends in the data, but some datasets are composed by distinct enough features and trends that most existing downsampling algorithms fail to preserve them. Pule was developed to downsample timeseries data for galvanostatic stack test data at the Idaho National Laboratory. These datasets were composed by approximately 4 million records, most of them being extremely uniform. However, during relatively brief time periods when the stack test changes state, for example when the test article is powered on, or a load is added, the data produce sparse asymptotes. No existing downsampling algorithm was capable of preserving the sparse asymptotes in electrolysis stack test data. Instead, we develop a downsampling algorithm that preserves important outliers in data, and otherwise aggressively downsamples uniform data. The algorithm has applications in other domains like seismology, in the measurement of earthquakes, or astronomy, in the measurement of quasars or transit photometry.

Woodruff, Nathan [Idaho National Laboratory (INL),↗

Deep Design Data Portal (D3P) v0.01

The Deep Design Data Portal (D3P) tool was developed to demonstrate how readily accessible data sources, such as building energy model reports for design and baseline energy performance data for projects, can provide the data required for reporting to an industry initiative (AIA 2030 commitment), as well as more detailed data that makes the industry dataset more valuable to all stakeholders, enabling project level analysis and analysis of BEM industry trends. D3P provides an easier and less time-consuming way for firms to auto-extract data from this data source, compared to the current reporting workflows of the firms. The BEM reports are the first of several data sources that D3P could integrate. D3P also provides the ability for firms to review, compare, and evaluate the performance of their projects to not only their portfolio, but also to the larger anonymized industry dataset created each time a project is added to D3P. The intent of D3P is to become part of a data-sharing ecosystem to assist creating large anonymized industry datasets that are accessible to industry.

Regnier, Cynthia [Lawrence Berkeley National Labor↗

Data for “Tree root nutrient uptake kinetics vary with nutrient availability, environmental conditions, and root traits: A global analysis”

This data package contains data and code used in the paper “Tree root nutrient uptake kinetics vary with nutrient availability, environmental conditions, and root traits: A global analysis”. The central product is a global dataset of root inorganic nutrient uptake rates and kinetics parameters covering temperate, boreal, and sub/tropical tree species, representing a collection of nutrient uptake data from published studies. This dataset enables tree investigation of root nutrient uptake rates across species, space, and experimental conditions. The data can also be combined with supplementary data on root and soil traits or with external datasets (e.g. R scripts contained within use data from FRED 3.0; (Iversen et al., 2021)). Contained within is the main nutrient data “uptake_data.csv” as well as 4 additional .csv files that link uptake data to supplementary measurements, source references, taxonomic information, and additional nutrient uptake measurements across nutrient gradients, and 1 .csv file that records meta-analysis results for plotting with the R scripts. There are seven R scripts that support data analysis and creation of the figures in the related publication.

54 ENVIRONMENTAL SCIENCES↗

A data integration framework of additive manufacturing based on FAIR principles

Abstract Laser-powder bed fusion (L-PBF) is a popular additive manufacturing (AM) process with rich data sets coming from both in situ and ex situ sources. Data derived from multiple measurement modalities in an AM process capture unique features but often have different encoding methods; the challenge of data registration is not directly intuitive. In this work, we address the challenge of data registration between multiple modalities. Large data spaces must be organized in a machine-compatible method to maximize scientific output. FAIR (findable, accessible, interoperable, and reusable) principles are required to overcome challenges associated with data at various scales. FAIRified data enables a standardized format allowing for opportunities to generate automated extraction methods and scalability. We establish a framework that captures and integrates data from a L-PBF study such as radiography and high-speed camera video, linking these data sets cohesively allowing for future exploration. Graphical abstract

36 MATERIALS SCIENCE↗

CLAS12 remote data-stream processing using ERSAP framework

Implementing a physics data processing application is relatively straightforward with the use of current containerization technologies and container image runtime services, which are prevalent in most high-performance computing (HPC) environments. However, the process is complicated by the challenges associated with data provisioning and migration, impacting the ease of workflow migration and deployment. Transitioning from traditional file-based batch processing to data-stream processing workflows is suggested as a method to streamline these workflows. This transition not only simplifies file provisioning and migration but also significantly reduces the necessity for extensive disk space. Data-stream processing is particularly effective for real-time processing during data acquisition, thereby enhancing data quality assurance. This paper introduces the integration of the JLAB CLAS12 event reconstruction application within the ERSAP data-stream processing framework that facilitates the execution of streaming event reconstruction at a remote data center and enables the return streaming of reconstructed events to JLAB while circumventing the need for temporary data storage throughout the process.

Gyurjyan, Vardan↗

MAPSTER: Automated Geospatial Data Sharing – Version 1.4.0

The US Department of Energy’s (DOE) Oak Ridge National Laboratory (ORNL) developed MAPSTER which is a geospatial data management tool that aggregates, organizes, and shares data from dispersed sources such as unmanned aerial systems (UAS). Built specifically for use in environments where communications may be limited, MAPSTER utilizes two key technologies to effectively manage data in the field and enable easy data sharing with authorized partners: Observer and Checkpoint. Observer is a lightweight software package on an edge device, such as a laptop, that automatically detects newly processed UAS data and sends to a central server called Checkpoint. Checkpoint is a centralized server at ORNL that receives and manages data from all Observer instances. Even in a very low bandwidth environment, Observer can still send information about the UAS data product almost instantly as it generates its own metadata package on the size of KB (kilobytes). MAPSTER is not only for UAS data but for any geospatial data collected at the austere edge and dispersed sources.

97 MATHEMATICS AND COMPUTING↗

Exploration of signal processing methods for superconducting magnet and quench data

Quenching is the phenomenon of a superconducting magnetic material carrying current transitioning into a regular conducting material. This may cause severe and irreparable damage to the superconductor due to Joule heating. The Magnet Department at Fermi National Accelerator Laboratory (FNAL) has acquired experimental data through quench antenna arrays that are recorded when the quench is detected. These data are in terms of voltage signals that are sampled at 100kHz for several minutes. There are multiple channels and each channel provides a data set of more than 20 million observations, while there is one channel, called the trigger channel which shows the time when quench is detected. Despite some advancements that were made including machine learning, data complexity still shadows the progress. In this work, we studied a multi-resolution analysis of the quench antenna data through the Haar wavelet transform. In particular, we applied the maximally overlapped discrete w avelet transform (MODWT) of a suitable level L to the given data and then projected it onto the wavelet basis. This decomposes a given signal (Original data) $x ϵ \mathbb{R}^N$ into $L + 1$ subspaces of $\mathbb{R}^N$. One of the subspaces called the approximation, captures the trend of the signal, and the others, called the details, capture the fluctuations at different frequency bands. This decomposition provides a clear trend of the data at a suitable level and also various activities (spikes) are seen in the details of the decomposition at every level. These spikes might reveal some information about the quench under investigation but in any case, give information about magnet behavior. Also, this decomposition is seen to be very useful in removing noise present in the data due to the source or mechanism of the experiment.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Qualification of Digitized Legacy Fast Reactor Data

The Integral Fast Reactor (IFR) fuel compatibility test program (1984-1994) included a variety of fuel pin examinations conducted at the Hot Fuel Examination Facility (HFEF) and the Alpha-Gamma Hot Cell Facility (AGHCF). Hard copy data records of these examinations have been recovered, scanned, and preserved in PDF format. Many hard copy records are now qualified in accordance with an NRC-approved Quality Assurance Program Plan (QAPP), and there is an ongoing effort to qualify additional legacy records. This legacy fuel performance data is vital to support design and licensing of fast reactors with validation of state-of-the-art codes and advanced methods for design and analysis. Stakeholders can most easily utilize this data when the PDF scans have been converted into digital data tables. However, qualification of the scanned hard copy data does not qualify the digital data file resulting from the digitization of the data contained in the record; the subject matter expert (SME) must make a review of the digitized data table as well before it can be designated as qualified. This report outlines a peer review process to qualify the digital data file(s), typically in CSV format, corresponding to hard copy records in accordance with the existing QAPP.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

In Situ Data Analysis Through Physics-informed Tensor Decompositions (LDRD Final Report)

We introduce a new low-dimensional model of high-dimensional numerical simulation data based on low-rank tensor decompositions. Our new model aims to minimize differences between the model data and simulation data as well as functions of the model data and functions of the simulation data. This novel approach to dimensionality reduction of simulation data provides a means of directly incorporating quantities of interests and invariants associated with conservation principles associated with the simulation data into the low-dimensional model, thus enabling more accurate analysis of the simulation without requiring access to the full set of high-dimensional data. Computational results of applying this approach to two standard low-rank tensor decompositions of data arising from simulation of combustion and plasma physics are presented.

97 MATHEMATICS AND COMPUTING↗

Quality Assurance Program Plan for SFR Metallic Fuel Data Qualification

This document contains an evaluation of the applicability of the current Quality Assurance Standards from the American Society of Mechanical Engineers Standard NQA-1 (NQA-1) criteria and identifies and describes the quality assurance process(es) by which attributes of historical, analytical, and other data associated with sodium-cooled fast reactor [SFR] metallic fuel will be evaluated. This process is being instituted to facilitate validation of data to the extent that such data may be used to support future licensing efforts associated with advanced reactor designs. The initial data to be evaluated under this program were generated during the US Integral Fast Reactor program between 1984-1994, where the data include, but are not limited to, research and development data and associated documents, test plans and associated protocols, operations and test data, technical reports, and information associated with past United States Nuclear Regulatory Commission reviews of SFR designs. It is recognized that managing the data generated by large research and development projects presents a significant challenge for retaining data integrity and availability. American Society of Mechanical Engineers Standard NQA-1 (NQA-1) 2008/2009a provides appropriate requirements for this plan.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Electromagnetic Transient Modeling of Large Data Centers for Grid-Level Studies

The magnitude and complexity of electricity usage patterns from large data centers are having significant impacts on the operation and dynamics of the power grid; grid operators and planners require a range of specialized data center models to properly evaluate these impacts and specify technical solutions as needed. Towards addressing this need, Pacific Northwest National Laboratory (PNNL) has developed a library of electromagnetic transient (EMT) models for grid-level studies of data centers called the data center model library (DML). This report describes how the DML was created and how it may properly be used. The models present in the DML are generic models; subject matter expertise and additional technical data are needed to modify these models before they can represent any real data center. However, they will significantly reduce the level of effort required to develop site-specific models and can serve as a common starting point to guide industry towards a more refined consensus. Most of the models within DML are dedicated to representing the power electronics interfaces commonly used in modern data centers, such as double-conversion uninterruptible power supplies and single-phase power factor correction converters. These models are intended for use in grid-level studies and are a simplified aggregation of many small components. That said, background material on the physical and electrical design of large data centers is provided as companion material so that users can be aware of many of the details which have been omitted or streamlined as a matter of practical necessity. Additionally, guidance on the application of EMT analysis for data center interconnection studies is provided, which aids users in identifying when the DML is necessary and what sort of additional model development may be necessary for conducting real-world studies.

24 POWER TRANSMISSION AND DISTRIBUTION↗

VA Community Determinants of Health Data Curation Documentation FY26-Q1

The U.S. Department of Veterans Affairs (VA) places the health and well-being of our nation’s veterans as its top priority. VA is dedicated to offering timely access to high-quality, evidence-based mental health care that meets the needs of veterans and supports their reintegration into society. One of our core missions is to prevent suicide among veterans through innovative approaches and resources. With funding from the VA Office of Mental Health and Suicide Prevention (OMHSP), the Community Determinants of Health (EDH) Data project has developed innovative datasets associated with specific health outcomes, a methodology for transforming spatiotemporal data from one spatial reference (e.g., a 1km grid) to another (e.g., US Census Tracts), and capabilities for modeling health outcomes. These datasets represent an enhancement of the Agency for Healthcare Research and Quality (AHRQ), addressing key gaps by introducing finer spatial resolution (Census Tract) and additional geographical covariates into existing data. The curation and standardization of these datasets is a complex task since they often originate from various sources and are measured at different spatial and temporal resolutions. For example, US Census data products typically use census blocks, block groups, or counties, while data like weather data are available on 1km grids. Some economic data may only be available at the zip code level. In this context, ‘standardized’ means that all datasets share the same spatial extent (e.g., US Census Tract and/or County), and ‘curated’ implies a repeatable process with data provenance and the use of appropriate methodologies for covariate conversion. The Community Determinants of Health datasets draw from multiple sources, resulting in variables with varying degrees of availability, patterns of missing data, and methodological considerations across different sources, geographies, and years.

99 GENERAL AND MISCELLANEOUS↗

VA Community Determinants of Health Data Curation Documentation FY26-Q2

The U.S. Department of Veterans Affairs (VA) places the health and well-being of our nation’s veterans as its top priority. VA is dedicated to offering timely access to high-quality, evidence-based mental health care that meets the needs of veterans and supports their reintegration into society. One of our core missions is to prevent suicide among veterans through innovative approaches and resources. With funding from the VA Office of Mental Health and Suicide Prevention (OMHSP), the Community Determinants of Health (EDH) Data project has developed innovative datasets associated with specific health outcomes, a methodology for transforming spatiotemporal data from one spatial reference (e.g., a 1 km grid) to another (e.g., U.S. Census Tracts), and capabilities for modeling health outcomes. These datasets represent an enhancement of the Agency for Healthcare Research and Quality (AHRQ), addressing key gaps by introducing finer spatial resolution (Census Tract) and additional geographical covariates into existing data. The curation and standardization of these datasets is a complex task since they often originate from various sources and are measured at different spatial and temporal resolutions. For example, U.S. Census data products typically use census blocks, block groups, or counties, while data like weather data are available on 1 km grids. Some economic data may only be available at the ZIP code level. In this context, ‘standardized’ means that all datasets share the same spatial extent (e.g., U.S. Census Tract and/or County), and ‘curated’ implies a repeatable process with data provenance and the use of appropriate methodologies for covariate conversion. The Community Determinants of Health datasets draw from multiple sources, resulting in variables with varying degrees of availability, patterns of missing data, and methodological considerations across different sources, geographies, and years.

99 GENERAL AND MISCELLANEOUS↗