Search NASA⌕ Search

SEARCH · Search NASA

Results for “data analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Author Correction: US oil and gas system emissions from nearly one million aerial site measurements

Correction to: Naturehttps://doi.org/10.1038/s41586-024-07117-5 Published online 13 March 2024 In the version of the article initially published, several errors were present and have been corrected in the HTML and PDF versions of the article and Supplementary Information. The main results, conclusions, and our interpretations of the data remain unchanged. See the new Supplementary Information Section S15 for a more detailed description of the errors corrected and the resulting effects on the analysis. Data processing and methods corrections Overflight count correction: We previously used pre-computed source coverage data for some Carbon Mapper campaigns that was computed differently than was required for our analysis. We have re-computed Carbon Mapper source coverage based on flightline polygons and source coordinates. Transition point computation, well sites: The updated version now correctly compares the cumulative emissions distribution of simulated well site emissions with that of aerially detected sources (rather than plumes) when computing the transition point. Transition point computation, midstream: Additionally, the transition point calculation has been corrected to exclude aerially detected midstream emissions below the transition point, which was previously leading to double counting of these emissions. This error was not present for upstream (well site) emissions. Calculation errors Unit error: We corrected a specific unit conversion error affecting well site emissions in the Kairos Fort Worth dataset. Across all datasets, we also correct the conversion factor for converting from standard volume to mass for midstream emissions. Sorting error: We correct code that was applying incorrect sorting when computing correction factors to account for partial detection at well sites. Small typographical corrections were made in Fig. 1b and SI Section S4.1. Data processing and methods corrections Overflight count correction: We previously used pre-computed source coverage data for some Carbon Mapper campaigns that was computed differently than was required for our analysis. We have re-computed Carbon Mapper source coverage based on flightline polygons and source coordinates. Transition point computation, well sites: The updated version now correctly compares the cumulative emissions distribution of simulated well site emissions with that of aerially detected sources (rather than plumes) when computing the transition point. Transition point computation, midstream: Additionally, the transition point calculation has been corrected to exclude aerially detected midstream emissions below the transition point, which was previously leading to double counting of these emissions. This error was not present for upstream (well site) emissions. Calculation errors Unit error: We corrected a specific unit conversion error affecting well site emissions in the Kairos Fort Worth dataset. Across all datasets, we also correct the conversion factor for converting from standard volume to mass for midstream emissions. Sorting error: We correct code that was applying incorrect sorting when computing correction factors to account for partial detection at well sites. Small typographical corrections were made in Fig. 1b and SI Section S4.1. The following practices may help researchers conducting similar analyses avoid making similar errors: 1, Clear, accessible documentation explaining the interpretation of all columns in data input tables and all internal variables within the model, 2, Simple cross-check calculations computed before and after unit conversions.

Sherwin, Evan D↗

Selection Algorithm Improvement for MicroBooNE

Data selection is an extremely important part of data analysis for any experiment. Finding a physics result is often the result of sifting through a massive amount of data, keeping data that we believe to be signal and throwing out data we do not. This process is called data selection. Creating a selection algorithm is an intensive process that must balance keeping enough data to have statistics and maximizing the signal purity of that data. In this study, we used three different reconstruction tools, Pandora, WireCell, and LANTERN, for the MicroBooNE experiment in conjunction to improve the selection algorithm for analysis. For the case of this study, we look into the charged current N proton 0 pions (CCNp0$\pi$) interaction channel. This is the dominant channel for the Short Baseline Neutrino (SBN) program and is expected to be a large contributor to the Deep Underground Neutrino Experiment (DUNE). We first investigated each of the three tools to find out more about their strengths and weaknesses as reconstructions. We then put together a direct comparison of the three methods to find which method or combination of methods would return the best result for us. While the study is ongoing, we have learned a lot about data selection for the experiment and the differences between the reconstruction tools.

Dillon, Brayden [Michigan State U.]↗

Deconvoluting thermomechanical effects in X-ray diffraction data using machine learning

X-ray diffraction is ideal for probing the sub-surface state during complex or rapid thermomechanical loading of crystalline materials. However, challenges arise as the size of diffraction volumes increases due to spatial broadening and because of the inability to deconvolute the effects of different lattice deformation mechanisms. Here, we present a novel approach that uses combinations of physics-based modeling and machine learning to deconvolve thermal and mechanical elastic strains for diffraction data analysis. The method builds on a previous effort to extract thermal strain distribution information from diffraction data. The new approach is applied to extract the evolution of the thermomechanical state during laser melting of an Inconel 625 wall specimen which produces significant residual stress upon cooling. A combination of heat transfer and fluid flow, elasto-plasticity and X-ray diffraction simulations is used to generate training data for machine-learning (Gaussian process regression, GPR) models that map diffracted intensity distributions to underlying thermomechanical strain fields. First-principles density functional theory is used to determine accurate temperature-dependent thermal expansion and elastic stiffness used for elasto-plasticity modeling. The trained GPR models are found to be capable of deconvoluting the effects of thermal and mechanical strains, in addition to providing information about underlying strain distributions, even from complex diffraction patterns with irregularly shaped peaks.

36 MATERIALS SCIENCE↗

Lipid droplet-associated proteins in alcohol-associated fatty liver disease: A proteomic approach

The earliest manifestation of alcohol-associated liver disease (ALD) is steatosis characterized by deposition of fat in specialized organelles called lipid droplets (LDs). While alcohol administration causes a rise in LD numbers in the hepatocytes, little is known regarding their characteristics that allow their accumulation and size to increase. The aim of the present study is to gain insights into underlying pathophysiological mechanisms by investigating the ethanol-induced changes in hepatic LD proteome as a function of LD size. Adult male Wistar rats (180–200 g BW) were fed with ethanol liquid diet for 6 weeks. At sacrifice, large-, medium-, and small-sized hepatic LD subpopulations (LD1, LD2, and LD3, respectively) were isolated and subjected to morphological and proteomic analyses. Morphological analysis of LD1-LD3 fractions of ethanol-fed rats clearly demonstrated that LD1 contained larger LDs compared with LD2 and LD3 fractions. Our preliminary results from principal component analysis showed that the proteome of different-sized hepatic LD fractions was distinctly different. Proteomic data analysis identified over 2000 proteins in each LD fraction with significant alterations in protein abundance among the three LD fractions. Among the altered proteins, several were related to fat metabolism, including synthesis, incorporation of fatty acid, and lipolysis. Ingenuity pathway analysis revealed increased fatty acid synthesis, fatty acid incorporation, LD fusion, and reduced lipolysis in LD1 compared to LD3. Overall, the proteomic findings indicate that the increased level of protein that facilitates fusion of LDs combined with an increased association of negative regulators of lipolysis dictates the generation of large-sized LDs during the development of alcohol-associated hepatic steatosis. Several significantly altered proteins were identified in different-sized LDs isolated from livers of ethanol-fed rats. Ethanol-induced increases in specific proteins that hinder LD lipid metabolism led to the accumulation and persistence of large-sized LDs in the liver.

60 APPLIED LIFE SCIENCES↗

Data Selection Improvement For MicroBooNE

Data selection is an extremely important part of data analysis for any experiment. Finding a physics result is often the result of sifting through a massive amount of data, keeping data that we believe to be signal, and throwing out data we do not. This process is called data selection. Creating a selection algorithm is an intensive process that must balance keeping enough data to have statistics and maximizing the signal purity of that data. We also need to choose the right reconstruction method, a tool to take raw data from the detector and convert it into physics results. In this study, we used three different reconstruction tools, Pandora, WireCell, and LANTERN, for the MicroBooNE experiment in conjunction to improve the selection algorithm for analysis. For the case of this study, we look into the charged current N proton 0 pions (CCNp0$\pi$) interaction channel. This is the dominant channel for the Short Baseline Neutrino (SBN) program and is expected to be a large contributor to the Deep Underground Neutrino Experiment (DUNE). We first investigated each of the three tools to find out more about their strengths and weaknesses as reconstructions, and compared them to the truth information directly from the MicroBooNE simulation pipeline. We then put together a direct comparison of the three methods to find which method or combination of methods would return the best result for us. While the study is ongoing, we have learned a lot about data selection for the experiment and the differences between the reconstruction tools.

Dillon, Brayden [Fermilab]↗

Electrification Analysis: Container Ports' Cargo Handling Equipment

This one-page highlight details the key takeaways from a project that utilized NREL's Fleet Research, Energy Data, and Insights (FleetREDI) data analysis pipeline, the Electrification Analysis of Container Ports' Cargo Handling Equipment project. This project created a scalable solution to model energy demand per shipping container moved (kWh/TEU) for an all-electric cargo handling equipment fleet located at a maritime port. The model allows stakeholders to understand energy demand at each electric vehicle (EV) equipment level and is easily scalable to container demand and EV adoption rate projections.

ADVANCED PROPULSION SYSTEMS↗

Challenges for monitoring and data analytics in a leadership public data repository

The availability and disposition of data has assumed increasing importance in large-scale computational science. Data repositories are evolving to meet new classes of requirements: compliance with government access guidelines, support for reproducibility of experimental results, and long-term availability of data products. The Constellation public data repository at the Oak Ridge Leadership Computing Facility faces these issues while being situated in one of the most productive data centers in the world. While monitoring and operational data analysis are ingrained in the operation of the OLCF’s large-scale high performance computing platforms, data repositories do not have this history of support. Problems faced by Constellation range from data size (over 7 petabytes in current holdings) to analytic complexity (detailed curation is both absolutely necessary for many data sets and absolutely impossible for humans to accomplish in any practical manner) to deployment environment (OLCF storage resources are oriented toward the needs of the compute platforms). In this paper we describe some of the challenges for collecting monitoring and analytic data from a leadership public data repository. We also discuss various strategies we are pursuing in order to address these challenges, from manual data collection to plans for introducing machine learning-based curatorial techniques.

Widener, Patrick [ORNL] (ORCID:0000000258820816)↗

Classification of events from α -induced reactions in the MUSIC detector via statistical and ML methods

The Multi-Sampling Ionization Chamber (MUSIC) detector is typically used to measure nuclear reaction cross sections relevant for nuclear astrophysics, fusion studies, and other applications. From the MUSIC data produced in one experiment scientists carefully extract an order of 10 3 events of interest from about 10 9 total events, where each event can be represented by an 18-dimensional vector. However, the standard data classification process is based on expert driven, manually intensive data analysis techniques that require several months to identify patterns and classify the relevant events from the collected data. Here, to address this issue, we present a method for the classification of events originating from specific α-induced reactions by combining statistical and machine learning methods that require significantly less input from the domain scientist, relative to the standard technique. Here, we applied the new method to two experimental data sets and compared our results with those obtained using traditional methods. With few exceptions, the number of events classified by our method agrees within ±20% with the results obtained using traditional methods. With the present method, which is the first of its kind for the MUSIC data, we have established the foundation for the automated extraction of physical events of interest from experiments using the MUSIC detector.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Imprinted Micelle Integration into a Commercial Platform (Progress Report)

PNNL has successfully integrated a commercial aerosol detector and the imprinted micelle technology. The integrated systems have been shown to have a limit of detection between 33-47 particles with several options for data analysis presented that vary on computational requirements. It is possible to integrate these systems and receive response data on the second time scale. While more work is needed, these technologies are compatible, which opens up a large field of air sampling looking for specific contaminates.

36 MATERIALS SCIENCE↗

Data and Code for: Observation-constrained agroecosystem model inversion reveals continental-scale variation of winter wheat traits

This repository contains the simulation outputs and processing scripts associated with the study of winter wheat traits across the United States, utilizing the Ecosys agroecosystem model. The dataset includes model results for both rainfed and irrigated winter wheat systems, supporting the findings presented in the manuscript titled "Observation-constrained agroecosystem model inversion reveals continental-scale variation of winter wheat traits." Data includes the original Ecosys simulation outputs (archived in .db format within the compressed .zip files) and extracted analysis data (stored in .pkl files for efficient processing). Python code for data processing and figure generation is provided in a Jupyter notebook. External Observational Datasets should refer to the following official repositories for the input and validation data used in this study. The eddy covariance data from the AmeriFlux network (https://ameriflux.lbl.gov/). Climate-forcing data of NLDAS-2 from NASA LDAS (https://ldas.gsfc.nasa.gov/nldas/nldas-2-forcing-data). Soil data from the Gridded Soil Survey Geographic Database (gSSURGO), available at (https://www.nrcs.usda.gov/resources/data-and-reports/gridded-soil-survey-geographic-gssurgo-database). Crop yields, planting and harvest dates from the USDA public databases (https://quickstats.nass.usda.gov/; https://webapp.rma.usda.gov/apps/actuarialinformationbrowser/CropCriteria.aspx). Satellite-derived SLOPE GPP data from ORNL DAAC (https://daac.ornl.gov/cgi-bin/dsviewer.pl?ds_id=1786). Land use and crop progress information from the USDA Crop Data Layer and Crop Progress and Condition Gridded Layers (https://www.nass.usda.gov/Research_and_Science/). The Ecosys model code is available online at https://github.com/jinyun1tang/ECOSYS.

Wheat↗

The Italian Summer Students Program at Fermilab and other US Laboratories: 40 years of education in particle physics and technology

Since 1983 the Italian groups collaborating with Fermilab (US) have been running a 2-month summer training program for Master students. While in the first year the program involved only 4 physics students, in the following years it was extended to engineering students. Many students have extended their collaboration with Fermilab with their Master Thesis and PhD. The program has involved more than 600 Italian students from more than 20 Italian universities. Each intern is supervised by a Fermilab Mentor responsible for the training program. Training programs spanned from Tevatron, CMS, Muon (g-2), Mu2e and SBN (MicroBooNE, Icarus, and SBND) and DUNE design and data analysis, development of particle detectors, design of electronic and accelerator components, development of infrastructures and software for tera-data handling, quantum computing and research on superconductive elements and accelerating cavities. In 2015 the University of Pisa included the program within its own educational programs. Summer Students are enrolled at the University of Pisa for the duration of the internship and at the end of the internship they write summary reports on their achievements. After positive evaluation by a University of Pisa Examining Board, interns are acknowledged 6 ECTS credits for their Diploma Supplement. The program was paused in 2020 and 2021 due to the COVID-19 pandemic, but it resumed in 2022. From 2022 to 2024, a total of 60 students participated in the nine-week training at Fermilab. We are currently organizing the 2025 program. This paper provides an overview of the program, which can serve as a model for other interested laboratories.

Barzi, Emanuela [Ohio State U.]↗

The Italian Summer Students Program at Fermilab and other US Laboratories: 40 years of education in particle physics and technology

Since 1983 the Italian groups collaborating with Fermilab (US) have been running a 2-month summer training program for Master students. While in the first year the program involved only 4 physics students, in the following years it was extended to engineering students. Many students have extended their collaboration with Fermilab with their Master Thesis and PhD. The program has involved almost 600 Italian students from more than 20 Italian universities. Each intern is supervised by a Fermilab Mentor responsible for the training program. Training programs spanned from Tevatron, CMS, Muon (g-2), Mu2e and SBN and DUNE design and data analysis, development of particle detectors, design of electronic and accelerator components, development of infrastructures and software for tera-data handling, quantum computing and research on superconductive elements and accelerating cavities. In 2015 the University of Pisa included the program within its own educational programs. Summer Students are enrolled at the University of Pisa for the duration of the internship and at the end of the internship they write summary reports on their achievements. After positive evaluation by a University of Pisa Examining Board, interns are acknowledged 6 ECTS credits for their Diploma Supplement. In the years 2020 and 2021 the program was canceled due to the sanitary emergency but in 2022 it was restarted and allowed a cohort of 21 students in 2022, and a cohort of 27 students in 2023 to be trained for nine weeks at Fermilab. We are now organizing the 2024 program.

Barzi, Emanuela↗

Web-based wide-area monitoring platform for ringdown and clustering analytics in power systems

This paper introduces an open-source research platform for monitoring the Mexican interconnected power grid, allowing real-time processing and information extraction of the grid’s dynamic condition. Moreover, the platform is a Python-based development that embeds different ringdown and clustering analytics tools. In the case of ringdown analysis, the modal information can be extracted using some of the most known algorithms, i.e., Prony analysis, eigensystem realization algorithm (ERA), and matrix pencil (MP). For clustering analysis, the coherent behaviour of generator and non-generator buses is provided by applying recent state-of-the-art techniques such as affinity propagation, K-means, hierarchical agglomerative clustering, and typicality data analysis. The results of up to 93 PMUs show that this open-source platform suits researchers’ and engineers’ power system dynamic analysis requirements.

Clustering↗

Performance Comparison of Machine Learning Models for Ultrasonic Nondestructive Evaluation of Alkali-Silica Reaction in Concrete

Alkali-silica reaction (ASR) causes concrete degradation, leading to cracking, rebar corrosion, and reduced structural integrity, which raises safety concerns. Ultrasonic nondestructive evaluation (NDE) effectively assesses concrete properties and monitors ASR progression. However, its deployment and analysis require specialized expertise and subjective interpretation. As computational power increases, artificial intelligence (AI) and machine learning (ML) algorithms are increasingly being used to automate NDE data analysis across various industries for AI-assisted automation. Regulatory agencies are adapting to this technological shift, prompting a need to evaluate current ML technologies’ capabilities and limitations in assessing concrete material properties and damage. This report presents a comparative analysis of four ML regression models for predicting concrete material damage induced by ASR expansion using long-term ultrasonic data monitoring. The models investigated include linear regression (LR), support vector regression (SVR), shallow neural networks (NN), and deep neural networks (DNN). LR, SVR, and shallow NN models use features extracted from ultrasonic signals, whereas the DNN model processes time-domain ultrasonic signals and frequency spectra directly. The study systematically compared the models’ performance from various perspectives, including model input, prediction performance, and generalization ability. The findings indicate significant variability in model performance, with some ML algorithms achieving very high or very low prediction accuracy depending on the preprocessing and feature engineering (extraction and selection) applied. Key insights include the observation that shallow ML models (LR, SVR, and shallow NNs) require meticulous preprocessing and feature extraction to achieve high accuracy. In contrast, the DNN model, although it bypasses the need for feature engineering, necessitates extensive preprocessing to mitigate noise and computational demands. The SVR model emerged as the top performer among the shallow models, and the DNN model exhibited superior performance on specific datasets but struggled with generalization across specimens from different batches. Additionally, the SVR model is sensitive to temperature variations, whereas the DNN model is robust in this regard. Using recurrent neural networks is recommended for future ASR expansion prediction studies. Recurrent neural networks’ inherent ability to capture temporal dependencies and long-term patterns makes them well suited for analyzing sequential ultrasonic monitoring data. Overall, the results and conclusions of this study could provide insights into the capabilities and effectiveness of ML when applied to ultrasonic NDE data and help identify best practices for using ML for ultrasonic NDE of concrete material properties.

36 MATERIALS SCIENCE↗

SpectraCodec: A Hilbert curve-based method for encoding metadata in mass spectra for machine learning applications (SpectraCodec) v1

Machine learning approaches to mass spectrometry (MS) data analysis require structured metadata for optimal performance. However, current MS file formats necessitate external metadata sources, creating integration challenges that impede analytical workflows. Here, we present a novel approach for encoding metadata directly within mzML files using one-hot encoding of ASCII characters mapped via Hilbert space-filling curves. This strategy embeds metadata in the first spectrum's m/z-intensity space, ensuring persistence with the primary data, eliminating the need for external metadata files, and maintaining compatibility with existing MS software. We demonstrate that the Hilbert curve mapping efficiently utilizes the two-dimensional spectral space while maintaining robust data recovery. This method offers a practical solution for machine learning applications in mass spectrometry by ensuring metadata and spectral data remain unified through all stages of analysis.

Bowen, Benjamin [Lawrence Berkeley National Labora↗

Identifying preferential flow from soil moisture time series: Review of methodologies

Abstract Identifying and quantifying preferential flow (PF) through soil—the rapid movement of water through spatially distinct pathways in the subsurface—is vital to understanding how the hydrologic cycle responds to climate, land cover, and anthropogenic changes. In recent decades, methods have been developed that use measured soil moisture time series to identify PF. Because they allow for continuous monitoring and are relatively easy to implement, these methods have become an important tool for recognizing when, where, and under what conditions PF occurs. The methods seek to identify a pattern or quantification that indicates the occurrence of PF. Most commonly, the chosen signature is either (1) a nonsequential response to infiltrated water, in which soil moisture responses do not occur in order of shallowest to deepest, or (2) a velocity criterion, in which newly infiltrated water is detected at depth earlier than is possible by nonpreferential flow processes. Alternative signatures have also been developed that have certain advantages but are less commonly utilized. Choosing among these possible signatures requires attention to their pertinent characteristics, including susceptibility to errors, possible bias toward false negatives or false positives, reliance on subjective judgments, and possible requirements for additional types of data. We review 77 studies that have applied such methods to highlight important information for readers who want to identify PF from soil moisture data and to inform those who aim to develop new methods or improve existing ones. Core Ideas Soil moisture data can be used to identify the occurrence of preferential flow (PF) and its initiating conditions. Various data‐analysis methods to identify PF differ in susceptibility to error, bias, and subjectivity. These methods can utilize vast amounts of data from soil moisture monitoring networks to develop understanding of when, where, and under what conditions PF occurs. Newly developed methods may lead to better accuracy and reliability, and reduce the need for subjective judgments. Plain Language Summary Preferential flow through soil occurs when a large amount of water is suddenly available, as during an intense storm. This type of flow moves rapidly through the soil in distinct narrow pathways rather than moving evenly throughout the body of soil, with major consequences for groundwater resources, ecosystems, spreading of contaminants, and other vital concerns. Methods of detecting preferential flow have been developed that utilize measurements of soil water content made by sensors installed at various depths. This measurement technology has been widely implemented, many locations now having datasets years in length, and various methods have been developed for using these to identify preferential flow. The various methods are based on different features in the soil moisture records and vary in their advantages and shortcomings. In this review, we explain and evaluate these methods, highlighting important information for their implementation to identify preferential flow from soil moisture data and for efforts to develop new methods or improve existing ones.

Nimmo, John R↗

Machine learning analysis of high-repetition-rate two-dimensional Thomson scattering spectra from laser-produced plasmas

With the emergence of high-repetition-rate two-dimensional Thomson scattering (TS) measurements, improving spectral data analysis is a key area of interest. Here, we present a new way to derive the electron temperature and density of laser-driven blast waves in plasmas from their TS spectra with machine learning (ML). This analysis occurs in both the non-collective (α < 1) and collective (α > 1) scattering regimes with the goal of autonomously and more accurately determining T c and n e both where spectral data has been collected and to give the ability to predict these attributes in regions where data has not been collected. We introduce three ML models, one trained only on experimental data, one only on synthetic data, and one using transfer learning, and compare their speed and accuracy with the conventional TS inversion algorithms in the open source PlasmaPy python package.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Nuclear Data Management and Analysis System Plan

The United States Department of Energy Advanced Reactor Technologies Program was formed in Fiscal Year 2015 and encompasses the Next Generation Nuclear Plant Project and Very High Temperature Reactor (VHTR) Program as they were known previously. The VHTR Program was created to support design and licensing of the first VHTR nuclear plant. Data created for and used by the program must be qualified for use, stored in a readily accessible electronic form, categorized to assure the correct data are used, and controlled to prevent data corruption or inadvertent changes. The Nuclear Data Management and Analysis System was designed to support the data needs of the VHTR Program, at the time and now the Advanced Reactor Technologies Program. Since its inception, use of the Nuclear Data Management and Analysis System has expanded to support additional projects and programs with similar requirements for control, analysis, and availability of large data sets.

99 GENERAL AND MISCELLANEOUS↗