Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data Analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Data Selection Improvement For MicroBooNE

Data selection is an extremely important part of data analysis for any experiment. Finding a physics result is often the result of sifting through a massive amount of data, keeping data that we believe to be signal, and throwing out data we do not. This process is called data selection. Creating a selection algorithm is an intensive process that must balance keeping enough data to have statistics and maximizing the signal purity of that data. We also need to choose the right reconstruction method, a tool to take raw data from the detector and convert it into physics results. In this study, we used three different reconstruction tools, Pandora, WireCell, and LANTERN, for the MicroBooNE experiment in conjunction to improve the selection algorithm for analysis. For the case of this study, we look into the charged current N proton 0 pions (CCNp0$\pi$) interaction channel. This is the dominant channel for the Short Baseline Neutrino (SBN) program and is expected to be a large contributor to the Deep Underground Neutrino Experiment (DUNE). We first investigated each of the three tools to find out more about their strengths and weaknesses as reconstructions, and compared them to the truth information directly from the MicroBooNE simulation pipeline. We then put together a direct comparison of the three methods to find which method or combination of methods would return the best result for us. While the study is ongoing, we have learned a lot about data selection for the experiment and the differences between the reconstruction tools.

Dillon, Brayden [Fermilab]↗

Electrification Analysis: Container Ports' Cargo Handling Equipment

This one-page highlight details the key takeaways from a project that utilized NREL's Fleet Research, Energy Data, and Insights (FleetREDI) data analysis pipeline, the Electrification Analysis of Container Ports' Cargo Handling Equipment project. This project created a scalable solution to model energy demand per shipping container moved (kWh/TEU) for an all-electric cargo handling equipment fleet located at a maritime port. The model allows stakeholders to understand energy demand at each electric vehicle (EV) equipment level and is easily scalable to container demand and EV adoption rate projections.

ADVANCED PROPULSION SYSTEMS↗

Challenges for monitoring and data analytics in a leadership public data repository

The availability and disposition of data has assumed increasing importance in large-scale computational science. Data repositories are evolving to meet new classes of requirements: compliance with government access guidelines, support for reproducibility of experimental results, and long-term availability of data products. The Constellation public data repository at the Oak Ridge Leadership Computing Facility faces these issues while being situated in one of the most productive data centers in the world. While monitoring and operational data analysis are ingrained in the operation of the OLCF’s large-scale high performance computing platforms, data repositories do not have this history of support. Problems faced by Constellation range from data size (over 7 petabytes in current holdings) to analytic complexity (detailed curation is both absolutely necessary for many data sets and absolutely impossible for humans to accomplish in any practical manner) to deployment environment (OLCF storage resources are oriented toward the needs of the compute platforms). In this paper we describe some of the challenges for collecting monitoring and analytic data from a leadership public data repository. We also discuss various strategies we are pursuing in order to address these challenges, from manual data collection to plans for introducing machine learning-based curatorial techniques.

Widener, Patrick [ORNL] (ORCID:0000000258820816)↗

Classification of events from α -induced reactions in the MUSIC detector via statistical and ML methods

The Multi-Sampling Ionization Chamber (MUSIC) detector is typically used to measure nuclear reaction cross sections relevant for nuclear astrophysics, fusion studies, and other applications. From the MUSIC data produced in one experiment scientists carefully extract an order of 10 3 events of interest from about 10 9 total events, where each event can be represented by an 18-dimensional vector. However, the standard data classification process is based on expert driven, manually intensive data analysis techniques that require several months to identify patterns and classify the relevant events from the collected data. Here, to address this issue, we present a method for the classification of events originating from specific α-induced reactions by combining statistical and machine learning methods that require significantly less input from the domain scientist, relative to the standard technique. Here, we applied the new method to two experimental data sets and compared our results with those obtained using traditional methods. With few exceptions, the number of events classified by our method agrees within ±20% with the results obtained using traditional methods. With the present method, which is the first of its kind for the MUSIC data, we have established the foundation for the automated extraction of physical events of interest from experiments using the MUSIC detector.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

The Scientific Case for Concurrent Neutron and X-ray Scattering and Spectroscopy

The interrogation of materials with X-rays or neutrons to determine the structure, energetics, and dynamics of materials is fundamental to advancing materials' physical and chemical science and developing innovative material technologies. A transcending challenge in developing novel materials is that progress hinges on understanding the structure and dynamics across multiple time and length scales in complex materials that feature multiple components, interfaces, and compositions. Despite the ever-growing demands on materials’ characterization, existing approaches are almost exclusively based on isolated X-ray or neutron scattering, i.e., an approach commensurate with the more narrowly defined needs of fifty years ago. A three-day workshop sponsored by the U.S. National Science Foundation (NSF) analyzed the demand for concurrent neutron and X-ray (NeX) experiments. It was held at the Spring Hill Suites, San Jose, California, from June 2 to 4, 2022. In this workshop, 70 national and international experts ascertained the crucial need to establish NeX capabilities to advance the science of complex materials and systems in the US. Here, we illustrate the need for NeX scattering and spectroscopy experiments by showcasing examples that span areas as diverse as biomaterials, energy science, soft matter, and nanomaterials. To provide NeX capability will require new instrumentation that enables concurrent experiments. Affected areas include chemistry, soft matter, quantum materials, pure and applied chemistry, bioscience, geoscience, and applied materials. NeX benefits research outcomes due to the complementarity of the two techniques, which is essential for better model refinement. While joint refinement of data from separate neutron and X-ray experiments is critical to avoid ambiguities, especially in multiphase-multicomponent materials, concurrent experiments overcome scientific and technical barriers associated with single measurements, separated by location and, thus, time. Among all the examples, these factors introduce uncertainties in the results that complicate data analysis. [1,2] [3] While models are strongly sample-dependent, the principles of joint refinement are generally applicable to these disciplines, including the development of advanced parameterization, modeling, and analysis techniques that also consider the temporal and spatial resolutions of the two methods, leading to unambiguous data interpretation. Solutions for technical barriers must be found to realize NeX experiments, including developing robust sample environments that meet the optical requirements of neutrons and X-rays.

36 MATERIALS SCIENCE↗

Imprinted Micelle Integration into a Commercial Platform (Progress Report)

PNNL has successfully integrated a commercial aerosol detector and the imprinted micelle technology. The integrated systems have been shown to have a limit of detection between 33-47 particles with several options for data analysis presented that vary on computational requirements. It is possible to integrate these systems and receive response data on the second time scale. While more work is needed, these technologies are compatible, which opens up a large field of air sampling looking for specific contaminates.

36 MATERIALS SCIENCE↗

Automated 3D cytoplasm segmentation in soft X-ray tomography

Cells’ structure is key to understanding cellular function, diagnostics, and therapy development. Soft X-ray tomography (SXT) is a unique tool to image cellular structure without fixation or labeling at high spatial resolution and throughput. Fast acquisition times increase demand for accelerated image analysis, like segmentation. Currently, segmenting cellular structures is done manually and is a major bottleneck in the SXT data analysis. This paper introduces ACSeg, an automated 3D cytoplasm segmentation model. ACSeg is generated using semi-automated labels and 3D U-Net and is trained on 43 SXT tomograms of immune T cells, rapidly converging to high-accuracy segmentation, therefore reducing time and labor. Furthermore, adding only 6 SXT tomograms of other cell types diversifies the model, showing potential for optimal experimental design. ACSeg successfully segmented unseen tomograms and is published on Biomedisa, enabling high-throughput analysis of cell volume and structure of cytoplasm in diverse cell types.

59 BASIC BIOLOGICAL SCIENCES↗

Data and Code for: Observation-constrained agroecosystem model inversion reveals continental-scale variation of winter wheat traits

This repository contains the simulation outputs and processing scripts associated with the study of winter wheat traits across the United States, utilizing the Ecosys agroecosystem model. The dataset includes model results for both rainfed and irrigated winter wheat systems, supporting the findings presented in the manuscript titled "Observation-constrained agroecosystem model inversion reveals continental-scale variation of winter wheat traits." Data includes the original Ecosys simulation outputs (archived in .db format within the compressed .zip files) and extracted analysis data (stored in .pkl files for efficient processing). Python code for data processing and figure generation is provided in a Jupyter notebook. External Observational Datasets should refer to the following official repositories for the input and validation data used in this study. The eddy covariance data from the AmeriFlux network (https://ameriflux.lbl.gov/). Climate-forcing data of NLDAS-2 from NASA LDAS (https://ldas.gsfc.nasa.gov/nldas/nldas-2-forcing-data). Soil data from the Gridded Soil Survey Geographic Database (gSSURGO), available at (https://www.nrcs.usda.gov/resources/data-and-reports/gridded-soil-survey-geographic-gssurgo-database). Crop yields, planting and harvest dates from the USDA public databases (https://quickstats.nass.usda.gov/; https://webapp.rma.usda.gov/apps/actuarialinformationbrowser/CropCriteria.aspx). Satellite-derived SLOPE GPP data from ORNL DAAC (https://daac.ornl.gov/cgi-bin/dsviewer.pl?ds_id=1786). Land use and crop progress information from the USDA Crop Data Layer and Crop Progress and Condition Gridded Layers (https://www.nass.usda.gov/Research_and_Science/). The Ecosys model code is available online at https://github.com/jinyun1tang/ECOSYS.

Wheat↗

The Italian Summer Students Program at Fermilab and other US Laboratories: 40 years of education in particle physics and technology

Since 1983 the Italian groups collaborating with Fermilab (US) have been running a 2-month summer training program for Master students. While in the first year the program involved only 4 physics students, in the following years it was extended to engineering students. Many students have extended their collaboration with Fermilab with their Master Thesis and PhD. The program has involved more than 600 Italian students from more than 20 Italian universities. Each intern is supervised by a Fermilab Mentor responsible for the training program. Training programs spanned from Tevatron, CMS, Muon (g-2), Mu2e and SBN (MicroBooNE, Icarus, and SBND) and DUNE design and data analysis, development of particle detectors, design of electronic and accelerator components, development of infrastructures and software for tera-data handling, quantum computing and research on superconductive elements and accelerating cavities. In 2015 the University of Pisa included the program within its own educational programs. Summer Students are enrolled at the University of Pisa for the duration of the internship and at the end of the internship they write summary reports on their achievements. After positive evaluation by a University of Pisa Examining Board, interns are acknowledged 6 ECTS credits for their Diploma Supplement. The program was paused in 2020 and 2021 due to the COVID-19 pandemic, but it resumed in 2022. From 2022 to 2024, a total of 60 students participated in the nine-week training at Fermilab. We are currently organizing the 2025 program. This paper provides an overview of the program, which can serve as a model for other interested laboratories.

Barzi, Emanuela [Ohio State U.]↗

The Italian Summer Students Program at Fermilab and other US Laboratories: 40 years of education in particle physics and technology

Since 1983 the Italian groups collaborating with Fermilab (US) have been running a 2-month summer training program for Master students. While in the first year the program involved only 4 physics students, in the following years it was extended to engineering students. Many students have extended their collaboration with Fermilab with their Master Thesis and PhD. The program has involved almost 600 Italian students from more than 20 Italian universities. Each intern is supervised by a Fermilab Mentor responsible for the training program. Training programs spanned from Tevatron, CMS, Muon (g-2), Mu2e and SBN and DUNE design and data analysis, development of particle detectors, design of electronic and accelerator components, development of infrastructures and software for tera-data handling, quantum computing and research on superconductive elements and accelerating cavities. In 2015 the University of Pisa included the program within its own educational programs. Summer Students are enrolled at the University of Pisa for the duration of the internship and at the end of the internship they write summary reports on their achievements. After positive evaluation by a University of Pisa Examining Board, interns are acknowledged 6 ECTS credits for their Diploma Supplement. In the years 2020 and 2021 the program was canceled due to the sanitary emergency but in 2022 it was restarted and allowed a cohort of 21 students in 2022, and a cohort of 27 students in 2023 to be trained for nine weeks at Fermilab. We are now organizing the 2024 program.

Barzi, Emanuela↗

Scaling the SciDAC QuantOm Workflow

As part of the Scientific Discovery through Advanced Computing (SciDAC) program, the Quantum Chromodynamics Nuclear Tomography (QuantOM) project aims to analyze data from Deep Inelastic Scattering (DIS) experiments conducted at Jefferson Lab and the upcoming Electron Ion Collider. The DIS data analysis is performed on an event-level by combining the input from theoretical and experimental nuclear physics into a single, composable workflow. The optimization itself (I.e. fitting the experimental data with theoretical predictions) is carried out by a machine / deep learning algorithm. The size of the acquired DIS data as well as the complexity of the workflow itself require that the analysis is performed across multiple GPUs on high performance computing systems, such as Polaris at Argonne National Laboratory. This presentation discusses the novelties and challenges that came along with parallelizing this workflow. Recent results are compared to common distributed training techniques.

Lersch, Daniel↗

Web-based wide-area monitoring platform for ringdown and clustering analytics in power systems

This paper introduces an open-source research platform for monitoring the Mexican interconnected power grid, allowing real-time processing and information extraction of the grid’s dynamic condition. Moreover, the platform is a Python-based development that embeds different ringdown and clustering analytics tools. In the case of ringdown analysis, the modal information can be extracted using some of the most known algorithms, i.e., Prony analysis, eigensystem realization algorithm (ERA), and matrix pencil (MP). For clustering analysis, the coherent behaviour of generator and non-generator buses is provided by applying recent state-of-the-art techniques such as affinity propagation, K-means, hierarchical agglomerative clustering, and typicality data analysis. The results of up to 93 PMUs show that this open-source platform suits researchers’ and engineers’ power system dynamic analysis requirements.

Clustering↗

Performance Comparison of Machine Learning Models for Ultrasonic Nondestructive Evaluation of Alkali-Silica Reaction in Concrete

Alkali-silica reaction (ASR) causes concrete degradation, leading to cracking, rebar corrosion, and reduced structural integrity, which raises safety concerns. Ultrasonic nondestructive evaluation (NDE) effectively assesses concrete properties and monitors ASR progression. However, its deployment and analysis require specialized expertise and subjective interpretation. As computational power increases, artificial intelligence (AI) and machine learning (ML) algorithms are increasingly being used to automate NDE data analysis across various industries for AI-assisted automation. Regulatory agencies are adapting to this technological shift, prompting a need to evaluate current ML technologies’ capabilities and limitations in assessing concrete material properties and damage. This report presents a comparative analysis of four ML regression models for predicting concrete material damage induced by ASR expansion using long-term ultrasonic data monitoring. The models investigated include linear regression (LR), support vector regression (SVR), shallow neural networks (NN), and deep neural networks (DNN). LR, SVR, and shallow NN models use features extracted from ultrasonic signals, whereas the DNN model processes time-domain ultrasonic signals and frequency spectra directly. The study systematically compared the models’ performance from various perspectives, including model input, prediction performance, and generalization ability. The findings indicate significant variability in model performance, with some ML algorithms achieving very high or very low prediction accuracy depending on the preprocessing and feature engineering (extraction and selection) applied. Key insights include the observation that shallow ML models (LR, SVR, and shallow NNs) require meticulous preprocessing and feature extraction to achieve high accuracy. In contrast, the DNN model, although it bypasses the need for feature engineering, necessitates extensive preprocessing to mitigate noise and computational demands. The SVR model emerged as the top performer among the shallow models, and the DNN model exhibited superior performance on specific datasets but struggled with generalization across specimens from different batches. Additionally, the SVR model is sensitive to temperature variations, whereas the DNN model is robust in this regard. Using recurrent neural networks is recommended for future ASR expansion prediction studies. Recurrent neural networks’ inherent ability to capture temporal dependencies and long-term patterns makes them well suited for analyzing sequential ultrasonic monitoring data. Overall, the results and conclusions of this study could provide insights into the capabilities and effectiveness of ML when applied to ultrasonic NDE data and help identify best practices for using ML for ultrasonic NDE of concrete material properties.

36 MATERIALS SCIENCE↗

SpectraCodec: A Hilbert curve-based method for encoding metadata in mass spectra for machine learning applications (SpectraCodec) v1

Machine learning approaches to mass spectrometry (MS) data analysis require structured metadata for optimal performance. However, current MS file formats necessitate external metadata sources, creating integration challenges that impede analytical workflows. Here, we present a novel approach for encoding metadata directly within mzML files using one-hot encoding of ASCII characters mapped via Hilbert space-filling curves. This strategy embeds metadata in the first spectrum's m/z-intensity space, ensuring persistence with the primary data, eliminating the need for external metadata files, and maintaining compatibility with existing MS software. We demonstrate that the Hilbert curve mapping efficiently utilizes the two-dimensional spectral space while maintaining robust data recovery. This method offers a practical solution for machine learning applications in mass spectrometry by ensuring metadata and spectral data remain unified through all stages of analysis.

Bowen, Benjamin [Lawrence Berkeley National Labora↗

Identifying preferential flow from soil moisture time series: Review of methodologies

Abstract Identifying and quantifying preferential flow (PF) through soil—the rapid movement of water through spatially distinct pathways in the subsurface—is vital to understanding how the hydrologic cycle responds to climate, land cover, and anthropogenic changes. In recent decades, methods have been developed that use measured soil moisture time series to identify PF. Because they allow for continuous monitoring and are relatively easy to implement, these methods have become an important tool for recognizing when, where, and under what conditions PF occurs. The methods seek to identify a pattern or quantification that indicates the occurrence of PF. Most commonly, the chosen signature is either (1) a nonsequential response to infiltrated water, in which soil moisture responses do not occur in order of shallowest to deepest, or (2) a velocity criterion, in which newly infiltrated water is detected at depth earlier than is possible by nonpreferential flow processes. Alternative signatures have also been developed that have certain advantages but are less commonly utilized. Choosing among these possible signatures requires attention to their pertinent characteristics, including susceptibility to errors, possible bias toward false negatives or false positives, reliance on subjective judgments, and possible requirements for additional types of data. We review 77 studies that have applied such methods to highlight important information for readers who want to identify PF from soil moisture data and to inform those who aim to develop new methods or improve existing ones. Core Ideas Soil moisture data can be used to identify the occurrence of preferential flow (PF) and its initiating conditions. Various data‐analysis methods to identify PF differ in susceptibility to error, bias, and subjectivity. These methods can utilize vast amounts of data from soil moisture monitoring networks to develop understanding of when, where, and under what conditions PF occurs. Newly developed methods may lead to better accuracy and reliability, and reduce the need for subjective judgments. Plain Language Summary Preferential flow through soil occurs when a large amount of water is suddenly available, as during an intense storm. This type of flow moves rapidly through the soil in distinct narrow pathways rather than moving evenly throughout the body of soil, with major consequences for groundwater resources, ecosystems, spreading of contaminants, and other vital concerns. Methods of detecting preferential flow have been developed that utilize measurements of soil water content made by sensors installed at various depths. This measurement technology has been widely implemented, many locations now having datasets years in length, and various methods have been developed for using these to identify preferential flow. The various methods are based on different features in the soil moisture records and vary in their advantages and shortcomings. In this review, we explain and evaluate these methods, highlighting important information for their implementation to identify preferential flow from soil moisture data and for efforts to develop new methods or improve existing ones.

Nimmo, John R↗

Machine learning analysis of high-repetition-rate two-dimensional Thomson scattering spectra from laser-produced plasmas

With the emergence of high-repetition-rate two-dimensional Thomson scattering (TS) measurements, improving spectral data analysis is a key area of interest. Here, we present a new way to derive the electron temperature and density of laser-driven blast waves in plasmas from their TS spectra with machine learning (ML). This analysis occurs in both the non-collective (α < 1) and collective (α > 1) scattering regimes with the goal of autonomously and more accurately determining T c and n e both where spectral data has been collected and to give the ability to predict these attributes in regions where data has not been collected. We introduce three ML models, one trained only on experimental data, one only on synthetic data, and one using transfer learning, and compare their speed and accuracy with the conventional TS inversion algorithms in the open source PlasmaPy python package.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Nuclear Data Management and Analysis System Plan

The United States Department of Energy Advanced Reactor Technologies Program was formed in Fiscal Year 2015 and encompasses the Next Generation Nuclear Plant Project and Very High Temperature Reactor (VHTR) Program as they were known previously. The VHTR Program was created to support design and licensing of the first VHTR nuclear plant. Data created for and used by the program must be qualified for use, stored in a readily accessible electronic form, categorized to assure the correct data are used, and controlled to prevent data corruption or inadvertent changes. The Nuclear Data Management and Analysis System was designed to support the data needs of the VHTR Program, at the time and now the Advanced Reactor Technologies Program. Since its inception, use of the Nuclear Data Management and Analysis System has expanded to support additional projects and programs with similar requirements for control, analysis, and availability of large data sets.

99 GENERAL AND MISCELLANEOUS↗

Anomaly Detection in Seismic Data with Deep Learning: Application for Instrument Failure Detection and Forecasting

Seismic data quality assessment (QA) is the first and one of the most important steps before conducting any further data analysis. Traditional methods involve checking various metrics, such as spike detection and power spectral density, by setting strict thresholds or comparing data against synthetic benchmarks. However, these approaches often rely on pre-existing knowledge and assumptions about data anomalies, leading to potential misclassification of unusual cases. Here, in this study, we propose a deep autoencoder model, an unsupervised learning approach that evaluates data quality without making assumptions about normal and anomalous data, which can be used to identify deviations in recorded data that may indicate nascent instrument failure. We test the model with the U.S. International Monitoring System (IMS) seismic stations and demonstrate the capability of detecting anomalies on a monthly scale. This could prompt station operators to examine potential problems early, allowing sufficient time for instrument maintenance to prevent data outages. In addition, we use a new manually selected testing dataset to compare our model performance against two supervised machine learning (ML) approaches and a standard QA package, as baseline models. When applied to the dataset containing known data anomalies, performance of the supervised and unsupervised ML approaches is similar, with an accuracy of 88.1% for our model compared to ∼90% for the supervised ML approach and 78.2% for the standard QA package. Our model outperforms the baseline models when applied to new stations, where new types of data anomalies can be station-specific and not included in the training dataset. Finally, we show model transferability by training the model with data from the Global Seismograph Network only and applying it to the IMS network data. The results suggest that our model is generalizable and can be applied to new stations with good accuracy.

Lin, Jiun-Ting [Lawrence Livermore National Labora↗