Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

LLM integration into EPICS

The utilization of large language models (LLMs) such as ChatGPT has seen a remarkable increase in various fields over the past few years. These models have demonstrated their versatility and capability in understanding and generating human-like text, making them invaluable tools in numerous applications. In this project, we explore the integration of a LLM into the Experimental Physics and Industrial Control System (EPICS). The primary focus of this integration is to employ the LLM for advanced image processing and spatial analysis on images obtained from the beamlines. By leveraging the capabilities of the LLM, we aim to enhance the accuracy and efficiency of image interpretation, enabling more precise data analysis and decision-making within the EPICS framework. This integration not only showcases the potential of LLMs in scientific and industrial applications but also sets the stage for future advancements in automated control systems.

Adams, Ethan↗

Blueprints for Training Information Bottlenecks for Collider Analyses

Dimensionality reduction is a crucial aspect of data analysis in high energy physics, even if accompanied by information loss. Several methods, including histogram- and kernel-based analyses, are only computationally feasible for low-dimensional data. Furthermore, simulation models used in HEP can often only be validated for low-dimensional data. We provide several blueprints for using machine learning to create low-dimensional data representations (continuous event variables and discrete classification labels) for use in signal discovery and parameter estimation tasks. We also describe how to design the learned representation to facilitate a) searches with unknown model parameters and b) validation of simulation models in data control regions.

43 PARTICLE ACCELERATORS↗

Parallel sorting algorithm classification: is manual instrumentation necessary?

Understanding parallel algorithms is crucial for accelerating scientific simulations on complex, distributed memory, high-performance computers. Modern algorithm classification approaches learn semantics directly from source code to differentiate between algorithms, however, accessing source code is not always possible. We can learn about parallel algorithms from observing their performance, as programs running the same algorithms and using the same hardware should exhibit similar performance characteristics. We present an approach to learn algorithm classes from parallel performance data directly in order to classify algorithms without access to the source code. We extend previous work to enable classifying parallel sorting algorithms using automatic instrumentation instead of requiring manual region annotations in the source code. In this work, we design and demonstrate a study for classification of parallel sorting algorithms using parallel performance data collected from automatic instrumentation, and evaluate the performance of our new methodology on classification. We leverage Caliper to collect the performance data, Thicket for our exploratory data analysis (EDA), and PyTorch and Scikit-learn to evaluate the effectiveness of random forests, support vector machines (SVMs), decision trees, neural networks, and logistic regressions on parallel performance data. Additionally, we study noise in parallel performance data, whether the removal of noise and pre-processing of the data is necessary to accurately classify parallel sorting algorithms, and determine the effectiveness of features created from performance data. In conclusion, we demonstrate classification accuracy for these five different models of up to 97.7% across four different parallel algorithm classes.

Algorithm Classification↗

G EANT 4 atomic relaxation data for transfermium nuclei (Z = 101–104)

Advanced theoretical methods can accurately calculate various atomic observables and predict electronic structure. Still, systematic computations of the radiative and non-radiative transition probabilities and energies are missing for the actinides and all the transfermium elements. However, these compilations are needed for comprehensive Monte-Carlo simulations (such as GEANT4) of the radioactive decay of transfermium nuclei. These simulations can forma basis for data analysis of experiments, especially with complex detection setups. Investigation of the transfermium nuclei is crucial for understanding the nature of the nuclear force. In this study, simulations based on data from the Jena Atomic Calculator (JAC) and the data from the Evaluated Atomic Data Library (EADL) present in GEANT4 were found compatible for the three elements Ba(Z = 56), U(Z = 92), and Fm(Z = 100), thus, validating the JAC calculations. For Z> 100, we also found sound agreement between simulations that used data generated with JAC and experimental results involving No(Z = 102) and Rf(Z = 104) isotopes. In conclusion, these results demonstrate that JAC can produce reliable atomic data sets for transfermium elements, which will assist in analyzing nuclear-decay-spectroscopy experiments.

GEANT4↗

A co-registered in-situ and ex-situ dataset from wire arc additive manufacturing process

Recent progress in sensing techniques and data analytics tools have significantly accelerated the development of Wire Arc Additive Manufacturing (WAAM) systems. This data-centric approach emphasizes leveraging sensor data available throughout the production process to optimize performance. Integration of extensive data analysis provides opportunities for improving precision, reducing waste, and enhancing the quality of produced parts. This method relies on AI/ML models and optimization techniques, which are developed using the data collected from various sources, including in-situ sensors, ex-situ imaging, and manufacturing process parameters. The quality and diversity of this data, along with the alignment between different data streams (achieved through spatiotemporal registration) are critical for the successful development of AI/ML and optimization models. In this work, we present a spatiotemporally registered dataset generated during the WAAM process of deposition of a rectangular block. The dataset includes a comprehensive description of the deposition process, process parameters, welding characteristics and acoustic data collected in-situ, and X-Ray Computed Tomography data of the build.

42 ENGINEERING↗

eDNAjoint: An R package for interpreting paired or semi‐paired environmental DNA and traditional survey data in a Bayesian framework

Abstract Environmental DNA (eDNA) sampling is increasingly used in surveys of species distribution as a potentially sensitive and efficient monitoring method. Yet access to modelling tools designed specifically for interpreting this new data type lags behind its ubiquity. While occupancy modelling software has dominated the analytical landscape for eDNA data analysis of single species, this type of model may not always be the most appropriate. The rate of eDNA detection often corresponds to species density, rather than just occupancy, and researchers often have access to observations from non‐genetic sampling methods at the same sites. To provide users access to a modelling framework designed to maximize the use of all available data, we developed an R package, eDNAjoint . The package provides an easy‐to‐use interface for fitting a ‘joint’ model that integrates data from paired or semi‐paired eDNA and traditional surveys in a Bayesian framework. The model can be used to estimate parameters like the probability of a false positive eDNA detection and mean catch rate at a site, and the package allows access to multiple model variations and Bayesian prior customization. Additional functionality can be used for model selection, summarising posteriors and comparing the relative sensitivities of the two survey methods. We demonstrate the use of eDNAjoint by fitting a variation of the model with site‐level covariates that scale the sensitivity of eDNA sampling relative to traditional sampling. The example workflow uses binary eDNA and seine count data for the endangered tidewater goby ( Eucyclogobius newberryi ) from a study by Schmelzle and Kinziger (2016). This use case includes a prior sensitivity analysis and an evaluation of the relationship between detection rates and environmental variables. eDNAjoint has the potential to greatly increase the range of users who will be able to rigorously analyse eDNA and traditional survey data in a Bayesian framework, understand if and how eDNA can improve monitoring practices, and gain confidence in the interpretability of eDNA data.

Keller, Abigail G. [Department of Environment Scie↗

Intrinsic Kinetics of Polyethylene Terephthalate Pyrolysis via Micropyrolysis and Multivariate Chromatographic Analysis

This study provides an in-depth investigation of the primary decomposition of polyethylene terephthalate (PET) via pyrolysis, employing an experimental-analytic workflow that integrates design of experiments (DoE), micropyrolysis coupled with comprehensive two-dimensional gas chromatography (GC×GC), and multivariate data analysis to verify intrinsic kinetic conditions and elucidate evolving product distributions for mapping key reaction pathways. Peaks that could not be identified using commercial spectral libraries were assigned using Mass Frontier simulations, enabling the identification of divinyl terephthalate, ethyl vinyl terephthalate, and 2-(benzoyloxy)ethyl vinyl terephthalate. A polar×polar (non-orthogonal) column set tailored for the detection of carboxylic acids enhanced the quantification of benzoic acid, 4-vinylbenzoic acid, 4-ethylbenzoic acid, and methylbenzoic acid by up to 6-fold relative to an orthogonal column combination (non-polar×mid-polar). Moreover, pyrolysis variables were systematically evaluated using a Box- Behnken design (BBD), encompassing pyrolysis temperature (500−600 °C), sample weight (50−150 μg), and carrier gas flow rate (100−300 mL min −1 ). Among these, pyrolysis temperature was the only statistically significant factor influencing product yields, ranging from 58.78 to 84.26 wt %. In contrast, neither the sample weight nor the carrier gas flow rate had a significant effect on product yields within the evaluated experimental space. At 600 °C, the major pyrolysis products were benzoic acid (up to 20.20 ± 1.46 wt %) and CO 2 (up to 21.28 ± 1.46 wt %), which can be produced through decarboxylation reactions. These findings underscore the critical importance of selecting appropriate analytical columns for the accurate quantification of heteroatomcontaining products such as carboxylic acids, which may otherwise be underestimated or undetected due to their reactivity with the stationary phase of non-polar and mid-polar columns, as well as other GC components. They also highlight the importance of selecting pyrolysis conditions for investigating the primary decomposition of PET under an isothermal kinetically limited regime.

aromatic compounds↗

RU-net for automatic characterization of TRISO fuel cross sections

During irradiation, phenomena such as kernel swelling and buffer densification may impact the performance of tristructural isotropic (TRISO) particle fuel. Post-irradiation microscopy is often used to identify these irradiation-induced morphologic changes. However, each fuel compact generally contains thousands of TRISO particles. Manually performing the work to get statistical information on these phenomena is cumbersome and subjective. Here, to reduce the subjectivity inherent in that process and to accelerate data analysis, we used convolutional neural networks (CNNs) to automatically segment cross-sectional images of microscopic TRISO layers. CNNs are a class of machine-learning algorithms specifically designed for processing structured grid data. They have gained popularity in recent years due to their remarkable performance in various computer vision tasks, including image classification, object detection, and image segmentation. In this research, we generated a large irradiated TRISO layer dataset with more than 2,000 microscopic images of cross-sectional TRISO particles and the corresponding annotated images. Based on these annotated images, we used different CNNs to automatically segment different TRISO layers. These CNNs include RU-Net (developed in this study), as well as three existing architectures: U-Net, Residual Network (ResNet), and Attention U-Net. The preliminary results show that the model based on RU-Net performs best in terms of Intersection over Union (IoU). Using CNN models, we can expedite the analysis of TRISO particle cross sections, significantly reducing the manual labor involved and improving the objectivity of the segmentation results.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Boundary Layer Structures Over the Northwest Atlantic Derived From Airborne High Spectral Resolution Lidar and Dropsonde Measurements During the ACTIVATE Campaign

The Planetary Boundary Layer height (PBLH) is essential for studying PBL and ocean-atmosphere interactions. Marine PBL is usually defined to include a mixed layer (ML) and a capping inversion layer. The ML height (MLH) estimated from the measurements of aerosol backscatter by a lidar was usually compared with PBLH determined from radiosondes/dropsondes in the past, as the PBLH is usually similar to MLH in nature. However, PBLH can be much greater than MLH for decoupled PBL. Here, in this study, we evaluate the retrieved MLH from an airborne lidar (HSRL-2) by utilizing 506 co-located dropsondes during the ACTIVATE field campaign over the Northwest Atlantic from 2020 to 2022. First, we define and determine the MLH and PBLH from the temperature and humidity profiles of each dropsonde, and find that the MLH values from HSRL-2 and dropsondes agree well with each other, with a coefficient of determination of 0.66 and median difference of 18 m. In contrast, the HSRL-2 MLH data do not correspond to dropsonde-derived PBLH, with a median difference of -47 m. Therefore, we modify the current operational and automated HSRL-2 wavelet-based algorithm for PBLH retrieval, decreasing the median difference significantly to -8 m. Further data analysis indicates that these conclusions remain the same for cases with higher or lower cloud fractions, and for decoupled PBLs. These results demonstrate the potential of using HSRL-2 aerosol backscatter data to estimate both marine MLH and PBLH and suggest that lidar-derived MLH should be compared with radiosonde/dropsonde-determined MLH (not PBLH) in general.

54 ENVIRONMENTAL SCIENCES↗

Compilation and utilization of a sorghum transcriptome compendium for gene regulatory network analysis and crop trait engineering

Sorghum bicolor (Sorghum) is a drought and heat tolerant C4 grass crop used to produce grain, forage, biofuels, and other bioproducts. Genetic improvement of sorghum hybrid crops is aided by a large and diverse germplasm, sorghum's diploid inbreeding genetics, and a relatively small genome that has facilitated genomic research. Over the past 20 years, the sorghum research community characterized the cytogenetic and recombinant landscapes of sorghum's 10 chromosomes, sequenced and annotated the sorghum genome, and used that information to identify genes/alleles that modulate flowering time, plant height, seed shattering, and other important traits. More recently, >1000 RNA-seq transcriptome profiles were collected from 15 sorghum genotypes to help understand the genetic basis of variation in growth and development of sorghum stems, tillers, roots, and leaves, and the regulation of biosynthetic pathways that produce epicuticular wax, dhurrin, and RFOs, compounds that contribute to sorghum's resilience. Transcriptome studies were designed to identify differentially expressed genes that are co-expressed during development or in response to a treatment to enable construction of gene regulatory networks. Co-expression and network analysis identified transcription factors and their cognate binding sites in target gene promoters and signaling pathways that modulate gene regulatory networks providing gene editing targets for further trait optimization. RNA-seq data from >20 experiments targeting sorghum organs, tissues, cell types, developmental stages, and responses to environmental conditions (i.e., diel, day-length, shading, water-deficit, temperature) has been compiled in a sorghum transcriptome compendium. The goal of this resource paper is to describe compendium content, accessibility, and a compendium data analysis pipeline and to illustrate the types of information that can be derived from the compendium with a focus on the elucidation of gene regulatory networks useful for guiding the improvement of sorghum traits through gene editing.

RNA-seq↗

Collaborative Research: Enhancing Laser-Based Ion Sources with High Data Rate Techniques

This collaborative research project focuses on leveraging advanced machine learning techniques to analyze and optimize data from high-repetition-rate laser experiments. The main goal is to apply modern computing hardware, customized data acquisition firmware/software, and machine learning approaches to improve data analysis and experimental control. The project also explores how methodology can be developed on smaller-scale experimental setups and then translated to larger facilities within DOE's LaserNetUS network. With extensive data collection and modeling, the research aims to predict and optimize experimental parameters to enhance performance and efficiency.

47 OTHER INSTRUMENTATION↗

MaPSA Quality Control and AI-Enhanced Grading For the CMS Phase-II Tracker Upgrade

The Compact Muon Solenoid (CMS) experiment will undergo changes as part of the Large Hadron Collider upgrade. The CMS tracker will be upgraded to cope with the new radiation environment and to provide tracking at the first level trigger. This upgrade features a new type of silicon module called PS Module, which combines a Pixel sensor and a Strip sensor in the same module. The pixel portion of the PS module has a sensor bump bonded to 16 Macro Pixel ASICs (MPA) to form a Macro Pixel Sub Assembly (MaPSA). At Fermilab, MaPSAs are tested for quality control before being assembled with the strip sensors, readout and service electronics to form a PS Module. All of this test data is stored in a centralized database, and is used to grade the final module to determine if it will be installed in the detector. The Phase II Outer Tracker Analyzer of Test Outputs (POTATO) is the software that processes this data and determines the module grades. Using recent technologies, an AI agent is being im plemented into POTATO in order to allow users to more efficiently sort through the large amounts of analysis data and ensure that only the user specified data is being considered. This poster will display the process of testing a MaPSA, how that test data is relevant to module assembly and grading, and how the POTATO grading tool is being improved with the use of an embedded AI agent.

Gzamouranis, Olivia [Purdue U.]↗

Direct inference of nuclear equation-of-state parameters from gravitational-wave observations

The observation of neutron star mergers with gravitational waves (GWs) has provided a new method to constrain the dense-matter equation of state (EOS) and to better understand its nuclear physics. However, inferring nuclear microphysics from GW observations necessitates the sampling of EOS model parameters that serve as input for each EOS used during the GW data analysis. The sampling of the EOS parameters requires solving the Tolman–Oppenheimer–Volkoff (TOV) equations a large number of times—a process that slows down each likelihood evaluation in the analysis on the order of a few seconds. Here, we employ emulators for the TOV equations built using multilayer perceptron neural networks to enable direct inference of nuclear EOS parameters from GW strain data. Our emulators allow us to rapidly solve the TOV equations, taking in EOS parameters and outputting the associated tidal deformability of a neutron star in only a few tens of milliseconds. We implement these emulators in PyCBC to directly infer the EOS parameters using the event GW170817, providing posteriors on these parameters informed solely by GWs. We benchmark these runs against analyses performed using the full TOV solver and find that the emulators achieve speed ups of nearly two orders of magnitude, with negligible differences in the recovered posteriors. Additionally, we constrain the slope and curvature of the symmetry energy at the 90% upper credible interval to be $L$ sym ≲ 106 MeV and $K$ sym ≲ 26 MeV.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

RU Net for Automatic Characterization of TRISO Fuel Cross Sections

TRistructural ISOtropic (TRISO) particle fuel is a type of nuclear fuel known for its high-temperature and high-burnup performance. Each sub-millimeter diameter TRISO particle consists of uranium-oxycarbide (UCO) or UO2 fuel kernel, coated with buffer, inner pyrolytic carbon (IPyC), silicon carbide (SiC), and outer pyrolytic carbon (OPyC) layers. The SiC layer acts as the main containment barrier for the TRISO particle to retain the fission products, while the IPyC and OPyC layers provide additional barriers to the release of fission products, especially fission gases. During irradiation, phenomena like kernel swelling, buffer densification, and IPyC fracture may impact fuel performance. Post-irradiation microscopy on entire compact cross sections or samples of individual particles deconsolidated from compacts is often used to identify these irradiation-induced changes in morphology. However, each fuel compact generally contains thousands of TRISO particles. To get statistical information on these phenomena, it is cumbersome work if done manually. For example, to get information about swelling/densification behaviors of different layers or kernels after irradiation, researchers previously manually measured the perimeter of each TRISO layer in hundreds of particles after four rounds of iterative grinding and polishing encompassing more than 2000 cross-section images for a total of four fuel compacts. To attempt to reduce the subjectivity inherent in that process and accelerate data analysis, we conducted a study on the automatic TRISO layer segmentation on cross-sectional microscopic images using Convolutional Neural Networks (CNNs). CNNs are a class of machine learning algorithms specifically designed for processing structured grid data that have gained popularity in recent years due to their remarkable performance in various computer vision tasks, including image classification, object detection, and image segmentation. In this research, we have generated the large irradiated TRISO layer dataset with more than 2000 cross-section TRISO microscopic images and the corresponding annotated images. Based on these annotated images, we have employed different CNNs for automatic segmentation of different TRISO layers. These include RU-Net (developed in this study), as well as three existing architectures: U-Net, Residual Network (ResNet), and Attention U-Net. The preliminary results show that the model based on RU-Net has the best performance in terms of intersection-over-union (IoU). Through the aid of these CNN models, we can expedite the analysis of TRISO particle cross-sections, significantly reducing the manual labor involved and improving the objectivity of the segmentation results.

Convolutional Neural Networks↗

Open-Source and FAIR Research Software for Proteomics

Scientific discovery relies on innovative software as much as experimental methods, especially in proteomics, where computational tools are essential for mass spectrometer setup, data analysis, and interpretation. Since the introduction of SEQUEST, proteomics software has grown into a complex ecosystem of algorithms, predictive models, and workflows, but the field faces challenges, including the increasing complexity of mass spectrometry data, limited reproducibility due to proprietary software, and difficulties integrating with other omics disciplines. Closed-source, platform-specific tools exacerbate these issues by restricting innovation, creating inefficiencies, and imposing hidden costs on the community. Open-source software (OSS), aligned with the FAIR Principles (Findable, Accessible, Interoperable, Reusable), offers a solution by promoting transparency, reproducibility, and community-driven development, which fosters collaboration and continuous improvement. In this manuscript, we explore the role of OSS in computational proteomics, its alignment with FAIR principles, and its potential to address challenges related to licensing, distribution, and standardization. Drawing on lessons from other omics fields, we present a vision for a future where OSS and FAIR principles underpin a transparent, accessible, and innovative proteomics community.

97 MATHEMATICS AND COMPUTING↗

SQuaD: Smart Quantum Detection for Photon Recognition and Dark Count Elimination

Quantum detectors of single photons are an essential component for quantum information processing across computing, communication and networking. Today's quantum detection system, which consists of single photon detectors, timing electronics, control and data processing software, is primarily used for counting the number of single photon detection events. However, it is largely incapable of extracting other rich physical characteristics of the detected photons, such as their wavelengths, polarization states, photon numbers, or temporal waveforms. This work, for the first time, demonstrates a smart quantum detection system, SQuaD, which integrates a field programmable gate array (FPGA) with a neural network model, and is designed to recognize the features of photons and to eliminate detector dark-count. The SQuaD is a fully integrated quantum system with high timing-resolution data acquisition, onboard multi-scale data analysis, intelligent feature recognition and extraction, and feedback-driven system control. Our \name experimentally demonstrates 1) reliable photon counting on par with the state-of-the art commercial systems; 2) high-throughput data processing for each individual detection events; 3) efficient dark count recognition and elimination; 4) up to 100% accurate feature recognition of photon wavelength and polarization. Additionally, we deploy the SQuaD to an atomic (erbium ion) photon emitter source to realize noise-free control and readout of a spin qubit in the telecom band, enabling critical advances in quantum networks and distributed quantum information processing.

Linne, Karl C. [U. Chicago (main)] (ORCID:00090009↗

Adsorptive denitrogenation of model aviation fuel using mesoporous silica in a packed bed adsorption system

This study aims to understand the effects of system process parameters such as flow rate, adsorbent particle size, and use of recycled adsorbent on denitrogenation performance of a model fuel using mesoporous silica gel. The goal is to reduce the nitrogen content of the model fuel from 1500 parts per million to single-digit ppm to meet ASTM specifications for drop-in fuels. This work was done with the intent of applying adsorptive denitrogenation to sustainable aviation fuel (SAF) product fractions produced via hydrothermal liquefaction (HTL). The adsorption performance of the silica is evaluated via packed column breakthrough data, with data generated from collecting from the column outlet and quantifying nitrogen content via gas chromatography. Select experiments use a significantly larger (2.5x column diameter and length) column to demonstrate linear scalability of the process. Thermogravimetric analysis data is collected to evaluate the effects of thermal calcination as a sorbent regeneration method. Effects of a more complex feed are also investigated using a known reference fuel with additional added nitrogen containing compounds. The results presented in this work successfully demonstrate up to 99.8 % removal of NCCs from a model fuel fraction at an original NCC concentration of approximately 1500 ppm to single-digit parts per million after treatment. We also examine calcination of sorbent materials to remove the adsorbed species to enable sorbent reuse and minimize waste generation and show that the calcined material can be reused up to 5 cycles with reduced adsorption capacity. Overall, this work indicates that adsorptive denitrogenation using silica gel is a viable solution to enable the integration of HTL-derived aviation fuels into existing fuel infrastructure.

Adsorption techniques↗

Tutorial: Extracting entanglement signatures from neutron spectroscopy

This tutorial is a pedagogical introduction to recent methods of computing quantum spin entanglement witnesses from spectroscopy, with a special focus on neutron scattering on quantum spin systems. We offer a brief introduction to the concepts and equations, define a data analysis protocol, and discuss the interpretation of three entanglement witnesses: one-tangle, two-tangle, and Quantum Fisher Information. We also discuss practical experimental considerations, and give three examples of extracting entanglement witnesses from experimental data: Copper Nitrate, KCuF 3 , and NiPS 3 .

47 OTHER INSTRUMENTATION↗