Search NASA⌕ Search

SEARCH · Search NASA

Results for “data similarity”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

A machine-learning-aided data recovery approach for predicting multi-material thermal behaviors in advanced test reactor capsules

Instrumented experiments conducted at test reactors are essential to the deployment of new advanced reactor systems. Designing new experiments and generating data on specific reactor conditions require significant investments in terms of both time and cost. Finite element analysis software can be used to create high-fidelity models of experiment environments in order to support the actual experiments, but computation time remains a concern in terms of applying outcomes to real-time usage of data (e.g., a digital twin [DT]). Here, the present research proposes a machine-learning (ML) aided approach to making temperature and displacement predictions based on the thickness of the outer gas gap on the experimental capsule used for in-pile demonstration of a novel new thermal conductivity probe in the Advanced Test Reactor (ATR). This capsule consisted of U10Zr fuel, a rodlet, sodium, and inner and outer capsules. Gas gaps existed between the fuel and the rodlet, and between the inner and the outer capsule. The learning data pertained to an experimental capsule's radial distributions of temperature and displacement, as obtained based on Abaqus and the physical features. For the first step of ML sequence, the temperature was predicted using three positional parameters. Next, the displacement was predicted using seven additional parameters. Each physical feature was normalized in order to be both nondimensional and standardized. The temperature and displacement predictions showed good agreement with the simulation results in all cases involving interpolation and extrapolation. Furthermore, data similarity enhancement increased the similarity between the training and the target data, thereby increasing the predictive accuracy of the ML models. In certain extrapolation cases involving limited original ML model accuracy, data similarity enhancement and data recovery was able to somewhat improve this accuracy.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Machine-Learning-aided Approach for Predicting the Thermal Expansion Behaviors in Advanced Test Reactor Capsules (NURETH-20 full paper)

Instrumented experiments at test reactors are essential to deploying new advanced reactor systems. Designing new experiments and generating data on specific conditions require both time and cost investment. A high-fidelity model of the experiment environment can be created using finite element analysis software to support the actual experiments, but computation time is still a concern in applying outcomes to real-time usage (e.g., a digital twin). This research proposes a machine-learning-aided approach to temperature and displacement predictions, based on the thickness of the outer gas gap on the experimental capsule used for the in-pile demonstration of a novel thermal conductivity probe in the Advanced Test Reactor. The capsule consisted of U10Zr fuel, a rodlet, sodium, and inner and outer capsules. There were gas gaps between the fuel and rodlet and between the inner and outer capsule. The learning data consisted of an experimental capsule’s radial distributions of temperature and displacement, as obtained from Abaqus and the physical features. For the first step, temperature was predicted using three positional parameters. Then the displacement was predicted using six different positional parameters. Each physical feature was normalized to be both nondimensional and standardized. The temperature and displacement predictions showed good agreement in all cases involving interpolation and extrapolation. Also, data similarity enhancement increased the similarity between training and target data increasing the predictive accuracy of machine-learning models. In some cases of extrapolation, the accuracy of the machine-learning model showed limited performance, but still data similarity enhancement improved the accuracy.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Comparison of pyrometry and thermography for thermal analysis of thermite reactions

This article examines the thermal behavior of a laser ignited thermite composed of aluminum and bismuth trioxide. Temperature data were collected during the reaction using a four-color pyrometer and a high-speed color camera modified for thermography. The two diagnostics were arranged to collect data simultaneously, with similar fields of view and with similar data acquisition rates, so that the two techniques could be directly compared. Results show that at initial and final stages of the reaction, a lower signal-to-noise ratio affects the accuracy of the measured temperatures. Both diagnostics captured the same trends in transient thermal behavior, but the average temperatures measured with thermography were about 750 K higher than those from the pyrometer. This difference was attributed to the lower dynamic range of the thermography camera’s image sensor, which was unable to resolve cooler temperatures in the field of view as well as the photomultiplier tube sensors in the pyrometer. Overall, while the camera could not accurately capture the average temperature of a scene, its ability to capture peak temperatures and spatial data make it the preferred method for tracking thermal behavior in thermite reactions.

36 MATERIALS SCIENCE↗

CMS Storage Performance with RNTuple

CMS is transitioning to use ROOT’s new RNTuple data storage format for the files CMS will write in the HL-LHC era. Based on initial tests, CMS expects faster I/O and smaller files compared to the present TTree storage format. This contribution will show a comprehensive performance comparison between RNTuple and TTree I/O using CMS AOD and MiniAOD data formats as test cases for both simulation and collision data corresponding to similar data taking conditions of LHC Run 3. Quantities such as the resulting file size, the memory usage of the I/O components, and the rate of events being read from a file or written to a file will be measured. CMS’ data processing relies heavily on reading files over the local or wide area networks. The file read patterns are important because the latencies have been seen to influence the total production job times. Therefore a study on the file read patterns will be conducted by recording traces of the offset, size, and timestamp of each read request for both RNTuple and TTree. The behavior of network reads will be mimicked by reading local files where artificial latency will be added to the read requests. The effect of different latency values on the job times will be studied.

Jones, Christopher D. [Fermilab]↗

Distinguishing isotropic and anisotropic signals for X-ray total scattering using machine learning

Understanding structure–property relationships is essential for advancing technologies based on thin films. X-ray pair distribution function (PDF) analysis can access relevant atomic structure details spanning local-, mid- and long-range structure. While X-ray PDF has been adapted for thin films on amorphous substrates, measurements on single-crystal substrates are necessary to accurately determine structure origins for some thin film materials, especially those for which the substrate changes the accessible structure and properties. However, when measuring films on single-crystal substrates, high-intensity anisotropic Bragg spots saturate 2D detector images, overshadowing the thin films' isotropic scattering signal. This renders previous data processing methods for films on amorphous substrates unsuitable for films on single-crystal substrates. To address this measurement need, we developed IsoDAT2D, an innovative data processing approach using unsupervised machine learning algorithms. The program combines dimensionality reduction and clustering algorithms to separate thin film and single-crystal substrate X-ray scattering signals. We use SimDAT2D , a program we developed to generate simulated thin film data, to validate IsoDAT2D . Here we also use IsoDAT2D to isolate X-ray total scattering signal from a thin film on a single-crystal substrate. The resulting PDF data are compared with similar data processed using previous methods, especially substrate subtraction for single-crystal and amorphous substrates. PDF data from IsoDAT2D -identified X-ray total scattering data are significantly better than from single-crystal substrate subtraction, but not as reliable as PDF data from amorphous substrate subtraction. With IsoDAT2D , there are new opportunities to expand PDF to a wider variety of thin films, including those on single-crystal substrates, with which new structure–property relationships can be elucidated to enable fundamental understanding and technological advances.

36 MATERIALS SCIENCE↗

Lightweight LSTM for CAN Signal Decoding

This paper describes an approach to identify undecoded Controller Area Network (CAN) data from one vehicle, based on the data similarity to previously decoded CAN data from another vehicle. Modern vehicles communicate data and signals from on-board sensors and controllers through the CAN bus. Networked sensors contain information such as wheel speeds, fuel gauges, turn signals, and radar signals. In the effort to use this information and make cars safer through human-in-the-loop CPS, signals on the CAN bus such as wheel speed and radar can be used to support the driver. However, data from the CAN bus are encoded and in some cases compressed, and different car manufacturers use different encoding schemes to represent data on the CAN bus. With hundreds of messages and thousands of possible encoding schemes to consider, it is laborious to identify the unique bits and encoding schemes that represent signals on each vehicle. In this study, we propose a method for training a Long Short-Term Memory (LSTM) neural network on known radar signals from one vehicle manufacturer, a Toyota, and successfully apply the network to identify the encoding for radar signals on a different vehicle, a Honda. By augmenting the training dataset with varied encoding bit boundaries, a small and lightweight LSTM network can learn to recognize radar data across different encoding schemes. The results are an improvement on exhaustive-search algorithms and other methods previously used in the search for such signals.

Ngo, Paul↗

Real-Time, Simultaneous Soil Water Content and Meteorological Data Measurement to Support TRACER over Harris County, Texas (Field Campaign Report: Part I)

The main purpose of this project was to provide ground-truth and satellite-based soil water content data in the Houston, Texas, area to support the U.S. Department of Energy Atmospheric Radiation Measurement (ARM) user facility’s 2022 field campaign, the Tracking Aerosol Convection Interactions ExpeRiment (TRACER). This involved installing four soil monitoring stations at different locations in the Houston area, two of which were part of the ARM facility and two were ancillary. The two stations associated directly with ARM were installed alongside other facilities managed and maintained by the TRACER research team during the intensive operational period (IOP), one located at the La Porte, Texas airport and the other near Guy, Texas. Data from these stations were assimilated with similar data collected by the Harris County Flood Control District (HCFCD) to improve the spatial coverage of the real-time monitoring network. The ground-truth data were compared to the satellite-based (National Aeronautics and Space Administration [NASA]’s Soil Moisture Active Passive [SMAP] mission) data in the TRACER area of interest, and analyzed further to nowcast gridded data over the Houston area. The nowcasted, gridded data product was made available to TRACER researchers for use in climate and land-atmosphere interaction modeling that can help understand formation and persistence of convective storms in dense urban areas, and predict possible environmental events such as floods.

54 ENVIRONMENTAL SCIENCES↗

Real-Time, Simultaneous Soil Water Content and Meteorological Data Measurement to Support TRACER over Harris County, Texas Field Campaign Report: Part II

The main purpose of this project was to provide ground-truth and satellite-based soil water content data in the Houston, Texas, area to support the U.S. Department of Energy Atmospheric Radiation Measurement (ARM) user facility’s 2022 field campaign, the Tracking Aerosol Convection Interactions ExpeRiment (TRACER). This involved installing four soil monitoring stations at different locations in the Houston area, two of which were part of the ARM facility and two were ancillary. The two stations associated directly with ARM were installed alongside other facilities managed and maintained by the TRACER research team during the intensive operational period (IOP), one located at the La Porte, Texas airport and the other located near Guy, Texas. Data from these stations were assimilated with similar data collected by the Harris County Flood Control District (HCFCD) to improve the spatial coverage of the real-time monitoring network. The ground-truth data were compared to the satellite-based (National Aeronautics and Space Administration [NASA]’s Soil Moisture Active Passive [SMAP] mission) data in the TRACER area of interest, and analyzed further to nowcast gridded data over the Houston area. The nowcasted, gridded data product was made available to TRACER researchers for use in climate and land-atmosphere interaction modeling that can help understand formation and persistence of convective storms in dense urban areas, and predict possible environmental events such as floods.

54 ENVIRONMENTAL SCIENCES↗

ESS-DIVE Unoccupied Aerial Systems (UAS) Reporting Format v1

Here we present documentation of the ESS-DIVE reporting format for Unoccupied Aerial System (UAS) data and metadata. This reporting format provides guidance to data contributors on how to store data to maximize their discoverability, facilitate their efficient reuse, and add value to individual datasets. For data users, the reporting format will better allow data repositories to optimize data search and extraction, and more readily integrate similar data into harmonized synthesis products. The reporting format provides templates and guidance for the reporting of metadata for UAS experimental campaigns, individual flights, platform and sensor description. To improve data access and discoverability, the reporting format proposes a data description scheme of Levels based on the degree of processing, where Level 0 includes raw data, through to Level 3 being derived data end products. A range of examples of data types for each Level are given, with suggested file naming schemes. The reporting format presented here is intended to form a foundation for future development that will accommodate new UAS technologies and approaches to data access and use in the future. The reporting format documentation is maintained and updated on the ESS-DIVE Community Space GitHub at https://github.com/ess-dive-community/essdive-uas. This data package is the first published version of this reporting format, and comprises a zip file of the complete content of https://github.com/ess-dive-community/essdive-uas v1.0. The zip contains the reporting format description, instructions and variable definitions in GitHub markdown language (*.md) and metadata templates in csv format. The reporting format is designed to be compatible with other ESS-DIVE formats, and it is specifically recommended that this reporting format be used in conjunction with the File-level metadata (FLMD) and comma separated values (csv) reporting formats for submission to the ESS-DIVE repository.

54 ENVIRONMENTAL SCIENCES↗

The Data Synergy Effects of Time-Series Deep Learning Models in Hydrology

When fitting statistical models to variables in geoscientific disciplines such as hydrology, it is a customary practice to stratify a large domain into multiple regions (or regimes) and study each region separately. Traditional wisdom suggests that models built for each region separately will have higher performance because of homogeneity within each region. However, each stratified model has access to fewer and less diverse data points. Here, through two hydrologic examples (soil moisture and streamflow), we show that conventional wisdom may no longer hold in the era of big data and deep learning (DL). We systematically examined an effect we call data synergy, where the results of the DL models improved when data were pooled together from characteristically different regions. The performance of the DL models benefited from modest diversity in the training data compared to a homogeneous training set, even with similar data quantity. Moreover, allowing heterogeneous training data makes eligible much larger training datasets, which is an inherent advantage of DL. A large, diverse data set is advantageous in terms of representing extreme events and future scenarios, which has strong implications for climate change impact assessment. The results here suggest the research community should place greater emphasis on data sharing.

54 ENVIRONMENTAL SCIENCES↗

Models for synthetic data generation

The software includes a suite of probabilistic statistical/machine learning models that can generate discrete synthetic data. Each model is trained on a set of real (private) data and then it can be used to generate synthetic but statistically similar data. Once ready, the model can generate as many samples as we want. Finally, in addition to the actual models, the software includes code to process data, evaluate results (based on cross validation), and produce reports.

De Oliveira Sales, Ana Paula↗

Neutral Loss Mass Spectral Data Enhances Molecular Similarity Analysis in $\mathrm{METLIN}$

We report Neutral loss (NL) spectral data presents a mirror of MS 2 data and is a valuable yet largely untapped resource for molecular discovery and similarity analysis. Tandem mass spectrometry (MS 2 ) data is effective for the identification of known molecules and the putative identification of novel, previously uncharacterized molecules (unknowns). Yet, MS 2 data alone is limited in characterizing structurally related molecules. To facilitate unknown identification and complement the METLIN-MS 2 fragment ion database for characterizing structurally related molecules, we have created a MS 2 to NL converter as a part of the METLIN platform. The converter has been used to transform METLIN’s MS 2 data into a neutral loss database (METLIN-NL) on over 860 000 individual molecular standards. The platform includes both the MS 2 to NL converter and a graphical user interface enabling comparative analyses between MS 2 and NL data. Examples of NL spectral data are shown with oxylipin analogues and two structurally related statin molecules to demonstrate NL spectra and their ability to help characterize structural similarity. Mirroring MS 2 data to generate NL spectral data offers a unique dimension for chemical and metabolite structure characterization.

59 BASIC BIOLOGICAL SCIENCES↗

Methodology for Digital Image Correlation and Infrared Measurement of Melting Aluminum Bars

Ultimately, our experiment measures two quantities on an aluminum bar: motion (which modeling must predict) and temperature (which sets thermal boundary conditions). For motion, stereo DIC is a technique to use imaging data to provide displacements relative to a reference image down to 1/100th of a pixel. We use a calibrated infrared imaging method for accurate temperature measurements. We will be capturing simultaneous data and then registering temperature data in space to the same coordinate system as the displacement data. While we will later show that our experiments are repeatable, indicating that separate experiments for motion and temperature would provide similar data, the simultaneous and registered data removes test to test variability as a source of uncertainty for model calibration and reduces the number of time-consuming tests that must be performed.

36 MATERIALS SCIENCE↗

Similarity Metric for Data Optimization and Efficient Training of Reactive Machine Learning Force Fields for Hydrocarbon Radiolysis

Radiolysis is a common approach to sterilize polymers, chemically modify them for upcycling, and accelerate their decomposition for recycling purposes. Reactive molecular dynamics (MD) simulations provide a powerful tool to generate atomic-level trajectories of the reactive processes and quantify radiolytic chemical degradation pathways. For this, machine learning (ML) surrogate models for reactive force fields with quantum mechanical accuracy are now widely used, which require ML training data sets that can provide information on atomic environments for target chemical systems. However, radiolysis chemistry can be highly complex and diverse, which poses significant challenges for generating training data to parametrize ML models. In this regard, we developed a method for optimizing the training data set using a cosine similarity metric to help guide training set selection for radiolysis of polyethylene, a model hydrocarbon polymer, as well as to enhance the transferability of our reactive ML force field (MLFF) to a variety of molecular and polymeric systems. Our approach performs atom-by-atom comparisons between local atomic environments to pinpoint important data points associated with rare and localized events, such as radiolysis damage within structures. We apply this approach to train the Chebyshev Interaction Model for Efficient Simulation (ChIMES) MLFF model, which expresses the atomic interaction potentials in terms of linear combinations of many-body Chebyshev polynomials. We first show that our method can reduce our training set size by ∼70% while improving overall accuracy compared to more standard MD model fitting approaches. We then validate our optimum model against diverse hydrocarbon simulation data, including simple alkanes and systems with unsaturated carbon bonds, over a wide range of thermodynamic conditions. Finally, we use our ChIMES model to perform MD simulations of radiolytic damage with large-scale systems that help avoid system size effects. Overall, our approach yields an MD force field that retains most of the accuracy of the underlying quantum method while yielding many orders of improvement in computational efficiency. In conclusion, our efforts will have impact on future hydrocarbon polymer radiolysis studies, where the chemical details of the polymer–radiation interactions can have a strong effect on the resulting products observed in experiments.

Hydrocarbons↗

Assessment of Potential Ergonomic Injury Risk in the Nuclear Material Processing Glovebox Environment [Capstone Project]

Gloveboxes are isolation barriers that are used within many different industries such as pharmaceuticals, electronic parts fabrication, nuclear, and biological. The research on glovebox ergonomics is currently limited with few ergonomic professionals that focusing exclusively on glovebox working environments. There is existing documentation on occupational injuries, such as musculoskeletal disorders (MSD), sustained as a direct result of working in gloveboxes. The main elements that pose ergonomic risks to glovebox workers are operational repetition, duration, force, vibration, lifting heavy (more than 15lbs with two hands) objects, and awkward postures. This paper examines a small sample of the potential causal or risk factors that lead to the ergonomic injuries. A meta-analysis utilizing a random effect model is used to examine data from several studies which focus on, dexterity and strength changes as result of glove thickness, and robotic assistive technology as a means to improve postural mechanics. With the meta-analysis technique, similar data sets from different research studies can be coalesced into a single weighted, statistically significant result for subsequent consideration. The results from the analysis show that increased glove thickness results in decreased dexterity for operators thus increasing ergonomic risk factors such as task duration. It is also shown that glove thickness decreases grip strength but a similar decrease in pinch strength is not definitively demonstrated. With decreased grip strength, operators will need to exert more force (a known ergonomic risk factor) on processing tools, etc. during operations thus increasing the risk of ergonomic injury. The use of robotic assistive technology as a means to improve operator posture (risk factor) was also examined. Although it may be intuitively assumed that human-robotic collaboration would be beneficial in reducing risk factors, the result from this study’s analysis was not statistically significant. It is inferred that with additional directed research on this topic another study/analysis could be statistically significant demonstrating the benefits of the technology.

99 GENERAL AND MISCELLANEOUS↗

FY23 Progress Report: SRNL Analysis of ICCWR LCM and WAMS data for Corrosion and Cracking

Algorithms for Machine Learning (ML) and data analysis for the 3013 Surveillance Program have been developed in an ongoing collaborative effort by the Savannah River National Laboratory (SRNL) and the University of South Carolina (USC). The objective of the algorithms is to automate the identification of corrosion and crack formation in the Inner Container Closure Weld Region (ICCWR) of the canister system used to store Pu-bearing material. Data for corrosion and cracking is collected from large binary files generated by a Laser Confocal Microscope (LCM), the Wide Area 3D Measurement System (WAMS), and in a recent proposal, by a Scanning Electron Microscope (SEM). The ML software uses the physical attributes in the data files (e.g., one or all of: height, color, and grayscale values as functions of position in a plane projection) to detect the presence of surface corrosion and cracking after being trained on similar data with the features to be detected labeled.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

FY24 Progress Report: SRNL Analysis of ICCWR LCM and WAMS data for Corrosion and Cracking

Algorithms for Machine Learning (ML) and data analysis for the 3013 Surveillance Program have been developed in an ongoing collaborative effort by the Savannah River National Laboratory (SRNL) and the University of South Carolina (USC). The objective of the algorithms is to automate the identification of corrosion and crack formation in the Inner Container Closure Weld Region (ICCWR) of the canister system used to store Pu-bearing material. Data for corrosion and cracking is collected from large binary files generated by a Laser Confocal Microscope (LCM), the Wide Area 3D Measurement System (WAMS), or,in a recent proposal, by a Scanning Electron Microscope (SEM). The ML software uses the physical attributes in the data files (e.g., one or more of: height, color, and 16-bit grayscale values as functions of position in a plane projection) to detect signs of surface corrosion and cracking after being trained on similar data, with the features to be detected. Although the initial scope included screening for broader indicators of corrosion, e.g., pitting, identification of potential cracks was prioritized for the past several years at the request of program leadership. Labeled training data is essential to developing the ML algorithm, and enhancements to data labeling capability have been developed to address this essential precursor to application of ML routines. Efficient labeling is particularly important in view of the large volume of data required to train ML algorithms and the relative rarity of cracks in the ICCWR data set. The updated program will read binary data from either LCM, WAMS or SEM files, interrogate data attributes, facilitate user labeling of data for training ML algorithms, execute ML algorithms, output parameters from trained ML algorithms, report ML model accuracy with respect to labeled data, and generate graphical representations for various analyses. In FY24, hourglass neural networks (HNNs) that were initiated in FY22 were further developed and tested using available LCM data, and their performance was tested against that of the alternative U-Net Neural Network algorithm structure. HNNs along with previously developed Convolutional Neural Networks (CNNs) and Deep Neural Networks (DNNs) comprise a suite of ML tools for identification of cracks in the ICCWR

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

FY25 Progress Report: SRNL Analysis of ICCWR LCM and WAMS data for Corrosion and Cracking

Algorithms for Machine Learning (ML) and image analysis for the 3013 Surveillance Program have been developed in an ongoing collaborative effort by the Savannah River National Laboratory (SRNL) and the University of South Carolina (USC). The objective of the algorithms is to automate the identification of corrosion and cracks in the Inner Container Closure Weld Region (ICCWR) of the canister system used to store Pu-bearing material. Data for corrosion and cracking is collected from large binary files generated by a Laser Confocal Microscope (LCM), the Wide Area 3D Measurement System (WAMS), or, in a recent proposal, by a Scanning Electron Microscope (SEM). The ML software uses the physical attributes in the data files (e.g., one or more of: height, color, and 16-bit grayscale values as functions of position in a plane projection) to detect signs of surface corrosion and cracking after being trained on similar data with the features to be detected. Although the initial scope included screening for broader indicators of corrosion, e.g., pitting, the identification of potential cracks was prioritized for the past several years at the request of program leadership.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗