Search NASA⌕ Search

SEARCH · Search NASA

Results for “training data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

Temporal Forecasting of Distributed Temperature Sensing in a Thermal Hydraulic System With Machine Learning and Statistical Models

We benchmark performance of long-short term memory (LSTM) network machine learning model and autoregressive integrated moving average (ARIMA) statistical model in temporal forecasting of distributed temperature sensing (DTS). Data in this study consists of fluid temperature transient measured with two co-located Rayleigh scattering fiber optic sensors (FOS) in a forced convection mixing zone of a thermal tee. We treat each gauge of a FOS as an independent temperature sensor. We first study prediction of DTS time series using Vanilla LSTM and ARIMA models trained on prior history of the same FOS that is used for testing. The results yield maximum absolute percentage error (MaxAPE) and root mean squared percentage error (RMSPE) of 1.58% and 0.06% for ARIMA, and 3.14% and 0.44% for LSTM, respectively. Next, we investigate zero-shot forecasting (ZSF) with LSTM and ARIMA trained on history of the co-located FOS only, which is advantageous when limited training data is available. The ZSF MaxAPE and RMSPE values for ARIMA are comparable to those of the Vanilla use case, while the error values for LSTM increase. We show that in ZSF, performance of LSTM network can be improved by training on most correlated gauges between the two FOS, which are identified by calculating the Pearson correlation coefficient. The improved ZSF MaxAPE and RMSPE for LSTM are 4.4% and 0.33%, respectively. Performance of ZSF LSTM can be further enhanced through transfer learning (TL), where LSTM is re-trained on a subset of the FOS that is the target of forecasting. We show that LSTM pre-trained on correlated dataset and re-trained on 30% of testing target dataset achieves MaxAPE and RMSPE values of 2.32% and 0.28%, respectively.

ARIMA↗

PySIDT: Subgraph Isomorphic Decision Trees for Molecular Property Prediction

Accurate molecular property prediction is important across all fields of chemistry. Deep neural networks (DNNs) have become increasingly popular due to their ability to train automatically, avoiding the incredibly tedious process of constructing and extending traditional property estimation schemes. However, DNNs require large amounts of training data, are challenging to interpret, require large amounts of memory to load even during inference, and have severe difficulties incorporating qualitative chemical knowledge, which are often desired for molecular property prediction tasks. Here, in this study, we present PySIDT (https://github.com/zadorlab/PySIDT), a software for training and running inference on Subgraph Isomorphic Decision Trees (SIDTs). SIDTs are graph-based decision trees made of nodes associated with molecular substructures. Inference is done by descending target molecular structures down the decision tree to nodes with matching subgraph isomorphic substructures and making predictions based on the final (most specific) nodes matched. SIDTs scale down well to dataset sizes much smaller than is feasible for DNNs. As trees of molecular substructures, SIDTs are inherently readable and easy to visualize, making them easy to analyze. They are also straightforward to extend and retrain, facilitate uncertainty estimation, and enable easy integration of expert knowledge. We demonstrate the SIDT approach discussing its application to a diverse range of molecular prediction tasks: rate coefficient estimation, diffusion coefficient estimation, thermochemistry estimation, transition state bond stretch prediction, p K a prediction, stability of molecular structures, stability of surface structures, and prediction of surface lateral interaction energetics. Additionally, we demonstrate the power of the SIDT algorithms in two direct learning curve vanilla comparisons with the popular DNN-based software Chemprop and the popular gradient boosted trees-based software XGBoost on enthalpy of formation and rate coefficient prediction tasks. In particular, in the enthalpy of formation case, vanilla PySIDT is able to outperform vanilla Chemprop and XGBoost across the full range of training/validation set sizes out to 11,560 data points.

Johnson, Matthew Sean [Sandia National Laboratorie↗

Quantum Annealing for Real-World Machine Learning Applications

Optimizing the training of a machine learning pipeline is important for reducing training costs and improving model performance. One such optimizing strategy is quantum annealing, which is an emerging computing paradigm that has shown potential in optimizing the training of a machine learning model. The implementation of a physical quantum annealer has been realized by D-Wave systems and is available to the research community for experiments. Recent experimental results on a variety of machine learning applications have shown interesting results especially under the conditions where the performance of classical machine learning techniques are limited such as limited training data and high dimensional features. This chapter explores the application of D-Wave’s quantum annealer for optimizing machine learning pipelines for real-world classification problems. We review the application domains on which a physical quantum annealer has been used to train machine learning classifiers. We discuss and analyze the experiments performed on the D-Wave quantum annealer for applications such as image recognition, remote sensing imagery, security, computational biology, biomedical sciences, and physics. We discuss the possible advantages and the problems for which quantum annealing is likely to be advantageous over classical computation.

Kumar nath, Rajdeep↗

Biomass Harmonization and SAR Analysis with the Multi-mission Algorithm and Analysis Platform (MAAP)

The Multi‐mission Algorithm and Analysis Platform (MAAP) is a collaborative effort between NASA and the European Space Agency (ESA) to support above ground biomass (AGB) research in an open science framework. MAAP brings together relevant data, algorithms, and computing capabilities in a common cloud environment to address the challenges of sharing and processing data from field, airborne and satellite measurements. MAAP was publicly released in October 2021, providing computing capabilities co-located with the data, a collaborative coding and analysis environment, and a set of interoperable tools and algorithms developed to support the estimation and visualization of data. MAAP has allowed scientists from both North America and Europe to collaborate on the generation and analysis/visualization of data derived from multiple, discipline-adjacent missions in an open, collaborative environment that has reached beyond traditional scientific investigation. MAAP has been used to support multiple scientific activities. To date, existing LiDAR data from multiple platforms has been calibrated with field measurements and combined for more comprehensive and accurate estimates of above ground biomass AGB; these LiDAR platforms include airborne (e.g. LVIS), the International Space Station (NASA’s Global Ecosystem Dynamics Investigation (GEDI), and satellites (e.g. ICESat-2). The current challenge is to effectively and seamlessly combine the aforementioned LiDAR-based data with new data sources such as P-band RADAR from ESA’s upcoming BIOMASS mission, existing ESA Sentinel-1 C-band SAR, and the 30 PB/yr of high cadence global coverage L-band SAR data from the upcoming NASA-ISRO SAR (NISAR) mission. Recent analysis using MAAP merged ICESat-2 and optical data (Harmonized Landsat Sentinel) produced the most comprehensively precise estimate of boreal-wide AGB to date. Another effort using MAAP is the production and open distribution of global comparisons of AGB map estimates, including from ICESat-2 and GEDI, to bolster stakeholder uptake for policy applications. These map estimates will feed into the Intergovernmental Panel on Climate Change (IPCC) database, likely aiding the next Global Carbon Stocktake of the UNFCCC. Furthermore, the biomass retrieval intercomparison exercise BRIX-2 could benefit from the MAAP providing standardized test cases (based on airborne campaign and spaceborne data) allowing the community to develop and apply retrieval algorithms based on these test cases, while forthcoming SAR data training curricula could also use the MAAP as a teaching and learning platform. The MAAP is meeting the challenges inherent in international, open science collaboration and large scale computing with a platform that is entirely open source and cloud native, using open standards for data access, manipulation, protocols, and formats. The MAAP data system consists of a dedicated data store whose data is indexed in an online catalog conforming to established metadata, application programmatic interfaces (APIs), and service interface standards, using an implementation of the open sourced NASA Common Metadata Repository. Federation of user identities allows users from either NASA or ESA to access and consume services from the other using a unified metadata catalog for the data utilized across the ESA and NASA MAAP platforms. Similarly, we are exploring how to increase interoperability to achieve a common approach to packaging, orchestrating and executing algorithms, with interoperable access to data for subsetting, fast browse, and cloud-optimized access, all using interoperable standards such as those from the Open Geospatial Consortium (OGC). Designed for interoperability, ESA and NASA utilize a common architecture for the software platform. It provides a cloud-based algorithm development environment (ADE) that enables scientists to develop algorithms collaboratively with access to the MAAP data catalog as well as other data archives. MAAP provides an Eclipse Che-based ADE supporting both Python and R languages, popular in this biomass community. Algorithms developed and containerized within the ADE can be deployed to run to thousands of computational nodes in the MAAP’s data processing system (DPS), dramatically speeding up processing and giving scientists a rapid, iterative turnaround of results. NASA’s implementation of the DPS is based on the Hybrid Science Data System (HySDS) framework, used by NASA flight projects to produce Earth science standard products.

cloud computing↗

Augmenting RANS Turbulence Models Guided by Field Inversion and Machine Learning

This report investigates the use of a data-driven approach, viz., Field Inversion and Machine Learning (FIML), to improve conventional RANS turbulence models like the Spalart-Allmaras model and the Menter SST k-ω model. One of the crucial aspects of using an ML-based approach with limited training data to produce corrections that are generalizable to a large range of flow configurations is to design appropriate “features” (inputs to the ML model). A model, based on guidance from the FIML methodology, is presented in analytical form. An additional list of potential features is provided. Although these were not used in the present correction, they were considered in the course of its development, and are included to fully document the complete process employed in the present work.

turbulence modeling↗

Soil, water, and vegetation conditions in south Texas

The author has identified the following significant results. The best wavelengths in the 0.4 to 2.5 micron interval were determined for detecting lead toxicity and ozone damage, distinguishing succulent from woody species, and detecting silverleaf sunflower. A perpendicular vegetation index, a measure of the distance from the soil background line, in MSS 5 and MSS 7 data space, of pixels containing vegetation was developed and tested as an indicator of vegetation development and crop vigor. A table lookup procedure was devised that permits rapid identification of soil background and green biomass or phenological development in LANDSAT scenes without the need for training data.

Wiegand, C. L.↗

Integrating very-high-resolution imagery, Sentinel-2 time-series data, and machine learning to map shrub fractional abundance across arid and semi-arid ecosystems in China

Shrub fractional abundance (SFA), the proportion of shrub cover per unit area, serves as a critical indicator of environmental aridity and ecosystem health in arid and semi-arid regions, particularly across the Mongolian steppe. However, large-scale SFA mapping in Mongolian steppe ecosystems remains challenging due to the small crown size of shrubs, their sparse distribution, and spectral overlap with coexisting low vegetation (e.g., grasses and herbs), which hinders accurate detection using coarser-resolution satellite data or traditional field surveys. To address these challenges, we developed a two-step approach that integrates very-high-resolution (VHR) imagery, time-series Sentinel-2 data, and deep learning techniques. First, we generated high-accuracy benchmark maps of individual shrub crowns from 0.5 m VHR imagery by combining manual segmentation with a hybrid deep learning framework (Dino V2 and convolutional neural networks). Second, we used these shrub crown maps as training data to build an XGBoost model for predicting SFA from 20 m Sentinel-2 time-series data, leveraging phenological information to improve estimation. We validated our approach across 70 sites (1km 2 each) in the Inner Mongolia Autonomous Region, which is representative of Mongolian steppe ecosystems. From VHR imagery, we mapped 1.31 million shrub crowns with an accuracy of R 2 = 0.92. Scaling up with Sentinel-2 data yielded regional SFA maps with an R 2 = 0.60. Further SHAP (SHapley Additive exPlanations) analysis on the developed XGBoost model revealed that phenological metrics (particularly observations in early-May, mid-July, and late-September), which distinguish shrub phenology from that of other land cover types (e.g., grasses and bare soil), were the most influential predictors of SFA. Finally, our regional SFA maps uncovered unimodal relationships between shrub distribution and climate variables, peaking at mean annual minimum temperatures near 0 °C and annual precipitation around 200 mm. Collectively, these findings demonstrate how the integration of multi-source remote sensing and machine learning can overcome historical limitations in SFA mapping, enabling accurate, spatially continuous assessments across vast Inner-Mongolian steppe ecosystems. Our framework has the potential to be applied to other steppe ecosystems and dryland ecosystems across the Mongolian steppe and beyond, offering a foundation for improved monitoring and ecological impact assessments in the face of global climate changes.

Arid and semi-arid landscapes↗

Dynamic data-driven multiscale modeling for predicting the degradation of a 316L stainless steel nuclear cladding material

Here, we have developed a long short-term memory stacked ensemble (LSTM-SE) surrogate modeling approach that can provide rapid predictions of microstructural evolution and the resultant mechanical properties of American Iron and Steel Institute (AISI) 316L series stainless steel (316LSS) fuel cladding under conditions of varying temperature and radiation dose rate. To acquire training data, we developed and implemented a kinetic Monte Carlo (KMC) model to simulate precipitation kinetics of M 23 C 6 , γ', and G phases within SS316L cladding. Experimentally reported precipitation kinetics of SS316L in literature were linked to the kinetic parameters of the simulated precipitation in our KMC model. The model was then used to simulate microstructure evolution under synthetically generated treatments of varying temperature and radiation dose rate, for periods of up to 3000 hours. Changes in volume fraction, number density, and particle size of precipitates were recorded, and particle area fractions were correlated using statistical methods to develop the surrogate model. Simultaneously, the mechanical properties of the simulated microstructures were evaluated using microstructure-based finite element method (FEM) analysis to determine the elastic modulus, yield stress, ultimate tensile strength, and elongation to failure of the aged microstructures. Using this approach, our surrogate model can predict precipitation behavior within 0.25% volume fraction and mechanical properties within 6% relative error from the values predicted by the KMC and FEM models using 50 training simulations as input. The trained recurrent neural network-based model can return estimations of precipitation kinetics and mechanical properties ~1000 times faster than the physics-based codes. This work demonstrates, as a proof of concept, that reactor material service lifetimes under variable service conditions can be predicted for a statistics-based model from a practicably obtainable dataset.

36 MATERIALS SCIENCE↗

Flow annealed importance sampling bootstrap meets differentiable particle physics

High-energy physics requires the generation of large numbers of simulated data samples from complex but analytically tractable distributions called matrix elements. Surrogate models, such as normalizing flows, are gaining popularity for this task due to their computational efficiency. We adopt an approach based on flow annealed importance sampling bootstrap (FAB) that evaluates the differentiable target density during training and helps avoid the costly generation of training data in advance. We show that FAB reaches higher sampling efficiency with fewer target evaluations in high dimensions in comparison to other methods.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Flood Mapping Using UAVSAR and Convolutional Neural Networks

We have mapped flooded areas in data collected by the NASA/JPL Uninhabited Aerial Vehicle Synthetic Aperture Radar (UAVSAR) using two convolutional neural network (CNN) image classifier architectures: U-Net and SegNet. Our study area was a region around Houston, TX, USA affected by widespread flooding in 2017 due to Hurricane Harvey. To train and test the classifiers, we manually labelled over 10000 image segments in two flight lines. Both U-Net and SegNet yielded higher accuracy than a previous non-machine learning classifier we used as a baseline. U-Net had slightly higher accuracy than SegNet. The classifiers performed better in areas with more homogeneous land cover. To independently validate the classifier accuracy we used NOAA aerial imagery, with overall accuracy around 80%. Future work includes assessing the classifier robustness in other study areas, assessing the classifier dependence on UAVSAR incidence angle, particularly for open water and bare ground, and collecting more training data, particularly in urban areas. This study demonstrates the potential of CNN image classifiers for mapping flooded areas in airborne polarimetric SAR imagery, and for land cover classification of polarimetric SAR imagery more generally.

Denbina, Michael W↗

High temperature nuclear data measurements of SiC, ZrC, and MgO [Slides]

Performed temperature dependent measurements of SiC and ZrC at ARCS instrument at SNS. Performed temperature dependent measurements of SiC, ZrC, and MgO at VISION instrument at SNS. Performed initial atomistic modeling of these materials using various techniques, including machine learned potentials. Future work includes temperature dependent transmission measurements of these materials, as well as improving the machine learned potentials with more training data and different machine learned frameworks.

ARCS↗

Improving vertical detail in simulated temperature and humidity data using machine learning

Atmospheric models used for weather forecasting and climate predictions discretise the atmosphere onto a vertical grid. There are however atmospheric phenomena that occur on scales smaller than the thickness of those model layers. The formation of low-level clouds due to temperature inversions is an example. This leads to atmospheric models underestimating, or even missing, these clouds and their radiative effects. Using radiosonde observations as training data, a machine learning model is used to improve the vertical detail of modelled profiles of temperature and specific humidity. In addition, a physics-informed machine learning model is developed and compared to the traditional approach; showing improvements in the cloud fraction profiles calculated from its predictions. The vertically enhanced profiles also improve the representation of layers of convective inhibition and anomalous refractivity gradients. This work facilitates targeted improvements to the representation of certain atmospheric processes without the burden of increased memory and computational cost from increasing vertical resolution throughout the whole model.

54 ENVIRONMENTAL SCIENCES↗

Classification Of Terrain In Polarimetric SAR Images

Two algorithms processing polarimetric synthetic-aperture-radar data found effective in assigning various parts of SAR images to classes representing different types of terrain. Partially automate interpretation of SAR imagery, reducing amount of photointerpretation needed and putting whole interpretation process on more quantitative and systematic basis. First algorithm implements Bayesian classification scheme "supervised" by use of training data. Second algorithm implements classification procedure unsupervised.

Van Zyl, Jakob J.↗

Deciphering the Scattering of Mechanically Driven Polymers Using Deep Learning

Here, we present a deep learning approach for analyzing two-dimensional scattering data of semiflexible polymers under external forces. In our framework, scattering functions are compressed into a three-dimensional latent space using a Variational Autoencoder (VAE), and two converter networks establish a bidirectional mapping between the polymer parameters (bending modulus, stretching force, and steady shear) and the scattering functions. The training data are generated using off-lattice Monte Carlo simulations to avoid the orientational bias inherent in lattice models, ensuring robust sampling of polymer conformations. The feasibility of this bidirectional mapping is demonstrated by the organized distribution of polymer parameters in the latent space. By integrating the converter networks with the VAE, we obtain a generator that produces scattering functions from given polymer parameters and an inferrer that directly extracts polymer parameters from scattering data. While the generator can be utilized in a traditional least-squares fitting procedure, the inferrer produces comparable results in a single pass and operates 3 orders of magnitude faster. This approach offers a scalable automated tool for polymer scattering analysis and provides a promising foundation for extending the method to other scattering models, experimental validation, and the study of time-dependent scattering data.

Ding, Lijie [Oak Ridge National Laboratory (ORNL),↗

Exploring 2D X-ray diffraction phase fraction analysis with convolutional neural networks: Insights from kinematic-diffraction simulations

Abstract Deep-learning models are effective for analyzing the complex information in 2D X-ray diffraction (XRD) patterns. Accurately collecting parameters of the material sample is crucial during model training, significantly impacting model performance. In this study, we employ a kinematic-diffraction simulator to generate simulated 2D XRD patterns for Ti–6Al–4V alloy, allowing precise control of sample parameters. These simulated patterns are used to train convolutional neural networks, predicting $$\upbeta$$ β -phase volume fractions. The training data set consists exclusively of 2D XRD patterns with pure $$\upalpha$$ α - or pure $$\upbeta$$ β -phase, while the testing set incorporates patterns with intermediate phase volume fraction. In particular, we investigate how the architectures of the model influence prediction reliability and computational performance. Experimental results reveal that, with appropriate training, the convolutional neural network accurately detects intermediate phase volume fractions even trained with only pure-phase patterns, achieving a mean square error accuracy of $$9.4 \times 10^{-4}$$ 9.4 × 10 - 4 . Graphical abstract

Yue, Weiqi↗

Shoulder Injuries in US Astronauts Related to EVA Suit Design

Introduction: For every one hour spent performing extravehicular activity (EVA) in space, astronauts in the US space program spend approximately six to ten hours training in the EVA spacesuit at NASA-Johnson Space Center's Neutral Buoyancy Lab (NBL). In 1997, NASA introduced the planar hard upper torso (HUT) EVA spacesuit which subsequently replaced the existing pivoted HUT. An extra joint in the pivoted shoulder allows increased mobility but also increased complexity. Over the next decade a number of astronauts developed shoulder problems requiring surgical intervention, many of whom performed EVA training in the NBL. This study investigated whether changing HUT designs led to shoulder injuries requiring surgical repair. Methods: US astronaut EVA training data and spacesuit design employed were analyzed from the NBL data. Shoulder surgery data was acquired from the medical record database, and causal mechanisms were obtained from personal interviews Analysis of the individual HUT designs was performed as it related to normal shoulder biomechanics. Results: To date, 23 US astronauts have required 25 shoulder surgeries. Approximately 48% (11/23) directly attributed their injury to training in the planar HUT, whereas none attributed their injury to training in the pivoted HUT. The planar HUT design limits shoulder abduction to 90 degrees compared to approximately 120 degrees in the pivoted HUT. The planar HUT also forces the shoulder into a forward flexed position requiring active retraction and extension to increase abduction beyond 90 degrees. Discussion: Multiple factors are associated with mechanisms leading to shoulder injury requiring surgical repair. Limitations to normal shoulder mechanics, suit fit, donning/doffing, body position, pre-existing injury, tool weight and configuration, age, in-suit activity, and HUT design have all been identified as potential sources of injury. Conclusion: Crewmembers with pre-existing or current shoulder injuries or certain anthropometric body types should conduct NBL EVA training in the pivoted HUT.

Scheuring, R. A.↗

Learning dynamical systems from data: An introduction to physics-guided deep learning

Modeling complex physical dynamics is a fundamental task in science and engineering. Traditional physics-based models are first-principled, explainable, and sample-efficient. However, they often rely on strong modeling assumptions and expensive numerical integration, requiring significant computational resources and domain expertise. While deep learning (DL) provides efficient alternatives for modeling complex dynamics, they require a large amount of labeled training data. Furthermore, its predictions may disobey the governing physical laws and are difficult to interpret. Physics-guided DL aims to integrate first-principled physical knowledge into data-driven methods. It has the best of both worlds and is well equipped to better solve scientific problems. Recently, this field has gained great progress and has drawn considerable interest across discipline Here, we introduce the framework of physics-guided DL with a special emphasis on learning dynamical systems. We describe the learning pipeline and categorize state-of-the-art methods under this framework. We also offer our perspectives on the open challenges and emerging opportunities.

97 MATHEMATICS AND COMPUTING↗

Improving Reliability of Large Language Models for Nuclear Power Plant Diagnostics [Poster]

Large Language Models (LLMs) struggle out of the box when answering factually about detailed questions, especially in domains that are sparsely represented in their training data. This causes hallucinations and reduces reliability making it difficult for them to be used in practice. This work shows that using RAG techniques can improve factual accuracy and reliability, allowing for the application of LLMs in specialized areas, even when those areas that aren’t extensively covered in their initial training.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗