Search NASA⌕ Search

SEARCH · Search NASA

Results for “limited data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Modular machine learning-based elastoplasticity: Generalization in the context of limited data

The development of highly accurate constitutive models for materials that undergo path-dependent processes continues to be a complex challenge in computational solid mechanics. Challenges arise both in considering the appropriate model assumptions and from the viewpoint of data availability, verification, and validation. Recently, data-driven modeling approaches have been proposed that aim to establish stress-evolution laws that avoid user-chosen functional forms by relying on machine learning representations and algorithms. However, these approaches not only require a significant amount of data but also need data that probes the full stress space with a variety of complex loading paths. Furthermore, they rarely enforce all necessary thermodynamic principles as hard constraints. Hence, they are in particular not suitable for low-data or limited-data regimes, where the first arises from the cost of obtaining the data and the latter from the experimental limitations of obtaining labeled data, which is commonly the case in engineering applications. In this work, we discuss a hybrid framework that can work on a variable amount of data by relying on the modularity of the elastoplasticity formulation where each component of the model can be chosen to be either a classical phenomenological or a data-driven model depending on the amount of available information and the complexity of the response. The method is tested on synthetic uniaxial data coming from simulations as well as cyclic experimental data for structural materials. The discovered material models are found to not only interpolate well but also allow for accurate extrapolation in a thermodynamically consistent manner far outside the domain of the training data. This ability to extrapolate from limited data was the main reason for the early and continued success of phenomenological models and the main shortcoming in machine learning-enabled constitutive modeling approaches. Training aspects and details of the implementation of these models into Finite Element simulations are discussed and analyzed.

42 ENGINEERING↗

Rapidly Changing Range Limits in a Warming World: Critical Data Limitations and Knowledge Gaps for Advancing Understanding of Mangrove Range Dynamics in the Southeastern USA

Climate change is altering species’ range limits and transforming ecosystems. For example, warming temperatures are leading to the range expansion of tropical, cold-sensitive species at the expense of their cold-tolerant counterparts. In some temperate and subtropical coastal wetlands, warming winters are enabling mangrove forest encroachment into salt marsh, which is a major regime shift that has significant ecological and societal ramifications. Here, we synthesized existing data and expert knowledge to assess the distribution of mangroves near rapidly changing range limits in the southeastern USA. We used expert elicitation to identify data limitations and highlight knowledge gaps for advancing understanding of past, current, and future range dynamics. Mangroves near poleward range limits are often shorter, wider, and more shrublike compared to their tropical counterparts that grow as tall forests in freeze-free, resource-rich environments. The northern range limits of mangroves in the southeastern USA are particularly dynamic and climate sensitive due to abundance of suitable coastal wetland habitat and the exposure of mangroves to winter temperature extremes that are much colder than comparable range limits on other continents. Thus, there is need for methodological refinements and improved spatiotemporal data regarding changes in mangrove structure and abundance near northern range limits in the southeastern USA. Advancing understanding of rapidly changing range limits is critical for foundation plant species such as mangroves, as it provides a basis for anticipating and preparing for the cascading effects of climate-induced species redistribution on ecosystems and the human communities that depend on their ecosystem services.

54 ENVIRONMENTAL SCIENCES↗

Utilization of Data Augmentation Techniques in Automated Inspection Systems for Defect Detection in Metals With Limited Data

Accurate identification of defects on metal surfaces is of great interest to many industry sectors, such as the automotive and aerospace industries. In contrast to conventional manual inspection techniques, recent automated inspection systems employ deep learning models trained to detect defects rapidly and precisely. The development of these models often requires a substantial image dataset to acquire adequate knowledge of defect features and enhance their predictive accuracy. When data is limited, augmentation techniques are often used to improve the precision and accuracy of defect detection systems. This study examined the prediction performance of two object detection models, namely Faster Region‐based Convolutional Neural Network (Faster R‐CNN) and You Only Look Once version 8 (YOLOv8), to identify dent defects in limited images of cast iron cylinder head surfaces. The original image set contains 46 images with 563 dents. To overcome limited data availability, common image augmentation techniques along with a copy‐paste method were applied. Results show that standard augmentation improved YOLOv8 accuracy by 8.00% and average precision (AP) by 3.00%. On the other hand, the copy‐paste technique achieved a 20.00% increase in accuracy and a 1% increase in AP with just 200 synthetic dents. Furthermore, these results provide support for using the copy‐paste augmentation strategy to enhance defect detection performance, with a limited dataset, contributing to more accurate defect identification in remanufacturing processes.

36 MATERIALS SCIENCE↗

Predicting weather impacts on corn production in a data-limited region using a transfer learning approach

The stability of food supply and prices may depend more on annual changes in yields from year-to-year variability in weather than on longer-term average changes from changing climatic conditions. However, the absence of high-quality data on crop yields at fine spatial resolutions in many regions of the world makes it challenging to statistically model their response to interannual variability in weather patterns. Therefore, there is a need for empirical methods that can project annual crop yield changes even in limited data regions. Here, we propose a transfer learning algorithm that uses high spatial resolution data from one region to project yields in another region with more limited data. The goal of our work is to understand what data types can be beneficial for transferring learning from a source region to a very different target region with more limited data. We utilize Long Short-Term Memory to develop a transfer learning model that is trained on historical county-level corn yield in the United States and predicts district-level corn yield variations in India. Even using smaller amounts of data in India, simulating a data-scarce region, we achieve an average root mean square error of 0.48 bu acre−1 in predicting interannual yield variations. Using Shapley values to interpret results, we explore the contribution of the different weather parameters to interannual yield variability and find a larger influence of precipitation-related variables. Our study demonstrates the usefulness of this method for transferring models of weather impacts on crop yields trained on a data-rich country to one with more limited data. It suggests the potential of applying the transfer learning model to mitigate the need for extensive raw data globally.

Vishwakarma, Srishti [ORNL] (ORCID:000000031674419↗

Deep learning enables nanoscale X-ray 3D imaging with limited data

Deep neural network can greatly improve tomography reconstruction with limited data. A recent effort of combining ptycho-tomography model with the 3D U-net demonstrated a significant reduction in both the number of projections and computation time, and showed its potential for integrated circuit imaging that requires high-resolution and fast measurement speed.

47 OTHER INSTRUMENTATION↗

Using GANs to predict milling stability from limited data

Milling is a key manufacturing process that requires the selection of operating parameters that provide efficient performance. However, the presence of chatter, a self-excited vibration causing poor surface finish and potential damage to the machine and cutting tool, makes it challenging to select the appropriate parameters. To predict chatter, stability maps are commonly used, but their generation requires expensive data, making it difficult to employ these maps in industry. Therefore, there is a pressing need for an approach that can accurately predict stability maps using limited experimental data. This study introduces the new Encoder GAN (EGAN) approach based on Generative Adversarial Networks (GANs) that predicts stability maps using limited experimental data. The approach consists of the encoder, generator, and discriminator subnetworks and uses the trained encoder and generator to predict the target stability map. This versatile method can be applied to various tool setups and can accurately predict stability maps with limited experimental data (five to 10 cutting tests) even when there is little information available for unknown parameters. In conclusion, the study evaluates the proposed approach using both numerical data and experiments and demonstrates its superior performance compared to state-of-the-art benchmarks.

42 ENGINEERING↗

Expanding the Domain of Applicability of Machine Learning Models with Limited Data for Drug Property Prediction

Accurate machine learning models for predicting small molecule interactions with biological targets are essential for therapeutic discovery, biothreat response, and computational drug design, but their performance is often limited for understudied targets with sparse experimental data. To address this challenge, we developed and evaluated methods to improve molecular property prediction under low-data conditions, using the NimA-related kinase (NEK) family as a proof-of-concept. This work focused on two complementary goals within the ATOM Modeling PipeLine (AMPL) and the Generative Molecular Design (GMD) loop: expanding model applicability through transfer learning, representation learning, feature scaling, sampling strategies, and active-learning-inspired compound selection; and enabling efficient virtual screening to prioritize compounds that balance predicted activity, design objectives, and synthetic accessibility.

organic↗

Transformer-powered surrogates close the ICF simulation-experiment gap with extremely limited data

Abstract Recent advances in machine learning, specifically transformer architecture, have led to significant advancements in commercial domains. These powerful models have demonstrated superior capability to learn complex relationships and often generalize better to new data and problems. This paper presents a novel transformer-powered approach for enhancing prediction accuracy in multi-modal output scenarios, where sparse experimental data is supplemented with simulation data. The proposed approach integrates transformer-based architecture with a novel graph-based hyper-parameter optimization technique. The resulting system not only effectively reduces simulation bias, but also achieves superior prediction accuracy compared to the prior method. We demonstrate the efficacy of our approach on inertial confinement fusion experiments, where only 10 shots of real-world data are available, as well as synthetic versions of these experiments.

97 MATHEMATICS AND COMPUTING↗

Data-Efficient Dimensionality Reduction and Surrogate Modeling of High-Dimensional Stress Fields

Tensor datatypes representing field variables like stress, displacement, velocity, etc., have increasingly become a common occurrence in data-driven modeling and analysis of simulations. Numerous methods [such as convolutional neural networks (CNNs)] exist to address the meta-modeling of field data from simulations. As the complexity of the simulation increases, so does the cost of acquisition, leading to limited data scenarios. Modeling of tensor datatypes under limited data scenarios remains a hindrance for engineering applications. Here, in this article, we introduce a direct image-to-image modeling framework of convolutional autoencoders enhanced by information bottleneck loss function to tackle the tensor data types with limited data. The information bottleneck method penalizes the nuisance information in the latent space while maximizing relevant information making it robust for limited data scenarios. The entire neural network framework is further combined with robust hyperparameter optimization. We perform numerical studies to compare the predictive performance of the proposed method with a dimensionality reduction-based surrogate modeling framework on a representative linear elastic ellipsoidal void problem with uniaxial loading. The data structure focuses on the low-data regime (fewer than 100 data points) and includes the parameterized geometry of the ellipsoidal void as the input and the predicted stress field as the output. The results of the numerical studies show that the information bottleneck approach yields improved overall accuracy and more precise prediction of the extremes of the stress field. Additionally, an in-depth analysis is carried out to elucidate the information compression behavior of the proposed framework.

artificial intelligence↗

A ModEx Framework for Watershed Subsurface Investigation With Limited Geophysical Data Using Machine Learning and Hydrologic Modeling

Abstract Subsurface heterogeneity influences watershed hydrology strongly but remains difficult to characterize at catchment scales with sparse and costly field data. Geophysical surveys such as electromagnetic induction (EMI) provide local spatial subsurface images yet scaling them to watershed scales and converting EMI‐derived resistivity into hydraulic properties remains a challenge. We present a Model–Experiment (ModEx) framework that integrates limited EMI data with machine learning (ML) and hydrologic modeling to improve process representation and guide field investigations. Sparse EMI surveys were scaled to the catchment scale using a Random Forest model, and the resulting resistivity fields were combined with nearby borehole constraints to parameterize a hydrologic model. The EMI‐informed hydrological simulations improved predictions of streamflow sustained by subsurface flow and shallow saturation patterns. By combining EMI data and ML with hydrologic modeling, the ModEx framework guides future subsurface surveys, providing a transferable and efficient strategy for data–model integration across diverse watersheds. Plain Language Summary Mapping the underground network of soil and rock that controls water is essential for predicting floods and droughts, but seeing underground is difficult and expensive. We cannot drill everywhere, so scientists use geophysical tools to scan broad areas. There are two key challenges: these geophysical scans are often sparse across the whole watershed, and the geophysical data is hard to translate into water‐related properties. We used artificial intelligence to solve these problems. We taught a computer to find patterns linking the limited geophysical data to the land surface properties. This allowed it to fill in the gaps and create a complete, useful subsurface map for the entire watershed. This new map improves hydrologic simulations, leading to more accurate predictions of water movement in the watershed. It also helps scientists build better models with less data and generates a priority map showing where to measure next, making future investigations more efficient. Key Points Limited EMI scaled with ML improves catchment‐scale subsurface parameterization for hydrologic models The framework integrates hydrologic modeling with limited geophysical data to support subsurface investigation design ModEx framework offers a transferable data–model integration strategy that quantifies and reduces uncertainty guiding watershed studies

Chen, Hang↗

Deployment of Traditional and Hybrid Machine Learning for Critical Heat Flux Prediction in the CTF Thermal-Hydraulics Code

Critical heat flux (CHF) marks the transition from nucleate to film boiling, where heat transfer to the working fluid can rapidly deteriorate. Accurate CHF prediction is essential for efficiency, safety, and preventing equipment damage, particularly in nuclear reactors. Although widely used, empirical correlations frequently exhibit discrepancies when compared to experimental data, limiting their reliability in diverse operational conditions. Traditional machine learning (ML) approaches have demonstrated potential for CHF prediction but often suffer from limited interpretability, data scarcity, and insufficient knowledge of physical principles. Hybrid model approaches, which combine data-driven ML with base models, mitigate these concerns by incorporating prior knowledge of the domain. This study integrates an externally trained purely data-driven ML model and two hybrid models (using the Biasi and Bowring CHF correlations) within the CTF subchannel code via a custom Fortran framework. Performance was evaluated using two validation cases: a subset of the Nuclear Regulatory Commission (NRC) CHF database and the Bennett dryout experiments. In both cases, the hybrid models demonstrated significantly lower error metrics compared to conventional empirical correlations, with the best models often reducing relative error by about 5 percentage points. The pure ML model achieved comparable accuracy, outperforming the hybrid Biasi model in the NRC test case (3.3% versus 5.5% relative error) but exhibiting slightly higher error against the hybrid Bowring model in the Bennett test case (7.7% versus 6.1%). Trend analysis of error parity indicated that ML-based models reduced the tendency for CHF overprediction, improving overall accuracy. These results demonstrate that ML-based CHF models can be effectively integrated into subchannel codes and could potentially increase performance compared to conventional methods.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Machine learning for reactor power monitoring with limited labeled data

Real-time reactor power monitoring is critical for a variety of nuclear applications, spanning safety, security, operations, and maintenance. While machine learning methods have shown promise in monitoring reactor power levels, there is limited research on their efficacy in label-starved environments. The goal of this work is to assess the feasibility of classifying nuclear reactor power level using multisource data in scenarios with limited labels. Data were collected using low-resolution multisensors at four nuclear reactor facilities: two large research reactors and two TRIGA reactors. Within each pair, one reactor dataset served as the source and the other as the target in a transfer learning paradigm. Twenty-three supervised models were trained on labeled sequences of magnetic field and acceleration data from each of the target sites. Self-learning and transfer learning methods were applied to the top performing models to assess their classification performance with increasing amounts of labeled data. While reactor power level classification was achieved with a Matthews Correlation Coefficient of up to 0.739 ± 0.003 and 0.622 ± 0.009 with only 400 sequences per power state for the large research reactor and TRIGA target sites, respectively, self-learning and transfer learning leveraging source site data did not improve target classification performance. These findings suggest that alternative methods, such as higher sensitivity sensors, digital twins, or the use of physics-informed models, are required to enable high-performance classification in machine learning approaches to reactor monitoring with a dearth of target ground truth.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Transfer Learning-Based Independent Component Analysis

Understanding the underlying component structure is crucial for multivariate signal analysis. Among all the techniques that try to learn the latent structure, independent component analysis (ICA) is one of the most important and popular methods, which aims to extract independent components from multivariate signals and enables further analysis. For example, in electroencephalogram (EEG) analysis, artifacts filtering and disease detection are conducted based on the independent components of the signals. One critical challenge in existing ICA approaches is that the component extraction accuracy may degrade when the available data of a unit are limited. To address this issue, this paper proposes a transfer learning-based ICA method by innovatively transferring component distribution from a source domain, so that accurate component extraction results can be achieved even when only limited data are available in the target domain. To the best of our knowledge, this is the first work that leverages transfer learning to improve ICA accuracy with limited available data. In particular, we first extract all the independent components from the source domain by maximizing the log-likelihood function with a Newton-like method on a smooth manifold. Then for the target domain, the component with the largest negentropy is extracted in each round. To effectively leverage the knowledge from the source domain and to prevent the negative transfer, we try to find a component in the source domain that matches the component we are extracting. The probability density function of the matched component will then be used to improve the component extraction accuracy if such matched component can be found; otherwise, no knowledge will be transferred. Finally, numerical simulations and a case study with electrocardiogram (ECG) data are conducted, showing the effectiveness of the proposed method in transferring knowledge and reducing negative transfer.

42 ENGINEERING↗

Quantitative Power System Resilience Metrics and Evaluation Approach

Power system resilience is an emerging topic and plays an essential role in helping the power industry understand and respond to the increasing threats of extreme weather events. The first step of power system resilience analysis is to introduce metrics to quantify the resilience reasonably. Existing resilience metrics are typically restrained by the limited data for extreme event modeling and fall short in terms of physical interpretation and comparability. This paper develops novel quantitative metrics to evaluate power system resilience in pre- and post-event contexts. The developed metrics illustrate clear physical meanings and can be effectively used to compare resilience across different systems under different extreme events. Moreover, the developed metrics can be applied to both transmission and distribution systems. Simulation on a distribution system is employed to validate the effectiveness of the proposed resilience metrics and resilience evaluation approach.

power system resilience↗