Search NASA⌕ Search

SEARCH · Search NASA

Results for “loss function regularization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Leveraging large language models to address data scarcity in machine learning for graphene synthesis

Machine learning in experimental materials science faces significant challenges due to the scarcity of data, which are costly and time-consuming to generate, particularly when relying on in-house experiments. Literature data mining offers a potential solution but introduces issues like mixed data quality, inconsistent formats, and non-uniform reporting of synthesis parameters, resulting in partially missing and heterogeneous features across the dataset. Here, we propose data imputation and feature engineering methods that employ pre-trained large language models (LLMs) to enhance machine learning performance on scarce, heterogeneous datasets, demonstrated on graphene CVD synthesis data and the ML-HydPARK hydrogen storage dataset. GPT models perform data imputation via tailored prompting and semantic normalization of inconsistently reported features through embeddings, for example, to harmonize the complex nomenclature of CVD substrates. Beyond yielding more diverse and richer feature representations than traditional methods such as K-nearest neighbors (KNN) and Multivariate Imputation by Chained Equations (MICE), LLM-based data imputation is evaluated against dataset characteristics and prompting strategies. We vary the level of autonomy granted to the LLM, from generic prompting that leverages pre-trained knowledge for autonomous data generation to data-informed prompting that constrains outputs using target-specific information, and demonstrate which level of autonomy yields superior imputation performance across datasets and feature types. The proposed data engineering methods markedly improve downstream performance; for example, in graphene layer number classification using a support vector machine (SVM), binary accuracy increases from 39% to 65% and ternary accuracy from 52% to 72%. Fine-tuning experiments on both datasets show that combining our proposed LLM-based data imputation and feature encoding methods with numerical machine learning predictors outperforms standalone fine-tuned LLM predictors in data-scarce settings. The proposed strategies emphasize data enhancement techniques rather than refining learning architectures or regularizing loss functions, offering a broadly applicable framework for improving machine learning performance on scarce, inhomogeneous datasets.

Chemical vapor deposition↗

Connect the Dots: In Situ 4-D Seismic Monitoring of CO 2 Storage With Spatio-Temporal CNNs

4-D seismic imaging has been widely used in CO 2 sequestration projects to monitor the fluid flow in the volumetric subsurface region that is not sampled by wells. Ideally, real-time monitoring and near-future forecasting would provide site operators with great insights to understand the dynamics of the subsurface reservoir and assess any potential risks. However, due to obstacles such as high deployment cost, availability of acquisition equipment, exclusion zones around surface structures, only very sparse seismic imaging data can be obtained during monitoring. That leads to an unavoidable and growing knowledge gap over time. The operator needs to understand the fluid flow throughout the project lifetime and the seismic data are only available at a limited number of times. This is insufficient for understanding reservoir behavior. To overcome those challenges, we have developed spatio-temporal neural-network-based models that can produce high-fidelity interpolated or extrapolated images effectively and efficiently. Specifically, our models are built on an autoencoder, and incorporate the long short-term memory (LSTM) structure with a new loss function regularized by optical flow. We validate the performance of our models using real 4-D post-stack seismic imaging data acquired at the Sleipner CO 2 sequestration field. We employ two different strategies in evaluating our models. Numerically, we compare our models with different baseline approaches using classic pixel-based metrics. We also conduct a blind survey and collect a total of 20 responses from domain experts to evaluate the quality of data generated by our models. Finally, via both numerical and expert evaluation, we conclude that our models can produce high-quality 2-D/3-D seismic imaging data at a reasonable cost, offering the possibility of real-time monitoring or even near-future forecasting of the CO 2 storage reservoir.

4-D seismic imaging↗

Inter-well connectivity detection in CO 2 WAG projects using statistical recurrent unit models

Routine well-wise injection and production measurements contain significant information on subsurface structure and properties. Data-driven technology that interprets surface data into subsurface structure or properties can assist operators in making informed decisions by providing a better understanding of field assets. Our machine-learning framework is built on the statistical recurrent unit (SRU) model and interprets well-based injection/production data into inter-well connectivity without relying on a geologic model. We test it on synthetic and field-scale CO 2 EOR projects utilizing the water-alternating-gas (WAG) process. SRU is a special type of recurrent neural network (RNN) that allows for better characterization of temporal trends, by learning various statistics of the input at different time scales. In our application, the complete states (injection rate, pressure and cumulative injection) at injectors and pressure states at producers are fed to SRU as the input and the phase rates at producers are treated as the output. Once the SRU is trained and validated, it is then used to assess the connectivity of each injector to any producer using permutation variable importance method, wherein inputs corresponding to an injector are shuffled and the increase in prediction error at a given producer is recorded as the importance (connectivity metric) of the injector to the producer. This method is tested in both synthetic and field-scale cases. The validation of the proposed data-driven inter-well connectivity assessment is performed using synthetic data from simulation models where inter-well connectivity can be easily measured using the streamline-based flux allocation. The SRU model is shown to offer excellent prediction performance on the synthetic case. Despite significant measurement noise and frequent well shut-ins imposed in the field-scale case, the SRU model offers good prediction accuracy, the overall relative error of the phase production rates at most producers ranges from 10% to 30%. It is shown that the dominant connections identified by the data-driven method and streamline method are in close agreement. This significantly improves confidence in our data-driven procedure. The novelty of this work is that it is purely data-driven method and can directly interpret routine surface measurements to intuitive subsurface knowledge. Furthermore, the streamline-based validation procedure provides physics-based backing to the results obtained from data analytics. This study results in a reliable and efficient data analytics framework that is well-suited for large field applications.

42 ENGINEERING↗

Using spatio-temporal graph neural networks to estimate fleet-wide photovoltaic performance degradation patterns

Accurate estimation of photovoltaic (PV) system performance is crucial for determining its feasibility as a power generation technology and financial asset. PV-based energy solutions offer a viable alternative to traditional energy resources due to their superior Levelized Cost of Energy (LCOE). A significant challenge in assessing the LCOE of PV systems lies in understanding the Performance Loss Rate (PLR) for large fleets of PV systems. Estimating the PLR of PV systems becomes increasingly important in the rapidly growing PV industry. Precise PLR estimation benefits PV users by providing real-time monitoring of PV module performance, while explainable PLR estimation assists PV manufacturers in studying and enhancing the performance of their products. However, traditional PLR estimation methods based on statistical models have notable drawbacks. Firstly, they require user knowledge and decision-making. Secondly, they fail to leverage spatial coherence for fleet-level analysis. Additionally, these methods inherently assume the linearity of degradation, which is not representative of real world degradation. To overcome these challenges, we propose a novel graph deep learning-based decomposition method called the Spatio-Temporal Graph Neural Network for fleet-level PLR estimation (PV-stGNN-PLR). PV-stGNN-PLR decomposes the power timeseries data into aging and fluctuation components, utilizing the aging component to estimate PLR. PV-stGNN-PLR exploits spatial and temporal coherence to derive PLR estimation for all systems in a fleet and imposes flatness and smoothness regularization in loss function to ensure the successful disentanglement between aging and fluctuation. We have evaluated PV-stGNN-PLR on three simulated PV datasets consisting of 100 inverters from 5 sites. Experimental results show that PV-stGNN-PLR obtains a reduction of 33.9% and 35.1% on average in Mean Absolute Percent Error (MAPE) and Euclidean Distance (ED) in PLR degradation pattern estimation compared to the state-of-the-art PLR estimation methods.

14 SOLAR ENERGY↗

Surrogate-based Analysis and Optimization

A major challenge to the successful full-scale development of modem aerospace systems is to address competing objectives such as improved performance, reduced costs, and enhanced safety. Accurate, high-fidelity models are typically time consuming and computationally expensive. Furthermore, informed decisions should be made with an understanding of the impact (global sensitivity) of the design variables on the different objectives. In this context, the so-called surrogate-based approach for analysis and optimization can play a very valuable role. The surrogates are constructed using data drawn from high-fidelity models, and provide fast approximations of the objectives and constraints at new design points, thereby making sensitivity and optimization studies feasible. This paper provides a comprehensive discussion of the fundamental issues that arise in surrogate-based analysis and optimization (SBAO), highlighting concepts, methods, techniques, as well as practical implications. The issues addressed include the selection of the loss function and regularization criteria for constructing the surrogates, design of experiments, surrogate selection and construction, sensitivity analysis, convergence, and optimization. The multi-objective optimal design of a liquid rocket injector is presented to highlight the state of the art and to help guide future efforts.

Queipo, Nestor V.↗

Addressing bias in bagging and boosting regression models

As artificial intelligence (AI) becomes widespread, there is increasing attention on investigating bias in machine learning (ML) models. Previous research concentrated on classification problems, with little emphasis on regression models. This paper presents an easy-to-apply and effective methodology for mitigating bias in bagging and boosting regression models, that is also applicable to any model trained through minimizing a differentiable loss function. Our methodology measures bias rigorously and extends the ML model's loss function with a regularization term to penalize high correlations between model errors and protected attributes. We applied our approach to three popular tree-based ensemble models: a random forest model (RF), a gradient-boosted model (GBT), and an extreme gradient boosting model (XGBoost). We implemented our methodology on a case study for predicting road-level traffic volume, where RF, GBT, and XGBoost models were shown to have high accuracy. Despite high accuracy, the ML models were shown to perform poorly on roads in minority-populated areas. Our bias mitigation approach reduced minority-related bias by over 50%.

97 MATHEMATICS AND COMPUTING↗

Data imbalance in drug response prediction: multi-objective optimization approach in deep learning setting

Abstract Drug response prediction (DRP) methods tackle the complex task of associating the effectiveness of small molecules with the specific genetic makeup of the patient. Anti-cancer DRP is a particularly challenging task requiring costly experiments as underlying pathogenic mechanisms are broad and associated with multiple genomic pathways. The scientific community has exerted significant efforts to generate public drug screening datasets, giving a path to various machine learning models that attempt to reason over complex data space of small compounds and biological characteristics of tumors. However, the data depth is still lacking compared to application domains like computer vision or natural language processing domains, limiting current learning capabilities. To combat this issue and improves the generalizability of the DRP models, we are exploring strategies that explicitly address the imbalance in the DRP datasets. We reframe the problem as a multi-objective optimization across multiple drugs to maximize deep learning model performance. We implement this approach by constructing Multi-Objective Optimization Regularized by Loss Entropy loss function and plugging it into a Deep Learning model. We demonstrate the utility of proposed drug discovery methods and make suggestions for further potential application of the work to achieve desirable outcomes in the healthcare field.

Biochemistry & Molecular Biology↗

Extension of the PINN diffusion model to k-eigenvalue problems

This paper extends our recent work on the Physics-Informed Neural Networks (PINN) approach for the fixed source diffusion models and applies it to the diffusion theory based k-eigenvalue problems. To make the PINN equitable for the eigenvalue problems, we introduce a novel integral regularization term to the loss function in the framework, and allow the direct inference of the principal eigenvalue and the associated eigenfunction. The regularization term enforces a pre-defined value on the integration of the model predictions, and this value can be directly related to a physical property of the system. We also introduce an additional learnable parameter to approximate the principal eigenvalue. As a proof of principle, we solve the one-group two-dimensional k-eigenvalue neutron diffusion equation in this work. We then provide two numerical examples to demonstrate the applicability of the PINN approach. In each example, we solve the k-eigenvalue diffusion equation in a multi-region configuration constrained with a set of Robin boundary conditions for generality. We use a FEM solution based on the power-iteration method to verify the results of the PINN solution. The results showed relative percentage error in the predicted eigenvalue of about 0.77% and about 1.2% for example 1 and example 2, respectively. The mean absolute error in the predicted flux for example 1 is ∼ 0.002 and for example 2 is ∼ 0.0024. These results indicate some preliminary successes of the PINN application to k-eigenvalue problems. (authors)

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

DENSECL: Haze Mitigation Using Dense Blocks and Contrastive Loss Regularization

Haze, which occurs as a result of the scattering of light in the atmosphere by small particles, diminishes the visibility of scene objects, inflicting important image applications such as object detection. To address the problem, this paper introduces a new physics-based end-to-end deep learning approach to haze mitigation in outdoor scenes, including those in airborne images. The proposed model named DenseCL is designed with dense blocks and adopts a contrastive loss function as an additional regularization. The model also maintains the cycle consistency by remapping the dehazed outputs into a hazy image using the physics-based light scattering function. DenseCL has been trained with publicly available outdoor images and demonstrates outstanding performance on outdoor, indoor, and remotely sensed nonhomogeneous haze satellite images.

Mitra, Somosmita↗

A physics-informed and hierarchically regularized data-driven model for predicting fluid flow through porous media

This paper presents a new deep learning data-driven model for predicting structure dependent pore-fluid velocity fields in rock. The model is based on a Convolutional Auto-Encoder (CAE) artificial neural network capable of learning from image data generated by direct numerical simulations of fluid flow through pore-structures, such as by Lattice Boltzmann or molecular dynamics methods. The main novelty of the model in comparison to previous CAE-based data-driven approaches consists of three parts. The first is a methodology for decomposing the full-domain of the porous media into sub-regions, or “sub-domains”, in order to reduce the overall size of the CAE, batch process the sub-domains in parallel, and enable the CAE to learn local and generalizable nonlinear mappings of pore-fluid velocities. The second consists of embedding the finite difference solutions of the incompressible Navier-Stokes and continuity equations into convolutional layers prior to the CAE in order to provide the CAE with knowledge of fluid dynamics physics (PhyFlow). The third main novelty is that the training of the CAE is regularized with a hierarchical loss function that encourages the learning of fluid flow patterns (in a way similar to ranked modes in principal component analysis), ranking from most to least important. This is shown to increase the stability in learning, reduce over-fitting, and promote interpretability of the CAE neural network layers (HierCAE). The comprehensive new data-driven model, which we call the PhyFlow-HierCAE model, is shown to exhibit improved accuracy and generalizability of flow field predictions over conventional CAE models, attributable to the embedded physical knowledge and the hierarchical regularization, as well as realize orders of magnitude speed-ups in computation times as a surrogate for the direct numerical simulations. Examples of training and forward predictions on unseen pore-structures are provided and evaluated for data from Lattice Boltzmann and molecular dynamics simulations of pore-fluid flow. The model is shown to be a fast and accurate emulator (or “surrogate”) for predicting effective permeability of unseen pore-structures based on learning from relatively small direct numerical simulation datasets.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Group-equivariant autoencoder for identifying spontaneously broken symmetries

We introduce the group-equivariant autoencoder (GE autoencoder), a deep neural network (DNN) method that locates phase boundaries by determining which symmetries of the Hamiltonian have spontaneously broken at each temperature. We use group theory to deduce which symmetries of the system remain intact in all phases, and then use this information to constrain the parameters of the GE autoencoder such that the encoder learns an order parameter invariant to these “never-broken” symmetries. This procedure produces a dramatic reduction in the number of free parameters such that the GE-autoencoder size is independent of the system size. We include symmetry regularization terms in the loss function of the GE autoencoder so that the learned order parameter is also equivariant to the remaining symmetries of the system. By examining the group representation by which the learned order parameter transforms, we are then able to extract information about the associated spontaneous symmetry breaking. We test the GE autoencoder on the 2D classical ferromagnetic and antiferromagnetic Ising models, finding that the GE autoencoder (1) accurately determines which symmetries have spontaneously broken at each temperature; (2) estimates the critical temperature in the thermodynamic limit with greater accuracy, robustness, and time efficiency than a symmetry-agnostic baseline autoencoder; and (3) detects the presence of an external symmetry-breaking magnetic field with greater sensitivity than the baseline method. Lastly, we describe various key implementation details, including a quadratic-programming-based method for extracting the critical temperature estimate from trained autoencoders and calculations of the DNN initialization and learning rate settings required for fair model comparisons.

42 ENGINEERING↗

O3BNN-R: An Out-Of-Order Architecture for High-Performance and Regularized BNN Inference

Binarized Neural Networks (BNN) have drawn tremendous attention due to significantly reduced computational complexity and memory demand. They have especially shown great potential in cost- and power-restricted domains, such as IoT and smart edge-devices, where reaching a certain accuracy bar is often sufficient, and real-time is highly desired.In this work, we demonstrate that the highly-condensed BNN model can be shrunk significantly further by dynamically pruning irregular redundant edges. Based on two new observations on BNN-specific properties, an out-of-order (OoO) architecture – O3BNN-R, can curtail edge evaluation in cases where the binary output of a neuron can be determined early. Similar to Instruction-Level-Parallelism(ILP), these fine-grained, irregular, runtime pruning opportunities are traditionally presumed to be difficult to exploit. In order to increase the pruning opportunities, we also optimize the training process by adding 2 regularization items in the loss function (1) for pooling pruning and (2) for threshold pruning. We evaluate our design on an FPGA platform using three well-known networks, including VggNet-16, AlexNet for ImageNet, and a VGG-like network for Cifar-10.

Geng, Tong↗

Keeping a Beat on the Heart

Feel the relief of a patient suffering from heart arrhythmia, who is able to return home while having her heart monitored by health professionals 24 hours a day, without the fear that she will miss an important indicator and suffer a fatal heart attack - using technology originally developed to conduct experiments on the Space Shuttle. Approximately 400,000 Americans die every year from sudden heart attacks . Medical research revealed that patterns of electrical activity in the heart can act as predictors of these lethal cardiac events known as arrhythmias. Fortunately, certain arrhythmias such as ventricular fibrillation (loss of regular heartbeat and subsequent loss of function) and ventricular tachycardia (rapid heartbeats), can be detected and appropriately treated. Today, patients at moderate risk of arrhythmias can benefit from technology that would permit long- term continuous monitoring of electrical cardiac rhythms outside the hospital environment in the comfort of their own homes. Medical telemetry systems, also known as telemedicine, are evolving rapidly as wireless communication technology advances, evidenced by the commercial products and research prototypes for remote health monitoring that have appeared in recent years. Wireless systems allow patients to move freely in their home and work environment while being monitored remotely by health care professionals.

Liszka, Kathy J.↗

Development of a Deep Learning Model for Predicting the Drag Coefficients of Spherical and Non-Spherical Particles,

There is yet to be a well-established drag model for non-spherical particles required in a particle-laden flow that could cover a wide range of sphericities. This talk will explore the development of a general drag model for non-spherical particles by applying deep learning using available experimental data available in the literature. The integration of several raw experimental measurements from different sources and research directions allows the training of robust Artificial Intelligence and Machine Learning (ML) models. Neural networks are an ML approach inspired by the inner biological workings of the brain. This work aims to develop a Deep Neural Network (DNN) that predicts drag coefficient values with the ability to adapt appropriately to unseen data. Given the limited number of data points available and the variance found within the data collected from various sources, challenges may arise when looking to train the model. Our study tests and implements various model regularization techniques and assesses different loss and activation functions for the proposed DNN. The proposed model considers a broader range of features other than sphericity and Reynold number. These features include density ratio, solid volume fraction, lengthwise and crosswise sphericity, and more. Furthermore, we present the features that play a significant role in predicting different drag coefficients through feature importance. Within the investigated parameter ranges in this study, the following conclusions can be achieved and summarized below: • An improved drag coefficient model can be developed by considering more features such as, aspect ratio, lengthwise sphericity, crosswise sphericity, and density ratio. • DNN model can predict better results compared to traditional methods using MAE metric. • The proposed model addresses data challenges such as limited data and extreme data points through expanded feature-set and regularization. • Three major features that mostly affect the drag coefficient were identified from a feature importance analysis.

Presa-Reyes, Maria↗

Magnetic field mapping of inaccessible regions using physics-informed neural networks

A difficult problem concerns the determination of magnetic field components within an experimentally inaccessible region when direct field measurements are not feasible. In this paper, we propose a new method of accessing magnetic field components using non-disruptive magnetic field measurements on a surface enclosing the experimental region. Magnetic field components in the experimental region are predicted by solving a set of partial differential equations (Ampere’s law and Gauss’ law for magnetism) numerically with the aid of physics-informed neural networks (PINNs). Prediction errors due to noisy magnetic field measurements and small number of magnetic field measurements are regularized by the physics information term in the loss function. We benchmark our model by comparing it with an older method. The new method we present will be of broad interest to experiments requiring precise determination of magnetic field components, such as searches for the neutron electric dipole moment.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Multi-resolution partial differential equations preserved learning framework for spatiotemporal dynamics

Traditional data-driven deep learning models often struggle with high training costs, error accumulation, and poor generalizability in complex physical processes. Physics-informed deep learning (PiDL) addresses these challenges by incorporating physical principles into the model. Most PiDL approaches regularize training by embedding governing equations into the loss function, yet this depends heavily on extensive hyperparameter tuning to weigh each loss term. To this end, we propose to leverage physics prior knowledge by “baking” the discretized governing equations into the neural network architecture via the connection between the partial differential equations (PDE) operators and network structures, resulting in a PDE-preserved neural network (PPNN). This method, embedding discretized PDEs through convolutional residual networks in a multi-resolution setting, largely improves the generalizability and long-term prediction accuracy, outperforming conventional black-box models. The effectiveness and merit of the proposed methods have been demonstrated across various spatiotemporal dynamical systems governed by spatiotemporal PDEs, including reaction-diffusion, Burgers’, and Navier-Stokes equations.

97 MATHEMATICS AND COMPUTING↗

Jensen–Shannon divergence based novel loss functions for Bayesian neural networks

Bayesian neural networks (BNNs) are state-of-the-art machine learning methods that can naturally regularize and systematically quantify uncertainties using their stochastic parameters. Kullback–Leibler (KL) divergence-based variational inference used in BNNs suffer from unstable optimization and challenges in approximating light-tailed posteriors due to the unbounded nature of the KL divergence. To resolve these issues, we formulate a novel loss function for BNNs based on a new modification to the generalized Jensen–Shannon (JS) divergence, which is bounded. In addition, we propose a Geometric JS divergence-based loss, which is computationally efficient since it can be evaluated analytically. We found that the JS divergence-based variational inference is intractable, and hence employed a constrained optimization framework to formulate these losses. Our theoretical analysis and empirical experiments on multiple regression and classification data sets suggest that the proposed losses perform better than the KL divergence-based loss, especially when the data sets are noisy or biased. Specifically, there are approximately 5% and 8% improvements in accuracy for a noise-added CIFAR-10 dataset and a regression dataset, respectively. There is about 13% reduction in false negative predictions of a biased histopathology dataset. Additionally, we quantify and compare the uncertainty metrics for the regression and classification tasks.

97 MATHEMATICS AND COMPUTING↗

Short term hearing loss in general aviation operations, phase 1, part 1

The effects of light aircraft noise on six subjects during flight operations were investigated. The noise environment in the Piper Apache light aircraft was found to be capable of producing hearing threshold shifts. The following are the principal findings and conclusions: (1) Through most of the frequency range for which measurements were taken (500 to 6000 Hz), there was a regular progression showing increased loss of auditory acuity as a function of increased exposure time. (2) Extensive variability was found in the results among subjects, and in the measured loss at discrete frequencies for each subject. (3) The principal loss of hearing occurred at the low frequencies, around 500 Hz.

Parker, J. F., Jr.↗