Search NASA⌕ Search

SEARCH · Search NASA

Results for “Knowledge constraint”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Posterior Regularized Bayesian Neural Network incorporating soft and hard knowledge constraints

Neural Networks (NNs) have been widely used in supervised learning due to their ability to model complex nonlinear patterns, often presented in high-dimensional data such as images and text. However, traditional NNs often lack the ability for uncertainty quantification. Bayesian NNs (BNNS) could help measure the uncertainty by considering the distributions of the NN model parameters. Besides, domain knowledge is commonly available and could improve the performance of BNNs if it can be appropriately incorporated. In this work, we propose a novel Posterior-Regularized Bayesian Neural Network (PR-BNN) model by incorporating different types of knowledge constraints, such as the soft and hard constraints, as a posterior regularization term. Furthermore, we propose to combine the augmented Lagrangian method and the existing BNN solvers for efficient inference. Furthermore, the experiments in simulation and two case studies about aviation landing prediction and solar energy output prediction have shown the knowledge constraints and the performance improvement of the proposed model over traditional BNNs without the constraints.

14 SOLAR ENERGY↗

Machine learning with knowledge constraints for process optimization of open-air perovskite solar cell manufacturing

Perovskite photovoltaics (PV) have achieved rapid development in the past decade in terms of power conversion efficiency of small-area lab-scale devices; however, successful commercialization still requires further development of low-cost, scalable, and high-throughput manufacturing techniques. One of the critical challenges of developing a new fabrication technique is the high-dimensional parameter space for optimization, but machine learning (ML) can readily be used to accelerate perovskite PV scaling. Herein, we present an ML-guided framework of sequential learning for manufacturing process optimization. We apply our methodology to the Rapid Spray Plasma Processing (RSPP) technique for perovskite thin films in ambient conditions. With a limited experimental budget of screening 100 process conditions, we demonstrated an efficiency improvement to 18.5% as the best-in-our-lab device fabricated by RSPP, and we also experimentally found 10 unique process conditions to produce the top-performing devices of more than 17% efficiency, which is 5 times higher rate of success than the control experiments with pseudo-random Latin hypercube sampling. Our model is enabled by three innovations: (a) flexible knowledge transfer between experimental processes by incorporating data from prior experimental data as a probabilistic constraint; (b) incorporation of both subjective human observations and ML insights when selecting next experiments; (c) adaptive strategy of locating the region of interest using Bayesian optimization first, and then conducting local exploration for high-efficiency devices. Furthermore, in virtual benchmarking, our framework achieves faster improvements with limited experimental budgets than traditional design-of-experiments methods (e.g., one-variable-at-a-time sampling). This framework shows the capability of incorporating researchers’ domain knowledge into the ML-guided optimization loop; therefore, it has the potential to facilitate the wider adoption of ML in scaling to perovskite PV manufacturing.

14 SOLAR ENERGY↗

IntraShuffler: A Privacy Preserving Framework for Heterogeneous DP Federated Learning

Heterogeneous Differential Privacy (HDP) in Federated Learning (FL) allows clients to select individual privacy budgets () according to institutional policies and data sensitivity. In practice, many HDP-FL systems employ -aware server aggregation to improve model utility by re-weighting client updates according to their declared privacy budgets. However, gradient updates in FL retain structural patterns induced by non-independent and identically-distributed (non-IID) data, and these additional signals exposed by -aware aggregation create new opportunities for inference by an honest-but-curious server. In this work, we first show that a server equipped with gradient denoising and surrogate modeling can mount a Privacy Inference Attack that infers distributional attributes of clients and links updates from the same client across training rounds, measured via surrogate inference accuracy and linkage success, under realistic knowledge constraints. The Shuffle-Model has been widely studied as a defense against such inference risks by anonymizing update sources, but it is fundamentally incompatible with HDP-FL -aware aggregation. To address this challenge, we propose IntraShuffler, a middleware defense framework designed for HDP-FL systems. IntraShuffler introduces a privacy-aware shuffling mechanism that groups clients into privacy-compatible buckets and performs parameter-level shuffling within each bucket to disrupt persistent gradient structure while preserving -aware aggregation. Experiments across four different datasets show that IntraShuffler reduces gradient recoverability by over 60% and decreases surrogate inference accuracy from 0.78 to 0.33 while maintaining comparable model utility across multiple FL aggregation rules.

Riya, Farhin Farhad [ORNL]↗

The Hubble Constant from Strongly Lensed Supernovae with Standardizable Magnifications

The dominant uncertainty in the current measurement of the Hubble constant (H 0 ) with strong gravitational lensing time delays is attributed to uncertainties in the mass profiles of the main deflector galaxies. Strongly lensed supernovae (glSNe) can provide, in addition to measurable time delays, lensing magnification constraints when knowledge about the unlensed apparent brightness of the explosion is imposed. We present a hierarchical Bayesian framework to combine a data set of SNe that are not strongly lensed and a data set of strongly lensed SNe with measured time delays. We jointly constrain (i) H 0 using the time delays as an absolute distance indicator, (ii) the lens model profiles using the magnification ratio of lensed and unlensed fluxes on the population level, and (iii) the unlensed apparent magnitude distribution of the SN population and the redshift–luminosity relation of the relative expansion history of the universe. We apply our joint inference framework on a future expected data set of glSNe and forecast that a sample of 144 glSNe of Type Ia with well-measured time series and imaging data will measure H 0 to 1.5%. We discuss strategies to mitigate systematics associated with using absolute flux measurements of glSNe to constrain the mass density profiles. Using the magnification of SN images is a promising and complementary alternative to using stellar kinematics. Future surveys, such as the Rubin and Roman observatories, will be able to discover the necessary number of glSNe, and with additional follow-up observations, this methodology will provide precise constraints on mass profiles and H 0 .

79 ASTRONOMY AND ASTROPHYSICS↗

Lyman continuum leaker candidates among highly ionised, low-redshift dwarf galaxies selected from He II

Contemporary research suggests that the reionisation of the intergalactic medium (IGM) in the early Universe was predominantly realised by star-forming (proto-)galaxies (SFGs). Due to observational constraints, our knowledge on the origins of sufficient amounts of ionising Lyman continuum (LyC) photons and the mechanisms facilitating their transport into the IGM remains sparse. Recent efforts have thus focussed on the study of local analogues to these high-redshift objects. We aim to acquire a set of very low-redshift SFGs that exhibit signs of a hard radiation field being present. A subsequent analysis of their emission line properties is intended to shed light on how the conditions prevalent in these objects compare to those predicted to be present in early SFGs that are thought to be LyC emitters (LCEs). We used archival spectroscopic SDSS DR12 data to select a sample of low-redshift He II 4686 emitters and restricted it to a set of SFGs with an emission line diagnostic sensitive to the presence of an active galactic nucleus, which serves as our only selection criterion. We performed a population spectral synthesis with FADO to reconstruct these galaxies’ star-formation histories (SFHs). Utilising the spectroscopic information at hand, we constrained the predominant ionisation mechanisms in these galaxies and inferred information on interstellar medium (ISM) conditions relevant for the escape of LyC radiation. Our final sample consists of eighteen ionised, metal-poor galaxies (IMPs). These low-mass (6.2 ≤ log(M * /M ⊙ ) ≤ 8.8), low-metallicity (7.54 ≤log(O/H) + 12 ≤ 8.13) dwarf galaxies appear to be predominantly ionised by stellar sources. We find large [OIII] 5007/[OII] 3727 ratios and [SII] 6717,6731/Hα deficiencies, which provide strong indications for these galaxies to be LCEs. At least 40% of these objects are candidates for featuring cosmologically significant LyC escape fractions ≳10%. The IMPs’ SFHs exhibit strong similarities and almost all galaxies appear to contain an old (> 1 Gyr) stellar component, while also harbouring a young, two-stage (~10 Myr and < 1 Myr) starburst, which we speculate might be related to LyC escape. The properties of the compact emission line galaxies presented here align well with those observed in many local LCEs. In fact, our sample may prove as an extension to the rather small catalogue of local LCEs, as the extreme ISM conditions we find are assumed to facilitate LyC leakage. Notably, all of our eighteen candidates are significantly closer (z < 0.1) than most established LCEs. If the inferred LyC photon loss is genuine, this demonstrates that selecting SFGs from He II 4686 is a powerful selection criterion in the search for LCEs.

79 ASTRONOMY AND ASTROPHYSICS↗

Physics-Guided, Physics-Informed, and Physics-Encoded Neural Networks and Operators in Scientific Computing: Fluid and Solid Mechanics

Abstract Advancements in computing power have recently made it possible to utilize machine learning and deep learning to push scientific computing forward in a range of disciplines, such as fluid mechanics, solid mechanics, materials science, etc. The incorporation of neural networks is particularly crucial in this hybridization process. Due to their intrinsic architecture, conventional neural networks cannot be successfully trained and scoped when data are sparse, which is the case in many scientific and engineering domains. Nonetheless, neural networks provide a solid foundation to respect physics-driven or knowledge-based constraints during training. Generally speaking, there are three distinct neural network frameworks to enforce the underlying physics: (i) physics-guided neural networks (PgNNs), (ii) physics-informed neural networks (PiNNs), and (iii) physics-encoded neural networks (PeNNs). These methods provide distinct advantages for accelerating the numerical modeling of complex multiscale multiphysics phenomena. In addition, the recent developments in neural operators (NOs) add another dimension to these new simulation paradigms, especially when the real-time prediction of complex multiphysics systems is required. All these models also come with their own unique drawbacks and limitations that call for further fundamental research. This study aims to present a review of the four neural network frameworks (i.e., PgNNs, PiNNs, PeNNs, and NOs) used in scientific computing research. The state-of-the-art architectures and their applications are reviewed, limitations are discussed, and future research opportunities are presented in terms of improving algorithms, considering causalities, expanding applications, and coupling scientific and deep learning solvers.

Computer Science↗

Comprehension of Spatial Constraints by Neural Logic Learning from a Single RGB-D Scan

Autonomous industrial assembly relies on the precise measurement of spatial constraints as designed by computer-aided design (CAD) software such as SolidWorks. This paper proposes a framework for an intelligent industrial robot to understand the spatial constraints for model assembly. An extended generative adversary network (GAN) with a 3D long short-term memory (LSTM) network was designed to composite 3D point clouds from a single RGB-D scan. The spatial constraints of the segmented point clouds are identified by a neural-logic network that incorporates general knowledge of spatial constraints in terms of first-order logic. The model was designed to comprehend a complete set of spatial constraints that are consistent with industrial CAD software, including left, right, above, below, front, behind, parallel, perpendicular, concentric, and coincident relations. The accuracy of 3D model composition and spatial constraint identification was evaluated by the RGB-D scans and 3D models in the ABC dataset. The proposed model achieved 57.23% intersection over union (IoU) in 3D model composition, and over 99% in comprehending all spatial constraints.

Wang, Dali↗

Combining data and theory for derivable scientific discovery with AI-Descartes

Abstract Scientists aim to discover meaningful formulae that accurately describe experimental data. Mathematical models of natural phenomena can be manually created from domain knowledge and fitted to data, or, in contrast, created automatically from large datasets with machine-learning algorithms. The problem of incorporating prior knowledge expressed as constraints on the functional form of a learned model has been studied before, while finding models that are consistent with prior knowledge expressed via general logical axioms is an open problem. We develop a method to enable principled derivations of models of natural phenomena from axiomatic knowledge and experimental data by combining logical reasoning with symbolic regression. We demonstrate these concepts for Kepler’s third law of planetary motion, Einstein’s relativistic time-dilation law, and Langmuir’s theory of adsorption. We show we can discover governing laws from few data points when logical reasoning is used to distinguish between candidate formulae having similar error on the data.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Extreme sparsification of physics-augmented neural networks for interpretable model discovery in mechanics

Data-driven constitutive modeling with neural networks has received increased interest in recent years due to its ability to easily incorporate physical and mechanistic constraints and to overcome the challenging and time-consuming task of formulating phenomenological constitutive laws that can accurately capture the observed material response. However, even though neural network-based constitutive laws have been shown to generalize proficiently, the generated representations are not easily interpretable due to their high number of trainable parameters. Sparse regression approaches exist that allow for obtaining interpretable expressions, but the user is tasked with creating a library of model forms which by construction limits their expressiveness to the functional forms provided in the libraries. Here, in this work, we propose to train regularized physics-augmented neural network-based constitutive models utilizing a smoothed version of $L^0$-regularization. This aims to maintain the trustworthiness inherited by the physical constraints, but also enables interpretability which has not been possible thus far on any type of machine learning-based constitutive model where model forms were not assumed a priori but were actually discovered. During the training process, the network simultaneously fits the training data and penalizes the number of active parameters, while also ensuring constitutive constraints such as thermodynamic consistency. We show that the method can reliably obtain interpretable and trustworthy constitutive models for compressible and incompressible hyperelasticity, yield functions, and hardening models for elastoplasticity, using synthetic and experimental data. This work aims to set a new paradigm for interpretable machine learning models in the broad area of solid mechanics where low and limited data is available along with prior knowledge of physical constraints that the learned maps need to obey. This paradigm can potentially be extended to a broader spectrum of scientific exploration.

Data-driven constitutive models↗

PolyODENet: Deriving mass-action rate equations from incomplete transient kinetics data

Kinetics of a reaction network that follows mass-action rate laws can be described with a system of ordinary differential equations (ODEs) with polynomial right-hand side. However, it is challenging to derive such kinetic differential equations from transient kinetic data without knowing the reaction network, especially when the data are incomplete due to experimental limitations. We introduce a program, PolyODENet, toward this goal. Based on the machine-learning method Neural ODE, PolyODENet defines a generative model and predicts concentrations at arbitrary time. As such, it is possible to include unmeasurable intermediate species in the kinetic equations. Importantly, we have implemented various measures to apply physical constraints and chemical knowledge in the training to regularize the solution space. Here, using simple catalytic reaction models, we demonstrate that PolyODENet can predict reaction profiles of unknown species and doing so even reveal hidden parts of reaction mechanisms.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Simulation Study of High-Precision Characterization of MeV Electron Interactions for Advanced Nano-Imaging of Thick Biological Samples and Microchips

The resolution of a mega-electron-volt scanning transmission electron microscope (MeV-STEM) is primarily governed by the properties of the incident electron beam and angular broadening effects that occur within thick biological samples and microchips. A precise understanding and mitigation of these constraints require detailed knowledge of beam emittance, aberrations in the STEM column optics, and energy-dependent elastic and inelastic critical angles of the materials being examined. This simulation study proposes a standardized experimental framework for comprehensively assessing beam intensity, divergence, and size at the sample exit. This framework aims to characterize electron-sample interactions, reconcile discrepancies among analytical models, and validate Monte Carlo (MC) simulations for enhanced predictive accuracy. Our numerical findings demonstrate that precise measurements of these parameters, especially angular broadening, are not only feasible but also essential for optimizing imaging resolution in thick biological samples and microchips. By utilizing an electron source with minimal emittance and tailored beam characteristics, along with amorphous ice and silicon samples as biological proxies and microchip materials, this research seeks to optimize electron beam energy by focusing on parameters to improve the resolution in MeV-STEM/TEM. This optimization is particularly crucial for in situ imaging of thick biological samples and for examining microchip defects with nanometer resolutions. Our ultimate goal is to develop a comprehensive mapping of the minimum electron energy required to achieve a nanoscale resolution, taking into account variations in sample thickness, composition, and imaging mode.

36 MATERIALS SCIENCE↗

PolyODENet

Kinetics of a reaction network that follows mass-action rate laws can be described with a system of ordinary differential equations (ODEs) with polynomial right-hand side. However, it is challenging to derive such kinetic differential equations from transient kinetic data without knowing the reaction network, especially when the data are incomplete due to experimental limitations. We introduce a program, PolyODENet, toward this goal. Based on the machine-learning method Neural ODE, PolyODENet defines a generative model and predicts concentrations at arbitrary time. As such, it is possible to include unmeasurable intermediate species in the kinetic equations. Importantly, we have implemented various measures to apply physical constraints and chemical knowledge in the training to regularize the solution space.

Wu, Qin [Brookhaven National Lab. (BNL), Upton, NY↗

Latent Mechanisms of Polarization Switching from In Situ Electron Microscopy Observations

In situ scanning transmission electron microscopy enables observation of the domain dynamics in ferroelectric materials as a function of externally applied bias and temperature. The resultant data sets contain a wealth of information on polarization switching and phase transition mechanisms. However, identification of these mechanisms from observational data sets has remained a problem due to a large variety of possible configurations, many of which are degenerate. Here, an approach based on a combination of deep learning-based semantic segmentation, rotationally invariant variational autoencoder (VAE), and non-negative matrix factorization to enable learning of a latent space representation of the data with multiple real-space rotationally equivalent variants mapped to the same latent space descriptors is introduced. By varying the size of training sub-images in the VAE, the degree of complexity in the structural descriptors is tuned from simple domain wall detection to the identification of switching pathways. Importantly, this yields a powerful tool for the exploration of the dynamic data in mesoscopic electron, scanning probe, optical, and chemical imaging. Moreover, this work adds to the growing body of knowledge of incorporating physical constraints into the machine and deep-learning methods to improve learned descriptors of physical phenomena.

36 MATERIALS SCIENCE↗

Dynamic Validation of CNN-Based Surrogate Models for Inverter-Based Resources in Open-Source Solvers

Traditionally, distribution system planning has focused on steady-state analyses, with limited consideration of dynamic behavior. However, as large or medium-scale inverter-based resources (IBRs), particularly grid-following (GFL) inverters in commercial or industry buildings, become more prevalent, understanding their dynamic impact is essential for grid planning and operation. This article presents an innovative deep-learning (DL)-approach using convolutional neural networks technique to model the GFL inverters. Developed from real grid-tied commercial IBR transient data, these dynamic DL models overcome proprietary constraints by requiring minimal knowledge of internal converter physics while maintaining high accuracy and flexibility. To demonstrate their applicability, the models were incorporated into GridLAB-D, an open-source, three-phase distribution analysis tool. This integration enables dynamic simulations of large-scale distribution networks with high IBR penetration stability analysis. Rigorous testing and validation, aligned with industry standards, confirmed the reliability and efficiency of this approach, paving the way for enhanced planning and operational assessments of modern power systems.

Deep-learning↗

Acceleration of Power System Dynamic Simulations Using a Deep Equilibrium Layer and Neural ODE Surrogate

The dominant paradigm for power system dynamic simulation is to build system-level simulations by combining physics-based models of individual components. The sheer size of the system along with the rapid integration of inverter-based resources exacerbates the computational burden of running time domain simulations. Here, in this paper, we propose a data-driven surrogate model based on implicit machine learningspecifically deep equilibrium layers and neural ordinary differential equationsto learn a reduced order model of a portion of the full underlying system. The data-driven surrogate achieves similar accuracy and reduction in simulation time compared to a physics-based surrogate, without the constraint of requiring detailed knowledge of the underlying dynamic models. This work also establishes key requirements needed to integrate the surrogate into existing simulation workflows; the proposed surrogate is initialized to a steady state operating point that matches the power flow solution by design.

Neural ordinary differential equations↗

EMMA: a new method for computing multiple sequence alignments given a constraint subset alignment

Abstract Background Adding sequences into an existing (possibly user-provided) alignment has multiple applications, including updating a large alignment with new data, adding sequences into a constraint alignment constructed using biological knowledge, or computing alignments in the presence of sequence length heterogeneity. Although this is a natural problem, only a few tools have been developed to use this information with high fidelity. Results We present EMMA (Extending Multiple alignments using MAFFT--add) for the problem of adding a set of unaligned sequences into a multiple sequence alignment (i.e., a constraint alignment). EMMA builds on MAFFT--add, which is also designed to add sequences into a given constraint alignment. EMMA improves on MAFFT--add methods by using a divide-and-conquer framework to scale its most accurate version, MAFFT-linsi--add, to constraint alignments with many sequences. We show that EMMA has an accuracy advantage over other techniques for adding sequences into alignments under many realistic conditions and can scale to large datasets with high accuracy (hundreds of thousands of sequences). EMMA is available at https://github.com/c5shen/EMMA . Conclusions EMMA is a new tool that provides high accuracy and scalability for adding sequences into an existing alignment.

Shen, Chengze↗

Path sampling of recurrent neural networks by incorporating known physics

Recurrent neural networks have seen widespread use in modeling dynamical systems in varied domains such as weather prediction, text prediction and several others. Often one wishes to supplement the experimentally observed dynamics with prior knowledge or intuition about the system. While the recurrent nature of these networks allows them to model arbitrarily long memories in the time series used in training, it makes it harder to impose prior knowledge or intuition through generic constraints. In this work, we present a path sampling approach based on principle of Maximum Caliber that allows us to include generic thermodynamic or kinetic constraints into recurrent neural networks. We show the method here for a widely used type of recurrent neural network known as long short-term memory network in the context of supplementing time series collected from different application domains. These include classical Molecular Dynamics of a protein and Monte Carlo simulations of an open quantum system continuously losing photons to the environment and displaying Rabi oscillations. Our method can be easily generalized to other generative artificial intelligence models and to generic time series in different areas of physical and social sciences, where one wishes to supplement limited data with intuition or theory based corrections.

59 BASIC BIOLOGICAL SCIENCES↗