Search NASASearch

SEARCH · Search NASA

Results for “Conformal prediction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

D–MOPH–25: diverse MOF–molecule pairs for Henry’s constants prediction

Computational methods like grand-canonical Monte Carlo simulations and machine learning (ML) have accelerated metal–organic frameworks (MOF) exploration but are typically limited to a narrow range of adsorbates due to data availability and force field constraints. In this study, we introduce a dataset of diverse MOF–molecule pairs for Henry’s constant prediction, D–MOPH–25, which systematically explores a diverse chemical space by combining 113 molecular adsorbates with over 5000 MOF structures through an active learning process. D–MOPH–25 constitutes the most diverse adsorbate dataset used in any ML study of molecular adsorption in MOFs to date. Our workflow builds a benchmark for predicting Henry’s constants at 300 K, leveraging conformal prediction for uncertainty quantification. Assessment through Shannon entropy and uniform manifold approximation and projection confirms the comprehensiveness of D–MOPH–25 while highlighting the importance of robust classification to filter out unphysical data points in regression tasks. Although future enhancements in model architecture and sampling criteria could improve predictive performance, our dataset already spans the target space using only 2.31% of total possibilities. This comprehensive dataset facilitates assessment of model generalizability across adsorbate species and can establish a foundation for high-throughput MOF screening and ML-driven separation processes.

active learning

An investigation on machine learning predictive accuracy improvement and uncertainty reduction using VAE-based data augmentation

The confluence of ultrafast computers with large memory, rapid progress in Machine Learning (ML) algorithms, and the availability of large datasets place multiple engineering fields at the threshold of dramatic progress. However, a unique challenge in nuclear engineering is data scarcity because experimentation on nuclear systems is usually more expensive and time-consuming than most other disciplines. One potential way to resolve the data scarcity issue is deep generative learning, which uses certain ML models to learn the underlying distribution of existing data and generate synthetic samples that resemble the real data. In this way, one can significantly expand the dataset to train more accurate predictive ML models. In this study, our objective is to evaluate the effectiveness of data augmentation using variational autoencoder (VAE)-based deep generative models. We investigated whether the data augmentation leads to improved accuracy in the predictions of a deep neural network (DNN) model trained using the augmented data. Additionally, the DNN prediction uncertainties are quantified using Bayesian Neural Networks (BNN) and conformal prediction (CP) to assess the impact on predictive uncertainty reduction. To test the proposed methodology, we used TRACE simulations of steady-state void fraction data based on the NUPEC Boiling Water Reactor Full-size Fine-mesh Bundle Test (BFBT) benchmark. Here, we found that augmenting the training dataset using VAEs has improved the DNN model’s predictive accuracy, improved the prediction confidence intervals, and reduced the prediction uncertainties.

Bayesian neural network

Conformalized-KANs: Uncertainty Quantification with Coverage Guarantees for Kolmogorov-Arnold Networks (KANs) in Scientific Machine Learning

This paper explores uncertainty quantification (UQ) methods in the context of Kolmogorov–Arnold Networks (KANs). We apply an ensemble approach to KANs to obtain a heuristic measure of UQ, enhancing interpretability and robustness in modeling complex functions. Building on this, we introduce Conformalized-KANs, which integrate conformal prediction, a distribution-free UQ technique, with KAN ensembles to generate calibrated prediction intervals with guaranteed coverage.} Extensive numerical experiments are conducted to evaluate the effectiveness of these methods, focusing particularly on the robustness and accuracy of the prediction intervals under various hyperparameter settings. We show that the conformal KAN predictions can be applied to recent extensions of KANs, including Finite Basis KANs (FBKANs) and multifideilty KANs (MFKANs). The results demonstrate the potential of our approaches to significantly improve the reliability and applicability of KANs in scientific machine learning.

• Artificial intelligence (AI) / machine learning

Uncertainty Quantification for Data-Driven Machine Learning Models in Nuclear Engineering Applications: Where We Are and What Do We Need?

Machine learning (ML) has been leveraged to tackle a diverse range of tasks in almost all branches of nuclear engineering. Many of the successes in ML applications can be attributed to the recent performance breakthroughs in deep learning, the growing availability of computational power, data, and easy-to-use ML libraries. However, these empirical successes have often outpaced our formal understanding of the ML algorithms. An important but under-rated area is uncertainty quantification (UQ) of ML. ML-based models are subject to approximation uncertainty when they are used to make predictions, due to sources including but not limited to, data noise, data coverage, extrapolation, imperfect model architecture and the stochastic training process. The goal of this paper is to clearly explain and illustrate the importance of UQ of ML. We will elucidate the differences in the basic concepts of UQ of physics-based models and data-driven ML models. Various sources of uncertainties in physical modeling and data-driven modeling will be discussed, demonstrated, and compared. We will also present and demonstrate a few techniques to quantify the ML prediction uncertainties, including Monte Carlo dropout, deep ensemble, Bayesian neural networks, Gaussian Processes and conformal prediction. Lastly, we will discuss the need for building a verification, validation and UQ framework to establish ML credibility.

22 GENERAL STUDIES OF NUCLEAR REACTORS

mphys-surrogate-model

This repository contains python scripts for building and studying reduced-order-modeling representations of droplet coalescence for eventual use in atmospheric models. The included data are generated from high-fidelity superdroplet methods and are utilized by machine learning pipelines to build data-driven models of droplet size distributions that evolve under coalescence. This repository further includes scripts to determine prediction (uncertainty) intervals on the data-driven model products based on conformal prediction.

Katona, JonasE [Lawrence Livermore National Labora

A Comparison of Metamodeling Techniques via Numerical Experiments

This paper presents a comparative analysis of a few metamodeling techniques using numerical experiments for the single input-single output case. These experiments enable comparing the models' predictions with the phenomenon they are aiming to describe as more data is made available. These techniques include (i) prediction intervals associated with a least squares parameter estimate, (ii) Bayesian credible intervals, (iii) Gaussian process models, and (iv) interval predictor models. Aspects being compared are computational complexity, accuracy (i.e., the degree to which the resulting prediction conforms to the actual Data Generating Mechanism), reliability (i.e., the probability that new observations will fall inside the predicted interval), sensitivity to outliers, extrapolation properties, ease of use, and asymptotic behavior. The numerical experiments describe typical application scenarios that challenge the underlying assumptions supporting most metamodeling techniques.

Crespo, Luis G.

Repetitive proteins that undergo large conformational changes evade structural prediction algorithms

Protein structure prediction algorithms, such as AlphaFold, have accelerated protein design and advanced the understanding of the relationship between amino acid sequence and protein structure. However, these algorithms are limited in their ability to predict the structures of conformationally dynamic, intrinsically disordered, and stimuli-responsive proteins. To evaluate sequence-to-structure predictions of such challenging proteins, we explored a class of conformationally dynamic, repeats-in-toxin (RTX) proteins. RTX proteins adopt intrinsically disordered conformations in the absence of calcium and undergo reversible folding into β-roll structures upon binding to calcium. RTX proteins are characterized by tandem repeats of the sequence GGXGXDXUX, in which X can be any amino acid and U is an aliphatic amino acid. We designed RTX sequence variants with global substitutions of nonconserved amino acids, tandem repeats of consensus sequences GGAGXDTLY, and tandem repeats of scrambled sequences GGAGXDTYL. AlphaFold2 and AlphaFold3 predicted that all of these RTX variants adopt β-roll structures, characteristic of wild-type RTX bound to calcium. However, modeling the predicted structures with molecular dynamics simulations and characterizing the protein variants with circular dichroism spectroscopy, small-angle x-ray scattering, and x-ray crystallography revealed that variants adopt diverse, sequence-dependent structures in the absence and presence of calcium. To better design proteins for applications in biotechnology and sustainability, it is critical to build predictive tools that consider intrinsically disordered protein states and validate these tools with multi-mode, multi-scale experimental data.

Chang, Marina P. [Stanford Univ., CA (United State

Functional protein mining with conformal guarantees

Molecular structure prediction and homology detection offer promising paths to discovering protein function and evolutionary relationships. However, current approaches lack statistical reliability assurances, limiting their practical utility for selecting proteins for further experimental and in-silico characterization. To address this challenge, we introduce a statistically principled approach to protein search leveraging principles from conformal prediction, offering a framework that ensures statistical guarantees with user-specified risk and provides calibrated probabilities (rather than raw ML scores) for any protein search model. Our method (1) lets users select many biologically-relevant loss metrics (i.e. false discovery rate) and assigns reliable functional probabilities for annotating genes of unknown function; (2) achieves state-of-the-art performance in enzyme classification without training new models; and (3) robustly and rapidly pre-filters proteins for computationally intensive structural alignment algorithms. Our framework enhances the reliability of protein homology detection and enables the discovery of uncharacterized proteins with likely desirable functional properties.

59 BASIC BIOLOGICAL SCIENCES

Regulation of body fluid volume and electrolyte concentrations in spaceflight

Despite a number of difficulties in performing experiments during weightlessness, a great deal of information has been obtained concerning the effects of spaceflight on the regulation of body fluid and electrolytes. Many paradoxes and questions remain, however. Although body mass, extracellular fluid volume, and plasma volume are reduced during spaceflight and remain so at landing, the changes in total body water are comparatively small. Serum or plasma sodium and osmolality have generally been unchanged or reduced during the spaceflight, and fluid intake is substantially reduced, especially during the first of flight. The diuresis that was predicted to be caused by weightlessness, has only rarely been observed as an increased urine volume. What has been well established by now, is the occurrence of a relative diuresis, where fluid intake decreases more than urine volume does. Urinary excretion of electrolytes has been variable during spaceflight, but retention of fluid and electrolytes at landing has been consistently observed. The glomerular filtration rate was significantly elevated during the SLS missions, and water and electrolyte loading tests have indicated that renal function is altered during readaptation to Earth's gravity. Endocrine control of fluid volumes and electrolyte concentrations may be altered during weightlessness, but levels of hormones in body fluids do not conform to predictions based on early hypotheses. Antidiuretic hormone is not suppressed, though its level is highly variable and its secretion may be affected by space motion sickness and environmental factors. Plasma renin activity and aldosterone are generally elevated at landing, consistent with sodium retention, but inflight levels have been variable. Salt intake may be an important factor influencing the levels of these hormones. The circadian rhythm of cortisol has undoubtedly contributed to its variability, and little is known yet about the influence of spaceflight on circadian rhythms. Atrial natriuretic peptide does not seem to play an important role in the control of natriuresis during spaceflight. Inflight activity of the sympathetic nervous system, assessed by measuring catecholamines and their metabolites and precursors in body fluids, generally seems to be no greater than on Earth, but this system is usually activated at landing. Collaborative experiments on the Mir and the International Space Station should provide more of the data needed from long-term flights, and perhaps help to resolve some of the discrepancies between U.S. and Russian data. The use of alternative methods that are easier to execute during spaceflight, such as collection of saliva instead of blood and urine, should permit more thorough study of circadian rhythms and rapid hormone changes in weightlessness. More investigations of dietary intake of fluid and electrolytes must be performed to understand regulatory processes. Additional hormones that may participate in these processes, such as other natriuretic hormones, should be determined during and after spaceflight. Alterations in body fluid volume and blood electrolyte concentrations during spaceflight have important consequences for readaptation to the 1-G environment. The current assessment of fluid and electrolyte status during weightlessness and at landing and our still incomplete understanding of the processes of adaptation to weightlessness and readaptation to Earth's gravity have resulted in the development of countermeasures that are only partly successful in reducing the postflight orthostatic intolerance experienced by astronauts and cosmonauts. More complete knowledge of these processes can be expected to produce countermeasures that are even more successful, as well as expand our comprehension of the range of adaptability of human physiologic processes.

NASA Discipline Number 18-10

From sequence to protein structure and conformational dynamics with artificial intelligence/machine learning

The 2024 Nobel Prize in Chemistry was awarded in part for de novo protein structure prediction using AlphaFold2, an artificial intelligence/machine learning (AI/ML) model trained on vast amounts of sequence and three-dimensional structure data. AlphaFold2 and related models, including RoseTTAFold and ESMFold, employ specialized neural network architectures driven by attention mechanisms to infer relationships between sequence and structure. At a fundamental level, these AI/ML models operate on the long-standing hypothesis that the structure of a protein is determined by its amino acid sequence. More recently, AlphaFold2 has been adapted for the prediction of multiple protein conformations by subsampling multiple sequence alignments. Herein, we provide an overview of the deterministic relationship between sequence and structure, which was hypothesized over half a century ago with profound implications for the biological sciences ever since. We postulate that protein conformational dynamics are also determined, at least in part, by amino acid sequence and that this relationship may be leveraged for construction of AI/ML models dedicated to predicting protein conformational ensembles. Accordingly, we describe a conceptual model architecture, which may be trained on sequence data in combination with conformationally sensitive structural information, coming primarily from nuclear magnetic resonance (NMR) spectroscopy. Notwithstanding certain limitations in this context, NMR offers abundant structural heterogeneity conducive to conformational ensemble prediction. As NMR and other data continue to accumulate, sequence-informed prediction of protein structural dynamics with AI/ML has the potential to emerge as a transformative capability across the biological sciences.

Artificial intelligence

Development of IR radiation simulator for spacecraft thermal testing

The aim was to simulate, in a ground test, the solar radiation environment to which the Ofeq satellite would be exposed in orbit. The solar simulator usually used is very expensive, as are its operation and maintenance, therefore an infrared (IR) simulator was used; its development involved the creation of uniform IR fluxes onto the irregular geometry of a spacecraft. The tests were carried out on a thermal model of the satellite and the model was verified in a solar simulator test at a European space center. Heat flux mapping software was developed to plan the positioning of the heating elements which generated the IR fluxes and a system was built for heat flux and temperature monitoring and control. A special heat flux sensor was developed to measure the energy absorbed by the satellite surfaces; its calibration had to be independent of wavelength, so that the measurements obtained at IR wavelengths would be equivalent to the actual solar radiation effects. The results of the verification experiment are presented; their conformity with predicted values indicates that the technique developed is a suitable tool for satellite design verification.

Shimrony, Yoram

Finding the global minimum: a fuzzy end elimination implementation

The 'fuzzy end elimination theorem' (FEE) is a mathematically proven theorem that identifies rotameric states in proteins which are incompatible with the global minimum energy conformation. While implementing the FEE we noticed two different aspects that directly affected the final results at convergence. First, the identification of a single dead-ending rotameric state can trigger a 'domino effect' that initiates the identification of additional rotameric states which become dead-ending. A recursive check for dead-ending rotameric states is therefore necessary every time a dead-ending rotameric state is identified. It is shown that, if the recursive check is omitted, it is possible to miss the identification of some dead-ending rotameric states causing a premature termination of the elimination process. Second, we examined the effects of removing dead-ending rotameric states from further considerations at different moments of time. Two different methods of rotameric state removal were examined for an order dependence. In one case, each rotamer found to be incompatible with the global minimum energy conformation was removed immediately following its identification. In the other, dead-ending rotamers were marked for deletion but retained during the search, so that they influenced the evaluation of other rotameric states. When the search was completed, all marked rotamers were removed simultaneously. In addition, to expand further the usefulness of the FEE, a novel method is presented that allows for further reduction in the remaining set of conformations at the FEE convergence. In this method, called a tree-based search, each dead-ending pair of rotamers which does not lead to the direct removal of either rotameric state is used to reduce significantly the number of remaining conformations. In the future this method can also be expanded to triplet and quadruplet sets of rotameric states. We tested our implementation of the FEE by exhaustively searching ten protein segments and found that the FEE identified the global minimum every time. For each segment, the global minimum was exhaustively searched in two different environments: (i) the segments were extracted from the protein and exhaustively searched in the absence of the surrounding residues; (ii) the segments were exhaustively searched in the presence of the remaining residues fixed at crystal structure conformations. We also evaluated the performance of the method for accurately predicting side chain conformations. We examined the influence of factors such as type and accuracy of backbone template used, and the restrictions imposed by the choice of potential function, parameterization and rotamer database. Conclusions are drawn on these results and future prospects are given.

NASA Program Exobiology

Cholesterol modulates membrane elasticity via unified biophysical laws

Cholesterol and lipid unsaturation underlie a balance of opposing forces that features prominently in adaptive cell responses to diet and environmental cues. These competing factors have resulted in contradictory observations of membrane elasticity across different measurement scales, requiring chemical specificity to explain incompatible structural and elastic effects. Here, we demonstrate that – unlike macroscopic observations – lipid membranes exhibit a unified elastic behavior in the mesoscopic regime between molecular and macroscopic dimensions. Using nuclear spin techniques and computational analysis, we find that mesoscopic bending moduli follow a universal dependence on the lipid packing density regardless of cholesterol content, lipid unsaturation, or temperature. Our observations reveal that compositional complexity can be explained by simple biophysical laws that directly map membrane elasticity to molecular packing associated with biological function, curvature transformations, and protein interactions. The obtained scaling laws closely align with theoretical predictions based on conformational chain entropy and elastic stress fields. These findings provide unique insights into the membrane design rules optimized by nature and unlock predictive capabilities for guiding the functional performance of lipid-based materials in synthetic biology and real-world applications.

Kumarage, Teshani [Virginia Polytechnic Inst. and

Beyond Point Estimates: Benchmarking Uncertainty Quantification Methods on the AION-1 Astronomical Foundation Model

Foundation models for astronomical surveys offer powerful learned representations that can be transferred to downstream regression tasks such as galaxy property estimation. However, point predictions alone are insufficient for scientific inference; reliable uncertainty quantification (UQ) is essential. We compare seven UQ methods on galaxy property regression using frozen AION-1 foundation-model embeddings, predicting redshift, stellar mass, stellar-population age, gas-phase metallicity, and specific star-formation rate, from Legacy Survey photometry/imaging and DESI spectra, with PROVABGS-derived labels. Distribution-free conformal methods achieve marginal coverage within $\sim$1 pp of the nominal 90% across all properties, while non-conformal baselines (Deep Ensembles, MC~Dropout) fail to calibrate reliably. Among conformal approaches, Conformalized Quantile Regression (CQR) delivers the best coverage in the bin with the poorest model predictions. More importantly, only the Locally Valid and Discriminative (LVD) framework -- particularly when operating on AION-1 embeddings -- also provides finite-sample \emph{local validity}, producing intervals that adapt to each galaxy's local prediction difficulty rather than relying on marginal guarantees alone. These results establish conformal prediction, and LVD in particular, as the preferred UQ framework for uncertainty-aware inference on foundation-model embeddings in astrophysics.

Tame-Narvaez, Karla [Fermilab] (ORCID:000000022249

Broken conformal window

We show that near the edges of the conformal window of supersymmetric SU(Nc) QCD, perturbed by Anomaly Mediated Supersymmetry Breaking (AMSB), chiral symmetry can be broken depending on the initial conditions of the RG flow. We do so by perturbatively expanding around Banks-Zaks fixed points and taking advantage of Seiberg duality. Interpolating between the edges of the conformal window, we predict that non-supersymmetric QCD breaks chiral symmetry up to N f ≤ 3N c − 1, while we cannot say anything definitive for N f ≥ 3N c at this moment.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

A kinematical/numerical analysis of rotor-stator interaction noise

In this study, the unsteady, thin-layer Navier-Stokes equations are solved using a system of patched grids for a rotor-stator configuration of an axial turbine. The study examines the plurality of spinning modes that are present in such an interaction. The propagation of these modes is analyzed and appropriate grid spacing chosen in the far upstream and downstream regions to attenuate reflections from the computational boundaries. In addition, radiating boundary conditions are implemented based on the farfield acoustical behavior of the flow field. Results in the form of pressure amplitudes and the spectra of turbine tones are presented. Numerical results and experimental data are compared wherever possible. The numerical results are also shown to conform with the predictions of a kinematical analysis of the flowfield.

Rangwalla, Akil A.