Search NASA⌕ Search

SEARCH · Search NASA

Results for “Model selection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

LASSO for CALPHAD Model Selection Enables Data-Efficient Thermodynamic Modeling: An Application in Thermochemical Hydrogen Production Materials

Phenomenological CALPHAD (CALculation of PHAse Diagrams) models, widely used for multicomponent materials, often contain a considerable number of parameters and require fitting using data from a relatively small number of experimental measurements or theoretical calculations. Sometimes these parameters are introduced for the purpose of improving model fits but without clear physical justification, which leads to overparametrized models with poor generalization performance. Automated approaches for optimal model selection based on the available data therefore become critical. Here, in this work, a least absolute shrinkage and selection operator (LASSO)-based approach is developed for model selection by leveraging the linearity of the CALPHAD model with respect to its parameters to convert the model selection and fitting to a LASSO minimization problem. We demonstrate its utility for thermodynamic modeling of thermochemical hydrogen (TCH) production materials using lanthanum strontium manganite (LSM) as an example. Various TCH-relevant properties, including oxygen stoichiometry as a function of oxygen partial pressure, enthalpy of reduction, and entropy of reduction, are successfully predicted with reasonable accuracy using a minimal set of model parameters. Importantly, the model selection and fitting involve minimal human decision; it can therefore be applied to high-throughput DFT defect calculations and yield efficient workflows for TCH material modeling and optimization.

CALPHAD↗

An empirical approach to model selection: weak lensing and intrinsic alignments

ABSTRACT In cosmology, we routinely choose between models to describe our data, and can incur biases due to insufficient models or lose constraining power with overly complex models. In this paper, we propose an empirical approach to model selection that explicitly balances parameter bias against model complexity. Our method uses synthetic data to calibrate the relation between bias and the χ2 difference between models. This allows us to interpret χ2 values obtained from real data (even if catalogues are blinded) and choose a model accordingly. We apply our method to the problem of intrinsic alignments – one of the most significant weak lensing systematics, and a major contributor to the error budget in modern lensing surveys. Specifically, we consider the example of the Dark Energy Survey Year 3 (DES Y3), and compare the commonly used non-linear alignment (NLA) and tidal alignment and tidal torque (TATT) models. The models are calibrated against bias in the Ωm–S8 plane. Once noise is accounted for, we find that it is possible to set a threshold Δχ2 that guarantees an analysis using NLA is unbiased at some specified level Nσ and confidence level. By contrast, we find that theoretically defined thresholds (based on, e.g. p-values for χ2) tend to be overly optimistic, and do not reliably rule out cosmological biases up to ∼1–2σ. Considering the real DES Y3 cosmic shear results, based on the reported difference in χ2 from NLA and TATT analyses, we find a roughly $30{{\ \rm per\ cent}}$ chance that were NLA to be the fiducial model, the results would be biased (in the Ωm–S8 plane) by more than 0.3σ. More broadly, the method we propose here is simple and general, and requires a relatively low level of resources. We foresee applications to future analyses as a model selection tool in many contexts.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

'Chain pooling' model selection as developed for the statistical analysis of a rotor burst protection experiment

A statistical decision procedure called chain pooling had been developed for model selection in fitting the results of a two-level fixed-effects full or fractional factorial experiment not having replication. The basic strategy included the use of one nominal level of significance for a preliminary test and a second nominal level of significance for the final test. The subject has been reexamined from the point of view of using as many as three successive statistical model deletion procedures in fitting the results of a single experiment. The investigation consisted of random number studies intended to simulate the results of a proposed aircraft turbine-engine rotor-burst-protection experiment. As a conservative approach, population model coefficients were chosen to represent a saturated 2 to the 4th power experiment with a distribution of parameter values unfavorable to the decision procedures. Three model selection strategies were developed.

Holms, A. G.↗

'Chain pooling' model selection for two-level fixed effects factorial experiments

As many as three iterated statistical model deletion procedures are considered for an experiment. Population model coefficients were chosen to simulate a saturated factorial experiment having an unfavorable distribution of parameter values. Using random number studies, three model selection strategies were developed, namely, (1) a strategy to be used in anticipation of large coefficients of variation (neighborhood of 65 percent), (2) strategy to be used in anticipation of small coefficients of variation (4 percent or less), and (3) a security regret strategy to be used in the absence of such prior knowledge.

Holms, A. G.↗

Improving Multi-Model Trajectory Simulation Estimators using Model Selection and Tuning

Multi-model Monte Carlo methods have been demonstrated to be an efficient and accurate alternative to standard Monte Carlo (MC) in the model-based propagation of uncertainty in entry, descent, and landing (EDL) applications. These multi-model MC methods fuse predictions from low-fidelity models with the high-fidelity EDL model of interest to produce unbiased statistics with a fraction of the computational cost. The accuracy and efficiency of the multi-model MC methods are dependent upon the magnitude of correlations of the low-fidelity models with the high-fidelity model, but also upon the correlation amongst the low-fidelity models, and their relative computational cost. Because of this layer of complexity, the question of how to optimally select the set of low-fidelity models has remained open. In this work, methods for optimal model construction and tuning are investigated as a means to increase the speed and precision of trajectory simulation for EDL. Specifically, the focus is on the inclusion of low-fidelity model tuning within the sample allocation optimization that accompanies multi-model MC methods. Preliminary results indicate that low-fidelity model tuning can significantly improve efficiency and precision of trajectory simulations and provide an increased edge to multi-model MC methods when compared to standard MC. The challenges and potential benefits to exploring a fully iterative and comprehensive optimization strategy in future work are highlighted.

uncertainty quantification↗

Improving Multi-Model Trajectory Simulation Estimators using Model Selection and Tuning

Multi-model Monte Carlo methods have been demonstrated to be an efficient and accurate alternative to standard Monte Carlo (MC) in the model-based propagation of uncertainty in entry, descent, and landing (EDL) applications. These multi-model MC methods fuse predictions from low-fidelity models with the high-fidelity EDL model of interest to produce unbiased statistics with a fraction of the computational cost. The accuracy and efficiency of the multi-model MC methods are dependent upon the magnitude of correlations of the low-fidelity models with the high-fidelity model, but also upon the correlation among the low-fidelity models, and their relative computational cost. Because of this layer of complexity, the question of how to optimally select the set of low-fidelity models has remained open. In this work, methods for optimal model construction and tuning are investigated as a means to increase the speed and precision of trajectory simulation for EDL. Specifically, the focus is on the inclusion of low-fidelity model tuning within the sample allocation optimization that accompanies multi-model MC methods. Preliminary results indicate that low-fidelity model tuning can significantly improve efficiency and precision of trajectory simulations and provide an increased edge to multi-model MC methods when compared to standard MC. The challenges and potential benefits to exploring a fully iterative and comprehensive optimization strategy in future work are highlighted.

uncertainty quantification↗

Chain Pooling modeling selection as developed for the statistical analysis of a rotor burst protection experiment

As many as three iterated statistical model deletion procedures were considered for an experiment. Population model coefficients were chosen to simulate a saturated 2 to the 4th power experiment having an unfavorable distribution of parameter values. Using random number studies, three model selection strategies were developed, namely, (1) a strategy to be used in anticipation of large coefficients of variation, approximately 65 percent, (2) a strategy to be sued in anticipation of small coefficients of variation, 4 percent or less, and (3) a security regret strategy to be used in the absence of such prior knowledge.

Holms, A. G.↗

Bayesian Model Selection for Reducing Bloat and Overfitting in Genetic Programming for Symbolic Regression

When performing symbolic regression using genetic programming, overfitting and bloat can negatively impact generalizability and interpretability of the resulting equations as well as increase computation times. A Bayesian fitness metric is introduced and its impact on bloat and overfitting during population evolution is studied and compared to common alternatives in the literature. The proposed approach was found to be more robust to noise and data sparsity in numerical experiments, guiding evolution to a level of complexity appropriate to the dataset. Further evolution of the population resulted not in overfitting or bloat, but rather in slight simplifications in model form. The ability to identify an equation of complexity appropriate to the scale of noise in the training data was also demonstrated. In general, the Bayesian model selection algorithm was shown to be an effective means of regularization which resulted in less bloat and overfitting when any amount of noise was present in the training data.

Uncertainty quantification↗

Machine-learning interatomic potentials for interfaces in all-solid-state batteries: Perspectives on training data, model selection, and validation

Interfaces play a pivotal role in dictating the performance and reliability of all-solid-state batteries (ASSBs), where complex electro-chemo-mechanical phenomena at grain boundaries (GBs) and interfaces can lead to degradation and failure. Traditional atomistic simulation methods, such as first-principles calculations and classical molecular dynamics, face limitations in modeling these interfaces due to either high computational cost or insufficient transferability to the diverse atomic environments evolving at interfaces. Machine-learning interatomic potentials (MLIPs) have emerged as a transformative approach, enabling large-scale, high-accuracy simulations of disordered and chemically complex systems by leveraging the predictability of machine learning models trained on first-principles data. Recent applications of MLIPs have demonstrated their ability to capture intricate behaviors at ASSB interfaces, including ion transport, interfacial evolution, and degradation mechanisms, with accuracy and efficiency unattainable by conventional methods. This prospective paper presents comprehensive analysis and practical guidance for MLIP development for GBs and interfaces in ASSBs, with a focus on three key pillars: data generation, model selection, and validation. Here, we review the current state of MLIP applications for GBs and interfaces in both general and ASSB-specific materials, highlighting best practices and challenges in constructing diverse and representative datasets, choosing appropriate machine learning architectures, and rigorously validating model performance. We also discuss emerging strategies and opportunities for improved reliability and efficiency of MLIPs to simulate realistic interfaces in ASSBs.

Energy - Storage↗

Bayesian Model Selection for Reducing Bloat and Overfitting in Genetic Programming for Symbolic Regression

When performing symbolic regression using genetic programming, overfitting and bloat can negatively impact generalizability and interpretability of the resulting equations as well as increase computation times. A Bayesian fitness metric is introduced and its impact on bloat and overfitting during population evolution is studied and compared to common alternatives in the literature. The proposed approach was found to be more robust to noise and data sparsity in numerical experiments, guiding evolution to a level of complexity appropriate to the dataset. Further evolution of the population resulted not in overfitting or bloat, but rather in slight simplifications in model form. The ability to identify an equation of complexity appropriate to the scale of noise in the training data was also demonstrated. In general, the Bayesian model selection algorithm was shown to be an effective means of regularization which resulted in less bloat and overfitting when any amount of noise was present in the training data.

G F Bomarito↗

Application of a Bayesian Framework for Plasticity Model Selection

Interpretable Machine Learning (IML) has performed well when tasked with deriving constitutive material models. However, IML has been shown to prefer models that overfit noise in data, which tends to lead to bloat and a decrease in interpretability. Due to these issues, the ability of IML to reliably derive models that fit the data and are both interpretable and generalizable is limited. A method developed recently has shown promise to improve upon traditional IML by using a Bayesian fitness definition for the evolution of free-form models with non-deterministic parameters. This framework was developed for genetic-programming-based symbolic regression(GPSR) and involves model parameter estimation using Sequential Monte Carlo sampling (SMC).The method has demonstrated a reduction in bloat when dealing with noisy data in comparison to conventional GPSR. The results of this framework applied to stress-strain data for copper show models that more effectively predict the experimental data better than was previously shown with GPSR.

plasticity↗

Modeling Selective Availability of the NAVSTAR Global Positioning System

As the development of the NAVSTAR Global Positioning System (GPS) continues, there will increasingly be the need for a software centered signal model. This model must accurately generate the observed pseudorange which would typically be encountered. The observed pseudorange varies from the true geometric (slant) range due to range measurement errors. Errors in range measurement stem from a variety of hardware and environment factors. These errors are classified as either deterministic or random and, where appropriate, their models are summarized. Of particular interest is the model for Selective Availability which is derived from actual GPS data. The procedure for the determination of this model, known as the System Identification Theory, is briefly outlined. The synthesis of these error sources into the final signal model is given along with simulation results.

Braasch, Michael↗

Hybrid Parameter Search and Dynamic Model Selection for Mixed-Variable Bayesian Optimization

Herein this article presents a new type of hybrid model for Bayesian optimization (BO) adept at managing mixed variables, encompassing both quantitative (continuous and integer) and qualitative (categorical) types. Our proposed new hybrid models (named hybridM) merge the Monte Carlo Tree Search structure (MCTS) for categorical variables with Gaussian Processes (GP) for continuous ones. hybridM leverages the upper confidence bound tree search (UCTS) for MCTS strategy, showcasing the tree architecture’s integration into Bayesian optimization. Our innovations, including dynamic online kernel selection in the surrogate modeling phase and a unique UCTS search strategy, position our hybrid models as an advancement in mixed-variable surrogate models. Numerical experiments underscore the superiority of hybrid models, highlighting their potential in Bayesian optimization.

97 MATHEMATICS AND COMPUTING↗

Development of a research project selection model: Application to a civil helicopter research program

A model is described for planning and decision making in research project selection. Evaluations of each project's direct and indirect benefits, uncertainty in achieving these benefits, and schedule priority with resource budget and program balance constraints are considered. The combination of the interactive effect of project selection, resource allocation and scheduling considerations into one model permits tradeoff alternatives to be studied. Clients' value judgments are used in evaluating the benefits from each proposed project. The model is applied to the NASA Civil Helicopter Technology Program. Research project priorities for this program are established, strengths and weaknesses of the model are discussed, and areas of future development are recommended.

Schoultz, M. B.↗