Search NASA⌕ Search

SEARCH · Search NASA

Results for “Model selection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

LASSO for CALPHAD Model Selection Enables Data-Efficient Thermodynamic Modeling: An Application in Thermochemical Hydrogen Production Materials

Phenomenological CALPHAD (CALculation of PHAse Diagrams) models, widely used for multicomponent materials, often contain a considerable number of parameters and require fitting using data from a relatively small number of experimental measurements or theoretical calculations. Sometimes these parameters are introduced for the purpose of improving model fits but without clear physical justification, which leads to overparametrized models with poor generalization performance. Automated approaches for optimal model selection based on the available data therefore become critical. Here, in this work, a least absolute shrinkage and selection operator (LASSO)-based approach is developed for model selection by leveraging the linearity of the CALPHAD model with respect to its parameters to convert the model selection and fitting to a LASSO minimization problem. We demonstrate its utility for thermodynamic modeling of thermochemical hydrogen (TCH) production materials using lanthanum strontium manganite (LSM) as an example. Various TCH-relevant properties, including oxygen stoichiometry as a function of oxygen partial pressure, enthalpy of reduction, and entropy of reduction, are successfully predicted with reasonable accuracy using a minimal set of model parameters. Importantly, the model selection and fitting involve minimal human decision; it can therefore be applied to high-throughput DFT defect calculations and yield efficient workflows for TCH material modeling and optimization.

CALPHAD↗

An empirical approach to model selection: weak lensing and intrinsic alignments

ABSTRACT In cosmology, we routinely choose between models to describe our data, and can incur biases due to insufficient models or lose constraining power with overly complex models. In this paper, we propose an empirical approach to model selection that explicitly balances parameter bias against model complexity. Our method uses synthetic data to calibrate the relation between bias and the χ2 difference between models. This allows us to interpret χ2 values obtained from real data (even if catalogues are blinded) and choose a model accordingly. We apply our method to the problem of intrinsic alignments – one of the most significant weak lensing systematics, and a major contributor to the error budget in modern lensing surveys. Specifically, we consider the example of the Dark Energy Survey Year 3 (DES Y3), and compare the commonly used non-linear alignment (NLA) and tidal alignment and tidal torque (TATT) models. The models are calibrated against bias in the Ωm–S8 plane. Once noise is accounted for, we find that it is possible to set a threshold Δχ2 that guarantees an analysis using NLA is unbiased at some specified level Nσ and confidence level. By contrast, we find that theoretically defined thresholds (based on, e.g. p-values for χ2) tend to be overly optimistic, and do not reliably rule out cosmological biases up to ∼1–2σ. Considering the real DES Y3 cosmic shear results, based on the reported difference in χ2 from NLA and TATT analyses, we find a roughly $30{{\ \rm per\ cent}}$ chance that were NLA to be the fiducial model, the results would be biased (in the Ωm–S8 plane) by more than 0.3σ. More broadly, the method we propose here is simple and general, and requires a relatively low level of resources. We foresee applications to future analyses as a model selection tool in many contexts.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Machine-learning interatomic potentials for interfaces in all-solid-state batteries: Perspectives on training data, model selection, and validation

Interfaces play a pivotal role in dictating the performance and reliability of all-solid-state batteries (ASSBs), where complex electro-chemo-mechanical phenomena at grain boundaries (GBs) and interfaces can lead to degradation and failure. Traditional atomistic simulation methods, such as first-principles calculations and classical molecular dynamics, face limitations in modeling these interfaces due to either high computational cost or insufficient transferability to the diverse atomic environments evolving at interfaces. Machine-learning interatomic potentials (MLIPs) have emerged as a transformative approach, enabling large-scale, high-accuracy simulations of disordered and chemically complex systems by leveraging the predictability of machine learning models trained on first-principles data. Recent applications of MLIPs have demonstrated their ability to capture intricate behaviors at ASSB interfaces, including ion transport, interfacial evolution, and degradation mechanisms, with accuracy and efficiency unattainable by conventional methods. This prospective paper presents comprehensive analysis and practical guidance for MLIP development for GBs and interfaces in ASSBs, with a focus on three key pillars: data generation, model selection, and validation. Here, we review the current state of MLIP applications for GBs and interfaces in both general and ASSB-specific materials, highlighting best practices and challenges in constructing diverse and representative datasets, choosing appropriate machine learning architectures, and rigorously validating model performance. We also discuss emerging strategies and opportunities for improved reliability and efficiency of MLIPs to simulate realistic interfaces in ASSBs.

Energy - Storage↗

Hybrid Parameter Search and Dynamic Model Selection for Mixed-Variable Bayesian Optimization

Herein this article presents a new type of hybrid model for Bayesian optimization (BO) adept at managing mixed variables, encompassing both quantitative (continuous and integer) and qualitative (categorical) types. Our proposed new hybrid models (named hybridM) merge the Monte Carlo Tree Search structure (MCTS) for categorical variables with Gaussian Processes (GP) for continuous ones. hybridM leverages the upper confidence bound tree search (UCTS) for MCTS strategy, showcasing the tree architecture’s integration into Bayesian optimization. Our innovations, including dynamic online kernel selection in the surrogate modeling phase and a unique UCTS search strategy, position our hybrid models as an advancement in mixed-variable surrogate models. Numerical experiments underscore the superiority of hybrid models, highlighting their potential in Bayesian optimization.

97 MATHEMATICS AND COMPUTING↗

Bayesian model selection for GRB 211211A through multiwavelength analyses

ABSTRACT Although GRB 211211A is one of the closest gamma-ray bursts (GRBs), its classification is challenging because of its partially inconclusive electromagnetic signatures. In this paper, we investigate four astrophysical scenarios as possible progenitors for GRB 211211A: a binary neutron star merger, a black hole–neutron star merger, a core-collapse supernova, and an r-process enriched core collapse of a rapidly rotating massive star (a collapsar). We perform a large set of Bayesian multiwavelength analyses based on different models describing these scenarios and priors to investigate which astrophysical scenarios and processes might be related to GRB 211211A. Our analysis supports previous studies in which the presence of an additional component, likely related to r-process nucleosynthesis, is required to explain the observed light curves of GRB 211211A, as it cannot be explained solely as a GRB afterglow. Fixing the distance to about $350~\rm Mpc$, namely the distance of the possible host galaxy SDSS J140910.47+275320.8, we find a statistical preference for a binary neutron star merger scenario.

(transients:) gamma-ray bursts↗

Boosting Noise2Inverse via enhanced model selection for denoising computed tomography data

Synchrotron-based x-ray tomographic imaging enables the examination of the internal structure of materials at high spatial and temporal resolution. Experimental constraints can impose dose and time limits on the measurements, introducing a higher level of noise and artifacts in the reconstructed images. Deep learning has emerged as a powerful tool to remove noise from reconstructed images. Recently, the Noise2Inverse method was designed specifically for denoising reconstructed images without requiring paired noisy and clean images. This method creates multiple statistically independent reconstructions used to pair the data in which training involves transforming one reconstruction into the other, and vice versa. Originally designed to be used after a fixed number of epochs, we see in practice that this approach may not produce the optimal model and may unnecessarily waste computational resources. Therefore, we propose an alternative method of identifying the best model during training that aligns with the Noise2Inverse method. During validation, we compare the model output of the multiple reconstructions among each other. We hypothesize that the best model is the one that produces images with the highest similarity, implying a convergence in the predicted material properties and absorption values. To compare model outputs, we consider the absolute error, square error, structural similarity index (SSIM), peak signal-to-noise ratio (PSNR), and cosine similarity. We evaluate our method on two simulated tomography datasets and two, real-world, low-contrast, high-energy x-ray tomography datasets. We show our approach is more effective at determining the best model, up to an increase of 12.50% and 12.53% in SSIM and PSNR, respectively, while only requiring a fifth of the training time compared to the original approach.

CT↗

Dataset for Top Model Decision Tree: Selecting Segmentation Models for Reliable Quantitative Analysis in Low- and Ultralow-Dose CryoEM

Motivation Multiple deep learning model architectures can be used to segment bacterial membranes in cryoEM images. However, an AI-based tool advancement is often presented with only a single segmentation model for broad use, and this single model may show inconsistent results across datasets from different users. Here, we present the Top Model Decision Tree, a model screening framework to screen for the best model to generate bacterial inner and outer membrane masks based on user priorities. We use pre-trained segmentation models from YOLOv11, YOLO26, U-Net, Detectron2 and SAM3 fine-tuned on bacterial inner and outer membranes imaged with cryoEM. Run the Framework This notebook must be opened in Google Colab. Mount Google Drive and run with a GPU-based runtime. Open the notebook and follow steps to git clone in folders and files within this repository. There will be a repeating top_model_decision_tree.ipynb (notebook clone) that will not be used. Save your .png binary mask files and .csv table outputs within your Google Drive or download before closing the notebook. The models and all analysis/training scripts are available at [GitHub: https://github.com/Lynnicia/CryoEM_membranes_top_model_decision_tree and https://github.com/Sireesiru/Semantic-Segmentation-of-bacterial-cell-envelope-using-U-Nets.

59 BASIC BIOLOGICAL SCIENCES↗

Model averaging approaches to data subset selection

Model averaging is a useful and robust method for dealing with model uncertainty in statistical analysis. Often, it is useful to consider data subset selection at the same time, in which model selection criteria are used to compare models across different subsets of the data. Two different criteria have been proposed in the literature for how the data subsets should be weighted. We compare the two criteria closely in a unified treatment based on the Kullback-Leibler divergence and conclude that one of them is subtly flawed and will tend to yield larger uncertainties due to loss of information. Here, analytical and numerical examples are provided.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Microbiome-enabled genomic selection improves prediction accuracy for nitrogen-related traits in maize

Root-associated microbiomes in the rhizosphere (rhizobiomes) are increasingly known to play an important role in nutrient acquisition, stress tolerance, and disease resistance of plants. However, it remains largely unclear to what extent these rhizobiomes contribute to trait variation for different genotypes and if their inclusion in the genomic selection protocol can enhance prediction accuracy. To address these questions, we developed a microbiome-enabled genomic selection method that incorporated host SNPs and amplicon sequence variants from plant rhizobiomes in a maize diversity panel under high and low nitrogen (N) field conditions. Our cross-validation results showed that the microbiome-enabled genomic selection model significantly outperformed the conventional genomic selection model for nearly all time-series traits related to plant growth and N responses, with an average relative improvement of 3.7%. The improvement was more pronounced under low N conditions (8.4–40.2% of relative improvement), consistent with the view that some beneficial microbes can enhance N nutrient uptake, particularly in low N fields. However, our study could not definitively rule out the possibility that the observed improvement is partially due to the amplicon sequence variants being influenced by microenvironments. Using a high-dimensional mediation analysis method, our study has also identified microbial mediators that establish a link between plant genotype and phenotype. Some of the detected mediator microbes were previously reported to promote plant growth. The enhanced prediction accuracy of the microbiome-enabled genomic selection models, demonstrated in a single environment, serves as a proof-of-concept for the potential application of microbiome-enabled plant breeding for sustainable agriculture.

60 APPLIED LIFE SCIENCES↗