Search NASA⌕ Search

SEARCH · Search NASA

Results for “ensemble method”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Ensemble methods for quantification of potassium oxide in ChemCam Mars and laboratory spectra

In this paper we test new approaches for predicting the amount of element oxides in rock samples from the ChemCam instrument suite onboard the NASA Curiosity rover by focusing on K 2 O. Using the expanded dataset compiled by Gasda et al. (2021) with and without the Earth to Mars (E2M and NoE2M) transformation discussed in Clegg et al. (2017) we trained blended submodels using the “double blending” technique and compared these to ensemble methods (Random Forest, ExtraTrees, and Gradient Boosting Regression). We found that ensemble methods performed similar to blended submodels when looking at RMSE-P on the laboratory spectra and provided significant advantages when looking at spectra coming from Mars. For the full model, blended submodels achieved an RMSE-P of 0.62 and 0.60 (E2M and NoE2M respectively) while Gradient Boosting Regression resulted in a slightly improved RMSE-P of 0.59 and 0.60. More importantly, by employing a local RMSE-P estimation technique where model performance is evaluated based on nearby test samples we found that using ensemble methods can lower the quantification limit for K 2 O from the current value of ≈0.6 wt% to ≈0.08 wt% using Extra Trees and Random Forest. This would allow for a much larger range of K 2 O values to be quantified on Mars with greater certainty given that most targets seen on Mars tend to have <1 wt% K2O. Finally, we used both Mean Decrease in Impurity (MDI) and permutation importance techniques to investigate the wavelengths used by the ensemble methods and found that they correspond to known potassium emission lines. This suggests that ensemble methods can provide an easier to train and improved alternative to blended submodels for predicting potassium compositions from Laser Induced Breakdown Spectroscopy (LIBS) data.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A Finite-Time Ensemble Method for Mixed Layer Model Comparison

Here, this work evaluates the fidelity of various upper-ocean turbulence parameterizations subject to realistic monsoon forcing and presents a finite-time ensemble vector (EV) method to better manage the design and numerical principles of these parameterizations. The EV method emphasizes the dynamics of a turbulence closure multimodel ensemble and is applied to evaluate 10 different ocean surface boundary layer (OSBL) parameterizations within a single-column (SC) model against two boundary layer large-eddy simulations (LES). Both LES include realistic surface forcing, but one includes wind-driven shear turbulence only, while the other includes additional Stokes forcing through the wave-average equations that generate Langmuir turbulence. The finite-time EV framework focuses on what constitutes the local behavior of the mixed layer dynamical system and isolates the forcing and ocean state conditions where turbulence parameterizations most disagree. Identifying disagreement provides the potential to evaluate SC models comparatively against the LES. Observations collected during the 2018 monsoon onset in the Bay of Bengal provide a case study to evaluate models under realistic and variable forcing conditions. The case study results highlight two regimes where models disagree 1) during wind-driven deepening of the mixed layer and 2) under strong diurnal forcing.

54 ENVIRONMENTAL SCIENCES↗

Feature Engineering and Ensemble Methods for Imbalanced ICS Intrusion Detection: Pipeline Audit and Constrained Evaluation

Industries are becoming increasingly connected and are more vulnerable to cyberattacks due to the widened attack surface. Industrial Control Systems (ICS) are among the most critical sectors that malicious actors can target, as such attacks can cause significant operational disruption and physical damage. It is imperative to detect such attacks as early as possible. This paper evaluates constraint-conditioned optimistic performance estimates for traditional ML models in ICS intrusion detection (i.e., estimates obtained under contiguous, non-shuffled temporal evaluation without test-set alteration, but with pre-split feature engineering that may introduce temporal leakage, due to dataset constraints). Our findings are threefold. First, we quantify how iterative feature engineering affects tree-based ensemble performance and examine how pipeline decisions (split strategy, sampling scope, and cleaning policy) can inflate or reduce reported IDS results under constraint-bound evaluation. Second, we compare intrinsic class-imbalance handling across ensemble models. Third, under our current pipeline constraints (including pre-split feature engineering), CatBoost achieves the best performance on Water Storage Tank (accuracy: 0.9831, class-1 F1: 0.9682), while Light- GBM achieves the best performance on Gas Pipeline (accuracy: 0.9618, class-1 F1: 0.9086).

97 MATHEMATICS AND COMPUTING↗

Enhanced Boundary Layer Height Detection Using Ceilometer, Surface Meteorology, and Radiation Products With a Random Forest Ensemble Method

This study develops and evaluates a Random Forest (RF) model for estimating planetary boundary layer height (PBLH) using 9 years of data from the Atmospheric Radiation Measurement Southern Great Plains (ARM SGP) user facility, with potential application in the NOAA Surface Radiation (SURFRAD) Network. The model integrates ceilometer, surface meteorology, and radiation measurements, and is trained using thermodynamic PBLH estimates derived from radiosondes. This approach aims to bridge gaps between aerosol-based and thermodynamic-based PBLH estimates. The RF model outperformed traditional methods during daytime and better captured transition periods, demonstrating improved accuracy and robustness. At ARM SGP, it showed a substantial reduction in both bias and RMSE, with a bias near zero (−4.9 m) compared with traditional Haar Wavelet (HW) (70.9 m) and Vaisala BL-View software (124.1 m), and an RMSE of 303.2 m, lower than both BL-View (566.9 m) and HW (404.6 m). During daytime hours, RF consistently outperformed both alternatives, maintaining lower bias and RMSE across all periods. At a second evaluation site, RF achieved the lowest overall RMSE (323.7 m), similar to HW (326.4 m) and significantly better than BL-View (738.3 m). However, all models showed reduced accuracy under stable nighttime conditions, limiting the reliability of PBLH estimates. Key predictors for the model included the lifting condensation level height (LCLH), aerosol gradients, and month for seasonal variability. The study underscores the potential of integrating machine learning with multiple data sets such as surface energy and thermodynamic data to advance PBLH estimation.

boundary layer height↗

Uncertainty quantification of a deep learning fuel property prediction model

Deep learning models are being widely used in the field of combustion. Given the black-box nature of typical neural network based models, uncertainty quantification (UQ) is critical to ensure the reliability of predictions as well as the training datasets, and for a principled quantification of noise and its various sources. Deep learning surrogate models for predicting properties of chemical compounds and mixtures have been recently shown to be promising for enabling data-driven fuel design and optimization, with the ultimate goal of improving efficiency and lowering emissions from combustion engines. In this study, UQ is performed for a multi-task deep learning model that simultaneously predicts the research octane number (RON), Motor Octane Number (MON), and Yield Sooting Index (YSI) of pure components and multicomponent blends. The deep learning model is comprised of three smaller networks: Extractor 1, Extractor 2, and Predictor, and a mixing operator. The molecular fingerprints of individual components are encoded via Extractor 1 and Extractor 2, the mixing operator generates fingerprints for mixtures/blends based on linear mixing operation, and the predictor maps the fingerprint to the target properties. Two different classes of UQ methods, Monte Carlo ensemble methods and Bayesian neural networks (BNNs), are employed for quantifying the epistemic uncertainty. Combinations of Bernoulli and Gaussian distributions with DropConnect and DropOut techniques are explored as ensemble methods. All the DropConnect, DropOut and Bayesian layers are applied to the predictor network. Aleatoric uncertainty is modeled by assuming that each data point has an independent uncertainty associated with it. The results of the UQ study are further analyzed to compare the performance of BNN and ensemble methods. Although this study is confined to UQ of fuel property prediction, the methodologies are applicable to other deep learning frameworks that are being widely used in the combustion community.

33 ADVANCED PROPULSION SYSTEMS↗

MINE: maximally informative next experiment—toward a new GWAS experimental design and methodology

Abstract The computational methodology of Genome Wide Association Studies (GWAS) currently has several limitations: (i) the number of observations (rows) on a quantitative trait tends to be smaller than the number of single nucleotide polymorphisms (SNPs) (columns) in the design matrix; (ii) each SNP is usually modeled separately, failing to acknowledge interaction between each other (ie epistasis); (iii) there is implicit linkage disequilibrium (LD) between neighboring SNPs due to their linkage. To overcome these issues, we developed a tool that uses ensemble methods to fit mixed linear models to GWAS data, and these ensemble methods include the development of a new experimental design approach in GWAS, which uses the resultant models and data to select the next informative experiment over time. This new adaptive and staged approach for GWAS experimental design was developed and tested in a 3 yr adaptive model-guided discovery experiment against a fixed classical design. In Sorghum bicolor a total of 79, 86, and 78 accessions were tested in years 1, 2, and 3, respectively out of 343 accessions available in the Bioenergy Association Panel (BAP) each identified for 232,303 SNPs, 1 every 2–3 kb in the genomes. We demonstrated the feasibility of MINE enacted with 8 people in the field per year over 3 yr vs in 1 large classical design enacted with 20 people in 1 yr. The MINE results for chromosomal regions identified controlling dry weight were confirmed against results from previous sorghum GWAS experiments and 1 large classical design for the BAP panel.

Genetics & Heredity↗

Comparison of Equilibrium and Nonequilibrium Approaches for Relative Binding Free Energy Predictions

Alchemical relative binding free energy calculations have recently found important applications in drug optimization. A series of congeneric compounds are generated from a preidentified lead compound, and their relative binding affinities to a protein are assessed in order to optimize candidate drugs. While methods based on equilibrium thermodynamics have been extensively studied, an approach based on nonequilibrium methods has recently been reported together with claims of its superiority. However, these claims pay insufficient attention to the basis and reliability of both methods. Here we report a comparative study of the two approaches across a large data set, comprising more than 500 ligand transformations spanning in excess of 300 ligands binding to a set of 14 diverse protein targets. Ensemble methods are essential to quantify the uncertainty in these calculations, not only for the reasons already established in the equilibrium approach but also to ensure that the nonequilibrium calculations reside within their domain of validity. If and only if ensemble methods are applied, we find that the nonequilibrium method can achieve accuracy and precision comparable to those of the equilibrium approach. Compared to the equilibrium method, the nonequilibrium approach can reduce computational costs but introduces higher computational complexity and longer wall clock times. There are, however, cases where the standard length of a nonequilibrium transition is not sufficient, necessitating a complete rerun of the entire set of transitions. This significantly increases the computational cost and proves to be highly inconvenient during large-scale applications. Our findings provide a key set of recommendations that should be adopted for the reliable implementation of nonequilibrium approaches to relative binding free energy calculations in ligand-protein systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Methods for Incorporating Model Uncertainty into Exoplanet Atmospheric Analysis

A key goal of exoplanet spectroscopy is to measure atmospheric properties, such as abundances of chemical species, in order to connect them to our understanding of atmospheric physics and planet formation. In this new era of high-quality JWST data, it is paramount that these measurement methods are robust. When comparing atmospheric models to observations, multiple candidate models may produce reasonable fits to the data. Typically, conclusions are reached by selecting the best-performing model according to some metric. This ignores model uncertainty in favor of specific model assumptions, potentially leading to measured atmospheric properties that are overconfident and/or incorrect. In this paper, we compare three ensemble methods for addressing model uncertainty by combining posterior distributions from multiple analyses: Bayesian model averaging, a variant of Bayesian model averaging using leave-one-out predictive densities, and stacking of predictive distributions. We demonstrate these methods by fitting the Hubble Space Telescope (HST) + Spitzer transmission spectrum of the hot Jupiter HD 209458b using models with different cloud and haze prescriptions. All of our ensemble methods lead to uncertainties on retrieved parameters that are larger but more realistic and consistent with physical and chemical expectations. Since they have not typically accounted for model uncertainty, uncertainties of retrieved parameters from HST spectra have likely been underreported. We recommend stacking as the most robust model combination method. Our methods can be used to combine results from independent retrieval codes and from different models within one code. They are also widely applicable to other exoplanet analysis processes, such as combining results from different data reductions.

79 ASTRONOMY AND ASTROPHYSICS↗

Progress in Normalizing Flows for 4d Gauge Theories

Normalizing flows have arisen as a tool to accelerate Monte Carlo sampling for lattice field theories. This work reviews recent progress in applying normalizing flows to 4-dimensional nonabelian gauge theories, focusing on two advancements: an architectural improvement referred to as learned active loops, and the application of correlated ensemble methods to QCD with N f = 2 dynamical fermions.

Abbott, Ryan [Massachusetts Institute of Technolog↗

MINE: a new way to design genetics experiments for discovery

Abstract The Maximally Informative Next Experiment or MINE is a new experimental design approach for experiments, such as those in omics, in which the number of effects or parameters p greatly exceeds the number of samples n (p > n). Classical experimental design presumes n > p for inference about parameters and its application to p > n can lead to over-fitting. To overcome p > n, MINE is an ensemble method, which makes predictions about future experiments from an existing ensemble of models consistent with available data in order to select the most informative next experiment. Its advantages are in exploration of the data for new relationships with n < p and being able to integrate smaller and more tractable experiments to replace adaptively one large classic experiment as discoveries are made. Thus, using MINE is model-guided and adaptive over time in a large omics study. Here, MINE is illustrated in two distinct multiyear experiments, one involving genetic networks in Neurospora crassa and a second one involving a genome-wide association study in Sorghum bicolor as a comparison to classic experimental design in an agricultural setting.

Biochemistry & Molecular Biology↗

Exploring Data Set Bias and Decision Support with Predictive Uncertainty Through Bayesian Approximations and Convolutional Neural Networks

Individual seismic catalogs can contain multiscale observations from fault level to global scales and associated waveforms from discrete events reflect crustal structure across many different scales and locations. Seismic network aperture, geographic location, and observation distance may not provide informative guidance or intuition on how different catalogs will behave across models trained under different conditions. We rely on uncertainty to provide guardrails for when to trust model decisions, but understanding when our uncertainty is trustworthy is an open challenge. Here, in this work, we explore Bayesian approximation methods for assigning predictive uncertainty in seismic event classification problems. We find that computationally expensive Bayesian approximations do not outperform simple ensemble methods. We also find that when exploiting multiple seismic event catalogs, joint training with data from all the catalogs combined with Bayesian approximations and supervised training for classification can obscure bias and result in less robust uncertainty while also not providing substantial performance benefits compared to training individual models for each catalog.

58 GEOSCIENCES↗

Large Ensemble Exploration of Global Energy Transitions Under National Emissions Pledges

Global climate goals require a transition to a deeply decarbonized energy system. Meeting the objectives of the Paris Agreement through countries' nationally determined contributions and long-term strategies represents a complex problem with consequences across multiple systems shrouded by deep uncertainty. Robust, large-ensemble methods and analyses mapping a wide range of possible future states of the world are needed to help policymakers design effective strategies to meet emissions reduction goals. This study contributes a scenario discovery analysis applied to a large ensemble of 5,760 model realizations generated using the Global Change Analysis Model. Eleven energy-related uncertainties are systematically varied, representing national mitigation pledges, institutional factors, and techno-economic parameters, among others. The resulting ensemble maps how uncertainties impact common energy system metrics used to characterize national and global pathways toward deep decarbonization. Results show globally consistent but regionally variable energy transitions as measured by multiple metrics, including electricity costs and stranded assets. Larger economies and developing regions experience more severe economic outcomes across a broad sampling of uncertainty. The scale of CO 2 removal globally determines how much the energy system can continue to emit, but the relative role of different CO 2 removal options in meeting decarbonization goals varies across regions. Previous studies characterizing uncertainty have typically focused on a few scenarios, and other large-ensemble work has not (to our knowledge) combined this framework with national emissions pledges or institutional factors. Our results underscore the value of large-ensemble scenario discovery for decision support as countries begin to design strategies to meet their goals.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Deep neural network uncertainty quantification for LArTPC reconstruction

We evaluate uncertainty quantification (UQ) methods for deep learning applied to liquid argon time projection chamber (LArTPC) physics analysis tasks. As deep learning applications enter widespread usage among physics data analysis, neural networks with reliable estimates of prediction uncertainty and robust performance against overconfidence and out-of-distribution (OOD) samples are critical for their full deployment in analyzing experimental data. While numerous UQ methods have been tested on simple datasets, performance evaluations for more complex tasks and datasets are scarce. Here we assess the application of selected deep learning UQ methods on the task of particle classification using the PiLArNet monte carlo 3D LArTPC point cloud dataset. We observe that UQ methods not only allow for better rejection of prediction mistakes and OOD detection, but also generally achieve higher overall accuracy across different task settings. We assess the precision of uncertainty quantification using different evaluation metrics, such as distributional separation of prediction entropy across correctly and incorrectly identified samples, receiver operating characteristic curves (ROCs), and expected calibration error from observed empirical accuracy. We conclude that ensembling methods can obtain well calibrated classification probabilities and generally perform better than other existing methods in deep learning UQ literature.

47 OTHER INSTRUMENTATION↗