Search NASA⌕ Search

SEARCH · Search NASA

Results for “Ensemble methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Efficient Unitary Designs from Random Sums and Permutations

A unitary k-design is an ensemble of unitaries that matches the first k moments of the Haar measure. In this work, we provide two efficient constructions of k-designs on n-qubits using new random matrix theory techniques. Our first construction is based on exponentiating sums of random i.i.d. Hermitian matrices and uses O(k2n2)-many gates. In the spirit of central limit theorems, we show that this random sum approximates the Gaussian Unitary Ensemble (GUE). We then show that the product of just two exponentiated GUE matrices is already approximately Haar random. Our second construction is based on products of exponentiated sums of random permutations and uses Õ(k poly (n)) many gates. The k dependence is optimal (up to polylogarithmic factors) and is inherited from the efficiency of existing k-wise independent permutations. Furthermore, replacing random permutations with quantum-secure pseudorandom permutations (PRPs), we also obtain a pseudorandom unitary (PRU) ensemble that is secure under nonadaptive queries. A central feature of both proofs is a new connection between the polynomial method in quantum query complexity and the large-dimension (N) expansion in random matrix theory. In particular, the first construction uses the polynomial method to control high moments of certain random matrix ensembles without requiring delicate Weingarten calculations. In doing so, we define and solve a moment problem on the unit circle, asking whether a finite number of equally weighted points can reproduce a given set of moments. In our second construction, the key step is to exhibit an orthonormal basis for irreducible representations of the partition algebra that has a low-degree large-N expansion. This allows us to show that the distinguishing probability is a low-degree rational polynomial of the dimension N.

algebra↗

An Ensemble Neural Network Model for Predicting Rare-Earth Oxide and Silicate Heat Capacities at High Temperature

In this work, a neural network model was developed to predict the constant pressure heat capacity for materials in the rare-earth oxide—silica material space. Several model architectures were trained and tested on heat capacity data generated from first-principles density functional theory calculations. Hyperparameter optimization was performed, and the optimal model was selected for heat capacity predictions. The optimal model architecture was found to have a root-mean-squared error of 5.12 ± 3.37 J/mol-K. The optimal model architecture was then used in a bagging ensemble model trained using the leave-one-group-out method to provide error estimates for model predictions. The out-of-bag score for the ensemble model was 0.997. The predicted heat capacities agree well with the DFT and experimental results and were computed orders of magnitude faster than DFT simulations. Machine learning shows the potential to provide a suitable surrogate model for thermochemical property predictions for candidate environmental barrier coating materials but refining of input material features and model architectures could further improve accuracy for these models.

environmental barrier coatings↗

Uncertainty-informed selection of CMIP6 Earth System Model subsets for use in multisectoral and impact models

Earth system models (ESMs) and general circulation models (GCMs) are heavily used to provide inputs to sectoral impact and multisector dynamic models, which include representations of energy, water, land, economics, and their interactions. Therefore, representing the full range of model uncertainty, scenario uncertainty, and interannual variability that ensembles of these models capture is critical to the exploration of the future co-evolution of the integrated human–Earth system. The pre-eminent source of these ensembles has been the Coupled Model Intercomparison Project (CMIP). With more modeling centers participating in each new CMIP phase, the size of the model archive is rapidly increasing, which can be intractable for impact modelers to effectively utilize due to computational constraints and the challenges of analyzing large datasets. In this work, we present a method to select a subset of the latest phase, CMIP6, featuring models for use as inputs to a sectoral impact or multisector dynamics models, while prioritizing preservation of the range of model uncertainty, scenario uncertainty, and interannual variability in the full CMIP6 ensemble results. This method is intended to help impact modelers select climate information from the CMIP archive efficiently for use in downstream models that require global coverage of climate information. This is particularly critical for large-ensemble experiments of multisector dynamic models that may be varying additional features beyond climate inputs in a factorial design, thus putting constraints on the number of climate simulations that can be used. We focus on temperature and precipitation outputs of CMIP6 models, as these are two of the most used variables among impact models, and many other key input variables for impacts are at least correlated with one or both of temperature and precipitation (e.g., relative humidity). Besides preserving the multi-model ensemble variance characteristics, we prioritize selecting CMIP6 models in the subset that preserve the very likely distribution of equilibrium climate sensitivity values as assessed by the latest Intergovernmental Panel on Climate Change (IPCC) report. This approach could be applied to other output variables of climate models and, possibly when combined with emulators, offers a flexible framework for designing more efficient experiments on human-relevant climate impacts. It can also provide greater insight into the properties of existing CMIP6 models.

Snyder, Abigail C.↗

DeepUQ: Assessing the Aleatoric Uncertainties from two Deep Learning Methods

Assessing the quality of aleatoric uncertainty estimates from uncertainty quantification (UQ) deep learning methods is important in scientific contexts, where uncertainty is physically meaningful and important to characterize and interpret exactly. We systematically compare aleatoric uncertainty measured by two UQ techniques, Deep Ensembles (DE) and Deep Evidential Regression (DER). Our method focuses on both zero-dimensional (0D) and two-dimensional (2D) data, to explore how the UQ methods function for different data dimensionalities. We investigate uncertainty injected on the input and output variables and include a method to propagate uncertainty in the case of input uncertainty so that we can compare the predicted aleatoric uncertainty to the known values. We experiment with three levels of noise. The aleatoric uncertainty predicted across all models and experiments scales with the injected noise level. However, the predicted uncertainty is miscalibrated to $\rm{std}(\sigma_{\rm al})$ with the true uncertainty for half of the DE experiments and almost all of the DER experiments. The predicted uncertainty is the least accurate for both UQ methods for the 2D input uncertainty experiment and the high-noise level. While these results do not apply to more complex data, they highlight that further research on post-facto calibration for these methods would be beneficial, particularly for high-noise and high-dimensional settings.

Nevin, Rebecca↗

Uncertainty based Online Ensemble on Non-Stationary Data for Fusion Science

Machine Learning (ML) is poised to play a pivotal role in the development and operation of next-generation fusion devices. Fusion data shows non-stationary behavior due to drifts in the data. The drifts can arise from both experimental evolution and machine wear-and-tear. ML models assume stationary distribution and fail to maintain performance when encountered with non-stationary data streams.Online learning can be used to continuously adapt the models with new data as it is acquired. However, traditional online learning can suffer from short-term performance degradation, as ground truth are not available before making the prediction. To address this challenge, we propose uncertainty aware ensemble approach for online learning. We use Deep Gaussian Process Approximation (DGPA) technique for calibrated uncertainty estimation and use the uncertainty values to guide a meta-algorithm that produces predictions based on ensemble of learners. Moreover, DGPA also provides uncertainty estimation along with the predictions for decision makers. This paper demonstrates that the proposed method outperforms traditional online learning approach, and a naive ensemble without uncertainty guidance by about 7% and 6%, respectively, on B-coil deflection prediction at DIII-D Fusion Facility.

Rajput, Kishansingh [Thomas Jefferson National Acc↗

Uncertainty based Online Ensemble on Non-Stationary Data for Fusion Science

Machine Learning (ML) is poised to play a pivotal role in the development and operation of next-generation fusion devices. Fusion data shows non-stationary behavior due to drifts in the data. The drifts can arise from both experimental evolution and machine wear-and-tear. ML models assume stationary distribution and fail to maintain performance when encountered with non-stationary data streams.Online learning can be used to continuously adapt the models with new data as it is acquired. However, traditional online learning can suffer from short-term performance degradation, as ground truth are not available before making the prediction. To address this challenge, we propose uncertainty aware ensemble approach for online learning. We use Deep Gaussian Process Approximation (DGPA) technique for calibrated uncertainty estimation and use the uncertainty values to guide a meta-algorithm that produces predictions based on ensemble of learners. Moreover, DGPA also provides uncertainty estimation along with the predictions for decision makers. This paper demonstrates that the proposed method outperforms traditional online learning approach, and a naive ensemble without uncertainty guidance by about 7% and 6%, respectively, on B-coil deflection prediction at DIII-D Fusion Facility.

Rajput, Kishansingh [Thomas Jefferson National Acc↗

Forced Component Estimation Statistical Method Intercomparison Project (ForceSMIP)

Anthropogenic climate change is unfolding rapidly, yet its regional manifestation can be obscured by internal variability. A primary goal of climate science is to identify the externally forced climate response from among the noise of internal variability. Separating the forced response from internal variability can be addressed in climate models by using a large ensemble to average over different possible realizations of internal variability. However, with only one realization of the real world, it is a major challenge to isolate the forced response directly in observations. In the Forced Component Estimation Statistical Method Intercomparison Project (ForceSMIP), contributors used existing and newly developed statistical and machine learning methods to estimate the forced response over 1950–2022 within individual realizations of the climate system. Participants used neural networks, linear inverse models, fingerprinting methods, and low-frequency component analysis, among other approaches. These methods were trained using large ensembles from multiple climate models and then applied to observations. Here, we evaluate method performance within large ensembles and investigate the estimates of the forced response in observations. Our results show that many different types of methods are skillful for estimating the forced response in climate models, though the relative skill of individual methods varies depending on the variable and evaluation metric. Methods with comparable skill in models can give a wide range of estimates of the forced response pattern in observations, illustrating the epistemic uncertainty in forced response estimates. ForceSMIP gives new insights into the forced response in observations, its uncertainty, and methods for its estimation.

Climate attribution↗

A comparison of measured and calculated optical properties of atmospheric aerosols at infrared wavelengths

Measurements of 10.6-micron lidar backscatter were compared with calculated backscatter based on nearly simultaneous observations of stratospheric and tropospheric aerosol size distributions. It was found that there is better agreement in the troposphere, even though the uncertainties of the calculation are greater for this region due to the variables in both the spatial concentration and the physical makeup of the aerosol. A second comparison study was made to test the consistency of the mean tropospheric extinction values at 1.02 micron (as reported by the SAGE satellite) with the values calculated from an ensemble of 400 measured size distributions thought to be representative of midcontinental tropospheric aerosol. The two methods produce consistent results within the expected degree of uncertainty. The ensemble of 400 'proven' size distributions is then used to calculate a statistical relationship between the 1.02-micron extinction and the 10.6-micron backscatter.

Rosen, James M.↗

A new method for simulating atmospheric turbulence for rotorcraft applications

Simulation of atmospheric turbulence as seen by a rotating blade element involves treatment of cyclostationary processes. Conventional filtering techniques do not lend themselves well to the generation of such turbulence sample functions as are required in rotorcraft flight dynamics simulation codes. A method to generate sample functions containing second order statistics of mean and covariance is presented. Compared to ensemble averaging involving excessive computer time, the novelty is to exploit cycloergodicity and thereby, replace ensemble averaging by averaging over a single path sample function of long duration. The method is validated by comparing its covariance results with the analytical and ensemble averaged results for a widely used 1-D turbulence approximation.

J. Riaz↗

Surface Terminations of LaAlO 3 Perovskite Nanoparticles as Viewed by Solid-State Nuclear Magnetic Resonance

Nanocrystal surfaces generally undergo reconstructions that differentiate them from the bulk structures, often in nontrivial ways. Understanding these terminations is critical across diverse fields, from heterogeneous catalysis to the formation of topological states and the synthesis of semiconductor nanomaterials. Determining surface structures is currently an interdisciplinary task, most often involving high-resolution electron microscopy and surface electron diffraction. These methods, however, do not provide a global view of the ensemble of structures present in a sample. Here, we show how surface-sensitive solid-state nuclear magnetic resonance (SSNMR) spectroscopy methods can bridge this gap. In this context, we investigated the surface structure of lanthanum aluminate (LaAlO 3 ) perovskite nanoparticles. Four distinct surface terminations have previously been observed for this material, but their relative abundances were unknown. Using an array of double- and triple-resonance SSNMR methods probing the relative proximities of surface 1 H, 27 Al, 17 O, and 139 La nuclei, we conclude the surface to be majority terminated (80%) by AlO x with substantial (20%) LaO x terminated regions.

Materials↗

Methods for Color Center Preserving Hydrogen‐Termination of Diamond

Abstract Chemical functionalization of diamond surfaces by hydrogen is an important method for controlling the charge state of near‐surface fluorescent color centers, an essential process in fabricating devices such as diamond field‐effect transistors and chemical sensors, and a required first step for realizing families of more complex terminations through subsequent chemical processing. In all these cases, termination is typically achieved using hydrogen plasma sources that can etch or damage the diamond, as well as deposited materials or embedded color centers. This work explores alternative methods for lower‐damage hydrogenation of diamond surfaces, specifically the annealing of diamond samples in high‐purity, non‐explosive mixtures of nitrogen and hydrogen gas, and the exposure of samples to microwave hydrogen plasmas in the absence of intentional stage heating. The effectiveness of these methods are characterized by x‐ray photoelectron spectroscopy (XPS), and comparison of the results to density‐functional modelling of the surface hydrogenation energetics implicates surface oxygen ligands as the primary factor limiting the termination quality of annealed samples. Finally, photoluminescence (PL) spectroscopy is used to verify that both the annealing and reduced sample temperature plasma methods are non‐destructive to near‐surface ensembles of nitrogen‐vacancy (NV) centers, in stark contrast to plasma treatments that use heated sample stages.

36 MATERIALS SCIENCE↗

Daily evapotranspiration changes during heatwaves at 32 NEON sites, 2019-2021

This dataset provides partitioned evapotranspiration (ET, the combined loss of water from soil and plant surfaces) anomalies during heatwave events—soil evaporation (E) and transpiration (T)—for 268 heatwave events across 32 National Ecological Observatory Network (NEON) flux sites in the contiguous United States from 2019–2021. Using an ensemble of four high-frequency turbulence methods (Flux-variance Similarity, Conditional Eddy Covariance [CEC], CEC with Water-Use Efficiency, and Conditional Eddy Accumulation; see Zahn and Bou-Zeid 2024), half-hourly transpiration-to-evapotranspiration (T/ET) ratios were derived from 20 hertz (Hz, cycles per second) eddy covariance measurements of carbon dioxide (CO₂) and water vapor (H₂O) concentrations. The dataset spans six vegetation types including evergreen and deciduous forests, grasslands, cultivated crops, shrublands, and emergent herbaceous wetlands. Data Package Contents: The dataset includes a single CSV (comma-separated values) file containing daily anomalies (deviations from baseline conditions) for transpiration (Delta_T), evaporation (Delta_E), total evapotranspiration (Delta_ET), and T/ET ratio (Delta_T_ET) during each day of identified heatwave events. The file also includes site codes, dates, heatwave event identifiers, and day-of-heatwave indicators. The CSV file can be opened with spreadsheet software (Microsoft Excel, Google Sheets) or programming environments (Python, R, MATLAB). This resource enables researchers to investigate ecosystem-specific responses to thermal extremes, validate land surface model partitioning of ET fluxes, and examine feedbacks between water cycling and surface energy balance during heatwaves. The dataset is particularly valuable for studies linking vegetation hydraulic strategies to climate resilience, as it captures the divergent responses of shallow-rooted versus deep-rooted ecosystems. Potential applications include improving drought early warning systems, informing irrigation management strategies, and advancing our mechanistic understanding of land-atmosphere interactions under extreme heat conditions.

Day of Heatwave↗

Linear Approximation to Optimal Control Allocation for Rocket Nozzles with Elliptical Constraints

In this paper we present a straightforward technique for assessing and realizing the maximum control moment effectiveness for a launch vehicle with multiple constrained rocket nozzles, where elliptical deflection limits in gimbal axes are expressed as an ensemble of independent quadratic constraints. A direct method of determining an approximating ellipsoid that inscribes the set of attainable angular accelerations is derived. In the case of a parameterized linear generalized inverse, the geometry of the attainable set is computationally expensive to obtain but can be approximated to a high degree of accuracy with the proposed method. A linear inverse can then be optimized to maximize the volume of the true attainable set by maximizing the volume of the approximating ellipsoid. The use of a linear inverse does not preclude the use of linear methods for stability analysis and control design, preferred in practice for assessing the stability characteristics of the inertial and servoelastic coupling appearing in large boosters. The present techniques are demonstrated via application to the control allocation scheme for a concept heavy-lift launch vehicle.

Orr, Jeb S.↗

Evaluation of a Regional Crop Model Implementation for Sub-National Yield Assessments in Kenya

CONTEXT: Cropping system models can be used to both assess regional food security and to monitor and predict agricultural drought. Agriculture in Kenya is extremely important to both the economy and food security of the country. OBJECTIVE: This study evaluated a regional implementation of a widely used crop model, the Decision Support System for Agrotechnology Transfer (DSSAT), within a coupled modeling framework, the Regional Hydrologic Extremes Assessment System (RHEAS), over Kenya. The goal of this study was to assess the ability of RHEAS to simulate the annual variability of maize yields at the county level and evaluate the uncertainty inherent in the model and inputs. METHODS: The RHEAS system implements a stochastic ensemble approach to account for field scale variabilities in crop management practices and underlying soil and weather conditions. Satellite-derived datasets were used to evaluate the land surface component of the system and seasonally disaggregated yield for 5 years was used to assess the performance of the cropping system model. RESULTS AND CONCLUSIONS: The median correlation between RHEAS and satellite-derived soil moisture and evapotranspiration estimates were 0.78, and 0.51, respectively, indicating that the model is able to capture the key drivers of the hydrological budget. Overall, RHEAS simulated yearly yield variations with a median correlation of 0.7 with reported yields, with the best performance in the short rains season. However, across both seasons, the RHEAS model was positively biased on the order of ~1.6 MT/ha. The overall median unbiased RMSE was 0.66 MT/ha. The RHEAS system shows skill at simulating extreme departures in anomalies, and a majority of the time (62.5%) the reported yields fall within the interquartile range of the simulations. SIGNIFICANCE: One of the most important areas of improvement for the next generation of agricultural data and models is to better understand and communicate the inherent uncertainties. This is especially critical in data-limited regions. Here we present a modeling system and its implementation that begins to address these concerns. We demonstrate the ability to simulate broad trends in yields at the county level for sub-annual yields with skills that commensurate previous national/annual level studies.

Crop model↗

Consistent Pixel Resolution Characterization of Deep Convective Clouds for Calibration

The NASA CERES project provides global TOA shortwave and longwave fluxes for climate monitoring and validation that spans over 20 years. CERES utilizes broadband fluxes derived from geostationary (GEO) imagers to estimate broadband fluxes between the CERES observations. In order for these fluxes to be viable, the GEO imager calibration must be stable over time. One such calibration method used by CERES is to evaluate ensemble sets of deep convective clouds (DCC) as invariant targets (IT) over time. DCC targets can also provide radiometric scaling between sensors. An international collaboration through GSICS is also evaluating the DCC-IT calibration methodology to provide consistent calibration coefficients across geostationary sensors. Tropical DCC are the coldest, brightest, and most Lambertian TOA Earth targets identified using a window channel brightness temperature threshold. The DCC-IT technique involves a large ensemble of TOA pixel-level reflectances, which are binned into probability density functions (PDF). The PDF structure dependency on sensor pixel resolution, which can vary greatly among sensors, is not well known. This study will identify the impact of pixel resolution on the DCC PDFs by aggregating VIIRS and Landsat OLI/TIRS pixel resolutions into various coarser pixel resolutions and comparing the shape and statistics of the resulting PDFs. This study should assist in mitigating the pixel resolution dependency in the DCC-IT approach for providing scaling factors between sensors.

Conor Haney↗

A new method for simulating atmospheric turbulence for rotorcraft applications

Simulation of atmospheric turbulence as seen by a rotating blade element involves treatment of cyclostationary processes. Conventional filtering techniques do not lend themselves well to the generation of such turbulence sample functions as are required in rotorcraft flight dynamics simulation codes. A method to generate sample functions containing second-order statistics of mean and covariance is presented. Compared to ensemble averaging involving excessive computer time, the novelty is to exploit cycloergodicity and thereby, replace ensemble averaging by averaging over a single-path sample function of long duration. The method is validated by comparing its covariance results with the analytical and ensemble-averaged results for a widely used one-dimensional turbulence approximation.

Prasad, J. V. R.↗

Adaptive Uncertainty Quantification for Stochastic Hyperbolic Conservation Laws

Here, we propose a predictor-corrector adaptive method for the study of hyperbolic partial differential equations (PDEs) under uncertainty. Constructed around the framework of stochastic finite volume (SFV) methods, our approach circumvents sampling schemes or simulation ensembles while also preserving fundamental properties, in particular hyperbolicity of the resulting systems and conservation of the discrete solutions. Furthermore, we augment the existing SFV theory with a priori convergence results for statistical quantities, in particular push-forward densities, which we demonstrate through numerical experiments. By linking refinement indicators to regions of the physical and stochastic spaces, we drive anisotropic refinements of the discretizations, introducing new degrees of freedom where deemed profitable. To illustrate our proposed method, we consider a series of numerical examples for nonlinear hyperbolic PDEs based on Burgers’ and Euler’s equations.

97 MATHEMATICS AND COMPUTING↗

An Ensemble-Based Smoother with Retrospectively Updated Weights for Highly Nonlinear Systems

Monte Carlo computational methods have been introduced into data assimilation for nonlinear systems in order to alleviate the computational burden of updating and propagating the full probability distribution. By propagating an ensemble of representative states, algorithms like the ensemble Kalman filter (EnKF) and the resampled particle filter (RPF) rely on the existing modeling infrastructure to approximate the distribution based on the evolution of this ensemble. This work presents an ensemble-based smoother that is applicable to the Monte Carlo filtering schemes like EnKF and RPF. At the minor cost of retrospectively updating a set of weights for ensemble members, this smoother has demonstrated superior capabilities in state tracking for two highly nonlinear problems: the double-well potential and trivariate Lorenz systems. The algorithm does not require retrospective adaptation of the ensemble members themselves, and it is thus suited to a streaming operational mode. The accuracy of the proposed backward-update scheme in estimating non-Gaussian distributions is evaluated by comparison to the more accurate estimates provided by a Markov chain Monte Carlo algorithm.

Monte Carlo↗