Search NASA⌕ Search

SEARCH · Search NASA

Results for “Ensemble methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

LTAU-FF: Loss Trajectory Analysis for Uncertainty in atomistic Force Fields

Model ensembles are effective tools for estimating prediction uncertainty in deep learning atomistic force fields. However, their widespread adoption is hindered by high computational costs and overconfident error estimates. In this work, we address these challenges by leveraging distributions of per-sample errors obtained during training and employing a distance-based similarity search in the model latent space. Our method, which we call LTAU (Loss Trajectory Analysis for Uncertainty), efficiently estimates the full probability distribution function of errors for any test point using the logged training errors, achieving speeds that are 2–3 orders of magnitudes faster than typical ensemble methods and allowing it to be used for tasks where training or evaluating multiple models would be infeasible. We apply LTAU towards estimating parametric uncertainty in atomistic force fields (LTAU-FF), demonstrating that it produces well-calibrated confidence intervals and predicts errors that correlate strongly with the true errors for data near the training domain. Furthermore, we show that the errors predicted by LTAU-FF can be used in practical applications for detecting out-of-domain data, tuning model performance, and predicting failure during simulations. We believe that LTAU will be a valuable tool for uncertainty quantification in atomistic force fields and is a promising method that should be further explored in other domains of machine learning.

97 MATHEMATICS AND COMPUTING↗

T-Matrix Computations of Light Scattering by Nonspherical Particles: A Review

We review the current status of Waterman's T-matrix approach which is one of the most powerful and widely used tools for accurately computing light scattering by nonspherical particles, both single and composite, based on directly solving Maxwell's equations. Specifically, we discuss the analytical method for computing orientationally-averaged light-scattering characteristics for ensembles of nonspherical particles, the methods for overcoming the numerical instability in calculating the T matrix for single nonspherical particles with large size parameters and/or extreme geometries, and the superposition approach for computing light scattering by composite/aggregated particles. Our discussion is accompanies by multiple numerical examples demonstrating the capabilities of the T-matrix approach and showing effects of nonsphericity of simple convex particles (spheroids) on light scattering.

Mischenko, Michael I.↗

Laplace-transform technique for deriving thermodynamic equations from the classical microcanonical ensemble

A direct and convenient method is presented for deriving expressions which equate any thermodynamic state function to averages of specific dynamical functions and their fluctuations over the classical microcanonical distribution. Specific expressions are obtained for a variety of thermodynamic quantities. The effect of various entropy definitions on the results are assessed, and the latter are compared to previous work in the literature. The derived formulas are applied to the analysis of molecular-dynamics computer simulations.

Pearson, E. M.↗

Scalable Generation of High-fidelity Synthetic Population Ensembles

Used within social simulations, synthetic population ensembles enable uncertainty quantification (UQ) methods for obtaining more robust model inference and prediction. A synthetic population ensemble is a series of plausible virtual reconstructions of an area’s population at the granularity of people and residences, generated stochastically to preserve privacy of the source population survey’s respondents. In this paper, we demonstrate the production of large synthetic population ensembles for the U.S. via Oak Ridge National Laboratory’s UrbanPop framework to support modeling of high spatial resolution energy affordability metrics from nationwide social surveys in collaboration with the fusionACS project. The study involves two scenarios: creating ensembles for (1) 17 U.S. metropolitan areas in 2019 and (2) full U.S. Census Divisions in 2023, with each scenario consisting of 41 population instances (a base realization and 40 replicates). To accomplish this task at scale, we configured an integrated system within a research cloud, comprised of virtual containerizations, GPU-enhanced functionality, and orchestrated deployments of UrbanPop’s maturing Likeness Python ecosystem. Results demonstrate we maintained high-fidelity approximations of residential totals by areas of interest and the demographic characteristics of neighborhoods while reducing manual workflow burdens. Finally, we discuss plans to fine-tune and further develop our automated workflows for truly distributed job orchestration to increase computational efficiency, as well as provide an outlook for broadening applications of the ensembles.

Cluster computing↗

Producing High-fidelity Synthetic Population Ensembles at Scale

Used within social simulations, synthetic population ensembles enable uncertainty quantification (UQ) methods for obtaining more robust model inference and prediction. A synthetic population ensemble is a series of plausible virtual reconstructions of an area’s population at the granularity of people and residences, generated stochastically to preserve privacy of the source population survey’s respondents. In this paper, we demonstrate the production of large synthetic population ensembles for the US via Oak Ridge National Laboratory’s UrbanPop framework to support modeling of high spatial resolution energy affordability metrics from nationwide social surveys in collaboration with the fusionACS project. Our initial task involves creating ensembles for 17 US metropolitan areas, each consisting of 41 population instances (a base realization and 40 replicates). To accomplish this task at scale, we configured an integrated system comprised of a research cloud, virtual containerization, GPU-enhanced functionality, and a dual API/CLI to interact with UrbanPop’s maturing Likeness Python ecosystem. We observe a reduction in theoretical execution time while maintaining high-fidelity approximations of residential totals by metropolitan area and the demographic characteristics of neighborhoods. We discuss expansion of our approach to produce synthetic population ensembles for the entire US, particularly plans to establish automated workflows for job orchestration to increase computational efficiency, as well as provide outlook for broadening applications of the ensembles.

Gaboardi, James [ORNL] (ORCID:0000000247766826)↗

Solar Rotational Modulations of Spectral Irradiance and Correlations with the Variability of Total Solar Irradiance

Aims: We characterize the solar rotational modulations of spectral solar irradiance (SSI) and compare them with the corresponding changes of total solar irradiance (TSI). Solar rotational modulations of TSI and SSI at wavelengths between 120 and 1600 nm are identified over one hundred Carrington rotational cycles during 2003-2013. Methods: The SORCE (Solar Radiation and Climate Experiment) and TIMED (Thermosphere Ionosphere Mesosphere Energetics and Dynamics)/SEE (Solar EUV Experiment) measured and SATIRE-S modeled solar irradiances are analyzed using the EEMD (Ensemble Empirical Mode Decomposition) method to determine the phase and amplitude of 27-day solar rotational variation in TSI and SSI. Results: The mode decomposition clearly identifies 27-day solar rotational variations in SSI between 120 and 1600 nm, and there is a robust wavelength dependence in the phase of the rotational mode relative to that of TSI. The rotational modes of visible (VIS) and near infrared (NIR) are in phase with the mode of TSI, but the phase of the rotational mode of ultraviolet (UV) exhibits differences from that of TSI. While it is questionable that the VIS to NIR portion of the solar spectrum has yet been observed with sufficient accuracy and precision to determine the 11-year solar cycle variations, the temporal variations over one hundred cycles of 27-day solar rotation, independent of the two solar cycles in which they are embedded, show distinct solar rotational modulations at each wavelength.

Spectral Solar Irradiance↗

Quantifying the Uncertainties in an Ensemble of Decadal Climate Predictions

Meaningful climate predictions should be accompanied by the corresponding uncertainty range. Common methods for estimating the uncertainty range are based on the spread of ensemble predictions. However, a simulation ensemble is not necessarily a proper sample of the real distribution of the climate, and therefore, the ensemble spread cannot be interpreted as the actual uncertainty. We propose a new method that links between the ensemble spread and the uncertainty without relying on any assumptions regarding the distribution of the ensemble predictions. The method is tested using CMIP5 1981-2010 decadal predictions and is shown to outperform other common methods.

Strobach, Ehud↗

Ensemble Data Assimilation Without Ensembles: Methodology and Application to Ocean Data Assimilation

Two methods to estimate background error covariances for data assimilation are introduced. While both share properties with the ensemble Kalman filter (EnKF), they differ from it in that they do not require the integration of multiple model trajectories. Instead, all the necessary covariance information is obtained from a single model integration. The first method is referred-to as SAFE (Space Adaptive Forecast error Estimation) because it estimates error covariances from the spatial distribution of model variables within a single state vector. It can thus be thought of as sampling an ensemble in space. The second method, named FAST (Flow Adaptive error Statistics from a Time series), constructs an ensemble sampled from a moving window along a model trajectory. The underlying assumption in these methods is that forecast errors in data assimilation are primarily phase errors in space and/or time.

Data Assimilation↗

Role of Forcing Uncertainty and Background Model Error Characterization in Snow Data Assimilation

Accurate specification of the model error covariances in data assimilation systems is a challenging issue. Ensemble land data assimilation methods rely on stochastic perturbations of input forcing and model prognostic fields for developing representations of input model error covariances. This article examines the limitations of using a single forcing dataset for specifying forcing uncertainty inputs for assimilating snow depth retrievals. Using an idealized data assimilation experiment, the article demonstrates that the use of hybrid forcing input strategies (either through the use of an ensemble of forcing products or through the added use of the forcing climatology) provide a better characterization of the background model error, which leads to improved data assimilation results, especially during the snow accumulation and melt-time periods. The use of hybrid forcing ensembles is then employed for assimilating snow depth retrievals from the AMSR2 (Advanced Microwave Scanning Radiometer 2) instrument over two domains in the continental USA with different snow evolution characteristics. Over a region near the Great Lakes, where the snow evolution tends to be ephemeral, the use of hybrid forcing ensembles provides significant improvements relative to the use of a single forcing dataset. Over the Colorado headwaters characterized by large snow accumulation, the impact of using the forcing ensemble is less prominent and is largely limited to the snow transition time periods. The results of the article demonstrate that improving the background model error through the use of a forcing ensemble enables the assimilation system to better incorporate the observational information.

assimilation↗

On the Confidence Limit of Hilbert Spectrum

Confidence limit is a routine requirement for Fourier spectral analysis. But this confidence limit is established based on ergodic theory: For stationary process, temporal average equals the ensemble average. Therefore, one can divide the data into n-sections and treat each section as independent realization. Most natural processes in general, and climate data in particular, are not stationary; therefore, there is a need for the Hilbert Spectral analysis for such processes. Here ergodic theory is no longer applicable. We propose to use various adjustable parameters in the shifting processes of the Empirical Mode Decomposition (EMD) method to obtain an ensemble of Intrinsic Mode Function 0 sets. Based on such an ensemble, we introduce a statistical measure in. a form of confidence limits for the Intrinsic Mode Functions, and consequently, the Hilbert spectra. The criterion of selecting the various adjustable parameters is based on the orthogonality test of the resulting M F sets. Length-of-day data from 1962 to 2001 will be used to illustrate this new approach. Its implication in climate data analysis will also be discussed.

Huang, Norden↗

Enhancing Automotive Intrusion Detection Through Multi-Modal Fusion: A CAN FD-LiDAR Approach

As vehicles become smarter and more autonomous, they increasingly depend on advanced sensors and communication technologies to operate securely. However, such growing dependence on technology—whether it’s CAN (Controller Area Network) for internal communication or LiDAR (Light Detection and Ranging) for sensing the world around them—also expands the attack surface for the types of cyber attacks. Traditional intrusion detection systems (IDS) typically monitor these systems in isolation, limiting their ability to detect sophisticated, crosssystem attacks. To address this, we propose a multi-modal fusion approach that combines real-world CAN FD signals (from the HCRL dataset) with LiDAR features (from the nuScenes dataset) to enhance attack detection. Our method employs a twostage ensemble approach. Calibrated XGBoost and LightGBM models initially process CAN FD (Fuzzing Data) and LiDAR data independently, detecting timing anomalies and space abnormalities. They are subsequently logarithmically combined with a logistic regression meta-model along with 17 engineered features capturing cross-modal behavior, prediction conflicts, and nonlinear interactions. This approach achieves an AUC of 0.87 and an F1-score of 0.82, surpassing single-modality baselines and early fusion methods, at merely 2 ms inference latency. Compared with deep learning competitors, it is 3 times more efficient, providing a lightweight, interpretable, and real time solution to automotive cybersecurity.

97 MATHEMATICS AND COMPUTING↗

Designing robust energy policy packages under deep uncertainty: A multi-metric decision support framework

The complexity of transitioning to sustainable energy systems requires policy frameworks capable of balancing multiple objectives while addressing deep uncertainty. However, existing approaches often lack systematic methods to identify combinations of policy levers that remain effective across a wide range of uncertain futures. This paper presents a novel decision support framework that guides the selection of robust policy packages based on their performance across multiple objectives under uncertainty. Our method leverages a large ensemble of scenarios and applies scenario discovery techniques to identify influential policy levers. Here, we introduce new indicators to assess the robustness of policies by evaluating their ability to mitigate adverse outcomes across metrics. These indicators support an iterative process to build a robust policy package. Finally, we map the technological and energy pathways associated with the robust policy package by leveraging an energy system optimization model. We illustrate the application of this framework to the Spanish energy system, providing insights into how specific combinations of policy levers shape decarbonization pathways under uncertainty.

Decision-support method↗

Antarctic ice sheet model comparison with uncurated geological constraints shows that higher spatial resolution improves deglacial reconstructions

Accurately reconstructing past changes to the shape and volume of the Antarctic ice sheet relies on the use of physically based and thus internally consistent ice sheet modeling, benchmarked against spatially limited geologic data. The challenge in model benchmarking against geologic data is diagnosing whether model-data misfits are the result of an inadequate model, inherently noisy or biased geologic data, and/or incorrect association between modeled quantities and geologic observations. In this work we address this challenge by (i) the development and use of a new model-data evaluation framework applied to an uncurated data set of geologic constraints, and (ii) nested high-spatial-resolution modeling designed to test the hypothesis that model resolution is an important limitation in matching geologic data. While previous approaches to model benchmarking employed highly curated datasets, our approach applies an automated screening and quality control algorithm to an uncurated public dataset of geochronological observations (specifically, cosmogenic-nuclide exposure-age measurements from glacial deposits in ice-free areas). This optimizes data utilization by including more geological constraints, reduces potential interpretive bias, and allows unsupervised assimilation of new data as they are collected. We also incorporate a nested model framework in which high-resolution domains are downscaled from a continent-wide ice sheet model. We highlight the application of this framework by applying these methods to a small ensemble of deglacial ice-sheet model simulations, and demonstrate that the nested approach improves the ability of model simulations to match exposure age data collected from areas of complex topography and ice flow. We develop a range of diagnostic model-data comparison metrics to provide more insight into model performance than possible from a single-valued misfit statistic, showing that different metrics capture different aspects of ice sheet deflation.

Geosciences↗

Analyses of Cometary Silicate Crystals: DDA Spectral Modeling of Forsterite

Comets are the Solar System's deep freezers of gases, ices, and particulates that were present in the outer protoplanetary disk. Where comet nuclei accreted was so cold that CO ice (approximately 50K) and other supervolatile ices like ethane (C2H2) were preserved. However, comets also accreted high temperature minerals: silicate crystals that either condensed (greater than or equal to 1400 K) or that were annealed from amorphous (glassy) silicates (greater than 850-1000 K). By their rarity in the interstellar medium, cometary crystalline silicates are thought to be grains that formed in the inner disk and were then radially transported out to the cold and ice-rich regimes near Neptune. The questions that comets can potentially address are: How fast, how far, and over what duration were crystals that formed in the inner disk transported out to the comet-forming region(s)? In comets, the mass fractions of silicates that are crystalline, f_cryst, translate to benchmarks for protoplanetary disk radial transport models. The infamous comet Hale-Bopp has crystalline fractions of over 55%. The values for cometary crystalline mass fractions, however, are derived assuming that the mineralogy assessed for the submicron to micron-sized portion of the size distribution represents the compositional makeup of all larger grains in the coma. Models for fitting cometary SEDs make this assumption because models can only fit the observed features with submicron to micron-sized discrete crystals. On the other hand, larger (0.1-100 micrometer radii) porous grains composed of amorphous silicates and amorphous carbon can be easily computed with mixed medium theory wherein vacuum mixed into a spherical particle mimics a porous aggregate. If crystalline silicates are mixed in, the models completely fail to match the observations. Moreover, models for a size distribution of discrete crystalline forsterite grains commonly employs the CDE computational method for ellipsoidal platelets (c:a:b=8.14x8.14xl in shape with geometrical factors of x:y:z=1:1:10, Fabian et al. 2001; Harker et al. 2007). Alternatively, models for forsterite employ statistical methods like the Distribution of Hollow Spheres (Min et al. 2008; Oliveira et al. 2011) or Gaussian Random Spheres (GRS) or RGF (Gielen et al. 200S). Pancakes, hollow spheres, or GRS shapes similar to wheat sheaf crystal habit (e.g., Volten et al. 2001; Veihelmann et al. 2006), however, do not have the sharp edges, flat faces, and vertices seen in images of cometary crystals in interplanetary dust particles (IDPs) or in Stardust samples. Cometary forsterite crystals often have equant or tabular crystal habit (J. Bradley). To simulate cometary crystals, we have computed absorption efficiencies of forsterite using the Discrete Dipole Approximation (DDA) DDSCAT code on NAS supercomputers. We compute thermal models that employ a size distribution of discrete irregularly shaped forsterite crystals (nonspherical shapes with faces and vertices) to explore how crystal shape affects the shape and wavelength positions of the forsterite spectral features and to explore whether cometary crystal shapes support either condensation or annealing scenarios (Lindsay et al. 2012a, b). We find forsterite crystal shapes that best-fit comet Hale-Bopp are tetrahedron, bricks or brick platelets, essentially equant or tabular (Lindsay et al. 2012a,b), commensurate with high temperature condensation experiments (Kobatake et al. 2008). We also have computed porous aggregates with crystal monomers and find that the crystal resonances are amplified. i.e., the crystalline fraction is lower in the aggregate than is derived by fitting a linear mix of spectral features from discrete subcomponents, and the crystal resonances 'appear' to be from larger crystals (Wooden et al. 2012). These results may indicate that the crystalline mass fraction in comets with comae dominated by aggregates may be lower than deduced by popular methods that only emoy ensembles of discrete crystals.

Wooden, Diane↗

The 27-Day Rotational Variations in Total Solar Irradiance Observations: from SORCE-TIM, ACRIMSAT-ACRIM III, and SOHO-VIRGO

During the last decade, observations from SORCE (Solar Radiation and Climate Experiment)/TIM (Total Irradiance Monitor), ACRIMSAT (Active Cavity Radiometer Irradiance Monitor Satellite)/ACRIM III, and SOHO (Solar and Heliospheric Observatory)VIRGO (Variability of IRradiance and Gravity Oscillations Sun PhotoMeter) provided the Total Solar Irradiance (TSI) measurements with unprecedented accuracy and stability to determine the amount of solar irradiance reaching the top of the atmosphere and how solar irradiance varies in different time scales. These three independent measurements are analyzed using the EEMD (Ensemble Empirical Mode Decomposition) method to characterize the phase and amplitude of 27-day solar rotational variation in TSI. The mode decomposition clearly identifies a 27-day solar rotational signature in TSI measurements. The rotational variations of TSI from the three independent observations are generally consistent with each other, despite different mean TSI values. During the declining phase of solar cycle 23, the amplitude of TSI 27-day variations is as high as 0.8 watts per square meter (approximately 0.05 percent), while during the rising phase of solar cycle 24, the amplitude is up to 0.4 watts per square meter (approximately 0.04 percent). During the minimum phase (2008-2009), the amplitude of the rotational mode is only 0.1 watts per square meter. The correlation of this rotational mode between TIM and ACRIM III is approximately 0.92 and the slope of the local peak values is approximately 0.98. The correlation between TIM and VIRGO is approximately 0.96 and the slope of the local peak values isapproximately 0.98, very similar to the slope with ACRIM III.

ACRIM III↗

Global Evolution of Solar Magnetic Fields and Prediction of Solar Activity Cycles

Prediction of solar activity cycles is challenging because the physical processes inside the Sun involve a broad range of multiscale dynamics that no model can reproduce, and the available observations are highly limited and cover mostly surface layers. Helioseismology makes it possible to probe solar dynamics in the convective zone, but variations in the differential rotation and meridional circulation are currently available for only two solar activity cycles. It has been demonstrated that sunspot observations, which cover over 400 years, can be used to calibrate the Parker-Kleeorin-Ruzmaikin model and that the Ensemble Kalman Filter (EnKF) method can be used to link the model magnetic fields to sunspot observations to make reliable predictions of a following cycle. However, for more accurate predictions, it is necessary to use actual observations of the solar magnetic fields, which are available for only four solar cycles. This raises the question of how limitations in observational data and model uncertainties affect predictive capabilities and implies the need for the development of new forecast methodologies and validation criteria. In this presentation, I will discuss the influence of the limited number of available observations on the accuracy of EnKF estimates of solar cycle parameters.

Kitiashvili, Irina N.↗

Application of Synoptic Magnetograms for Prediction of Solar Activity Using Ensemble Kalman Filter

Solar activity predictions using the data assimilation approach have demonstrated great potential to build reliable long-term forecasts of solar activity. In particular, it has been shown that the Ensemble Kalman Filter (EnKF) method applied to a non-linear dynamo model is capable of predicting solar activity up to one sunspot cycle ahead in time, as well as estimating the properties of the next cycle a few years before it begins. These developments assume an empirical relationship between the mean toroidal magnetic field flux and the sunspot number. Estimated from the sunspot number series, variations of the toroidal field have been used to assimilate the data into the Parker-Kleeorin-Ruzmakin (PKR) dynamo model by applying the EnKF method. The dynamo model describes the evolution of the toroidal and poloidal components of the magnetic field and the magnetic helicity. Full-disk magnetograms provide more accurate and complete input data by constraining both the toroidal and poloidal global field components, but these data are available only for the last four solar cycles. In this presentation, using the available magnetogram data, we discuss development of the methodology and forecast quality criteria (including forecast uncertainties and sources of errors). We demonstrate the influence of limited time series observations on the accuracy of solar activity predictions. We present EnKF predictions of the upcoming Solar Cycle 25 based on both the sunspot number series and observed magnetic fields and discuss the uncertainties and potential of the data assimilation approach.

Kitiashvili, Irina N.↗

Solar Activity Modeling: From Subgranular Dynamical Scales to the Solar Cycles

Dynamical effects of solar magnetoconvection span a wide range spatial and temporal scales that extends from the interior to the corona and from fast turbulent motions to the global-Sun magnetic activity. To study the solar activity on short temporal scales (from minutes to hours), we use 3D radiative MHD simulations that allow us to investigate complex turbulent interactions that drive various phenomena, such as plasma eruptions, spontaneous formation of magnetic structures, funnel-like structures and magnetic loops in the corona, and others. In particular, we focus on multi-scale processes of energy exchange across the different layers, which contribute to the corona heating and eruptive dynamics, as well as interlinks between different layers of the solar interior and atmosphere. For modeling the global-scale activity we use the data assimilation approach that has demonstrated great potential for building reliable long-term forecasts of solar activity. In particular, it has been shown that the Ensemble Kalman Filter (EnKF) method applied to the Parker-Kleeorin-Ruzmakin dynamo model is capable of predicting solar activity up to one sunspot cycle ahead in time, as well as estimating the properties of the next cycle a few years before it begins. In this presentation, using the available magnetogram data, we discuss development of the methodology and forecast quality criteria (including forecast uncertainties and sources of errors). We demonstrate the influence of observational limitation on the prediction accuracy. We present the EnKF predictions of the upcoming Solar Cycle 25 based on both the sunspot number series and observed magnetic fields, and discuss the uncertainties and potential of the data assimilation approach for modeling and forecasting the solar activity.

Kitiashvili, I. N.↗