Search NASA⌕ Search

SEARCH · Search NASA

Results for “Ensemble methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Direct Simulation of Extinction in a Slab of Spherical Particles

The exact multiple sphere superposition method is used to calculate the coherent and incoherent contributions to the ensemble-averaged electric field amplitude and Poynting vector in systems of randomly positioned nonabsorbing spherical particles. The target systems consist of cylindrical volumes, with radius several times larger than length, containing spheres with positional configurations generated by a Monte Carlo sampling method. Spatially dependent values for coherent electric field amplitude, coherent energy flux, and diffuse energy flux, are calculated by averaging of exact local field and flux values over multiple configurations and over spatially independent directions for fixed target geometry, sphere properties, and sphere volume fraction. Our results reveal exponential attenuation of the coherent field and the coherent energy flux inside the particulate layer and thereby further corroborate the general methodology of the microphysical radiative transfer theory. An effective medium model based on plane wave transmission and reflection by a plane layer is used to model the dependence of the coherent electric field on particle packing density. The effective attenuation coefficient of the random medium, computed from the direct simulations, is found to agree closely with effective medium theories and with measurements. In addition, the simulation results reveal the presence of a counter-propagating component to the coherent field, which arises due to the internal reflection of the main coherent field component by the target boundary. The characteristics of the diffuse flux are compared to, and found to be consistent with, a model based on the diffusion approximation of the radiative transfer theory.

spheres↗

Project trades model for complex space missions

A Project Trades Model (PTM) is a collection of tools/simulations linked together to rapidly perform integrated system trade studies of performance, cost, risk, and mission effectiveness. An operating PTM captures the interactions between various targeted systems and subsystems through an exchange of computed variables of the constituent models. Selection and implementation of the order, method of interaction, model type, and envisioned operation of the ensemble of tools rpresents the key system engineering challenge of the approach. This paper describes an approach to building a PTM and using it to perform top-level system trades for a complex space mission. In particular, the PTM discussed here is for a future Mars mission involving a large rover.

web services↗

Advances in Land Data Assimilation at the NASA Goddard Space Flight Center

Research in land surface data assimilation has grown rapidly over the last decade. In this presentation we provide a brief overview of key research contributions by the NASA Goddard Space Flight Center (GSFC). The GSFC contributions to land assimilation primarily include the continued development and application of the Land Information System (US) and the ensemble Kalman filter (EnKF). In particular, we have developed a method to generate perturbation fields that are correlated in space, time, and across variables and that permit the flexible modeling of errors in land surface models and observations, along with an adaptive filtering approach that estimates observation and model error input parameters. A percentile-based scaling method that addresses soil moisture biases in model and observational estimates opened the path to the successful application of land data assimilation to satellite retrievals of surface soil moisture. Assimilation of AMSR-E surface soil moisture retrievals into the NASA Catchment model provided superior surface and root zone assimilation products (when validated against in situ measurements and compared to the model estimates or satellite observations alone). The multi-model capabilities of US were used to investigate the role of subsurface physics in the assimilation of surface soil moisture observations. Results indicate that the potential of surface soil moisture assimilation to improve root zone information is higher when the surface to root zone coupling is stronger. Building on this experience, GSFC leads the development of the Level 4 Surface and Root-Zone Soil Moisture (L4_SM) product for the planned NASA Soil-Moisture-Active-Passive (SMAP) mission. A key milestone was the design and execution of an Observing System Simulation Experiment that quantified the contribution of soil moisture retrievals to land data assimilation products as a function of retrieval and land model skill and yielded an estimate of the error budget for the SMAP L4_SM product. Terrestrial water storage observations from GRACE satellite system were also successfully assimilated into the NASA Catchment model and provided improved estimates of groundwater variability when compared to the model estimates alone. Moreover, satellite-based land surface temperature (LST) observations from the ISCCP archive were assimilated using a bias estimation module that was specifically designed for LST assimilation. As with soil moisture, LST assimilation provides modest yet statistically significant improvements when compared to the model or satellite observations alone. To achieve the improvement, however, the LST assimilation algorithm must be adapted to the specific formulation of LST in the land model. An improved method for the assimilation of snow cover observations was also developed. Finally, the coupling of LIS to the mesoscale Weather Research and Forecasting (WRF) model enabled investigations into how the sensitivity of land-atmosphere interactions to the specific choice of planetary boundary layer scheme and land surface model varies across surface moisture regimes, and how it can be quantified and evaluated against observations. The on-going development and integration of land assimilation modules into the Land Information System will enable the use of GSFC software with a variety of land models and make it accessible to the research community.

Reichle, Rolf↗

Observational Constraints Reduce Model Spread but Not Uncertainty in Global Wetland Methane Emission Estimates

The recent rise in atmospheric methane (CH 4 ) concentrations accelerates climate change and offsets mitigation efforts. Although wetlands are the largest natural CH 4 source, estimates of global wetland CH 4 emissions vary widely among approaches taken by bottom-up (BU) process-based biogeochemical models and top-down (TD) atmospheric inversion methods. Here, we integrate in situ measurements, multi-model ensembles, and a machine learning upscaling product into the International Land Model Benchmarking system to examine the relationship between wetland CH 4 emission estimates and model performance. We find that using better-performing models identified by observational constraints reduces the spread of wetland CH 4 emission estimates by 62% and 39% for BU- and TD-based approaches, respectively. However, global BU and TD CH 4 emission estimate discrepancies increased by about 15% (from 31 to 36 TgCH 4 year −1 ) when the top 20% models were used, although we consider this result moderately uncertain given the unevenly distributed global observations. Our analyses demonstrate that model performance ranking is subject to benchmark selection due to large inter-site variability, highlighting the importance of expanding coverage of benchmark sites to diverse environmental conditions. We encourage future development of wetland CH 4 models to move beyond static benchmarking and focus on evaluating site-specific and ecosystem-specific variabilities inferred from observations.

Kuang-Yu Chang↗

Bayesian reduced-order deep learning surrogate model for dynamic systems described by partial differential equations

We propose a reduced-order deep-learning surrogate model for dynamic systems described by time-dependent partial differential equations. This method employs space–time Karhunen–Loève expansions (KLEs) of the state variables and space-dependent KLEs of space-varying parameters to identify the reduced (latent) dimensions. Subsequently, a deep neural network (DNN) is used to map the parameter latent space to the state variable latent space. An approximate Bayesian method is developed for uncertainty quantification (UQ) in the proposed KL-DNN surrogate model. The KL-DNN method is tested for the linear advection–diffusion and nonlinear diffusion equations, and the Bayesian approach for UQ is compared with the deep ensembling (DE) approach, commonly used for quantifying uncertainty in DNN models. It was found that the approximate Bayesian method provides a more informative distribution of the PDE solutions in terms of the coverage of the reference PDE solutions (the percentage of nodes where the reference solution is within the confidence interval predicted by the UQ methods) and log predictive probability. The DE method is found to underestimate uncertainty and introduce bias. For the nonlinear diffusion equation, we compare the KL-DNN method with the Fourier Neural Operator (FNO) method and find that KL-DNN is 10% more accurate and needs less training time than the FNO method.

97 MATHEMATICS AND COMPUTING↗

Physics-Based Machine Learning Methods for U-235 Forensics Signatures

Signatures of low-intensity U-235 sources have been recently studied by utilizing a variety of machine learning (ML) classifiers using features derived from gamma spectral measurements collected under structured campaigns. Several ML classifiers, such as ensemble of tress and classification trees, revealed misleadingly-optimistic training error due to over-fitting, and furthermore, their performance is not directly relatable to the physical properties due to their data-driven, opaque designs. We present a regression-based ML method that first estimates the inverse distance to the source and then utilizes a threshold to infer its presence, by representing the background as a source located at an infinite distance. For the inverse distance estimation, we study the ensemble of trees and Gaussian process regression methods, and a hyper parameter auto-tuning and selection method that employs five regression estimators. These methods avoid the over-fitting observed in several ML classifiers, while providing the classification error nearly comparable to them based on independent test data. Their error is directly related to estimates of the inverse physical distance to source, and the precision of error determines the seperability property that determines the false alarm and missed detection rates. The property of monotonic decrease of the source strength with increasing detector distance combined with Poisson distribution of measurements is utilized to analytically validate these methods by deriving the generalization equations of underlying regression methods.

Rao, Nageswara↗

Large Eddy simulation of turbulence: A subgrid scale model including shear, vorticity, rotation, and buoyancy

The Reynolds numbers that characterize geophysical and astrophysical turbulence (Re approximately equals 10(exp 8) for the planetary boundary layer and Re approximately equals 10(exp 14) for the Sun's interior) are too large to allow a direct numerical simulation (DNS) of the fundamental Navier-Stokes and temperature equations. In fact, the spatial number of grid points N approximately Re(exp 9/4) exceeds the computational capability of today's supercomputers. Alternative treatments are the ensemble-time average approach, and/or the volume average approach. Since the first method (Reynolds stress approach) is largely analytical, the resulting turbulence equations entail manageable computational requirements and can thus be linked to a stellar evolutionary code or, in the geophysical case, to general circulation models. In the volume average approach, one carries out a large eddy simulation (LES) which resolves numerically the largest scales, while the unresolved scales must be treated theoretically with a subgrid scale model (SGS). Contrary to the ensemble average approach, the LES+SGS approach has considerable computational requirements. Even if this prevents (for the time being) a LES+SGS model to be linked to stellar or geophysical codes, it is still of the greatest relevance as an 'experimental tool' to be used, inter alia, to improve the parameterizations needed in the ensemble average approach. Such a methodology has been successfully adopted in studies of the convective planetary boundary layer. Experienc e with the LES+SGS approach from different fields has shown that its reliability depends on the healthiness of the SGS model for numerical stability as well as for physical completeness. At present, the most widely used SGS model, the Smagorinsky model, accounts for the effect of the shear induced by the large resolved scales on the unresolved scales but does not account for the effects of buoyancy, anisotropy, rotation, and stable stratification. The latter phenomenon, which affects both geophysical and astrophysical turbulence (e.g., oceanic structure and convective overshooting in stars), has been singularly difficult to account for in turbulence modeling. For example, the widely used model of Deardorff has not been confirmed by recent LES results. As of today, there is no SGS model capable of incorporating buoyancy, rotation, shear, anistropy, and stable stratification (gravity waves). In this paper, we construct such a model which we call CM (complete model). We also present a hierarchy of simpler algebraic models (called AM) of varying complexity. Finally, we present a set of models which are simplified even further (called SM), the simplest of which is the Smagorinsky-Lilly model. The incorporation of these models into the presently available LES codes should begin with the SM, to be followed by the AM and finally by the CM.

Canuto, V. M.↗

Large-Scale High-Resolution Coastal Mangrove Forests Mapping Across West Africa With Machine Learning Ensemble and Satellite Big Data

Coastal mangrove forests provide important ecosystem goods and services, including carbon sequestration, biodiversity conservation, and hazard mitigation. However, they are being destroyed at an alarming rate by human activities. To characterize mangrove forest changes, evaluate their impacts, and support relevant protection and restoration decision making, accurate and up-to-date mangrove extent mapping at large spatial scales is essential. Available large-scale mangrove extent data products use a single machine learning method commonly with 30 m Landsat imagery, and significant inconsistencies remain among these data products. With huge amounts of satellite data involved and the heterogeneity of land surface characteristics across large geographic areas, finding the most suitable method for large-scale high-resolution mangrove mapping is a challenge. The objective of this study is to evaluate the performance of a machine learning ensemble for mangrove forest mapping at 20 m spatial resolution across West Africa using Sentinel-2 (optical) and Sentinel-1 (radar) imagery. The machine learning ensemble integrates three commonly used machine learning methods in land cover and land use mapping, including Random Forest (RF), Gradient Boosting Machine (GBM), and Neural Network (NN). The cloud-based big geospatial data processing platform Google Earth Engine (GEE) was used for pre-processing Sentinel-2 and Sentinel-1 data. Extensive validation has demonstrated that the machine learning ensemble can generate mangrove extent maps at high accuracies for all study regions in West Africa (92%–99% Producer’s Accuracy, 98%–100% User’s Accuracy, 95%–99% Overall Accuracy). This is the first-time that mangrove extent has been mapped at a 20 m spatial resolution across West Africa. The machine learning ensemble has the potential to be applied to other regions of the world and is therefore capable of producing high-resolution mangrove extent maps at global scales periodically.

coastal environment↗

Acoustic Testing of Flight Hardware Using Loudspeakers: How Much do We Know About This Method of Testing?

Loudspeakers have been used for acoustic qualification of spacecrafts, reflectors, solar panels, and other acoustically responsive structures for more than a decade. Even though a lot of hardware has been acoustic tested using this method, the nature of the acoustic field generated by controlling an ensemble of speakers with and without the hardware in the test volume has not been thoroughly investigated. Limited measurements from some of the recent speaker tests used to qualify flight hardware have indicated significant spatial variation of the acoustic field within the test volume. Also structural responses have been reported to differ when similar tests were performed using reverberant chambers. Unlike the reverberant chamber acoustic test, for which the acoustic field in most chambers is known to be diffuse except below several tens of Hz where acoustic standing waves and large spatial variations exist, the characteristics of the acoustic field within the speaker test volume has not been quantified. It has only been recently that a detailed acoustic field characterization of speaker testing has been made at Jet Propulsion Laboratory (JPL) with involvement of various organizations. To address the impact of non-uniform acoustic field on structures, a series of acoustic tests were performed using a flat panel and a 3-ft cylinder exposed to the field controlled by speakers and repeated in a reverberant chamber. The analysis of the data from this exercise reveals that there are significant differences both in the acoustic field and in the structural responses. In this paper the differences between the two methods are reviewed in some detail and the over- or under-testing of articles that could pose un-anticipated structural and flight qualification issues are discussed. A framework for discussing the validity of the speaker acoustic testing method with the current control system and a path forward for improving it will be provided.

Acoustics↗

Robust Strategy for Rocket Engine Health Monitoring

Monitoring the health of rocket engine systems is essentially a two-phase process. The acquisition phase involves sensing physical conditions at selected locations, converting physical inputs to electrical signals, conditioning the signals as appropriate to establish scale or filter interference, and recording results in a form that is easy to interpret. The inference phase involves analysis of results from the acquisition phase, comparison of analysis results to established health measures, and assessment of health indications. A variety of analytical tools may be employed in the inference phase of health monitoring. These tools can be separated into three broad categories: statistical, rule based, and model based. Statistical methods can provide excellent comparative measures of engine operating health. They require well-characterized data from an ensemble of "typical" engines, or "golden" data from a specific test assumed to define the operating norm in order to establish reliable comparative measures. Statistical methods are generally suitable for real-time health monitoring because they do not deal with the physical complexities of engine operation. The utility of statistical methods in rocket engine health monitoring is hindered by practical limits on the quantity and quality of available data. This is due to the difficulty and high cost of data acquisition, the limited number of available test engines, and the problem of simulating flight conditions in ground test facilities. In addition, statistical methods incur a penalty for disregarding flow complexity and are therefore limited in their ability to define performance shift causality. Rule based methods infer the health state of the engine system based on comparison of individual measurements or combinations of measurements with defined health norms or rules. This does not mean that rule based methods are necessarily simple. Although binary yes-no health assessment can sometimes be established by relatively simple rules, the causality assignment needed for refined health monitoring often requires an exceptionally complex rule base involving complicated logical maps. Structuring the rule system to be clear and unambiguous can be difficult, and the expert input required to maintain a large logic network and associated rule base can be prohibitive.

Santi, L. Michael↗

Permafrost Region Greenhouse Gas Budgets Suggest a Weak CO 2 Sink and CH 4 and N 2 O Sources, But Magnitudes Differ Between Top-Down and Bottom-Up Methods

Large stocks of soil carbon (C) and nitrogen (N) in northern permafrost soils are vulnerable to remobilization under climate change. However, there are large uncertainties in present-day greenhouse gas (GHG) budgets. We compare bottom-up (data-driven upscaling and process-based models) and top-down (atmospheric inversion models) budgets of carbon dioxide (CO 2 ), methane (CH 4 ) and nitrous oxide (N 2 O) as well as lateral fluxes of C and N across the region over 2000–2020. Bottom-up approaches estimate higher land-to-atmosphere fluxes for all GHGs. Both bottom-up and top-down approaches show a sink of CO 2 in natural ecosystems (bottom-up: -29 (-709, 455), top-down: -587 (-862, -312) Tg CO 2 -C yr -1 ) and sources of CH 4 (bottom-up: 38 (22, 53), top-down: 15 (11, 18) Tg CH 4 -C y -1 ) and N 2 O (bottom-up: 0.7 (0.1, 1.3), top-down: 0.09 (-0.19, 0.37) Tg N 2 O-N yr -1 ). The combined global warming potential of all three gases (GWP-100) cannot be distinguished from neutral. Over shorter timescales (GWP-20), the region is a net GHG source because CH 4 dominates the total forcing. The net CO 2 sink in Boreal forests and wetlands is largely offset by fires and inland water CO 2 emissions as well as CH 4 emissions from wetlands and inland waters, with a smaller contribution from N 2 O emissions. Priorities for future research include the representation of inland waters in process-based models and the compilation of process-model ensembles for CH 4 and N 2 O. Discrepancies between bottom-up and top-down methods call for analyses of how prior flux ensembles impact inversion budgets, more and well-distributed in situ GHG measurements and improved resolution in upscaling techniques.

54 ENVIRONMENTAL SCIENCES↗

Standardising the “Gregory method” for calculating equilibrium climate sensitivity

The equilibrium climate sensitivity (ECS) – the equilibrium global mean temperature response to a doubling of atmospheric CO 2 – is a high-profile metric for quantifying the Earth system's response to human-induced climate change. A widely applied approach to estimating the ECS is the “Gregory method” (Gregory et al., 2004), which uses an ordinary least squares (OLS) regression between the net radiative flux, N, and surface air temperature anomalies, ΔT, from a 150 year experiment in which atmospheric CO 2 concentrations are quadrupled. The ECS is determined by extrapolating the linear fit to N=0, i.e. the ΔT-intercept, indicating the point at which the system is back in equilibrium. This method has been used to compare ECS estimates across the CMIP5 and CMIP6 ensembles and will likely be a key diagnostic for CMIP7. Despite its widespread application, there is little consistency or transparency between studies in how the climate model data is processed prior to the regression, leading to potential discrepancies in ECS estimates. We identify 32 alternative data processing pathways, varying by differences in global mean weighting, net radiative flux variable, anomaly calculation method, and linear regression fit. Using 44 CMIP6 models, we systematically assess the impact of these choices on ECS estimates and calculate uncertainty ranges using two bootstrap approaches. While the inter-model ECS range is insensitive to the data processing pathway, individual outlier models exhibit notable differences. Approximating a model's native grid cell area (if irregular) with cosine of the latitude can decrease the ECS by 11 %, the choice of N-variable can change the ECS by 6 %, and some anomaly calculation methods can introduce spurious temporal correlations in the processed data. Beyond data processing choices, we also evaluate an alternative linear regression method – total least squares (TLS) – which has a more statistically robust basis than OLS. However, for consistency with previous literature, and given TLS may reduce the ECS compared to OLS (by up to 24 %), thereby making a known bias in the Gregory method worse, we do not feel there is sufficient clarity to recommend a transition to TLS in all cases. To improve reproducibility and comparability in future studies, we recommend a standardised Gregory method: weighting the global mean by cell area, using the top of the atmosphere (as opposed to the top of model) N-variable, and calculating anomalies by first applying a rolling average to the preindustrial control timeseries then subtracting from the raw CO 2 quadrupling experiment. This approach accounts for model drift while reducing noise in the data to best meet the pre-conditions of the linear regression. While CMIP6 results of the multi-model mean ECS appear insensitive to these processing choices, similar assumptions may not hold for CMIP7, underscoring the need for standardised data preparation in future climate sensitivity assessments.

Geosciences↗

Enhanced accuracy through ensembling of randomly initialized auto-regressive models for dynamical systems

Computational mechanics simulations using traditional finite element methods (FEM) require prohibitively expensive computational resources for real-time engineering applications, design optimization, and digital twin implementations. While machine learning (ML) surrogate models offer significant computational speedups, autoregressive ML models for time-dependent mechanical systems suffer from error accumulation that compromises long-term prediction reliability - a critical concern for engineering applications where accuracy over extended time horizons is essential for safety and performance assessments. Here, we propose a deep ensemble framework specifically designed to address this challenge in computational mechanics applications, where multiple ML surrogate models with random weight initializations are trained in parallel and their predictions aggregated during inference. This approach leverages statistical diversity to maximize information gain from a fixed set of training data and to mitigate error propagation, while maintaining the computational efficiency that makes ML surrogates attractive for engineering practice. We validate the framework on three representative problems spanning critical areas of computational mechanics: stress field evolution in heterogeneous microstructures under complex loading (relevant to advanced materials design and composite analysis), planetary-scale shallow water dynamics (applicable to environmental and geotechnical engineering), and Gray-Scott reaction-diffusion systems (relevant to mass transport and chemical process engineering). Across all test cases, the ensemble approach demonstrates consistent error reduction of 15-33% compared to individual models. The codes for this work are available on GitHub (https://github.com/Graham-Brady-Research-Group/AutoregressiveEnsemble_SpatioTemporal_Evolution).

autoregressive prediction↗

Reweighting Monte Carlo predictions and automated fragmentation variations in Pythia 8

This work reports on a method for uncertainty estimation in simulated collider-event predictions. The method is based on a Monte Carlo-veto algorithm, and extends previous work on uncertainty estimates in parton showers by including uncertainty estimates for the Lund string-fragmentation model. This method is advantageous from the perspective of simulation costs: a single ensemble of generated events can be reinterpreted as though it was obtained using a different set of input parameters, where each event now is accompanied with a corresponding weight. This allows for a robust exploration of the uncertainties arising from the choice of input model parameters, without the need to rerun full simulation pipelines for each input parameter choice. Such explorations are important when determining the sensitivities of precision physics measurements. Accompanying code is available at https://gitlab.com/uchep/mlhad-weights-validation.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Cloud Fusion of Big Data and Multi-Physics Models using Machine Learning for Discovery, Exploration, and Development of Hidden Geothermal Resources

The primary goals of this project are identifying hidden geothermal resources in the USA and designing profitable enhanced geothermal systems (EGS). Many non-obvious processes and parameters could characterize geothermal resources and could control the ultimate energy potential of geothermal fields. Diverse datasets (e.g., geology, geochemistry, geophysics, satellite, airborne geophysics) are available to help characterize geothermal resources, but this data is sparse and multi-scale. This has hindered attempts to leverage the datasets for geothermal exploration and profitable EGS design. Recent advancements in machine learning (ML) give promise to overcome these issues. Modern ML methods and tools can (1) analyze large datasets, (2) assimilate model ensembles that include a multitude of inputs and outputs, (3) process sparse datasets, (4) perform transfer learning between sites with different data quality, (5) extract hidden geothermal signatures from field and simulation data, (6) label geothermal resources and processes, (7) identify high-value data acquisition targets, and (8) guide geothermal exploration and production by selecting optimal exploration, production, and drilling strategies. In this work, we implement ML-based geothermal exploration and an enhanced geothermal systems (EGS) design tool to achieve the above goals. Our exploration tool is GeoThermalCloud (GTC) EGS design tool is GeoDT-ML. GTC (github.com/SmartTensors/GeoThermalCloud.jl) utilizes a LANL unsupervised ML platform called SmartTensors (https://tensors.lanl.gov/) to automate data analyses and interpretations by extracting hidden signatures to identify geothermal prospects. It enables the identification of critical measurements needed to identify geothermal resource signatures. GeoDT-ML (github.com/SmartTensors/GeoThermalCloud.jl/tree/master/) adds coupling to GeoDT (https://github.com/GeoDesignTool/GeoDT.git) for stochastic EGS design optimization and performance prediction. GeoDT-ML leverages recent advances in deep learning and high-performance computing. Contributors to this effort include LANL, PNNL, Google, Stanford, and Julia Computing.

15 GEOTHERMAL ENERGY↗

Automatic Determination of the Conic Coronal Mass Ejection Model Parameters

Characterization of the three-dimensional structure of solar transients using incomplete plane of sky data is a difficult problem whose solutions have potential for societal benefit in terms of space weather applications. In this paper transients are characterized in three dimensions by means of conic coronal mass ejection (CME) approximation. A novel method for the automatic determination of cone model parameters from observed halo CMEs is introduced. The method uses both standard image processing techniques to extract the CME mass from white-light coronagraph images and a novel inversion routine providing the final cone parameters. A bootstrap technique is used to provide model parameter distributions. When combined with heliospheric modeling, the cone model parameter distributions will provide direct means for ensemble predictions of transient propagation in the heliosphere. An initial validation of the automatic method is carried by comparison to manually determined cone model parameters. It is shown using 14 halo CME events that there is reasonable agreement, especially between the heliocentric locations of the cones derived with the two methods. It is argued that both the heliocentric locations and the opening half-angles of the automatically determined cones may be more realistic than those obtained from the manual analysis

Pulkkinen, A.↗

Forecasting generative amplification

Generative networks are perfect tools to enhance the speed and precision of LHC simulations. Especially when generating events beyond the size of the training dataset, it is important to understand their statistical precision. We present two complementary methods to estimate the amplification factor without large holdout datasets. Averaging amplification uses Bayesian networks or ensembling to estimate amplification from the precision of integrals over given phase-space volumes. Differential amplification uses hypothesis testing to quantify amplification without any resolution loss. Applied to state-of-the-art event generators, both methods indicate that amplification is already possible in specific regions of phase space.

Bahl, Henning [Heidelberg Univ. (Germany)] (ORCID:↗

Machine Learning the COSMO Model for Predicting Thermodynamics of Electrolyte Mixtures

Bottom-up design of electrolyte mixtures for battery systems requires predicting macro thermodynamic properties from molecular constituents. For instance, molten salt electrolyte batteries require conditions far above room temperature to operate. Therefore, discovering mixtures with increasingly lower eutectic melting points is desirable. A model that can approximate chemical activity is a valuable tool to search through the vast compositional design space. Machine learning can predict properties of materials such as vibrational free energies, electronic energy gaps, and thermal conductivities. Moreover, they can learn physical models such as interatomic potentials. The COSMO-SAC model uses theory and empirical parameterization to predict liquid-vapor and liquid-solid properties using first-principles calculations. However, obtaining activity coefficients required for parameterizing the COSMO-SAC model is costly and limited to a select chemical space. In this work, we explored if machine learning methods could improve the COSMO-SAC model and bridge density functional theory calculations to liquid phase thermodynamic properties. Our data-driven approach uses existing databases for sigma-profiles of organic solvents and reconciles their methodological differences via ensemble averaging. First, an optimal machine learning model is constructed for each dataset. Our machine learning algorithms use the sigma-profile as an input feature to predict binary mixtures' activity coefficients using multi-output regression. Each dataset uses different choices of functionals, methods, and basis sets. Therefore, our ensemble model attempts to predict corrected activity coefficients given the combination of all the model outputs. The activity coefficients used for training are generated using the COSMO-SAC model. This approach enables the extraction of meaningful information from the existing datasets to improve the COSMO-SAC model for obtaining thermodynamic properties of electrolyte mixtures. With the liquid phase activities, we can identify electrolyte mixtures that meet desired phase equilibria conditions.

Thermodynamics↗