Random forest regression feature importance for climate impact pathway detection
Not Available
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Not Available
Explore the source record for details and available documents.
All-solid-state structural lithium-ion batteries are sought to enable all-electric propulsion in next generation aerospace concepts through improved safety and systems level weight savings. In this work, the influence of processing conditions on microstructural evolution was evaluated for anode composites of strain-free Li4Ti5O12 and metallic nickel current collector. Beyond size distributions, this study explored methods of quantifying microstructural features that describe changes in the spatial distribution and coalescence of nickel particles as a function of sample composition and sintering conditions. Processing-microstructure-property relationships were described by microstructure quantifiers including nickel particle count per area, nearest neighbor distance distribution, and edge-to-edge distance distribution. Machine learning methods were applied to compare the relative influence of processing conditions and microstructural features on electrical conductivity and mechanical strength to optimize for simultaneous energy storage and load bearing performance. Insights gained from this work inform future evaluation of alternative energy storage materials and microstructures for multifunctional performance, and generation of microstructural descriptors strengthens modeling across length scales.
Hub-height turbulence intensity is essential for a variety of wind energy applications. However, simulating it is a challenging task. Simple analytical models have been proposed in the literature, but they all come with significant limitations. Even state-of-the-art numerical weather prediction models, such as the Weather Research and Forecasting model, currently struggle to predict hub-height turbulence intensity. Here, we propose a machine-learning-based approach to predict hub-height turbulence intensity from other hub-height and ground-level atmospheric measurements, using observations from the Perdigao field campaign and the Southern Great Plains atmospheric observatory. We consider a random forest regression model, which we validate first at the site used for training and then under a more robust round-robin approach, and compare its performance to a multivariate linear regression. The random forest successfully outperforms the linear regression in modeling hub-height turbulence intensity, with a normalized root-mean-square error as low as 0.014 when using 30-minute average data. In order to achieve such low root-mean-square error values, the knowledge of hub-height turbulence kinetic energy (which can instead be modeled in the Weather Research and Forecasting model) is needed. Interestingly, we find that the performance of the random forest generalizes well when considering a round-robin validation (i.e., when the algorithm is trained at one site such as Perdigao or Southern Great Plains) and then applied to model hub-height turbulence intensity at the other location.
These 20-meter spatial resolution gridded products provide per-pixel fractional cover (%) of Next Generation Ecosystem Experiments (NGEE) Arctic Plant Functional Types (PFTs) Tier 3 across Alaska, north of the boreal treeline. The products were developed for the NGEE Arctic project, which is improving Arctic vegetation representation and parameterization of the E3SM Land Model. This dataset includes 8 files containing fractional cover for NGEE Tier 3 PFTs (https://data.ess-dive.lbl.gov/view/doi:10.15485/2529470): (1) bryophytes; (2) lichens; (3) non-vascular plants, i.e., the sum of lichens and bryophytes; (4) deciduous shrubs, (5) evergreen shrubs, (6) forbs, (7) graminoids, and a non-PFT class, (8) litter. Each pixel contains the percent cover (expressed as a fraction of total ground cover) that was predicted by random-forest regression models. The random-forest models were trained on cover data collected at 978 plots from 2010 to 2021, of which are archived in the Pan-Arctic Vegetation Cover (PAVC) database (https://data.ess-dive.lbl.gov/datasets/doi:10.15485/2483557). The plot cover was linked to 20-meter spatial resolution, satellite-derived predictor variables: Sentinel-2 spectra and Sentinel-1 polarizations averaged over the 2019 growing season, as well as topographical features derived from ArcticDEM. Then, spatio-temporally anomalous plot data that introduced large variability to the regression outcomes were dropped using the Cook’s distance outlier detection method, and the models were re-created using high-quality plots and their associated satellite derived explanatory variables per each PFT. The correlations between plot-observed and satellite-derived fractional cover for all PFTs were well correlated (R2 = 0.69–0.95 and 0.5 for litter) and had low RMSE bias (0.02–0.11). This research was performed as a part of the NGEE Arctic project. The NGEE Arctic project was a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research.
This article presents the OpenMP parallelization of the preconditioning Jacobian assembly and right‐hand side residual evaluation in UEDGE. A continuation algorithm, utilizing the internal NKSOL implicit Jacobian‐Free Newton‐Krylov solver to efficiently scan physical parameters, is also presented. The implemented parallelization reduces the computational time for a benchmark scan run on 32 threads by compared to the serial version when using trained random forest regression models to identify the optimal decomposition of the system of equations. Random forest regression models applied to the UEDGE time‐dependent and continuation solver algorithms did not yield meaningful improvement in computational performance. A benchmark DIII‐D gas injection rate scan in the 0.35–0.75 kA interval, performed on a test cluster using the parallelized code and continuation solver, produced 1066 steady‐state solutions with a 22 s average wall‐clock computational time per steady‐state solution.
Agent-based modeling (ABM) is a powerful simulation technique which describes a complex dynamic system based on its interacting constituent entities. While the flexibility of ABM enables broad application, the complexity of real-world models demands intensive computing resources and computational time; however, a metamodel may be constructed to gain insight at less computational expense. Here, we developed a model in NetLogo to describe the growth of a microbial population consisting of Pantoea . We applied 13 parameters that defined the model and actively changed seven of the parameters to modulate the evolution of the population curve in response to these changes. We efficiently performed more than 3,000 simulations using a Python wrapper, NL4Py . Upon evaluation of the correlation between the active parameters and outputs by random forest regression, we found that the parameters which define the depth of medium and glucose concentration affect the population curves significantly. Subsequently, we constructed a metamodel, a dense neural network, to predict the simulation outputs from the active parameters and found that it achieves high prediction accuracy, reaching an R 2 coefficient of determination value up to 0.92. Our approach of using a combination of ABM with random forest regression and neural network reduces the number of required ABM simulations. The simplified and refined metamodels may provide insights into the complex dynamic system before their transition to more sophisticated models that run on high-performance computing systems. The ultimate goal is to build a bridge between simulation and experiment, allowing model validation by comparing the simulated data to experimental data in microbiology.
Enhanced characterization and phase identification of post-PUREX Pu Oxalates (PuOXA) are pivotal for nonproliferation and pre-detonation nuclear forensics. Despite significant advances in the characterization of PuO 2 samples, little is known about the impact of both the chemical structure and oxidation states of PuOXA (i.e., Pu(III) and Pu(IV)) have on optical emission signatures. Here, we demonstrate the analytical capabilities of laser-induced breakdown spectroscopy (LIBS) applied to Fe(II) and Fe(III) oxalate samples as surrogates for PuOXA, highlighting the discriminating features in the LIBS emission spectra arising from differences in the oxidation states within mixed FeOXA samples. We report the enhancement of spectral feature selection using Principal Component Analysis (PCA), which enables the analytical superiority of machine learning algorithms such as Linear Discriminant Analysis (LDA), Quadratic Discriminant Analysis (QDA), Partial Least Squares Regression (PLSR), Support Vector Regression (SVR), and Random Forest Regression (RFR) over conventional univariate techniques for phase discrimination and chemometric analysis. Cluster analysis revealed how both matrix effects and laser ablation influence cluster separability by introducing spectral artifacts that misdirect the maximization of variance. PCA-selected emission lines were used in the regression models, demonstrating that both univariate and multivariate linear regression models (i.e., PLSR and SVR) can achieve acceptable performance, with machine learning models outperforming conventional calibration regressions. Furthermore, the application of non-linearly activated PCA-selected emission lines illustrates how simplifying the data while retaining captured variance enables the use of less complex and more computationally efficient models. Furthermore, this is particularly evident in the underperformance of RFR, which suffers from increased computational costs and overfitting owing to its high complexity.
Low-cost sensors (LCSs) for air quality monitoring have enormous potential to improve air quality data coverage in resource-limited parts of the world such as sub-Saharan Africa. LCSs, however, are affected by environment and source conditions. To establish high-quality data, LCSs must be collocated and calibrated with reference grade PM2.5 monitors. From March 2020, a low-cost PurpleAir PM2.5 monitor was collocated with a Met One Beta Attenuation Monitor 1020 in Accra, Ghana. While previous studies have shown that multiple linear regression (MLR) and random forest regression (RF) can improve accuracy and correlation between PurpleAir and reference data, MLR and RF yielded suboptimal improvement in the Accra collocation (R2 = 0.81 and R2 = 0.81, respectively). We present the first application of Gaussian mixture regression (GMR) to air quality data calibration and demonstrate improvement over traditional methods by increasing the collocated PM2.5 correlation and accuracy to R2 = 0.88 and MAE = 2.2 μg/cu. m. Gaussian mixture models (GMMs) are a probability density estimator and clustering method from which nonlinear regressions that tolerate missing inputs can be derived. We find that even when given missing inputs, GMR provides better correlation than MLR and RF performed with complete data. GMR also allows us to estimate calibration certainty. When evaluated, 95% confidence intervals agreed with reference PM2.5 data 96% of the time, suggesting that the model accurately assesses its own confidence. Additionally, clustering within the GMM is consistent with climate characteristics, providing confidence that the calibration approach can learn underlying relationships in data.
Particulate matter air pollution is a leading cause of global mortality, particularly in Asia and Africa. Addressing the high and wide-ranging air pollution levels requires ambient monitoring, but many low- and middle-income countries (LMICs) remain scarcely monitored. To address these data gaps, recent studies have utilized low-cost sensors. These sensors have varied performance, and little literature exists about sensor intercomparison in Africa. By colocating 2 QuantAQ Modulair-PM, 2 PurpleAir PA-II SD, and 16 Clarity Node-S Generation II monitors with a reference-grade Teledyne monitor in Accra, Ghana, we present the first intercomparisons of different brands of low-cost sensors in Africa, demonstrating that each type of low-cost sensor PM2.5 is strongly correlated with reference PM2.5, but biased high for ambient mixture of sources found in Accra. When compared to a reference monitor, the QuantAQ Modulair-PM has the lowest mean absolute error at 3.04 μg/m3, followed by PurpleAir PA-II (4.54 μg/m3) and Clarity Node-S (13.68 μg/m3). We also compare the usage of 4 statistical or machine learning models (Multiple Linear Regression, Random Forest, Gaussian Mixture Regression, and XGBoost) to correct low-cost sensors data, and find that XGBoost performs the best in testing (R2: 0.97, 0.94, 0.96; mean absolute error: 0.56, 0.80, and 0.68 μg/m3 for PurpleAir PA-II, Clarity Node-S, and Modulair-PM, respectively), but tree-based models do not perform well when correcting data outside the range of the colocation training. Therefore, we used Gaussian Mixture Regression to correct data from the network of 17 Clarity Node-S monitors deployed around Accra, Ghana, from 2018 to 2021. We find that the network daily average PM2.5 concentration in Accra is 23.4 μg/m3, which is 1.6 times the World Health Organization Daily PM2.5 guideline of 15 μg/m3. While this level is lower than those seen in some larger African cities (such as Kinshasa, Democratic Republic of the Congo), mitigation strategies should be developed soon to prevent further impairment to air quality as Accra, and Ghana as a whole, rapidly grow.
Atmospheric chemistry is a high-dimensionality, large-data problem and thus may be suited to machine-learning algorithms. We show here the potential of a random forest regression algorithm to replace the gas-phase chemistry solver in the GEOS-Chem chemistry model. In this proof-of-concept study, we used one month of model output to train random forest regression models to predict the concentrations of each long-lived chemical species after integration based upon the physical and chemical conditions before the chemical integration. The choice of prediction type has a strong impact on the skill of the regression model. We find best results from predicting the change in concentration for very long-lived species and the absolute concentration for shorter lived species. The skill of the machine learning algorithm is further improved by using a family approach for NO and NO2 rather than treating them independently.By replacing the numerical integrator with the random forest algorithm and running this model for one month, we find that the model is able to reproduce many of the features of the reference chemistry simulation. Replacing the integration methodology with a machine learning algorithm has the potential to be substantially faster. There are a wide range of applications for such an approach, e.g. to generate boundary conditions, for use in air quality forecasts or chemical data assimilation systems, etc.
Fully defined physics-based building energy models can accurately represent building systems; however, generating models based on high-level parameters is time consuming and simulation time of complex models can be slow. This article discusses the development of a Metamodelling Framework to create metamodels from a building energy modelling dataset. The framework generates metamodels using either linear regression, random forests, or support vector regressions. A fifth-generation district heating and cooling system analysis use case was used to motivate the development of the framework. The use case required quick and accurate representations of annual building loads reported hourly. Typical annual building modelling approaches can result in a runtime of 10 min. The metamodels runtime was reduced to less than 10 s to load and run an annual simulation with user-defined covariates. The results of the metamodel performance and an abbreviated topology analysis based on the motivating use case will be presented.
Astronauts are exposed to a unique set of stressors in spaceflight. Microgravity, isolation, confinement, and environmental and operational hazards: all of these can impact sleep, vigilant attention, and alertness, which are critical to mission success. In this paper, we seek to understand the most important predictors of alertness over the course of a space mission, using self-reported, cognitive, and environmental data collected from 24 astronauts on 6-month missions to the International Space Station (ISS). Alertness was repeatedly and objectively assessed on the ISS with a brief 3-minute Psychomotor Vigilance Test (PVT) that is highly sensitive to sleep deprivation. To relate PVT performance to time-varying and sparsely-measured environmental, operational, and psychological covariates, we propose a n ensemble prediction model comprising of linear mixed effects regression, random forest, and functional concurrent regression models. An extensive cross-validation procedure reveals that this ensemble outperforms any one of its components alone. We also discover that a participant’s past performance, reported fatigue and stress, and temperature and radiation exposure were among the most important variables associated with alertness. This method is broadly applicable to environmental studies where the main goal is accurate, individualized prediction involving a mixture of person-level traits and irregularly measured time series.
The Mars Entry, Descent, and Landing Instrumentation (MEDLI2) sensor suite collected data during entry of the Mars 2020 Perseverance rover into Mars’ atmosphere. This suite included a network of MEDLI2 Instrumented Sensor Plugs (MISPs). Each MISP was comprised of a cylinder made of Thermal Protection System (TPS) material with 1-3 embedded thermocouples (TCs), and it was flush mounted into the heatshield or backshell. Data from these in-depth TCs were used to reconstruct the aeroheating environment of the vehicle throughout entry. Surface heating was posed as an inverse problem, with the goal of estimating the surface heating by minimizing an objective function of the difference between MISP temperature measurements during flight and the temperature predictions derived from the Fully Implicit Ablation and Thermal response (FIAT) program. Given an aerothermal environment, FIAT calculates the material response and provides in-depth temperatures throughout the TPS material. To achieve the reverse, an internal tool called FIAT_Opt runs through multiple different environments until the output temperature at the TC depth closely matches the flight data. 95% confidence intervals on the reconstructed surface heating were obtained using Monte Carlo analysis, in which uncertainties in the thermocouple depth and the TPS material properties (e.g., density, thermal conductivity, heat capacity, emissivity) based on flight-lot material testing were included. A variance decomposition method using Sobol indices was employed to assess the sensitivity of the reconstructed peak heating to the TC placement and material property uncertainties. Variance decomposition was found to require tens of thousands of FIAT_Opt runs in order for the Sobol indices to converge. With a single FIAT_Opt run taking on the order of 40 minutes, the required number of computations would take months to complete, even if using multiple CPUs. To mitigate this problem, three machine learning models (ridge regression with cross-validation, random forest regression, and a deep neural network) were trained and tested using the 2000 Monte Carlo runs that were already completed. A subset of 1600 runs were used to train the model (i.e., training set), while the remaining 400 runs were used as the test set. The predictions from the deep neural network (DNN) on the test set showed nearly perfect agreement to the actual values computed with FIAT_Opt (R2 > 0.99). Using the DNN as a surrogate model, the variance decomposition using 50,000 runs was completed within minutes. The resulting Sobol indices showed that the reconstructed peak surface heating was most sensitive to the uncertainties in the thermal conductivity (ST = 0.37) and heat capacity (ST = 0.26). This method can be leveraged to provide requirements for material property measurements needed to improve the accuracy of surface heating prediction and ultimately lead to the reduction of design margins in the future. This presentation will include background on the MEDLI2 suite; the method used for inverse heating estimation; the way that material property uncertainties were accounted for using Monte Carlo analysis; a brief background on variance decomposition; the motivation for using machine learning in this context; how a neural network was trained on the data to enable variance decomposition in a fraction of the time; and the variance decomposition results for one of the MISPs.
We use artificial intelligence (AI) to learn and infer the physics of higher order gravitational wave modes of quasi-circular, spinning, non precessing binary black hole mergers. We trained AI models using 14 million waveforms, produced with the surrogate model NRHybSur3dq8, that include modes up to $\ell$ ≤ 4 and (5,5), except for (4,0) and (4,1), that describe binaries with mass-ratios $\textit{q}$ ≤ 8, individual spins $s^z_{\{1,2\}} \in$[–0.8,0.8], and inclination angle $θ \in$ [0,π]. Our probabilistic AI surrogates can accurately constrain the mass-ratio, individual spins, effective spin, and inclination angle of numerical relativity waveforms that describe such signal manifold. We compared the predictions of our AI models with Gaussian process regression, random forest, k-nearest neighbors, and linear regression, and with traditional Bayesian inference methods through the PyCBC Inference toolkit, finding that AI outperforms all these approaches in terms of accuracy, and are between three to four orders of magnitude faster than traditional Bayesian inference methods. Our AI surrogates were trained within 3.4 hours using distributed training on 1,536 NVIDIA V100 GPUs in the Summit supercomputer.
Void growth plays a central role in ductile fracture, yet the specific mechanisms that control this remain obscure. Classical models, such as those proposed by Rice and Tracey in 1969, are able to capture average rates of void growth, but cannot capture the heterogeneity of individual void growth. Building on recent work, the present study employs laboratory-based diffraction contrast tomography and in-situ x-ray computed tomography to investigate the effect of grain structure and other microstructural factors on void growth in an Al-2219 alloy. Crystal plasticity finite element (CP-FE) modeling is used alongside experimental data to evaluate the contributions of local mechanical states, grain orientation, grain size, and neighboring microstructural features. No strong linear relationships are found with any of the considered descriptors and void growth rate. Potential complex nonlinear relationships are explored with the use of a random forest regression model, which identifies initial void volume, void aspect ratio, local normal stress state, local shear stress state, and local equivalent plastic strain (EQPS) as features that most improve void growth rate predictions. The combination of these analyses suggests that these features should be prioritized to improve models of void growth.
We present a novel data-driven approach for prediction of the estimated time of arrival (ETA) of aircraft in the terminal area via the implementation of a Random Forest regression model. The model uses data fused from a number of sources (flight track, weather, flight plan information, etc.) and provides predictions for the remaining flight time for aircraft landing at Dallas/Fort Worth (DFW) International Airport. The predictions are made when the aircraft is at a distance of 200-miles from the airport. The results show that the model is able to predict estimated time of arrival to within ± 5 min for 90% of the flights in the test data with the mean absolute error being lower at 145 seconds. This paper covers the entire pipeline of data collection, preprocessing, setup and training of the ML model, and the results obtained for DFW.
The radiative forcing of anthropogenic aerosols associated with aerosol–cloud interactions (RF(sub aci)) remains the largest source of uncertainty in climate prediction. The calculation of particle number concentration (PNC), one of the critical parameters affecting RF(sub aci), is generally simplified in climate models. Here we employ outputs from long-term (30-years) simulations of a global size-resolved (sectional) aerosol microphysics model and a machine-learning tool to develop a Random Forest Regression Model (RFRM) for PNC. We have implemented the PNC RFRM in GISS-ModelE2.1 with a mass-based One-Moment Aerosol module, which is one of CMIP6 models. Compared to the default setting, the GISS-ModelE2.1 simulation based on RFRM reduces the changes of cloud droplet number concentration associated with anthropogenic emissions, and decreases the RF(sub aci) from −1.46 W⋅m(exp −2) to −1.11 W⋅m(exp −2). This work highlights a promising approach based on machine learning to reduce uncertainties of climate models in predicting PNC and RF(sub aci) without compromising their computing efficiency.