Optical Turbulence Profile Modeling in the Atmospheric Boundary Layer: A Random Forest Regression Approach
Explore the source record for details and available documents.
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Constraining cloud feedback in global climate models (GCMs) using observations is important for establishing accurate predictions of future climate. Uncertainty in shortwave cloud feedback (SW FB ) dominates uncertainty in total cloud feedback. Recent studies show a shift toward more positive extratropical SW FB in the latest generations of GCMs leading to the emergence of very high equilibrium climate sensitivity (ECS). In this study, we use precipitation efficiency and albedo susceptibility to constrain liquid water path (LWP) response to warming and SW FB in the Southern Ocean (SO; 50°–80°S). We analyze precipitation in extratropical cyclones (ECs) to learn about extratropical condensed water sink processes, combined with observations of clouds and moisture convergence, and use the analysis to better understand and constrain SW FB . We utilize a perturbed parameter ensemble (PPE) hosted in the Community Atmosphere Model, version 6 (CAM6), to provide a constraint on SW FB based on observations from Clouds and the Earth’s Radiant Energy System (CERES) and Multisensor Advanced Climatology of LWP (MAC-LWP). We apply Gaussian process regression to emulate the model response to all parameters perturbed in the PPE. Confronting the emulator output with observations provides a new estimated response of Earth to global warming. Furthermore, our new estimates of SO LWP reduce the PPE range by 66%–72%, which results in a shortwave cloud radiative effect estimated range that is 27%–34% less than the PPE range. Observations suggest a more positive SO SW FB than the Community Earth System Model, version 2 (CESM2), and consequently do not reject the high climate sensitivity GCMs emerging from the Coupled Model Intercomparison Project phase 6 (CMIP6).
Abstract A real-time capable core Ion Cyclotron Range of Frequencies (ICRF) heating model on NSTX and WEST is developed. The model is based on two nonlinear regression algorithms, the random forest ensemble of decision trees and the multilayer perceptron neural network. The algorithms are trained on TORIC ICRF spectrum solver simulations of the expected flat-top operation scenarios in NSTX and WEST assuming Maxwellian plasmas. The surrogate models are shown to successfully capture the multi-species core ICRF power absorption predicted by the original model for the high harmonic fast wave and the ion cyclotron minority heating schemes while reducing the computational time by six orders of magnitude. Although these models can be expanded, the achieved regression scoring, computational efficiency and increased model robustness suggest these strategies can be implemented into integrated modeling frameworks for real-time control applications.
Although machine learning (ML) has emerged as a powerful tool for rapidly assessing grid contingencies, prior studies have largely considered a static grid topology in their analyses. This limits their application, since they need to be re-trained for every new topology. Here, this paper explores the development of generalizable graph convolutional network (GCN) models by pre-training them across a range of grid topologies and contingency types. We found that a GCN model with auto-regressive moving average (ARMA) layers with a line graph representation of the grid offered the best predictive performance in predicting voltage magnitudes (VM) and voltage angles (VA). We introduced the concept of phantom nodes to consider disparate grid topologies with a varying number of nodes and lines. For pre-training the GCN ARMA model across a variety of topologies, distributed graphics processing unit (GPU) computing afforded us significant training scalability. The predictive performance of this model on grid topologies that were part of the training data is substantially better than the direct current (DC) approximation. Although direct application of the pre-trained model to topologies that are not part of the grid is not particularly satisfactory, fine-tuning with small amounts of data from a specific topology of interest significantly improves predictive performance. In general, this paper highlights the feasibility of training large-scale GNN models to assess the reliability of power grids by considering a wide variety of grid topologies and contingency types. With the advent of foundational models in ML and the exponential increase in GPU computing clusters, generalizable ML models will significantly enhance how utilities manage power systems and make decisions in real-time or near-real-time.
A machine learning-enabled multiscale framework is developed for modeling the mechanical response of both pure metal and nanoparticle-reinforced metal matrix nanocomposites (MMNCs). Using aluminum–silicon carbide (Al-SiC) as an example MMNC, atomistic simulations reveal three distinct deformation mechanisms (i.e., defect-free, dislocation-based, and interface separation) governed by the interfaces between the Al matrix and SiC nanoparticles. As compared with single crystal Al, the lattice undergoes a more abrupt failure once the dislocation network becomes extensive and void nucleation initiates, whereas in Al-SiC, nanoparticle interfaces enable a more gradual progression of damage. These mechanisms are captured through a combined classification-regression neural network surrogate model that bridges atomic-scale insights with continuum-scale finite element analysis. Machine learning-enabled multiscale modeling of pure Al accurately predicted strain localization and confirmed by in-situ scanning electron microscopic tensile testing on perforated Al specimens. This study underscores the promise of integrating physics-informed machine learning with hierarchical modeling to capture the interface dominated phenomena and guide the design of advanced MMNCs.
This research program established a transformative framework for the discovery and design of mechanical metamaterials, which are architected structures engineered to control physical phenomena like sound and vibration in ways natural materials cannot. To overcome the traditional reliance on trial-and-error, the project developed an interpretable Artificial Intelligence (AI) framework that moves beyond "black box" models to reveal the specific geometric patterns—such as "unit-cell templates"—that govern a material’s performance. A major breakthrough was the development of a hierarchical design method, which allows a single material to block vibrations across multiple frequency ranges simultaneously by layering patterns at different scales without them interfering with one another. This was further expanded to include irregular, graph-based designs that use spanning tree algorithms to ensure structural connectivity while allowing for customized, direction-dependent properties like stiffness and acoustic impedance. Beyond design, the project addressed the practicalities of real-world production by developing uncertainty quantification techniques that account for manufacturing defects and material variability, reducing the need for expensive physical testing by orders of magnitude. To speed up the discovery process, the team implemented Gaussian Process Regression and other surrogate models that provide accurate performance predictions at a fraction of the traditional computational cost. The AI-generated designs were successfully validated through fabrication of physical samples and wave propagation experiments, confirming their ability to accurately guide or reflect waves as predicted. By contributing these tools and high-quality FAIR benchmark datasets to the wider scientific community, this work provides a scalable foundation for advancing technologies in aerospace vibration control, medical imaging, and noise reduction.
Chiral 2D metal halide perovskites (MHPs) are promising for spin-optoelectronic applications, yet their absorption dissymmetry factor (g abs ) exhibits significant variability due to complex, co-dependent structural and experimental factors. Here, we established a data-driven framework using Pearson’s correlation, ANOVA, and Gaussian process regression to identify and model key synthesis “knobs” governing these properties. The analysis revealed that solvent choice is the primary factor driving variability. For acetonitrile-based films, g abs was maximized by optimizing annealing temperature and film thickness. Conversely, films from higher boiling point solvents showed complex dependencies on annealing temperature, excitonic integral intensity, and film texture. These statistical correlations provide a roadmap for the rational design of high-performance chiral MHPs and establish a foundation for future machine learning-driven material exploration.
A model-independent search for low-mass resonances decaying into pairs of oppositely charged muons is presented. The analysis uses proton-proton collision data corresponding to an integrated luminosity of 140 fb−1, recorded by the ATLAS detector at the Large Hadron Collider between 2015 and 2018. The search targets hypothetical dimuon resonances in the invariant mass range from 35 GeV to 75 GeV. The modelling of this mass region is particularly challenging for conventional analytic background parameterisations. To address this, a Gaussian process regression technique is used to model the background. The dimuon mass spectrum is analysed for potential signals, and no statistically significant excess is observed. Upper limits at the 95% confidence level are set on the fiducial production cross-section of new resonances decaying promptly into muons, ranging from 20 fb to 110 fb, depending on the resonance mass. These results are further interpreted in the context of dark-photon and dark-matter-mediator models, leading to new constraints on their parameter spaces.
Not Available
This study applies machine learning methods to analyze natural gas pipeline incidents in the United States using the Pipeline and Hazardous Materials Safety Administration (PHMSA) Gas Distribution Incident Dataset (2010–2024). The dataset includes over 600 variables describing incident characteristics, infrastructure attributes, and contributing factors associated with unintentional gas releases. The objective is to assess whether these features can reliably predict the underlying cause of pipeline failures. Multinomial logistic regression and Random Forest models were developed to classify incident causes, including excavation damage, corrosion, equipment failure, and natural forces. Results show that excavation damage is both the most frequent and most predictable cause, with models achieving strong performance for this category. However, when excavation damage is excluded, model accuracy declines significantly, with some models performing near random levels. Across all approaches, severe class imbalance and limited variability in key predictors constrain predictive performance. Pipeline age and diameter emerge as the most influential variables, but they provide insufficient discriminatory power to distinguish among less frequent failure types. These findings indicate that non-excavation-related incidents are rare, heterogeneous, and weakly represented in the dataset, limiting the effectiveness of machine learning classification. Overall, this study highlights the structural limitations of the PHMSA dataset for predictive modeling and underscores the need for improved data balance and feature enrichment. The results reinforce excavation damage prevention as the most impactful strategy for reducing pipeline incidents.
Introduction: Crops are vulnerable to precipitation and heat extremes during late spring through summer. Methods: We analyzed for a north-central U.S. region short-term drought and agricultural heat stress during April-May-June-July. We used the 4-km Parameter Elevation Regression on Independent Slopes Model (PRISM) for observations, aggregated to a 25-km grid, and two 25-km Regional Climate Model version 4 (RegCM4) simulns used either GFDL- or MPI-GCM boundary conditions. We chose 1981-2000 as our contemporary time period, and 2041- 2060 as our scenario time period, which used the Representative Concentration Pathway 8.5 emissions scenario. We used object-oriented analysis to identify events of interest in observations and simulations by identifying objects in a space-time domain that meet specified criteria, such as exceeding a heat-stress temperature threshold. The event diagnosis allowed analysis of compound events, occurring when temperature and drought objects overlap. Results: Identified objects yielded events that can undermine agricultural productivity and which are thus relevant to decision makers, making them building blocks for possible climate storylines. The observations and simulations showed similar spatial distributions of event frequencies across the analysis region. However, the simulations attained this distribution by having fewer events that tend to cover larger areas compared to observed events, suggesting that the effective resolution of the simulations was coarser than their 25-km grids. Short-term drought frequency increased and heat-stress frequency decreased in transitioning to the scenario climate. When compounding occurred heat-stress events generally preceded the short-term drought events. The overlapping, compound events tended to be more extreme compared to non-overlapping events of either type. Discussion: The information yielded projected changes in these agriculturally motivated events. One prominent conditional behavior emerging from the work was that a heat-stress event should be a warning to watch for potential drought, as both could compound each other to more intense levels.
We present a new Bayesian model for the problem of multiclass classification. In this model, the probabilities of class membership of a given observation are determined by the mean of a latent Gaussian distribution. The mean functions of this latent distribution consist of combinations of highly flexible basis functions of the inputs: multivariate adaptive regression splines (MARS), first developed for multiple regression. We use reversible jump Markov chain Monte Carlo to make inference on the classification model, including the number of basis functions. We compare the probabilistic classification performance of our proposed approach to existing methods on simulated and benchmark data, and compare uncertainty estimates on simulated data. Our proposed method compares favorably with existing Bayesian and frequentist multiclass classification methods in out-of-sample probabilistic classification, and uncertainty estimation of these probabilistic classifications. We examine the fit of the proposed method to a data set of hurricane storm surge levels near Delaware Bay, US, and conclude that sea level rise is a key contributor to damage delivered by storm surge.
The strength of materials is influenced by a range of external conditions, such as temperature and deformation rate. Consequently, materials that demonstrate substantial variations in their mechanical behavior due to fluctuations in temperature and strain rate require complex strength models to accurately predict material performance in real-world applications. To predict such complex behavior, a robust and flexible strength model is necessary. In this work, we utilize genetic programming-based symbolic regression (GPSR) to develop data-driven strength models that accurately represent the measured stress–strain responses of tin across a wide range of strain, strain rate and temperature regimes. The GPSR models are constrained by physically-informed conditions, which leads to significant improvement in extrapolation. The best model is integrated into a multi-physics code to perform Taylor impact simulations, validating the model’s accuracy and robustness. In conclusion, the model predictions showed excellent agreement with experimental results, particularly when compared to predictions using traditional strength models.
Here, the goal of the present work was to provide the necessary reaction emulation information to enable detailed process simulation of a chemical looping H 2 production system from fossil fuels using CaFe 2 O 4 . This specifically pertained to the necessary kinetic data, reaction model development, and model rate parameters required for reaction emulation in both reducing and oxidizing environments. A logical methodology was defined, which included discretization of the reaction network, establishing a core model for reaction emulation that could be adapted based on the system phenomena, and development of a rate parameter regression tool designed around the core model. An extensive array of data sets was acquired by which parametric regressions were performed. The work presented and tabulated a comprehensive set of rate parameters for the reduction and oxidation reactions of CaFe 2 O 4 and descendent phases of Ca 2 Fe 2 O 5 , FeO, Fe 3 O 4 , Fe, and CaO to emulate reaction behavior in a looping-based process environment. This included direct reduction using CH 4 , H 2 , and CO, and direct oxidation reactions with steam, CO 2 and O 2 . Dynamic equilibrium was quantified for reactions that could utilize H 2 O and CO 2 as soft oxidants to re-saturate lattice oxygen in the depleted structure/phases. The kinetics associated with the oxidative mechanisms with the soft oxidants were quantified and compared to those of the reducing counterparts. The analysis provided critical insight to emulate reactions for a process that seeks to use natural gas (NG) or other fossil fuels as a direct reductant for the end goal of H 2 production.
We propose a new set of nuclear mass predictions based on multiple theoretical mass models. By employing Gaussian process regression with the Matérn kernel, we achieved root-mean-square (rms) deviations below 100 keV for the training dataset. The best-performing mass models achieved rms deviations below 150 keV for the new precise mass data from AME2020, whereas the ensemble average showed robust performance across the nuclear chart. Our approach uniquely combines: (1) systematic refinement of eight mass models through their residuals, (2) physics-informed features, including magic numbers, nucleon parity numbers, neutron excess, and nuclear collectivity, and (3) theory-to-theory validation demonstrating robust extrapolation capability. We find that the Matérn kernel provides superior uncertainty quantification compared to the RBF kernel, with a length-scale analysis revealing enhanced inter-nuclei correlations. We provide complete mass predictions for all unknown nuclides in AME2020, offering valuable constraints for nuclear structure studies and astrophysical modeling when used with proper uncertainty propagation.
Foundation models for astronomical surveys offer powerful learned representations that can be transferred to downstream regression tasks such as galaxy property estimation. However, point predictions alone are insufficient for scientific inference; reliable uncertainty quantification (UQ) is essential. We compare seven UQ methods on galaxy property regression using frozen AION-1 foundation-model embeddings, predicting redshift, stellar mass, stellar-population age, gas-phase metallicity, and specific star-formation rate, from Legacy Survey photometry/imaging and DESI spectra, with PROVABGS-derived labels. Distribution-free conformal methods achieve marginal coverage within $\sim$1 pp of the nominal 90% across all properties, while non-conformal baselines (Deep Ensembles, MC~Dropout) fail to calibrate reliably. Among conformal approaches, Conformalized Quantile Regression (CQR) delivers the best coverage in the bin with the poorest model predictions. More importantly, only the Locally Valid and Discriminative (LVD) framework -- particularly when operating on AION-1 embeddings -- also provides finite-sample \emph{local validity}, producing intervals that adapt to each galaxy's local prediction difficulty rather than relying on marginal guarantees alone. These results establish conformal prediction, and LVD in particular, as the preferred UQ framework for uncertainty-aware inference on foundation-model embeddings in astrophysics.
We explore the application of machine learning algorithms specifically to enhance the selection process of Russet potato (Solanum tuberosum L.) clones in breeding trials by predicting their suitability for advancement. This study addresses the challenge of efficiently identifying high-yield, disease-resistant, and climate-resilient potato varieties that meet processing industry standards. Leveraging manually collected data from trials in the state of Oregon, we investigate the potential of a wide variety of state-of-the-art binary classification models. The dataset includes 1086 clones, with data on 38 attributes recorded for each clone, focusing on yield, size, appearance, and frying characteristics, with several control varieties planted consistently across four Oregon regions from 2013 to 2021. We conduct a comprehensive analysis of the dataset that includes preprocessing, feature engineering, and imputation to address missing values. We focus on several key metrics such as accuracy, F1-score, and Matthews correlation coefficient (MCC) for model evaluation. The top-performing models, namely a feedforward neural network classifier (Neural Net), a histogram-based gradient boosting classifier (HGBC), and a support vector machine classifier (SVM), demonstrate consistent and significant results. To further validate our findings, we conducted a simulation study using the aims, data-generating mechanisms, estimands, methods, and performance measures (ADEMP) framework, simulating different data-generating scenarios to assess model robustness and performance through true positive, true negative, false positive, and false negative distributions, area under the receiver operating characteristic curve (AUC-ROC) and MCC. The simulation results highlight that non-linear models like SVM and HGBC consistently show higher AUC-ROC and MCC than logistic regression, thus outperforming the traditional linear model across various distributions, and emphasizing the importance of model selection and tuning in agricultural trials. Variable selection further enhances model performance and identifies influential features in predicting trial outcomes. The findings emphasize the potential of machine learning in streamlining the selection process for potato varieties, offering benefits such as increased efficiency, substantial cost savings, and judicious resource utilization. Our study contributes insights into precision agriculture and showcases the relevance of advanced technologies for informed decision-making in breeding programs.
ABSTRACT In this study, we consider three different machine‐learning methods—a three‐hidden‐layer neural network, support vector regression, and Gaussian process regression—and compare how well they can learn from a synthetic data set for proton acceleration in the Target Normal Sheath Acceleration regime. The synthetic data set was generated from a previously published theoretical model by Fuchs et al. 2005 that we modified. Once trained, these machine‐learning methods can assist with efforts to maximize the peak proton energy, or with the more general problem of configuring the laser system to produce a proton energy spectrum with desired characteristics. In our study, we focus on both the accuracy of the machine‐learning methods and the performance on one GPU including memory consumption. Although it is arguably the least sophisticated machine‐learning model we considered, support vector regression performed very well in our tests.