Search NASA⌕ Search

SEARCH · Search NASA

Results for “random forests”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Machine learning approaches for influenza A virus risk assessment identifies predictive correlates using ferret model in vivo data

In vivo assessments of influenza A virus (IAV) pathogenicity and transmissibility in ferrets represent a crucial component of many pandemic risk assessment rubrics, but few systematic efforts to identify which data from in vivo experimentation are most useful for predicting pathogenesis and transmission outcomes have been conducted. To this aim, we aggregated viral and molecular data from 125 contemporary IAV (H1, H2, H3, H5, H7, and H9 subtypes) evaluated in ferrets under a consistent protocol. Three overarching predictive classification outcomes (lethality, morbidity, transmissibility) were constructed using machine learning (ML) techniques, employing datasets emphasizing virological and clinical parameters from inoculated ferrets, limited to viral sequence-based information, or combining both data types. Among 11 different ML algorithms tested and assessed, gradient boosting machines and random forest algorithms yielded the highest performance, with models for lethality and transmission consistently better performing than models predicting morbidity. Comparisons of feature selection among models was performed, and highest performing models were validated with results from external risk assessment studies. Our findings show that ML algorithms can be used to summarize complex in vivo experimental work into succinct summaries that inform and enhance risk assessment criteria for pandemic preparedness that take in vivo data into account.

59 BASIC BIOLOGICAL SCIENCES↗

A framework to evaluate machine learning crystal stability predictions

The rapid adoption of machine learning in various scientific domains calls for the development of best practices and community agreed-upon benchmarking tasks and metrics. We present Matbench Discovery as an example evaluation framework for machine learning energy models, here applied as pre-filters to first-principles computed data in a high-throughput search for stable inorganic crystals. We address the disconnect between (1) thermodynamic stability and formation energy and (2) retrospective and prospective benchmarking for materials discovery. Alongside this paper, we publish a Python package to aid with future model submissions and a growing online leaderboard with adaptive user-defined weighting of various performance metrics allowing researchers to prioritize the metrics they value most. To answer the question of which machine learning methodology performs best at materials discovery, our initial release includes random forests, graph neural networks, one-shot predictors, iterative Bayesian optimizers and universal interatomic potentials. We highlight a misalignment between commonly used regression metrics and more task-relevant classification metrics for materials discovery. Accurate regressors are susceptible to unexpectedly high false-positive rates if those accurate predictions lie close to the decision boundary at 0 eV per atom above the convex hull. The benchmark results demonstrate that universal interatomic potentials have advanced sufficiently to effectively and cheaply pre-screen thermodynamic stable hypothetical materials in future expansions of high-throughput materials databases.

Riebesell, Janosh↗

Computational multiphysics modeling of radioactive aerosol deposition in diverse human respiratory tract geometries

The evaluation of aerosol exposure relies on generic mathematical models that assume uniform particle deposition profiles over the human respiratory tract and do not account for subject-specific characteristics. Here we introduce a hybrid-automated computational workflow that generates personalized particle deposition profiles in 3D reconstructed human airways from computed tomography scans using Computational Fluid and Particle Dynamics simulations. This is the first large-scale study to consider realistic airways variability, where 380 lower and 40 upper human respiratory tract 3D geometries are reconstructed and parameterized. The data is clustered into nine groups using random forest regression. Computational fluid and particle dynamics simulations are conducted on these representative geometries using a realistic heavy-breathing respiratory cycle and radioactive iodine-131 as a source term. Monte Carlo radiation transport simulations are performed to obtain detailed energy deposition maps. Our findings emphasize the importance of personalized studies, as minor respiratory tract variations notably influence deposition patterns rather than global parameters of the lower airways, observing more than 30% variance in the mass deposition fraction.

62 RADIOLOGY AND NUCLEAR MEDICINE↗

Decoding substrate specificity determining factors in glycosyltransferase-B enzymes – insights from machine learning models

Substrate specificity is an essential characteristic of any enzyme's function and an understanding of the factors that determine this specificity is crucial for enzyme engineering. Unlike the structure of an enzyme which is directly impacted by its sequence, substrate specificity as an enzyme attribute involves a rather indirect relationship with sequence as it also depends on structural aspects that dictate substrate accessibility and active site dynamics. In this study, we explore the performance of classifier-based machine learning models trained on curated sequence and structural data for a class of glycosyltransferases (GTs), namely GT-Bs, to understand their substrate specificity determining factors. GTs enable the transfer of sugar moieties to other biomolecules such as oligosaccharides or proteins and are found in all kingdoms of life. In plants, GTs participate in the biosynthesis of plant cell wall biopolymers (e.g.: hemicelluloses and pectins) and are an integral part of the enzymatic machinery that enables the storage of carbon and energy as plant biomass. To elucidate the substrate specificity of uncharacterized GT-Bs, we constructed multi-label machine learning models (Support Vector Classifier, K-Nearest Neighbors, Gaussian Naïve-Bayes, Random Forest) that incorporate both sequence and structural features. These models achieve good predictive accuracies on test datasets. However, despite our use of structural information, we highlight that there is further scope for improvement in training these models to draw interpretable relationships between sequence, structure and substrate specificity determining motifs in GT-Bs.

97 MATHEMATICS AND COMPUTING↗

Machine learning-enabled discovery of ionic liquid–solvent electrolytes exhibiting high ionic conductivity

Ionic liquids (ILs), which are a class of materials with versatile nature and growing popularity, are facing impediments toward widespread usage as electrolytes due to various factors such as low ionic conductivity, high viscosity, high market price etc. One of the ways these limitations can be addressed is by mixing ILs with a molecular solvent. In a combinatorial sense, there exists an immense number of specific IL–solvent combinations. An exhaustive experimental or even simulation-based investigation of the chemical space spanned by such combinations can be extremely time-consuming, expensive, and nearly impossible. An alternative approach is to employ machine learning-based models developed from available databases. Although there exists prior literature that integrates machine learning to investigate mixtures of specific solvents with ILs, these models lack generalization necessitating development of a large number of ML models to handle various solvents. To remedy this shortcoming, as a part of designing green electrolytes with high ionic conductivity that can have potential applications in next-generation batteries and solar cells, this work aims to develop a unified machine learning model to predict ionic conductivity of any IL–solvent mixture system. In this regard, three models, namely, Random Forest, extreme gradient boosting (XGBoost), and artificial neural network (ANN) were formulated using the NIST ILThermo database. The dataset contained 549 unique ionic liquids from 16 cation families and 81 unique solvents, representing a total of 23 712 datapoints. SHAPLEY additive explanation (SHAP) method was used to assess the impact of various features on model prediction and their significance was compared with literature to gain physical insight about the model behavior. Finally, using the developed models, approximately 2.5 million IL–solvent mixtures at five different compositions were screened at room temperature. The high-throughput screening yielded nearly 19 000 IL–solvent mixtures for which ionic conductivity was found to exceed the ionic conductivity of conventional Li-ion battery electrolyte.

25 ENERGY STORAGE↗

Charting the chemical space of Zintl phases with graph neural networks and bonding insights

A large number of Zintl phases have been discovered by solid-state chemists driven by empirical knowledge, chemical intuition and in some cases, through serendipitous accidents. These discoveries have only scratched the surface, given the vast compositional and structural diversity that Zintl phases can accommodate. The large chemical space of Zintl phases, as well as intermetallic compounds in general, remain under-explored. Here, we use graph neural networks and the upper bound energy minimization approach to efficiently scan a large chemical space of >90 000 hypothetical Zintl phases and accurately discover 1810 new thermodynamically stable phases with 90% precision, as validated with first-principles calculations. We show that our approach is more than 2× more accurate in predicting DFT stability than M3GNet (40% precision) on the same dataset. Using a random forest model and SHAP analysis, we demonstrate the critical role of ionic bonding in the thermodynamic stability of Zintl phases. Our results not only expand the known chemical landscape of Zintl phases but also highlight the efficacy of machine learning frameworks combined with domain knowledge in uncovering chemically meaningful insights across complex intermetallics.

36 MATERIALS SCIENCE↗

Ensemble Federated Machine Learning‐Based Cybersecurity Situational Awareness in Microgrid Network

Cyber-physical microgrids are vulnerable to stealthy cybersecurity threats that disguise their actions through the exploitation of system knowledge. Such actions can severely impacts microgrids deployed in defense bases, slowing the response time of military forces during national emergencies. Several machine-learning algorithms have been proposed to detect intrusions in the grid networks; however, these traditional machine-learning algorithms lack data privacy and are subject to several adversarial machine-learning threats. This paper proposes a novel federated machine learning (FML)-based three-model framework to detect and identify stealthy data-integrity attacks while ensuring data privacy in microgrid networks. The proposed architecture uses a variational mode decomposition technique to extract derived features from incoming measurement and control datasets. The extraction of these derived features allows FML models to learn minute variations in data patterns that allow them to perform significantly better than the models trained with generic datasets consisting of raw features. Our experimental results show the efficient performance of the proposed methodology against different types of data integrity attacks while considering primary and secondary controllers in microgrids. Further, the applied FML-integrated random forest ensemble algorithm outperforms the existing generic FML algorithms during noisy and noise-free datasets with prediction latencies of only 91–134 µs per sample within the 0.1 s sampling interval and requires communication bandwidth of around ∼8.25 KB/s at the control center and ∼2.7 KB/s per edge client for communication.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Travelling wave‐based fault detection and location in a real low‐voltage DC microgrid

Abstract This paper discusses a device‐level implementation of a travelling wave (TW) protection device (PD) designed for a real low‐voltage DC microgrid. The TWPD fault detection and location algorithm is executed on a commercial digital signal processor (DSP) board, involving signal sampling at 1 MHz via the DSP board's analog‐to‐digital converter (ADC). The analogue input card measures positive pole, negative pole and pole‐to‐pole voltages at the TWPD location. Upon a successful fault detection using a second‐order high‐pass filter, the voltage data is normalised and multi‐resolution analysis (MRA) is performed on a 128‐sample buffer around the TW arrival time. MRA employs the discrete wavelet transform (DWT) to capture high‐frequency voltage patterns, and then the Parseval's energy theorem quantifies these TW characteristics by computing the energy of reconstructed wavelet coefficients. These energy values per decomposed frequency band are the basis for training a random forest classifier that predicts fault location and type. The TWPD is fully implemented and connected to a real DC microgrid in Albuquerque, NM, USA, for validation, and results are shown for field tests verifying the performance under faults.

Paruthiyil, Sajay Krishnan [Department of Electric↗

ATAT: Astronomical Transformer for time series and Tabular data

Context. The advent of next-generation survey instruments, such as theVera C. RubinObservatory and its Legacy Survey of Space and Time (LSST), is opening a window for new research in time-domain astronomy. The Extended LSST Astronomical Time-Series Classification Challenge (ELAsTiCC) was created to test the capacity of brokers to deal with a simulated LSST stream. Aims. Our aim is to develop a next-generation model for the classification of variable astronomical objects. We describe ATAT, the Astronomical Transformer for time series And Tabular data, a classification model conceived by the ALeRCE alert broker to classify light curves from next-generation alert streams. ATAT was tested in production during the first round of the ELAsTiCC campaigns. Methods. ATAT consists of two transformer models that encode light curves and features using novel time modulation and quantile feature tokenizer mechanisms, respectively. ATAT was trained on different combinations of light curves, metadata, and features calculated over the light curves. We compare ATAT against the current ALeRCE classifier, a balanced hierarchical random forest (BHRF) trained on human-engineered features derived from light curves and metadata. Results. When trained on light curves and metadata, ATAT achieves a macro F1 score of 82.9 ± 0.4 in 20 classes, outperforming the BHRF model trained on 429 features, which achieves a macro F1 score of 79.4 ± 0.1. Conclusions. The use of transformer multimodal architectures, combining light curves and tabular data, opens new possibilities for classifying alerts from a new generation of large etendue telescopes, such as theVera C. RubinObservatory, in real-world brokering scenarios.

Astronomy & Astrophysics↗

Barium stars as tracers of s -process nucleosynthesis in AGB stars

Barium (Ba) stars help to verify asymptotic giant branch (AGB) star nucleosynthesis models since they experienced pollution from an AGB binary companion and thus their spectra carry the signatures of the slow neutron capture process (s process). For a large number (180) of Ba stars, we searched for AGB stellar models that match the observed abundance patterns. We aim to uncover any systematic deviations of the sample abundances from the predictions of the nucleosynthesis models. We employed three machine learning algorithms as classifiers: a Random Forest method, developed for this work, and the two classifiers used in our previous study. Compared to that work, we also expanded our observational sample with 11 Ba stars available in the supersolar metallicity range. We studied the statistical behaviour of the different s-process elements in the observational sample to investigate if the AGB models systematically under- or overpredict the abundances observed in the Ba stars and show the results in the form of violin plots of the residuals between spectroscopic abundances and model predictions. We inspected the correlations between the observed [Fe/H], the s-process elemental abundances, and the residuals. We employed the [Zr/Fe] and [Nb/Fe] abundances as a thermometer to constrain the operational temperature that rules the production of these elements in the sample stars, assuming a steady-state s process. We also investigated the mass distribution of the identified polluter AGB stars and the behaviour of the δ parameter, which describes the fraction of accreted AGB material relative to the Ba star envelope. We find a significant trend in the residuals that implies an underproduction of the elements just after the first s-process peak (Nb, Mo, and Ru) in the models relative to the observations. This may originate from a neutron-capture process (e.g. the intermediate neutron-capture process, i process) not yet included in the AGB models of metallicity from solar to roughly 1/5 solar, corresponding to the range of the Ba stars. Correlations are found between the residuals of these peculiar elements, suggesting a common origin for the deviations from the models. In addition, there is a weak metallicity dependence of the residuals of these elements. The s-process temperatures derived with the [Zr/Fe] – [Nb/Fe] thermometer have an unrealistic value for the majority of our stars. The most likely explanation is that at least a fraction of these elements are not produced in a steady-state s process, and instead may be due to processes not included in the AGB models. The mass distribution of the identified models confirms that our sample of Ba stars was polluted by low-mass AGB stars (< 4 M ⊙ ). Most of the matching AGB models require low accreted mass, but a few systems with high accreted mass are needed to explain the observations.

79 ASTRONOMY AND ASTROPHYSICS↗

Machine learning enhanced predictions of ICRF heating: Overcoming numerical limitations via data curation

In this work, we present the development of robust surrogate models for Ion Cyclotron Range of Frequencies (ICRF) and High-Harmonic Fast Wave (HHFW) heating predictions in fusion plasmas. Building upon our previous efforts to achieve real-time capable models, we identify the cause of the outliers found using TORIC in certain HHFW heating scenarios. The outliers are observed to be spurious ion Bernstein wave (IBW)-like modes caused by a wavelength control algorithm designed to address challenging scenarios with high perpendicular wavenumbers. The effect arises from the modulation in the perpendicular susceptibility, which can induce sign reversal and IBW-like propagation for scenarios featuring normalized ion Larmor radius λ i ≫ 1. We use TORIC with this algorithm disabled to generate a novel HHFW-NSTX database that is free of outliers. Surrogate models trained on this database, including Random Forest Regressor (RFR), Multi-Layer Perceptrons, and Gaussian Process Regressors (GPR), demonstrate the ability to accurately predict HHFW heating profiles, with regression scores of R 2 ∈[0.93−0.99]. Additionally we demonstrate that it is possible to generalize predictions beyond training data by the use of both RFR and GPR models, enabling the prediction of scenarios previously limited to the original model. GPR models also provide uncertainty quantification, offering insights into model confidence. This work introduces a comprehensive Verification, Validation, and Uncertainty Quantification methodology for surrogate modeling, applicable not only to ICRF heating but also to other RF heating challenges and fusion physics problems. Beyond accelerated inference, these models show effective extrapolation capabilities, providing an alternative for addressing numerical challenges.

Artificial neural networks↗

Detection of Diversion in a Realistic Heat Pipe Microreactor Using Supervised Machine Learning

Microreactors (MRs) pose new challenges for international safeguards. Here, their small size and mass reproducibility make them ideal for deployment in greater numbers and in remote locations, making the job of safeguards inspectors more challenging. Machine learning (ML) is currently being applied to many fields to augment human performance and increase automation; in particular, ML could be used to provide insight for international inspectors to help detect the diversion of nuclear fuel from MR cores. Four ML model types (k-nearest neighbors, decision tree, random forest, and histogram-based gradient boosted ensemble) were trained on integrated flux and critical control drum angle data generated with Serpent 2 for a realistic heat pipe MR design, achieving nearly 100% binary classification accuracy of nominal and diversion core configurations by the end of 1 full power year for three of the four model types. Regression model variants were also trained, using the same input data, for predicting the number of fuel pins diverted. Root-mean-square errors below 5% of the total number of fuel pins were achieved by the 1 full power year mark for all models.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Machine Learning–Augmented Laser-Induced Breakdown Spectroscopy for Spectral Discrimination of Iron Oxalates

Enhanced characterization and phase identification of post-PUREX Pu Oxalates (PuOXA) are pivotal for nonproliferation and pre-detonation nuclear forensics. Despite significant advances in the characterization of PuO 2 samples, little is known about the impact of both the chemical structure and oxidation states of PuOXA (i.e., Pu(III) and Pu(IV)) have on optical emission signatures. Here, we demonstrate the analytical capabilities of laser-induced breakdown spectroscopy (LIBS) applied to Fe(II) and Fe(III) oxalate samples as surrogates for PuOXA, highlighting the discriminating features in the LIBS emission spectra arising from differences in the oxidation states within mixed FeOXA samples. We report the enhancement of spectral feature selection using Principal Component Analysis (PCA), which enables the analytical superiority of machine learning algorithms such as Linear Discriminant Analysis (LDA), Quadratic Discriminant Analysis (QDA), Partial Least Squares Regression (PLSR), Support Vector Regression (SVR), and Random Forest Regression (RFR) over conventional univariate techniques for phase discrimination and chemometric analysis. Cluster analysis revealed how both matrix effects and laser ablation influence cluster separability by introducing spectral artifacts that misdirect the maximization of variance. PCA-selected emission lines were used in the regression models, demonstrating that both univariate and multivariate linear regression models (i.e., PLSR and SVR) can achieve acceptable performance, with machine learning models outperforming conventional calibration regressions. Furthermore, the application of non-linearly activated PCA-selected emission lines illustrates how simplifying the data while retaining captured variance enables the use of less complex and more computationally efficient models. Furthermore, this is particularly evident in the underperformance of RFR, which suffers from increased computational costs and overfitting owing to its high complexity.

Oxalates↗

The Role of Nuclear Data Sensitivities in Prompt α-Eigenvalue Predictions of Delayed Critical Benchmarks

Alpha (α) eigenvalues, which describe the logarithmic time derivative of the neutron population in a multiplying system, are integral to time-dependent behavior and diagnostic applications. However, uncertainties in the evaluated nuclear data can significantly impact the accuracy of transport simulations for such quantities. This work explores the use of machine learning models to predict two key outputs, α-eigenvalues and keff bias, using input features derived from α-eigenvalue sensitivities to nuclear data. The criticality safety benchmark models used in this study come from the International Handbook of Evaluated Criticality Safety Benchmark Experiments. Three models, random forest, XGBoost, and NGBoost, are trained on both energy-resolved and energy-summed α sensitivities. For the α-eigenvalue bias prediction, NGBoost achieved the highest R 2 (0.9476) using energy-resolved features, while XGBoost performed best using summed sensitivities. In contrast, when predicting the keff bias, all the models showed moderate predictive capability (best R 2 ≈ 0.72), as the mapping from the static α-sensitivities to the static keff bias was less direct. SHAP (SHapley Additive exPlanations) analysis was used to interpret the model predictions. Across both prediction tasks, the features associated with neutron capture [H-1 (n, γ)], uranium scattering reactions (such as 235 U elastic/inelastic), and actinide capture/fission reactions (such as 239 Pu and 234 U) were consistently identified as the most impactful. This highlights the key role of specific nuclear reactions and energy ranges in shaping both time-dependent and steady-state criticality behavior. These results demonstrated that α-sensitivities, despite being computed for time-dependent metrics, can provide valuable insights for predicting both α-eigenvalues and the keff bias. Moreover, machine learning models offer a promising pathway for uncovering important nuclear data dependencies and guiding future data evaluation efforts.

Nuclear data↗

Are there differences in public interest toward automated vehicles’ ownership in burdened and non-burdened communities?

This study examines differences in public interest toward automated vehicle (AV) adoption between burdened communities (BCs) and non-burdened communities (non-BCs). The study uniquely captures travel behavior by integrating the census tract-level community indicators with the 2017–2019 Puget Sound Household Travel Survey. Public interest in AV ownership is modeled as a five-level ordered outcome, ranging from ‘Not interested at all’ to ‘Very interested.’ Segmented ordered logit models and Random Forest are used to identify key drivers and barriers to AV adoption across BCs and non-BCs. Results show that longer commute times, lack of personal vehicles, and concerns about AV performance in adverse weather conditions are significantly associated with higher interest levels toward owning AVs in BCs. In conclusion, the findings offer insights for developing targeted policies to promote AV adoption and address differences between BCs and non-BCs in accessing emerging transportation technologies.

Automated vehicles↗

Real-time capable modeling of ICRF heating on NSTX and WEST via machine learning approaches

Abstract A real-time capable core Ion Cyclotron Range of Frequencies (ICRF) heating model on NSTX and WEST is developed. The model is based on two nonlinear regression algorithms, the random forest ensemble of decision trees and the multilayer perceptron neural network. The algorithms are trained on TORIC ICRF spectrum solver simulations of the expected flat-top operation scenarios in NSTX and WEST assuming Maxwellian plasmas. The surrogate models are shown to successfully capture the multi-species core ICRF power absorption predicted by the original model for the high harmonic fast wave and the ion cyclotron minority heating schemes while reducing the computational time by six orders of magnitude. Although these models can be expanded, the achieved regression scoring, computational efficiency and increased model robustness suggest these strategies can be implemented into integrated modeling frameworks for real-time control applications.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

A new data-driven map predicts substantial undocumented peatland areas in Amazonia

Tropical peatlands are among the most carbon-dense terrestrial ecosystems yet recorded. Collectively, they comprise a large but highly uncertain reservoir of the global carbon cycle, with wide-ranging estimates of their global area (441 025–1700 000 km 2 ) and below-ground carbon storage (105–288 Pg C). Substantial gaps remain in our understanding of peatland distribution in some key regions, including most of tropical South America. Here we compile 2413 ground reference points in and around Amazonian peatlands and use them alongside a stack of remote sensing products in a random forest model to generate the first field-data-driven model of peatland distribution across the Amazon basin. Our model predicts a total Amazonian peatland extent of 251 015 km 2 (95th percentile confidence interval: 128 671–373 359), greater than that of the Congo basin, but around 30% smaller than a recent model-derived estimate of peatland area across Amazonia. The model performs relatively well against point observations but spatial gaps in the ground reference dataset mean that model uncertainty remains high, particularly in parts of Brazil and Bolivia. For example, we predict significant peatland areas in northern Peru with relatively high confidence, while peatland areas in the Rio Negro basin and adjacent south-western Orinoco basin which have previously been predicted to hold Campinarana or white sand forests, are predicted with greater uncertainty. Similarly, we predict large areas of peatlands in Bolivia, surprisingly given the strong climatic seasonality found over most of the country. Very little field data exists with which to quantitatively assess the accuracy of our map in these regions. Data gaps such as these should be a high priority for new field sampling. This new map can facilitate future research into the vulnerability of peatlands to climate change and anthropogenic impacts, which is likely to vary spatially across the Amazon basin.

54 ENVIRONMENTAL SCIENCES↗

Short-Term Energy and Meteorological Impacts on Thanksgiving CO2 in Salt Lake City

Abstract Long-term, high-frequency atmospheric CO2 measurements at multiple sites in the Salt Lake City (SLC), Utah, reveal that annual and monthly CO2 variability aligns with a priori estimates of emissions from anthropogenic and biological sources. In this study, we investigate whether short-term fluctuations in anthropogenic emissions, as captured in the Vulcan3 dataset for the United States, can be detected in atmospheric CO2 observations. Specifically, we focus on Thanksgiving holidays, when traffic and energy usage patterns differ from the rest of November. Onroad CO2 emissions exhibit a double peak during weekday morning and evening rush hours but remain relatively low on weekends and Thanksgiving. Interestingly, CO2 mole fractions during Thanksgiving were higher than the rest of November at all SLC monitoring sites, particularly from 2008 to 2013. This increase is partially attributed to elevated energy-related emissions — especially residential sources — and meteorological factors such as weak wind speeds, cold temperature, and a low planetary boundary layer height (PBLH).&#xD;&#xD; While CO₂ emissions and mole fraction patterns align over time, notable spatial differences exist. For instance, the near-highway site in Murray shows the highest CO₂ mole fractions despite low local emissions, suggesting pollution transport via highways and wind advection. Random Forest model-based SHapley Additive exPlanations (SHAP) analysis reveals that onroad emissions dominate CO2 contributions on weekdays and weekends, while energy-related emissions play a larger role during Thanksgiving, alongside meteorological drivers such as wind speed and PBLH. Across six urban cities, CO2 emissions display a consistent pattern: residential and commercial (onroad) emissions peak during Thanksgiving (weekday) with substantial (minimal) year-to-year variability. These findings highlight that urban CO₂ variability is driven by the combined influence of emissions and meteorology, underscoring the need for integrated mitigation strategies. Additionally, multi-site measurements are essential for accurate source attribution and the development of effective policy interventions. &#xD;

Ryoo, Ju-Mee (ORCID:0000000234256296)↗