Search NASA⌕ Search

SEARCH · Search NASA

Results for “Association Learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Arm and shoulder muscle segmentation in axial MRI with UNet deep learning model

Quantifying individual upper-limb muscle volumes from MRI provides key insight into muscle-specific strength, deficits, and adaptations. Manual delineation is the gold standard but time‑intensive, and the performance of current deep learning approaches, particularly for small or anatomically complex muscles, remains incompletely characterized. We evaluated a state‑of‑the‑art deep learning framework across the entire upper limb and analyzed factors governing segmentation performance, with attention to the forearm. Three previously published MRI datasets (1.5 T, 3D GRE T1‑weighted; total n = 39) spanning young, middle‑aged, and older adults were curated and quality‑checked, including expert manual segmentations for 31 muscles. Following multiclass mask reconstruction, we trained three 3D nnU‑Net multiclass models matched to the muscle subsets present across datasets, using five‑fold cross‑validation and a composite Dice Similarity Coefficient (DSC) + cross entropy loss. Segmentation accuracy was assessed with DSC. Performance varied across muscles (mean DSC = 0.806 ± 0.098), ranging from 0.920 (Deltoid) to 0.461 (Extensor pollicis brevis). In uncertainty‑weighted regressions, muscle volume was positively associated with DSC (R2 = 0.36, p < 0.001), whereas training segmentation count and muscle orientation showed negligible associations (R2 ≤ 0.06). A weighted mixed‑effects model identified volume as the strongest evaluated predictor, explaining 23.9% of variance in DSC; orientation and training count each contributed <1%, leaving 61.5% unexplained. These results indicate that deep learning–based segmentation can accurately quantify muscle volume for many upper‑limb muscles but remains constrained for small, low‑contrast forearm muscles.

Gillespie, Samuel↗

Modeling injection-induced fault slip using long short-term memory networks

Stress changes due to changes in fluid pressure and temperature in a faulted formation may lead to the opening/shearing of the fault. This can be due to subsurface (geo)engineering activities such as fluid injections and geologic disposal of nuclear waste. Such activities are expected to rise in the future making it necessary to assess their short- and long-term safety. Here, a new machine learning (ML) approach to model pore pressure and fault displacements in response to high-pressure fluid injection cycles is developed. The focus is on fault behavior near the injection borehole. To capture the temporal dependencies in the data, long short-term memory (LSTM) networks are utilized. To prevent error accumulation within the forecast window, four critical measures to train a robust LSTM model for predicting fault response are highlighted: (i) setting an appropriate value of LSTM lag, (ii) calibrating the LSTM cell dimension, (iii) learning rate reduction during weight optimization, and (iv) not adopting an independent injection cycle as a validation set. Several numerical experiments were conducted, which demonstrated that the ML model can capture peaks in pressure and associated fault displacement that accompany an increase in fluid injection. The model also captured the decay in pressure and displacement during the injection shut-in period. Further, the ability of an ML model to highlight key changes in fault hydromechanical activation processes was investigated, which shows that ML can be used to monitor risk of fault activation and leakage during high pressure fluid injections.

58 GEOSCIENCES↗

Non-destructive structural characterization of graphite components using mechanical resonance and deep learning

As compared to conventional nuclear reactors, microreactors have the potential to significantly reduce construction timelines and capital costs, decreasing the barriers for advanced nuclear reactor technologies. However, the lower power output of these microreactors (typically < 20 MWe) creates challenging economics if operation and maintenance costs cannot be sufficiently reduced. The compact size of these designs presents an opportunity for comprehensive in-situ structural health monitoring to provide real-time feedback in order to reduce operational costs associated with maintenance and downtime. Many microreactor concepts use graphite for both in-core neutron moderation and as a structural material, which has typically required some form of periodic and laborious inspection. This report provides a description and assessment of recent work with graphite to couple acoustic-based experimental measurements and characterization with machine learning models to mature structural health monitoring capabilities and generate benefits for the nuclear microreactor industry. With resilient embedded sensors in development in other programs funded by the US Department of Energy’s Office of Nuclear Energy and elsewhere, the work described herein builds upon previously funded efforts to mature non-destructive testing technology that relates measured vibrational signatures to structural changes, using a combination of new experimental measurements and machine learning processing. Building on past successful demonstrations of predictive workflows to identify structural changes in a hexagonal stainless steel test article with excellent acoustic propagation, we first performed baseline characterization on graphite samples with canonical geometries to ensure compatibility and confidence in the applied techniques for a material with distinctly different mechanical properties. In contrast to efforts in previous years, we worked exclusively with unidirectional vibration data that is more comparable to those expected from the existing embedded sensor technologies which are suitable for deployment in a reactor setting. Established acoustic and modern machine-learning-based characterization approaches were applied to the resulting datasets from these simple geometries. Both approaches were found to be highly capable of detecting even small geometric irregularities amongst nominally identical samples. As such, we then moved to testing these approaches for detection of artificial local stress perturbations introduced into a more complex geometry: a hexagonal block with drilled holes. A main outcome of this work is that a generalizable ML workflow can be used to detect and predict the characteristics of small artificial anomalies in a graphite component with a relevant geometry. While this work was performed using surficial vibration data, we expect the approach to be flexible and viable for other monitoring scenarios, such as those with different arrangements or types of sensor arrays. As compared to previously funded efforts, an existing ML workflow based on neural networks was enhanced through the addition of recently developed Fourier neural operators. As applied to previously collected and new vibration datasets, prediction accuracies of anomaly characterizations were greatly improved with minimal added computational cost. As trained on small durations of vibration data (tens of seconds) collected over a realistic number of locations, the model was able to reliably determine the presence of a subtle stress anomaly and begin to provide location estimates. Such an approach is likely to be viable for more relevant reactor damage scenarios for graphite components, such as progressive crack growth or creep.

36 MATERIALS SCIENCE↗

Ferroelectric phase transition in group-IV monochalcogenides from an equivariant machine learned force field

Group-IV monochalcogenides are a class of layered ferroelectric semiconductors that have demonstrated spontaneous intrinsic polarization above room temperature. Here, in this study, we use the multi-atomic cluster expansion (MACE) machine learning architecture to train and test a force field capable of modeling the structural properties and second-order ferroelectric-to-paraelectric phase transition in a Group-IV monochalcogenide, GeSe. The model captures the double-well potential energy surface associated with the onset of macroscopic polarization in bulk GeSe within 12.5 meV/atom, as well as near-equilibrium properties like the phonon dispersion. The development of this quantitatively accurate force field enables long-time molecular dynamics simulations, which predict the critical temperature of the ferroelectric-to-paraelectric phase transition in bulk GeSe to be T c = 600 K. This study demonstrates the capabilities of equivariant force-fields to accurately describe phenomena associated with structural symmetry breaking.

ferroelectricity↗

Bayesian Entropy Neural Networks for physics-aware prediction

This article addresses the need for deep learning models to integrate well-defined constraints into their outputs, driven by their application in surrogate models, learning with limited data and partial information, and scenarios requiring flexible model behavior to incorporate non-data sample information. We introduce Bayesian Entropy Neural Networks (BENN), a framework grounded in Maximum Entropy (MaxEnt) principles, designed to impose constraints on Bayesian Neural Network (BNN) predictions. BENN is capable of constraining not only the predicted values but also their derivatives and variances, ensuring a more robust and reliable model output. To achieve simultaneous uncertainty quantification and constraint satisfaction, we employ the method of multipliers approach. This allows for the concurrent estimation of neural network parameters and the Lagrangian multipliers associated with the constraints. Our experiments, spanning diverse applications such as beam deflection modeling and microstructure generation, demonstrate the effectiveness of BENN. The results highlight significant improvements over traditional BNNs and showcase competitive performance relative to contemporary constrained deep learning methods.

14 SOLAR ENERGY↗

Simultaneous Probe of the Charm and Bottom Quark Yukawa Couplings Using $t\bar{t}$𝐻 Events

A search for the standard model Higgs boson decaying to a charm quark-antiquark pair, 𝐻→$c\bar{c}$, produced in association with a top quark-antiquark pair ($t\bar{t}$𝐻) is presented. The search is performed with data from proton-proton collisions at √𝑠 =13 TeV, corresponding to an integrated luminosity of 138 fb−1. Advanced machine learning techniques are employed for jet flavor identification and event classification. The Higgs boson decay to a bottom quark-antiquark pair is measured simultaneously and the observed $t\bar{t}$𝐻(𝐻→$b\bar{b}$) event rate relative to the standard model expectation is 0.91$^{+0.26}_{−0.22}$. The observed (expected) upper limit on the product of production cross section and branching fraction 𝜎⁡($t\bar{t}$𝐻)⁢ℬ⁡(𝐻→$c\bar{c}$) is 0.11 (0.13) pb at 95% confidence level, corresponding to 7.8 (8.7) times the standard model prediction. When combined with the previous search for 𝐻 →$c\bar{c}$ via associated production with a 𝑊 or 𝑍 boson, the observed (expected) 95% confidence interval on the Higgs-charm Yukawa coupling modifier, 𝜅 𝑐 , is |𝜅 𝑐 | < 3.5 (2.7), the most stringent constraint to date.

Bottom quark↗

Crowdsourcing the Frontier: Advancing Hybrid Physics‐ML Climate Simulation via a $\$$50,000 Kaggle Competition

Subgrid machine-learning (machine learning [ML]) parameterizations have the potential to introduce a new generation of climate models that incorporate the effects of higher-resolution physics without incurring the prohibitive computational cost associated with more explicit physics-based simulations. However, important issues, ranging from online instability to inconsistent online performance, have limited their operational use for long-term climate projections. To more rapidly drive progress in solving these issues, domain scientists and ML researchers opened up the offline aspect of this problem to the broader ML and data science community with the release of ClimSim, a NeurIPS Data sets and Benchmarks publication, and an associated Kaggle competition. This paper reports on the downstream results of the Kaggle competition by coupling emulators inspired by the winning teams' architectures to an interactive climate model (including full cloud microphysics, a regime historically prone to online instability) and systematically evaluating their online performance. Our results demonstrate that online stability in the low-resolution real-geography setting is reproducible across multiple diverse architectures, which we consider a key milestone. All tested architectures exhibit strikingly similar offline and online biases, though their responses to architecture-agnostic design choices (e.g., expanding the list of input variables) can differ significantly. Multiple Kaggle-inspired architectures achieve state-of-the-art results on certain metrics such as zonal mean bias patterns and global Root Mean Squared Error, indicating that crowdsourcing the essence of the offline problem is one path to improving online performance in hybrid physics-AI climate simulation.

Environmental sciences↗

Accuracy versus precision in boosted top tagging with the ATLAS detector

The identification of top quark decays where the top quark has a large momentum transverse to the beam axis, known as top tagging , is a crucial component in many measurements of Standard Model processes and searches for beyond the Standard Model physics at the Large Hadron Collider. Machine learning techniques have improved the performance of top tagging algorithms, but the size of the systematic uncertainties for all proposed algorithms has not been systematically studied. This paper presents the performance of several machine learning based top tagging algorithms on a dataset constructed from simulated proton-proton collision events measured with the ATLAS detector at $\sqrt{s}$ = 13 TeV. The systematic uncertainties associated with these algorithms are estimated through an approximate procedure that is not meant to be used in a physics analysis, but is appropriate for the level of precision required for this study. The most performant algorithms are found to have the largest uncertainties, motivating the development of methods to reduce these uncertainties without compromising performance. To enable such efforts in the wider scientific community, the datasets used in this paper are made publicly available.

47 OTHER INSTRUMENTATION↗

Biosynthesis of bioprivileged, linear molecules via novel carboligase reactions

Over the award period, we made progress on the three aims. We screened twenty-five carboligases for activity coupling twenty-one possible -keto acids (Aim 1). The carboligases were selected across a diverse set of protein sequences. Using Q-Exactive UHPLC-MS, we tested a total of 210 coupled products per enzyme and generated a dataset of 5250 enzyme-substrate activity relationships. We identified multiple enzymes that had activity for synthesizing suberic acid and heptanoic acid (Aim 2). We built a random forest model for predicting the activity of each enzyme toward substrates on which it was not tested using the data from Aim 1. Finally, we evaluated growth defects that occurred due to expression of different carboligases in E. coli (Aim 3). We were able to identify specific metabolites and putative pathways that, when supplemented in the media, recovered the growth defect associated with the presence of specific carboligases. We are in the process of publishing two manuscript describing the methods for high-throughput screening of enzyme promiscuity, using machine learning to predict activity on untested substrates, and enzyme activity data we collected. This project has produced enabling data for biosynthesis of a range of new-to-nature compounds to support biomanufacturing.

60 APPLIED LIFE SCIENCES↗

Tuning water dissociation at oxide–electrolyte interfaces with electric fields

Understanding how electric fields influence water dissociation at heterogeneous interfaces is crucial for controlling interfacial chemical reactions and advancing next-generation energy technologies. Herein, ab initio–based machine learning simulations show that even small electric field changes can significantly alter the water dissociation fraction at planar TiO 2 –electrolyte interfaces. The resulting free energy difference between undissociated and dissociated interfacial water exhibits a linear dependence on the field change with a slope of 1.97 eÅ, which far exceeds the dissociation-induced dipole change of a water molecule. Employing a machine-learned collective variable to investigate the reaction statistics of thousands of water dissociation/recombination events, we find that small electric field changes exert minor effects on individual reaction energy barriers but significantly influence the populations of local configurations associated with initial states that are most favorable for reactions. These findings elucidate the pronounced impact of electric fields on interfacial water dissociation and reveal a mechanism for electric-field-controlled chemical reactions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Genetic variations and their interaction with thirdhand smoke exposure on anxiety and memory in Collaborative Cross mice

Thirdhand smoke (THS) is linked to adverse health effects, but the effect of genetic variations on behavioral outcomes is poorly understood. To investigate this, we assessed anxiety- and memory-related behaviors in 820 mice from 21 strains of the genetically diverse Collaborative Cross (CC) mouse that were exposed to THS from 4 through 10 weeks of age. Anxiety was evaluated with a light/dark box assay with a previously established risk score system. Females were generally more sensitive: THS reduced anxiety risk in strains CC013, CC019, and CC051, but increased risk in CC036 and CC061, while males showed no significant effects. Memory was tested using passive avoidance: impairments were observed in both sexes in CC016 and CC019, with sex-dependent effects in CC002 and CC051. A genome-wide association study identified 2,347 SNPs associated with anxiety and 1,568 SNPs with memory, with 32 and 85 SNPs, respectively, interacting with THS exposure. Enrichment analyses revealed distinct biological processes underlying susceptibility, including axonogenesis, synapse organization, cognition, and learning and memory. KEGG pathway analysis identified distinct genetic pathways, including GTPase binding and GTPase regulatory activity, that act as critical molecular switches in the brain that regulate synaptic plasticity, dendritic spine structure, and neuronal signaling, directly influencing anxiety-like behaviors and memory formation. These findings show that THS exposure affects neurobehavioral outcomes in a sex- and genotype-dependent manner, highlighting critical gene-environment interactions and providing a foundation for mechanistic insights into THS neurotoxicity

Anxiety↗

Reweighting configurations generated by transferable, machine learned models for protein sidechain backmapping

Multiscale modeling requires the linking of models at different levels of detail, with the goal of gaining accelerations from lower fidelity models while recovering fine details from higher resolution models. Communication across resolutions is particularly important in modeling soft matter, where tight couplings exist between molecular-level details and mesoscale structures. While multiscale modeling of biomolecules has become a critical component in exploring their structure and self-assembly, backmapping from coarse-grained to fine-grained, or atomistic, representations presents a challenge, despite recent advances through machine learning. A major hurdle, especially for strategies utilizing machine learning, is that backmappings can only approximately recover the atomistic ensemble of interest. We demonstrate conditions for which backmapped configurations may be reweighted to exactly recover the desired atomistic ensemble. By training separate decoding models for each sidechain type, we develop an algorithm based on normalizing flows and geometric algebra attention to autoregressively propose backmapped configurations for any protein sequence. Critical for reweighting with modern protein force fields, our trained models include all hydrogen atoms in the backmapping and make probabilities associated with atomistic configurations directly accessible. We also demonstrate, however, that reweighting is extremely challenging despite state-of-the-art performance on recently developed metrics and generation of configurations with low energies in atomistic protein force fields. Through detailed analysis of configurational weights, we show that machine-learned backmappings must not only generate configurations with reasonable energies, but also correctly assign relative probabilities under the generative model. These are broadly important considerations in generative modeling of atomistic molecular configurations.

Monroe, Jacob I. [Univ. of Arkansas, Fayetteville,↗

Multigene engineering in plants: Technologies, applications, and future prospects

The emerging bioeconomy presents a promising solution to both economic and environmental challenges. Within the bioeconomy, plants serve as a renewable, sustainable, and cost-effective source of foods, fuels, chemicals, and materials. However, traditional breeding and single-gene engineering approaches fall short in addressing complex traits (e.g., drought tolerance, disease resistance, yield, nutrient use efficiency) which are controlled by multiple genes. The complexity of plant biology often necessitates the use of multigene engineering (MGE), which involves simultaneous ectopic expression, up/down-regulation, or editing of multiple genes, to enhance plant traits relevant to the bioeconomy. These genes may be associated with distinct traits or function as components of specific metabolic and regulatory pathways. This review summarizes current technologies for MGE within the synthetic biology-driven Design-Build-Test-Learn (DBTL) framework, detailing its four key stages: Design – gene construct development; Build – DNA assembly and plant transformation; Test – the molecular, biochemical, and physiological characterization of engineered plants; and Learn – computational modeling to refine, multiplex and iterate the process. Despite good progress in the applications of MGE in biofortification, metabolic engineering, and stress resilience, challenges remain in construct stability, coordinated gene expression, and regulatory predictability. We identified optimization paths and future directions to accelerate MGE deployment in sustainable agriculture, with possible societal benefits including reduced production costs, increased yield, and improved food and nutritional security.

AI-aided plant engineering↗

Computational Analysis of the Energetic Stability of High-Entropy Structures of a Prototypical Lanthanide-Based Metal–Organic Framework

High-entropy materials are characterized by their complex compositions, typically comprising five or more elements in near-equiatomic proportions. Applying this concept to metal ions in metal−organic frameworks (MOFs) has paved the way for exploring a new class of high-entropy MOFs. While the compositional strategy of high-entropy materials leverages configurational entropy to aid thermodynamic stability, it also poses significant analytical challenges due to the vast compositional landscape and diverse phases that these materials can adopt. We present a computational study of several complexities associated with selecting potential high-entropy versions of a prototype lanthanidebased MOF. We compute the energetics of metal mixing of these heterometallic MOFs using density functional theory (DFT) and machine learning interatomic potential (MLIP) methods. The use of MLIP methods allows a systematic exploration of the convex hull of thermodynamically stable MOF structures containing up to 5 distinct metals.

Chemical structure↗

Renewable Energy Integration in Remote Alaska Communities

This guide was developed through the U.S. Department of Energy (DOE) Energy Transitions Initiative Partnership Project (ETIPP) technical assistance (TA) projects for Nikolski and St. George. Both communities had wind turbine projects that failed due to integration and maintenance issues. Due to these failures, Nikolski and St. George sought help to investigate renewable energy alternatives and associated integration strategies, with a focus on technologies that could be easily maintained within the community. The primary objective of this guide was to interview subject matter experts and document lessons learned to address potential renewable energy and storage integration issues in remote Alaskan communities. This guide is intended to provide information to rural Alaskan communities considering renewables with the purpose of sharing best practices and case studies to facilitate the successful implementation of future renewable energy projects.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Dominant Controls on Preferential Flow and Their Implications for Future Soil Water Fluxes

Abstract Soil water flow, particularly preferential flow (PF), is a critical control on hydrological and biogeochemical processes, including groundwater recharge, contaminant transport, and carbon cycling. However, it remains challenging to predict PF occurrence across large environmental gradients. Here, we developed a deep learning (DL) model to estimate event‐scale soil water flow velocity and the probability of PF occurrence using high‐frequency soil moisture and precipitation data from 33 sites across the National Ecological Observatory Network. The model demonstrated high skill in predicting the binary occurrence of PF (91% F1‐score; 85% accuracy) but the performance was limited in predicting soil water velocity ( R 2 = 0.31). We found that precipitation characteristics (duration, volume, and intensity) were the most important predictors for soil water velocity. Among the non‐precipitation event variables, sand content showed relatively high predictive skill, though differences among non‐event climate variables were generally modest. Lower sand content was associated with increased predicted soil water velocity, a finding that highlights the role of soil structure in producing more non‐uniform flow, which contrasts with traditional uniform flow models. Projecting a reduced DL model under both moderate and high‐emissions future climate scenarios (2060–2099 Representative Concentration Pathways 4.5 and 8.5), we found ∼7.3% increase under RCP4.5 and ∼15% under RCP8.5 of soil water velocities compared to the historical simulation, while modeled likelihood of PF changed little. These findings suggest climate change is not making PF more frequent, but it is making existing PF pathways more efficient with important consequences for associated nutrient and contaminant transport under climate change. Plain Language Summary Water movement in soil is critical for water quality. While often modeled as a uniform flow process, in reality water moves rapidly through cracks and burrows in what is called “preferential flow” (PF), which limits natural filtration and can transport pollutants. We developed a deep learning model, trained on data from 33 U.S. sites, to predict when and how fast this PF occurs based on precipitation, soil, and climate data. The model showed that precipitation characteristics (duration, intensity, volume) were the most important predictors of PF. Lower soil sand content/higher clay content was associated with faster water flow, likely due to clay soils forming aggregates and cracks that water moves through rather than infiltrating uniformly. Further analyses based on climate projections suggest that the speed at which PF occurs will become more rapid under future climate scenarios compared to historical simulation. This highlights the need to represent PF in soil water models when assessing future water quality. Key Points The effect of precipitation peak intensity on soil water velocities declined with increasing precipitation intensity Antecedent soil moisture failed to predict preferential flow (PF), contrasting the high predictive power of sand content Climate predictions suggest that soil water velocities through PF paths will increase ∼15% by 2099

Li, Bonan↗

Advancing International Integration and Strengthening Responsible Peaceful Uses of Nuclear Applications Through Specialized Curriculums

Nuclear technology has been pivotal in addressing some of the most pressing global challenges, ranging from energy production to combatting infectious diseases to agricultural security. Oak Ridge National Laboratory (ORNL), as a leader in nuclear research and development including in the field of radioisotopes, is well situated to share lessons learned and enhance global collaboration from years of discoveries in the field. Accordingly, the U.S. Department of Energy’s National Nuclear Security Administration (NNSA) sponsors specialized educational programs at ORNL, with support from the IAEA, focused on advancing peaceful nuclear applications and associated industries while upholding strong nuclear safety, security, and safeguards standards. The Joint U.S./IAEA International School on Peaceful Uses of Nuclear Applications, launched in 2024, is a cornerstone of this collaboration. The school provides an opportunity for early-career professionals from around the globe to gain practical knowledge and skills in utilizing nuclear technologies for peaceful purposes. This initiative demonstrates the United States’ commitment to its obligations under the Treaty on the Non-Proliferation of Nuclear Weapons (NPT), specifically Article IV, which calls on nuclear-weapon states to facilitate access to the peaceful uses of nuclear energy while guarding against the proliferation of nuclear weapons. In this paper, we explore the curriculum, objectives, and global impact of the program. Participants of the ICARST-2025 conference are invited to learn more about this program, contribute to its development, and explore opportunities for collaboration. This school exemplifies how strategic partnerships, and educational initiatives can drive the peaceful and beneficial use of nuclear technology worldwide

Raffo Caiado, Ana [ORNL] (ORCID:0009000239304805)↗

Statistical and Machine Learning Approaches to Analyzing Pipeline Incidents in the United States (2010–2024)

This study applies machine learning methods to analyze natural gas pipeline incidents in the United States using the Pipeline and Hazardous Materials Safety Administration (PHMSA) Gas Distribution Incident Dataset (2010–2024). The dataset includes over 600 variables describing incident characteristics, infrastructure attributes, and contributing factors associated with unintentional gas releases. The objective is to assess whether these features can reliably predict the underlying cause of pipeline failures. Multinomial logistic regression and Random Forest models were developed to classify incident causes, including excavation damage, corrosion, equipment failure, and natural forces. Results show that excavation damage is both the most frequent and most predictable cause, with models achieving strong performance for this category. However, when excavation damage is excluded, model accuracy declines significantly, with some models performing near random levels. Across all approaches, severe class imbalance and limited variability in key predictors constrain predictive performance. Pipeline age and diameter emerge as the most influential variables, but they provide insufficient discriminatory power to distinguish among less frequent failure types. These findings indicate that non-excavation-related incidents are rare, heterogeneous, and weakly represented in the dataset, limiting the effectiveness of machine learning classification. Overall, this study highlights the structural limitations of the PHMSA dataset for predictive modeling and underscores the need for improved data balance and feature enrichment. The results reinforce excavation damage prevention as the most impactful strategy for reducing pipeline incidents.

03 NATURAL GAS↗