Search NASA⌕ Search

SEARCH · Search NASA

Results for “machine learning classification”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

Explainable AI classification for parton density theory

Quantitatively connecting properties of parton distribution functions (PDFs, or parton densities) to the theoretical assumptions made within the QCD analyses which produce them has been a longstanding problem in HEP phenomenology. To confront this challenge, we introduce an ML-based explainability framework, XAI4PDF, to classify PDFs by parton flavor or underlying theoretical model using ResNet-like neural networks (NNs). By leveraging the differentiable nature of ResNet models, this approach deploys guided backpropagation to dissect relevant features of fitted PDFs, identifying x-dependent signatures of PDFs important to the ML model classifications. By applying our framework, we are able to sort PDFs according to the analysis which produced them while constructing quantitative, human-readable maps locating the x regions most affected by the internal theory assumptions going into each analysis. This technique expands the toolkit available to PDF analysis and adjacent particle phenomenology while pointing to promising generalizations.

Artificial Intelligence↗

Discriminative versus generative approaches to simulation-based inference

Most of the fundamental, emergent, and phenomenological parameters of particle and nuclear physics are determined through parametric template fits. Simulations are used to populate histograms which are then matched to data. This approach is inherently lossy, since histograms are binned and low-dimensional. Deep learning has enabled unbinned and high-dimensional parameter estimation through neural likelihood(-ratio) estimation. We compare two approaches for neural simulation-based inference (NSBI): one based on discriminative learning (classification) and one based on generative modeling. These two approaches are directly evaluated on the same datasets, with a similar level of hyperparameter optimization in both cases. In addition to a Gaussian dataset, we study NSBI using a Higgs boson dataset from the FAIR Universe Challenge. We find that both the direct likelihood and likelihood ratio estimation are able to effectively extract parameters with reasonable uncertainties. For the numerical examples and within the set of hyperparameters studied, we found that the likelihood ratio method is more accurate and/or precise. Both methods have a significant spread from the network training and would require ensembling or other mitigation strategies in practice.

high energy physics↗

Multi-contrast machine learning improves schistosomiasis diagnostic performance

Schistosomiasis currently affects over 250 million people and remains a public health burden despite ongoing global control efforts. Conventional microscopy is a practical tool for diagnosis and screening ofSchistosoma haematobium, but identification of eggs requires a skilled microscopist. Here we present a machine learning (ML)-based strategy for automated detection ofS. haematobiumthat combines two imaging contrasts, brightfield (BF) and darkfield (DF), to improve diagnostic performance. We collected BF and DF images of urine samples, many of them containingS. haematobiumeggs, during two different field studies in Côte d’Ivoire using a mobile phone-based microscope, the SchistoScope. We then trained separate egg-detection ML models and compared the patient-level performance of BF and DF models alone to combinations of BF and DF models, using annotations from trained microscopists as the gold standard. We found that models trained on DF images, and almost all BF and DF combinations, performed significantly better than models trained on BF images only. When models were trained on images from the first field study (n = 349 patients, 748 images of each contrast), patient-level classification performance on patient images from the second study (n = 375 patients, 752 images of each contrast) met the WHO Diagnostic Target Product Profile (TPP) sensitivity and specificity for the monitoring and evaluation use case (sensitivity for all models and combinations was >75% when evaluated at a confidence score threshold that resulted in specificity >96.5%). When we used images from both field studies for the training set, performance of the models was improved. Overall, this work shows that the use of DF and BF increases the performance of ML models on images from devices with low-cost optics, while retaining the portability, power, and time-to-results of the WHO’s diagnostic TPP. DF requires no additional sample preparation and does not increase the complexity of the imaging system. It thus offers a practical means to improve performance of automated diagnostics forS. haematobiumas well as other microscopy-based diagnostics.

Infectious Diseases↗

Demonstrating Advanced Sensors for In-Situ Monitoring Towards Qualification of Nuclear Relevant Components

The U.S. Department of Energy’s Office of Nuclear Energy Advanced Materials and Manufacturing Technologies (AMMT) program is pursuing qualification of laser powder bed fusion (LPBF) components for nuclear applications. A major focus of this effort is the use of in situ process monitoring and machine learning–based tools to establish real-time quality assurance. The primary objective of this report is to identify and evaluate the most relevant in situ sensor systems for LPBF, and to document the deployment of these systems across platforms critical to the AMMT program. This work demonstrates how in situ monitoring can detect process anomalies, track geometry-dependent flaws, and identify limiting combinations of processing parameters—particularly those related to energy density and complex geometries (e.g., overhanging structures). To support this goal, a diverse suite of sensor modalities was evaluated across LPBF platforms, including visible and near-infrared (NIR) imaging, fringe projection profilometry, long-wavelength infrared (LWIR) thermography, and high-speed photodiode/pyrometry systems. These sensor streams were integrated with Peregrine, a machine-agnostic software platform that, among other capabilities, can generate real-time process anomaly classification. This report documents sensor deployments on multiple AMMT flagship platforms, including the Concept Laser M2 and Renishaw AM400/AM250 systems. Calibration builds with complex, flaw-prone geometries such as unsupported overhangs, stepped features, and thin walls, were used to evaluate how well Peregrine and its associated sensors could detect process anomalies and other instabilities under varied energy densities. It will be shown how Peregrine reliably identifies common process anomalies such as recoater streaking, superelevation, etc., and can be used in post-build analysis for anomaly spatial distributions throughout the build height to better understand the impact of geometry and processing parameter choice on the build. This work demonstrates measurable progress toward the vision that components can be born-qualified by establishing a real-time monitoring framework, identifying limiting process conditions, and laying the foundation for sensor fusion–enabled prediction pipelines that are scalable across platforms and applicable to nuclear-relevant components.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Data-Driven Clustering and Classification of Outage Patterns with Insights into their Links to Extreme Events

At a global level extreme events have increased in both scale and impact. These events have the potential to affect the electrical grid infrastructure and cause a wide range of outages, which can lead to a disruption in daily patterns, cost millions of dollars and also the loss of life. Currently, to track these outage events there have been various approaches developed ranging from regional to national level quantifications for what defines an outage. However, this variation in methods can potentially lead to subjective decision-making and a lack of proper management in relation to the event. While previous work has made strides in determining spatio-temporal patterns, minimal attention has been given to the type and number of outages an area may be exposed to. The differences in incurred cost and the overall severity of an event between a transformer box malfunction and a hurricane are drastic, and by finding historical signals, we can allow for more efficient management, potentially saving lives and millions of dollars. Here, we leverage unsupervised machine learning techniques to delineate outage patterns among 22 counties within the United States and find that there are clear, segregated clusters (0.93 silhouette) of data which are related by event behavior and underlying cause. This finding will allow for energy stakeholders, policy makers, and researchers to gain a deeper understanding of the extent and severity of historic events and to better prepare for electrical grid infrastructure planning and management.

Koob, Benjamin [ORNL]↗

Learned adaptive properties for mitigation of weight perturbations in embedded spiking networks

Recent years have seen an increased importance of neural network inference in edge-based scenarios, which impose size and power constraints requiring novel computing devices. These same edge scenarios may require operating over long periods of time, or exposure to extreme environments, resulting in a drift of neural network weights that cause degraded performance. In searching for ways to develop neural network approaches that perform robustly under these conditions, we propose a biologically-inspired mechanism for the dynamic adaptation of within-neuron parameters that is guided by a global context signal carrying information about perturbations and variability in incoming stimuli. Specifically, we demonstrate that adaptive voltage thresholds or neuronal time constants, when informed by a global context signal, can enable network-level mechanisms to recover from perturbed synaptic weights. Consistent with prior literature, the context-modulated approach is effective for recurrent, but not feedforward networks, by modulating network level dynamics. We demonstrate this approach successfully recovers performance in image classification tasks and spatiotemporal tracking tasks under idealized and Gaussian noise as well as for realistic perturbations from a memristive device when exposed to ionizing radiation. Finally, we discuss how this approach enables the design of robust and energy-efficient neuromorphic systems that perform well, even in resource-constrained scenarios with extreme environments such as edge processing.

context modulation↗

Optimal transport for 𝑒/𝜋 0 particle classification in LArTPC neutrino experiments

The efficient classification of electromagnetic activity from 𝜋 0 and electrons remains an open problem in the reconstruction of neutrino interactions in liquid argon time projection chamber (LArTPC) detectors. We address this problem using the mathematical framework of optimal transport (OT), which has been successfully employed for event classification in other high energy physics contexts and is ideally suited to the high-resolution calorimetry of LArTPCs. Using a publicly available simulated dataset from the MicroBooNE Collaboration, we show that OT methods achieve state-of-the-art reconstruction performance in 𝑒/𝜋 0 classification. The success of this first application indicates the broader promise of OT methods for LArTPC-based neutrino experiments.

Neutrino detection↗

Micro-photoluminescence mapping and Chemometrics for the rapid classification of rare earth materials

This article introduces advancements in chemically mapping rare earth materials using photoluminescence (PL) and chemometrics. By leveraging the high sensitivity and selectivity of PL compared to alternative optical techniques, as well as its compatibility with microscopy, we present enhanced capabilities for noninvasive material screening and characterization. Exemplary PL spectra of samarium(III) and europium(III) in oxide, nitrate, and chloride forms demonstrated the ability to extract detailed chemical information of diverse rare earth particles. Additionally, we introduced efficient PL mapping sequences capable of covering a 9 mm diameter carbon tab within minutes, which highlighted the benefits of rapid, large-area imaging. Furthermore, an integrated approach combining PL mapping with principal component analysis and a random forest classifier enabled the resolution of overlapping spectral peaks from different chemistries and provided accurate material classification. In conclusion, these advancements underscored the versatility and robustness of PL for chemically mapping rare earth materials, with the potential to support applications in mining, energy, environmental monitoring, isotope production and beyond.

Chemometrics↗

FPGA-accelerated SpeckleNN with SNL for real-time X-ray single-particle imaging

We present the implementation of a specialized version of our previously published unified embedding model, SpeckleNN, for real-time speckle pattern classification in X-ray Single-Particle Imaging (SPI), using the SLAC Neural Network Library (SNL) on an FPGA platform. This hardware realization transitions SpeckleNN from a prototypic model into a practical edge solution, optimized for running inference near the detector in high-throughput X-ray free-electron laser (XFEL) facilities, such as those found at the Linac Coherent Light Source (LCLS). To address the resource constraints inherent in FPGAs, we developed a more specialized version of SpeckleNN. The original model, which was designed for broader classification across multiple biological samples, comprised ~5.6 million parameters. The new implementation, while reducing the parameter count to 64.6K (a 98.8% reduction), focuses on maintaining the model's essential functionality for real-time operation, achieving an accuracy of 90%. Furthermore, we compressed the latent space from 128 to 50 dimensions. This implementation was demonstrated on the KCU1500 FPGA board, utilizing 71% of available DSPs, 75% of LUTs, and 48% of FFs, with an average power consumption of 9.4W according to the Vivado post-implementation report. The FPGA performed inference on a single image with a latency of 45.015 microseconds at a 200 MHz clock rate. In comparison, running the same inference on an NVIDIA A100 GPU resulted in an average power consumption of ~73W and an image processing latency of around 400 microseconds. Our FPGA-accelerated version of SpeckleNN demonstrated significant improvements, achieving an 8.9 × speedup and a 7.8 × reduction in power consumption compared to the GPU implementation. Key advancements include model specialization and dynamic weight loading through SNL, which eliminates the need for time-consuming FPGA design re-synthesis, allowing fast and continuous deployment of models (re)trained online. These innovations enable real-time adaptive classification and efficient vetoing of speckle patterns, making SpeckleNN more suited for deployment in XFEL facilities. This implementation has the potential to significantly accelerate SPI experiments and enhance adaptability to evolving experimental conditions.

47 OTHER INSTRUMENTATION↗

Assessing decision boundaries under uncertainty

In order to make design decisions, engineers may seek to identify regions of the design domain that are acceptable in a computationally efficient manner. A design is typically considered acceptable if its reliability with respect to parametric uncertainty exceeds the designer’s desired level of confidence. Despite major advancements in reliability estimation and in design classification via decision boundary estimation, the current literature still lacks a design classification strategy that incorporates parametric uncertainty and desired design confidence. To address this gap, this paper offers a novel interpretation of the acceptance region by defining the decision boundary as the hypersurface which isolates the designs that exceed a user-defined level of confidence given parametric uncertainty. This work addresses the construction of this novel decision boundary using computationally efficient algorithms that were developed for reliability analysis and decision boundary estimation. The approach proposed in this paper is verified on two physical examples from structural and thermal analysis using Support Vector Machines and Efficient Global Optimization-based contour estimation.

97 MATHEMATICS AND COMPUTING↗

Predicting Dynamic-to-Static Correction Factor from Petrophysical Data and Chemostratigraphy using Unsupervised Machine Learning

Estimating static mechanical properties of stratigraphic layers is critical for optimizing subsurface engineering applications. To estimate dynamic-to-static correction factor F ds (static-to-dynamic Young’s modulus ratio) across the Caney shale interval in Oklahoma, USA, we integrated triaxial test measurements and petrophysical data, including well logs and X-ray fluorescence (XRF) using unsupervised machine learning (ML). We used a novel workflow that includes principal component analysis (PCA) to reduce data set dimensionality of well logs and XRF data sets—both separately and combined—creating three scenarios, and later applied inverse distance weighting (IDW) to derive F ds profiles for these scenarios. Furthermore, we applied K-means clustering on each scenario to predict depositional facies, and built a stiffness zonation profile through chemostratigraphic analysis of the terrigenous elements to validate the predicted F ds . The predicted F ds profile from each scenario using the PCA-IDW method was compared with the constant F ds approach from our previous study by calculating the root mean square error (RMSE). The combined data sets scenario yielded the lowest RMSE value of 0.113, while the RMSE values for the well logs and XRF scenarios were 0.131 and 0.129, respectively. In addition, the predicted F ds from the XRF scenario well-matched the stiffness zonation from the chemostratigraphic analysis that was built using the optimized K-means clustering of nine clusters for that scenario. These methods and findings offer a valuable tool for refining lithological classification and improving the F ds profile, potentially enhancing drilling and stimulation strategies for subsurface energy engineering applications.

clastic rock↗

Machine Learning Correlation of Electron Micrographs and ToF-SIMS for the Analysis of Organic Biomarkers in Mudstone

The spatial distribution of organics in geological samples can be used to determine when and how these organics were incorporated into the host rock. Mass spectrometry (MS) imaging can rapidly collect a large amount of data, but ions produced are mixed without discrimination, resulting in complex mass spectra that can be difficult to interpret. Here, we apply unsupervised and supervised machine learning (ML) to help interpret spectra from time-of-flight-secondary ion mass spectrometry (ToF-SIMS) of an organic-carbon-rich mudstone of the Middle Jurassic of England (UK). It was previously shown that the presence of sterane molecular biomarkers in this sample can be detected via ToF-SIMS (Pasterski, M. J. et al., Astrobiology 2023, 23, 936). We use unsupervised ML on scanning electron microscopy–electron dispersive spectroscopy (SEM-EDS) measurements to define compositional categories based on differences in elemental abundances. We then test the ability of four ML algorithms─k-nearest neighbors (KNN), recursive partitioning and regressive trees (RPART), eXtreme gradient boost (XGBoost), and random forest (RF)─to classify the ToF-SIM spectra using (1) the categories assigned via SEM-EDS, (2) organic and inorganic labels assigned via SEM-EDS, and (3) the presence or absence of detectable steranes in ToF-SIMS spectra. In terms of predictive accuracy and balanced accuracy, KNN was the best performing model and RPART the worst. The feature importance, or the specific features of the ToF-SIM spectra used by the models to make classifications, cannot be determined for KNN, preventing posthoc model interpretation. Nevertheless, the feature importance extracted from the other models was useful for interpreting spectra. In conclusion, we determined that some of the organic ions used to classify biomarker containing spectra may be fragment ions derived from kerogen which is abundant in this mudstone sample.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Scenario Storyline Discovery for Planning in Multi‐Actor Human‐Natural Systems Confronting Change

Scenarios have emerged as valuable tools in managing complex human-natural systems, but the traditional approach of limiting focus on a small number of predetermined scenarios can inadvertently miss consequential dynamics, extremes, and diverse stakeholder impacts. Exploratory modeling approaches have been developed to address these issues by exploring a wide range of possible futures and identifying those that yield consequential vulnerabilities. However, vulnerabilities are typically identified based on aggregate robustness measures that do not take full advantage of the richness of the underlying dynamics in the large ensembles of model simulations and can make it hard to identify key dynamics and/or storylines that can guide planning or further analyses. This study introduces the FRamework for Narrative Storylines and Impact Classification (FRNSIC; pronounced “forensic”): a scenario discovery framework that addresses these challenges by organizing and investigating consequential scenarios using hierarchical classification of diverse outcomes across actors, sectors, and scales, while also aiding in the selection of scenario storylines, based on system dynamics that drive consequential outcomes. We present an application of this framework to the Upper Colorado River Basin, focusing on decadal droughts and their water scarcity implications for the basin's diverse users and its obligations to downstream states through Lake Powell. We show how FRNSIC can explore alternative sets of impact metrics and drought dynamics and use them to identify drought scenario storylines, that can be used to inform future adaptation planning.

54 ENVIRONMENTAL SCIENCES↗

Data Science Shows that Entropy Correlates with Accelerated Zeolite Crystallization in Monte Carlo Simulations

We have performed a data science study of Monte Carlo simulation trajectories to understand factors that can accelerate formation of zeolite nanoporous crystals, a process that can take days or even weeks. In previous work, Monte Carlo simulations predicted and experiments confirmed that using a secondary organic structure-directing agent (OSDA) accelerates crystallization of all-silica LTA zeolite, with experiments finding a three-fold speedup [PCCP 24, 142-148 (2022)]. However, it remains unclear what physical factors cause the speed-up. Here, we apply data science to analyze the simulation trajectories to discover what drives accelerated zeolite crystallization in Monte Carlo going from a one-OSDA synthesis (1OSDA) to a two-OSDA version (2OSDA). We encoded simulation snapshots using the Smooth Overlap of Atomic Positions approach, which represents all 2- and 3-body correlations within a given cutoff distance. Principal component analyses failed to discriminate datasets of structures from 1OSDA and 2OSDA simulations, while the Support Vector Machine (SVM) approach succeeded at classifying such structures with an area-under-curve (AUC) score of 0.99 (where AUC = 1 is a perfect classification) with all 3-body correlations, and as high as 0.94 with only 2-body correlations. SVM decision functions reveal relatively broad / narrow histograms for 1OSDA / 2OSDA datasets, suggesting that the two simulations differ strongly in information heterogeneity. Informed by these results, we performed pair (2-body) entropy calculations during crystallization, resulting in entropy differences that semi-quantitatively account for the speedup observed in the previous Monte Carlo simulations. We conclude that altering synthesis conditions in ways that substantially changes the entropy of labile silica networks may accelerate zeolite crystallization, and we discuss possible approaches for achieving such acceleration.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

Using active learning to improve quasar identification for the DESI spectra processing pipeline

The Dark Energy Spectroscopic Instrument (DESI) survey uses an automatic spectral classification pipeline to classify spectra. QuasarNET is a convolutional neural network used as part of this pipeline originally trained using data from the Baryon Oscillation Spectroscopic Survey (BOSS). In this paper we implement an active learning algorithm to optimally select spectra to use for training a new version of the QuasarNET weights file using only DESI data, with the goal of improving classification accuracy. This active learning algorithm includes a novel outlier rejection step using a Self-Organizing Map to ensure we label spectra representative of the larger quasar sample observed in DESI. We perform two iterations of the active learning pipeline, assembling a final dataset of 5600 labeled spectra, a small subset of the approximately 1.3 million quasar targets in DESI's Data Release 1. When splitting the spectra into training and validation subsets we achieve similar performance to the previously trained weights file in completeness and purity calculated on the validation dataset but do so with less than one tenth of the amount of training data. The new weights also more consistently classify objects in the same way when used on unlabeled data compared to the old weights file. In the process of improving QuasarNET's classification accuracy we discovered a systemic error in QuasarNET's redshift estimation and used our findings to improve our understanding of QuasarNET's redshifts.

Machine learning↗

Smart Pixels: In-pixel AI for on-sensor data filtering

We present a smart pixel prototype readout integrated circuit (ROIC) designed in CMOS 28 nm bulk process, with in-pixel implementation of an artificial intelligence (AI) / machine learning (ML) based data filtering algorithm designed as proof-of-principle for a Phase III upgrade at the Large Hadron Collider (LHC) pixel detector. The first version of the ROIC consists of two matrices of 256 smart pixels, each 25$\times$25 $\mu$m$^2$ in size. Each pixel consists of a charge-sensitive preamplifier with leakage current compensation and three auto-zero comparators for a 2-bit flash-type ADC. The frontend is capable of synchronously digitizing the sensor charge within 25 ns. Measurement results show an equivalent noise charge (ENC) of $\sim$30e$^-$ and a total dispersion of $\sim$100e$^-$ The second version of the ROIC uses a fully connected two-layer neural network (NN) to process information from a cluster of 256 pixels to determine if the pattern corresponds to highly desirable high-momentum particle tracks for selection and readout. The digital NN is embedded in-between analog signal processing regions of the 256 pixels without increasing the pixel size and is implemented as fully combinatorial digital logic to minimize power consumption and eliminate clock distribution, and is active only in the presence of an input signal. The total power consumption of the neural network is $\sim$ 300 $\mu$W. The NN performs momentum classification based on the generated cluster patterns and even with a modest momentum threshold, it is capable of 54.4% - 75.4% total data rejection, opening the possibility of using the pixel information at 40MHz for the trigger. The total power consumption of analog and digital functions per pixel is $\sim$ 6 $\mu$W per pixel, which corresponds to $\sim$ 1 W/cm$^2$ staying within the experimental constraints.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Smart Pixels: In-pixel AI for on-sensor data filtering

We present a smart pixel prototype readout integrated circuit (ROIC) designed in CMOS 28 nm bulk process, with in-pixel implementation of an artificial intelligence (AI) / machine learning (ML) based data filtering algorithm designed as proof-of-principle for a Phase III upgrade at the Large Hadron Collider (LHC) pixel detector. The first version of the ROIC consists of two matrices of 256 smart pixels, each 25$\times$25 µm\textsuperscript{2} in size. Each pixel consists of a charge-sensitive preamplifier with leakage current compensation and three auto-zero comparators for a 2-bit flash-type ADC. The frontend is capable of synchronously digitizing the sensor charge within 25 ns. Measurement results show an equivalent noise charge (ENC) of $\sim$30e\textsuperscript{-} and a total dispersion of $\sim$100e\textsuperscript{-} The second version of the ROIC uses a fully connected two-layer neural network (NN) to process information from a cluster of 256 pixels to determine if the pattern corresponds to highly desirable high-momentum particle tracks for selection and readout. The digital NN is embedded in-between analog signal processing regions of the 256 pixels without increasing the pixel size and is implemented as fully combinatorial digital logic to minimize power consumption and eliminate clock distribution, and is active only in the presence of an input signal. The total power consumption of the neural network is $\sim$ 300 $\mu$W. The NN performs momentum classification based on the generated cluster patterns and even with a modest momentum threshold, it is capable of 54.4\% – 75.4\% total data rejection, opening the possibility of using the pixel information at 40MHz for the trigger. The total power consumption of analog and digital functions per pixel is $\sim$ 6 $\mu$W per pixel, which corresponds to $\sim$ 1 W/cm\textsuperscript{2} staying within the experimental constraints.

Parpillon, Benjamin↗

Harnessing Machine Learning and Data Fusion for Accurate Undocumented Well Identification in Satellite Images

This study utilizes satellite data to detect undocumented oil and gas wells, which pose significant environmental concerns, including greenhouse gas emissions. Three key findings emerge from the study. Firstly, the problem of imbalanced data is addressed by recommending oversampling techniques like Rotation–GaussianBlur–Solarization data augmentation (RGS), the Synthetic Minority Over-Sampling Technique (SMOTE), or ADASYN (an extension of SMOTE) over undersampling techniques. The performance of borderline SMOTE is less effective than that of the rest of the oversampling techniques, as its performance relies heavily on the quality and distribution of data near the decision boundary. Secondly, incorporating pre-trained models trained on large-scale datasets enhances the models’ generalization ability, with models trained on one county’s dataset demonstrating high overall accuracy, recall, and F1 scores that can be extended to other areas. This transferability of models allows for wider application. Lastly, including persistent homology (PH) as an additional input improves performance for in-distribution testing but may affect the model’s generalization for out-of-distribution testing. A careful consideration of PH’s impact on overall performance and generalizability is recommended. Overall, this study provides a robust approach to identifying undocumented oil and gas wells, contributing to the acceleration of a net-zero economy and supporting environmental sustainability efforts.

SMOTE↗