Search NASA⌕ Search

SEARCH · Search NASA

Results for “Resource prediction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Pavement condition and climatic data in southeast Texas: A dataset for evaluating flood impacts on pavement performance

Effective pavement maintenance is essential for economic stability, optimal network performance, and roadway safety. Achieving this requires thorough evaluation of pavement conditions, including structural integrity, surface roughness, and distress characteristics. Pavement performance indicators play a critical role in influencing vehicle safety and ride quality. Recent advances have emphasized the use of data-driven modeling to anticipate pavement behavior, with the goal of optimizing resource allocation and refining Maintenance and Rehabilitation (M&R) strategies through accurate condition assessment. A foundational requirement for these modeling efforts is the availability of standardized, high-quality datasets that can support robust and reproducible infrastructure analysis. This data article presents a comprehensive dataset assembled to facilitate pavement performance prediction, with a geographic focus on Southeast Texas, particularly the flood-vulnerable area of Beaumont. The dataset encompasses pavement and traffic attributes, meteorological records, flood simulation outputs, ground deformation measurements, and topographic indices, enabling detailed examination of both load-associated and non-load-associated degradation mechanisms. Data preprocessing was performed using ArcGIS Pro, Microsoft Excel, and Python to ensure consistency and usability in data-driven modeling applications, including machine learning workflows. Key contributions of this dataset include its utility in analyzing the climatic and environmental factors affecting pavement conditions, identifying critical predictive features, and enabling in-depth correlation analysis across diverse variables. By filling existing gaps in input variable selection resources, this dataset supports the development of predictive tools for estimating future maintenance demand and enhancing the resilience of pavement networks in flood-impacted areas. The resource highlights the importance of standardized datasets for advancing pavement management practices and provides a robust foundation for ongoing infrastructure performance modeling.

42 ENGINEERING↗

Denoising Autoencoder for Reconstructing Sensor Observation Data and Predicting Evapotranspiration: Noisy and Missing Values Repair and Uncertainty Quantification

Abstract Machine learning (ML) methods applied in scientific research often deal with interrelated features in high‐dimensional data. Reducing data noise and redundancy is needed to increase prediction accuracy and efficiency especially when dealing with data from field sensors. We explored an unsupervised learning method, the denoising autoencoder (DAE), to extract the underlying data structure from noisy raw data in the context of predicting hydrologic quantities from multiple field sensors. These sensors have intrinsic instrumental noise and occasional malfunctions that cause missing values. Our DAE neural network reconstructed meteorological sensor data containing noise and missing values to predict evapotranspiration in a mountainous watershed. The DAE reconstructed the sensor variables with a mean coefficient of determination value of 0.77 across 15 dimensions representing individual sensors. It reduced variance and bias uncertainties compared to a classical autoencoder model. The reconstruction quality varied across dimensions depending on their cross‐correlation and alignment with the underlying data structure. Uncertainties arising from the model structure were overall higher than those resulting from data corruption. We attached the DAE structure to a downstream ET‐prediction neural network in three formats and achieved reasonably accurate ET predictions . The use of the DAE notably reduced variance uncertainty in ET prediction. However, excessive variance reduction may be accompanied by an increase in bias due to the intrinsic bias‐variance tradeoff. Our method of evaluating and reducing uncertainties in aggregated data from different sources can be used to improve predictive models, process understanding, and uncertainty quantification for better water resource management. Plain Language Summary We present a machine learning method, namely the denoising autoencoder, which reduces the effects of data noise and missing values typically present in scientific data sets collected through sensor measurements. This method selects the most relevant information from noisy raw data collected by the instruments and fills in missing values. To demonstrate the effectiveness of our method, we applied it to predict evapotranspiration, a hydrologic variable that represents the water moved from the land surface to the atmosphere through a combination of evaporation and plant water use (transpiration). We also used a random sampling technique (the Monte Carlo method) to compare the uncertainty in the predictions when using the raw and noisy data versus the reconstructed data. The denoising process produced more accurate predictions of evapotranspiration with less uncertainty. Improved predictions of evapotranspiration can lead to a better understanding and accounting of water budgets. This ML approach is broadly suitable for a wide variety of applications that involve noisy sensor data with missing values. Key Points We used a denoising autoencoder (DAE) neural network to reduce noise in meteorological and soil sensor observations by on average We used Monte Carlo sampling to estimate the bias and variance of all model outputs, including uncertainty sources from data and the model We attached the DAE component to a downstream neural network to predict ET with the variance reduced by , compared to that without the DAE

denoising autoencoder↗

A Comparison of Pre‐Construction and Operational Wake Loss Estimates for Land‐Based Wind Plants

The overall bias between pre‐construction energy yield assessment (EYA) estimates of wind plant energy production and the achieved operational production is improving in the wind industry, but uncertainty remains high for individual wind plants. Wake effects within wind plants are one of the largest sources of energy loss considered in the EYA process, and previous work shows wake loss estimates to be a major source of disagreement among wind energy consultants who perform EYAs. To better understand the accuracy of wake loss predictions, we compare overall operational wake loss estimates based on supervisory control and data acquisition data to pre‐construction estimates provided by six wind energy consultants for five land‐based wind plants in North America. By augmenting existing approaches for quantifying operational wake losses, we estimate wake losses during the period of record for which operational data are available as well as the expected long‐term wake losses, based on historical reanalysis weather data, to which the EYA estimates are compared. To account for power variations at different turbine locations caused by terrain‐induced wind resource heterogeneity, we correct the operational wake loss estimates using predicted freestream wind speed variations from the Wind Systems Engineering Reynolds‐averaged Navier–Stokes (RANS) tool. We identify long‐term corrected operational wake losses between 1.9% and 6.4% for the five plants, with a mean loss of 4%. For the project deemed most acceptable for operational wake loss assessment, which is located in the simplest terrain and isolated from neighboring plants, the mean EYA wake loss estimate is within 0.7 percentage points of the operational value of 6.4%. For most of the remaining plants, results suggest that wake losses are generally overpredicted by 2.6–6.3 percentage points. However, operational wake losses may be underestimated for many of these projects because of spatial wind resource variations not captured by the RANS model, external wake effects that are unaccounted for in the estimation process, and wind plant blockage effects. To better understand factors that contribute to the observed wake losses, we investigate operational wake losses as a function of wind direction and wind speed. As expected, wake losses are generally concentrated near wind directions that are aligned with rows of closely spaced turbines and at below‐rated wind speeds; however, for some projects, the energy produced by the wind plant exceeds the estimated potential energy of the plant without wake interactions for certain wind directions and wind speeds, suggesting inaccurate assumptions in the wake loss estimation method for those plants. Lastly, we compare predicted and operational wake losses for individual wind turbines, finding that even when overall wake losses are predicted accurately, large uncertainty exists at the turbine level.

17 WIND ENERGY↗

From soil to sequence: filling the critical gap in genome-resolved metagenomics is essential to the future of soil microbial ecology

Abstract Soil microbiomes are heterogeneous, complex microbial communities. Metagenomic analysis is generating vast amounts of data, creating immense challenges in sequence assembly and analysis. Although advances in technology have resulted in the ability to easily collect large amounts of sequence data, soil samples containing thousands of unique taxa are often poorly characterized. These challenges reduce the usefulness of genome-resolved metagenomic (GRM) analysis seen in other fields of microbiology, such as the creation of high quality metagenomic assembled genomes and the adoption of genome scale modeling approaches. The absence of these resources restricts the scale of future research, limiting hypothesis generation and the predictive modeling of microbial communities. Creating publicly available databases of soil MAGs, similar to databases produced for other microbiomes, has the potential to transform scientific insights about soil microbiomes without requiring the computational resources and domain expertise for assembly and binning.

59 BASIC BIOLOGICAL SCIENCES↗

Controls From Above and Below: Snow, Soil, and Steepness Drive Diverging Trends of Subsurface Water and Streamflow Dynamics

ABSTRACT The importance of subsurface water dynamics, such as water storage and flow partitioning, is well recognised. Yet, our understanding of their drivers and links to streamflow generation has remained elusive, especially in small headwater streams that are often data‐limited but crucial for downstream water quantity and quality. Large‐scale analyses have focused on streamflow characteristics across rivers with varying drainage areas, often overlooking the subsurface water dynamics that shape streamflow behaviour. Here we ask the question: What are the climate and landscape characteristics that regulate subsurface dynamic storage, flow path partitioning, and dynamics of streamflow generation in headwater streams? To answer this question, we used streamflow data and a widely‐used hydrological model (HBV) for 15 headwater catchments across the contiguous United States. Results show that climate characteristics such as aridity and precipitation phase (snow or rain) and land attributes such as topography and soil texture are key drivers of streamflow generation dynamics. In particular, steeper slopes generally promoted more streamflow, regardless of aridity. Streams in flat, rainy sites (< 30% precipitation as snow) with finer soils exhibited flashier regimes than those in snowy sites (> 30% precipitation as snow) or sites with coarse soils and deeper flow paths. In snowy sites, less weathered, thinner soils promoted shallower flow paths such that discharge was more sensitive to changes in storage, but snow dampened streamflow flashiness overall. Results here indicate that land characteristics such as steepness and soil texture modify subsurface water storage and shallow and deep flow partitioning, ultimately regulating streamflow response to climate forcing. As climate change increases uncertainty in water availability, understanding the interacting climate and landscape features that regulate streamflow will be essential to predict hydrological shifts in headwater catchments and improve water resources management.

Kerins, Devon [Department of Civil and Environment↗

The Role of Bedrock Circulation Depth and Porosity in Mountain Streamflow Response to Prolonged Drought

Quantitative understanding is lacking on how the depth of active groundwater circulation in bedrock affects mountain streamflow response to a multi-year drought. We use an integrated hydrological model to explore the sensitivity of a variety of streamflow metrics to bedrock circulation depth and porosity under a plausible extreme drought scenario lasting up to 5 years. Endmember depth versus hydraulic conductivity relationships and porosity values for fractured crystalline rock are simulated. With drought, a deeper circulation system with higher drainable porosity more effectively buffers minimum flow and significantly limits perennial stream loss in comparison to a shallow circulation system. Streamflow buffering is accomplished through extensive groundwater storage loss. However, deeper circulation systems experience prolonged recovery from drought in comparison to storage-limited shallow systems. Research highlights the importance of characterizing the deeper bedrock hydrogeology in mountainous watersheds to better understand and predict drought impacts on stream ecosystem health and water resource sustainability.

54 ENVIRONMENTAL SCIENCES↗

Electronic structure prediction of multi-million atom systems through uncertainty quantification enabled transfer learning

The ground state electron density — obtainable using Kohn-Sham Density Functional Theory (KS-DFT) simulations — contains a wealth of material information, making its prediction via machine learning (ML) models attractive. However, the computational expense of KS-DFT scales cubically with system size which tends to stymie training data generation, making it difficult to develop quantifiably accurate ML models that are applicable across many scales and system configurations. Here, we address this fundamental challenge by employing transfer learning to leverage the multi-scale nature of the training data, while comprehensively sampling system configurations using thermalization. Our ML models are less reliant on heuristics, and being based on Bayesian neural networks, enable uncertainty quantification. We show that our models incur significantly lower data generation costs while allowing confident — and when verifiable, accurate — predictions for a wide variety of bulk systems well beyond training, including systems with defects, different alloy compositions, and at multi-million-atom scales. Moreover, such predictions can be carried out using only modest computational resources.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Quantum Tensor-Product Decomposition from Choi-State Tomography

The Schmidt decomposition is the go-to tool for measuring bipartite entanglement of pure quantum states. Similarly, it is possible to study the entangling features of a quantum operation using its operator-Schmidt or tensor-product decomposition. While quantum technological implementations of the former are thoroughly studied, entangling properties on the operator level are harder to extract in the quantum computational framework because of the exponential nature of sample complexity. Here, we present an algorithm for unbalanced partitions into a small subsystem and a large one (the environment) to compute the tensor-product decomposition of a unitary the effect of which on the small subsystem is captured in classical memory, while the effect on the environment is accessible as a quantum resource. This quantum algorithm may be used to make predictions about operator nonlocality and effective open quantum dynamics on a subsystem, as well as for finding low-rank approximations and low-depth compilations of quantum circuit unitaries. We demonstrate the method and its applications on a time-evolution unitary of an isotropic Heisenberg model in two dimensions. Published by the American Physical Society 2024

Mansuroglu, Refik (ORCID:000000017352513X)↗

Water use of co‐occurring loblolly ( Pinus taeda ) and shortleaf ( Pinus echinata ) in a loblolly pine plantation in the Piedmont

Abstract Measuring water use in co‐occurring loblolly pine (Pinus taedaL.) and shortleaf pine (Pinus echinataMill.) enhances our understanding of their competitive water use and aids in refining watershed water budget model parameters. This study was conducted in a 12‐ha forested headwater catchment in the Piedmont of North Carolina, southeastern U.S., from 2018 to 2019 (pre‐thinning) to 2020 (post‐thinning). Sap flux density (J s ), species‐level transpiration (T s ), and watershed‐level transpiration (T w ) were quantified. Water use efficiency (WUE) in loblolly and shortleaf pines was compared, alongside an investigation into how both species'J s andT s responded to atmospheric vapor pressure deficit (VPD). Loblolly pine had 19%–36% higherJ s than shortleaf pine. DailyT s for loblolly pine ranged from 15.0 to 29.0 L/day whileT s in shortleaf pine ranged from 3.0 to 6.8 L/day. TheT s was significantly higher in loblolly pine when compared to shortleaf pine likely due to higher canopy position and higher growth rates of the former. WUE, defined by annual tree biomass growth per tree water use, was not significantly different between the two. DailyJ s andT s in both species responded nonlinearly to VPD, with loblolly pine being more sensitive and variable. Species‐specific water use should be considered when quantifyingT w and developing reliable models to predict the effects of forest management practices on water resources.

Engineering↗

A reproducible study design for the MIMIC-IV in-hospital mortality task

Open, tabular electronic health record (EHR) datasets such as MIMIC-III and MIMIC-IV have become critical resources for developing machine learning (ML) models addressing clinical prediction tasks, including hospital readmission, length of stay, and in-hospital mortality (IHM). While MIMIC-III has benefited from well-established preprocessing pipelines and standardized feature sets, MIMIC-IV remains comparatively challenging to work with because there are no standardized benchmarks to support reproducibility and comparability across studies. To address this limitation, we present a rigorously curated MIMIC-IV custom feature set optimized for IHM prediction, constructed through a reproducible preprocessing pipeline and feature selection strategy.

97 MATHEMATICS AND COMPUTING↗

Validating Greater Sage-Grouse Individual-based Model (IBM) Tool (Final Report)

The project focused on validating the previously developed Greater Sage-Grouse Individual-based Model (GrSG IBM; LaGory et al. 2012, 2021). The objective was to transform this predictive, spatially and temporally explicit model into a portable resource to assist siting/resource managers in proactively assessing the cumulative impacts of wind energy development on the greater sage-grouse. Utilizing a bottom-up, individual-based approach, the GrSG IBM accounts for landscape context and species behavior, aiming to reduce uncertainty in estimating development impacts and support ecologically mindful land-based wind energy development. The validation effort covered approximately 6,540 km 2 near the Seven Mile Hill Wind Project in Wyoming. The GrSG IBM tool, built on the NetLogo platform (Tisue and Wilensky 2004), was executed over a 50-year period, with the analysis focusing on years following a 10-year initialization phase. Key results demonstrated the tool’s biological soundness across five key biological metrics: non-chick age class distribution (older than 10 weeks), adult sex ratio, life expectancy, population size, and overall population growth. For instance, the tool estimated that 58.6% of the non-chick population was reproductively immature, while the reference ranges from 51.4% to 57.8% (Patterson 1952, Rogers 1964). Experts confirmed the tool’s estimate was within a reasonable range for the species. The tool estimated average life expectancy of 1.43 years, while the reference ranges from 0.9 years to 1.1 years (Ammann 1957, Hamerstrom 1949). Experts also supported the model’s life-expectancy estimate as ecologically sound for the species in the study area. In terms of population change, the model estimated an annual shift between a 0.6% decline and a 1.0% increase over 50 years. While the reference suggests 2.9% annual decline in range-wide populations (Cortes et al. 2023), that includes many at-risk populations in South Dakota and Washington, for example. Our study area—in the northeastern part of Carbon County and western-edge of Albany County, Wyoming—is one of the remaining greater sage-grouse habitats supporting some of the most stable populations. Experts confirmed that the range of the annual population change spanning from a 0.6% decline to a 1.0% increase estimated by the tool was reasonable for our study area for this reason and confirmed that aligned with population estimates from existing studies on the greater sage-grouse and wind energy development in the study area (LeBeau et al. 2017a, Smith et al. 2024). Furthermore, the project showed that temporally explicit biological metrics generated by the GrSG IBM tool can complement the USGS’ Prioritizing Restoration of Sagebrush Ecosystems Tool (PReSET; Duchardt et al. 2021) by incorporating habitat restoration strategies into seasonal habitat suitability models to visualize population responses over time.

17 WIND ENERGY↗

CrossMP: Enabling Cross-Modality Translation between Single-Cell RNA-Seq and Single-Cell ATAC-Seq through Web-Based Portal

In recent years, there has been a growing interest in profiling multiomic modalities within individual cells simultaneously. One such example is integrating combined single-cell RNA sequencing (scRNA-seq) data and single-cell transposase-accessible chromatin sequencing (scATAC-seq) data. Integrated analysis of diverse modalities has helped researchers make more accurate predictions and gain a more comprehensive understanding than with single-modality analysis. However, generating such multimodal data is technically challenging and expensive, leading to limited availability of single-cell co-assay data. Here, we propose a model for cross-modal prediction between the transcriptome and chromatin profiles in single cells. Our model is based on a deep neural network architecture that learns the latent representations from the source modality and then predicts the target modality. It demonstrates reliable performance in accurately translating between these modalities across multiple paired human scATAC-seq and scRNA-seq datasets. Additionally, we developed CrossMP, a web-based portal allowing researchers to upload their single-cell modality data through an interactive web interface and predict the other type of modality data, using high-performance computing resources plugged at the backend.

59 BASIC BIOLOGICAL SCIENCES↗

Climatic imprint on interfacially-controlled platinum-palladium resources

Abstract Iron oxide-rich laterites, soils, and regolith formed from the weathering of ultramafic rocks represent untapped unconventional resources for the critical minerals platinum and palladium, but the fundamental surficial geochemistry of these elements remains poorly understood. Depletion of Pd relative to Pt occurs in some weathering zones in semi-arid climates. The accepted model attributes this platinum-palladium chemical fractionation to preferential complexation of Pd by dissolved chloride. However, similar fractionation is not observed in laterites of humid equatorial regions despite substantial wet deposition of chloride. The established mechanistic model for Pt and Pd behavior during weathering thus inaccurately predicts the distribution of these critical minerals in many settings, hindering global resource assessment. We show through mineral-fluid partitioning experiments coupled to element-specific spectroscopy that this canonical explanation for platinum-palladium fractionation is invalid: chloride complexation does not differentially mobilize Pd versus Pt. Instead, mineral-specific interfacial reactions control Pd and Pt accumulation. Modeling of platinum-palladium fractionation in representative weathering zone profiles demonstrates sub-equal retention in goethite-rich settings and Pd depletion in hematite-rich zones, accurately predicting trends observed in soils and laterites. Iron oxide mineralogy, reflecting modern and past regional climate conditions, is likely the primary determinant of Pt and Pd endowment in weathering zone resources. This new model for Pt and Pd mobilization and accumulation behavior provides a mechanistic foundation for exploration and recovery of platinum group elements from novel ultramafic regolith deposits.

58 GEOSCIENCES↗

Queue wait time prediction in high performance computing (HPC) systems

High Performance Computing (HPC) systems are critical enablers for groundbreaking scientific research across various domains. Efficient resource allocation, facilitated by job scheduling, is paramount for maximizing the utilization of HPC systems. However, the variability in wait times for queued jobs poses challenges for users, necessitating accurate job wait time estimation. This paper explores the influence of job characteristics, including job size (the number of nodes requested and walltime), the queue to which the job is submitted and other resource requirements, on job wait times in leadership-class HPC systems. Focusing on the Theta Cray XC40 and Polaris machines at Argonne National Laboratory, the study evaluates the performance of different supervised learning algorithms in predicting job wait times. It also evaluates the impact of data preprocessing, including outlier detection, Principal Component Analysis (PCA), and feature selection, on the performance of wait time prediction models. The findings reveal insights into the relationship between job characteristics and wait times, offering a foundation for optimizing resource allocation and enhancing user experience. The methodologies and tools developed in this study are adaptable to other leadership-class HPC systems, providing a valuable contribution to the broader HPC community aiming to improve job scheduling efficiency and user satisfaction.

Okafor, Nwamaka↗

LHC EFT WG note: SMEFT predictions, event reweighting, and simulation

This note provides a comprehensive overview of tools for predicting observables in the Standard Model effective field theory (SMEFT) at both tree level and one loop using event generators. We evaluate three primary methodologies–event reweighting, separate simulation of squared matrix elements, and full SMEFT process simulation–focusing on their statistical performance, computational efficiency, and potential biases. Each approach is assessed in terms of its accuracy, highlighting trade-offs between precision and resource demands. Practical insights into their applicability for high-energy physics analyses are offered, with particular attention to processes where SMEFT effects are significant. Additionally, we discuss the role of helicity in reweighting strategies and its impact on the quality of predictions. By comparing the methods across various LHC processes, this note provides guidance for selecting the most effective strategy for various SMEFT studies, ensuring robust predictions while optimizing computational resources.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A global soil plasmidome resource unveils functional and ecological roles of plasmids in soil microbiomes

Plasmids play significant roles in microbial adaptation to ecosystems, yet their dynamics remain poorly understood due to identification challenges. We present the Global Soil Plasmidome Resource (GSPR), a comprehensive dataset of 98,728 plasmid sequences amassed from 6860 terrestrial microbial communities and isolates. We explore this resource through various computational approaches, including phylogenetic diversity analysis, host prediction, and extensive functional annotation, to understand the contribution of plasmids to the genetic and functional diversity in soil, correlating these findings with sample type, as well as the soil habitat they were retrieved from. Our analysis reveals insights into plasmid-encoded functions such as effector modules, quorum sensing, and stress resistance, which may contribute to their persistence and microbial adaptation in soil. Furthermore, CRISPR analysis suggests a prevalent role of these elements related to intra-plasmid competition. By contrasting plasmids from cultivated and uncultivated organisms, we identify important functions that expand existing knowledge of plasmid roles in these habitats. This study represents a notable step forward in elucidating plasmid diversity and function within soil microbiomes and establishes a foundational framework for exploring their roles in natural environments.

Fiamenghi, Mateus B↗

An accelerated framework for predicting creep rupture lifetimes in engineering alloys

Confidently predicting high-temperature deformation, including creep and creep rupture, is paramount for the design and commercialization of candidate materials for advanced nuclear energy systems. To accelerate creep quantification, we introduce a framework that enables rapid, cost-effective, and reliable prediction of creep rupture lifetimes, minimizing reliance on time-intensive bulk creep testing. Unlike conventional creep analysis, which requires extensive time and resources, our method leverages a maximum of four short-term bulk creep tests as training data for prediction. This framework combines high-throughput nanoindentation up to 700 °C with these targeted bulk tests to inform our creep rupture model in order to predict rupture lifetimes. The strong agreement between our predictions and conventional experimental data demonstrates the effectiveness of our approach for accelerated creep analysis and lifetime prediction of structural components in high-temperature applications. Our multi-pronged approach motivates further integration of computational tools and advanced instrumentation to establish a universal framework for understanding high-temperature material responses.

36 MATERIALS SCIENCE↗

Machine Learning Models for Mapping Groundwater Pollution Risk: Advancing Water Security and Sustainable Development Goals in Georgia, USA

The widespread use of pesticides, such as atrazine and malathion, in agricultural systems raises significant concerns regarding the contamination of groundwater, which serves as a critical resource for drinking water. This study applies machine learning techniques to predict the concentrations of atrazine and malathion in groundwater across Georgia, USA, using 2019 data. A Random Forest classifier was employed to integrate various environmental and demographic factors, including pesticide application rates, precipitation, lithology, and population density, to predict pesticide contamination in groundwater. The models demonstrated high training accuracies of 100% and moderate average testing accuracy of 55% for atrazine and 60% for malathion across five iterations. The low test accuracy of the model, ranging from 50% to 75%, is likely due to overfitting, which can be attributed to the small dataset size and the complex nature of pesticide-contamination patterns, making it challenging for the model to generalize to unseen data. Feature importance analysis revealed that average pesticide usage emerged as the most influential factor for atrazine, while aquifer lithology and precipitation played crucial roles in both models. These results provide valuable insights into the dynamics of pesticide contamination, highlighting areas at greater risk of contamination. The findings underscore the importance of integrating environmental, geological, and agricultural variables for more effective groundwater management and sustainable agricultural practices, contributing to the protection of water resources and public health.

54 ENVIRONMENTAL SCIENCES↗