Search NASASearch

SEARCH · Search NASA

Results for “Well Log Correlation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Machine Learning for Well Log Analysis in Uranium Mining

This project explores the use of Artificial Intelligence (AI) and Machine Learning (ML) techniques to automate well log analysis for uranium mining. Geophysical log data—spontaneous potential, resistivity, and gamma ray—were used to classify lithology, correlate well logs and identify roll front zonation patterns, which are critical for locating uranium ore bodies. Supervised ML algorithms such as eXtreme Gradient Boosting (XGBoost), Categorical Boosting (CatBoost), and Random Forest were trained to classify lithology with high accuracy. Gradient Boosting Machines (GBM), XGBoost, Random Forest, and Neural Networks were also used for role front zone identification. Moreover, a Fast Dynamic Time Warping (FastDTW) algorithm was employed for well log correlation. Additionally, sample lag was addressed using dynamic programming. Results demonstrate the potential of AI and ML to streamline well log analysis and enhance uranium exploration workflows.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS

Facies Analysis of the Prairie Du Chien Group in the Illinois Basin and Analogous Rocks in Missouri and Kentucky

Funded in 2023 by the U.S. Department of Energy’s Phase II Carbon Storage Assurance Facility Enterprise (CarbonSAFE) initiative, a Heidelberg Materials cement plant in Mitchell, Indiana, is currently being evaluated as a potential Carbon Capture and Storage (CCS) subsurface injection site. The Heidelberg CCS project targets the middle to upper Prairie du Chien Group (Early Ordovician) in southwestern Indiana. Assessment of reservoir feasibility requires collection of field data, seismic surveys, well-log correlation, geologic modeling, characterization well drilling, well testing, and reservoir simulation. However, the proposed Heidelberg CCS site is in a data-limited region, lacking both outcrop analogs and deep wells penetrating the target interval, which makes geologic modelling difficult prior to drilling a characterization well. To directly address this problem, the present study was undertaken to understand the sedimentologic composition and stratigraphic architecture of the Prairie du Chien Group from analogous outcrops and cores in the Illinois Basin and adjacent regions.

Ali, Shah Bilawal [Univ. of Illinois at Urbana-Cha

Evaluation of Drilling Performance at The Geysers with Machine Learning Methods Using Geologic Data

A recent well, GDC-36, was drilled in The Geysers Geothermal Field served in a Department of Energy-industry to demonstrate improved drilling performance with polycrystalline diamond compact (PDC) bits. Both PDC and roller cone drill bits were used to drill this well. Key challenges encountered during drilling included lost circulation in the mud-drilled section, and bit damage interfacial severity in the deeper, air-drilled section. The objective of this study is to evaluate the drilling performance in relation to the local geological characteristics using machine learning methods. By applying K-clustering to the sonic log data, we were able to identify areas correlated with measured lost circulation. Also, the boundaries defined by clustering of the mineralogical and lithological data from the mud logs correlate well with interfacial severity during drilling. A random forest model was employed to build correlation between drilling data and rock strength. The confined compressive strength (CCS) of the rock in the training of the machine learning model was inferred from the dipole sonic log. The R-squared of the testing data is 0.78, and the RMSE (Root Mean Squared Error) is 0.06. The trained model was used to forecast rock strength for the section where sonic log data are not available. CCS could also be inferred from mud logs provided the relationship between mineralogy and rock strength is established through core testing data.

15 GEOTHERMAL ENERGY

Structural Evolution of the Hogback Monocline and Its Tectonic Significance in the San Juan Basin

The San Juan Basin is recognized as a Laramide foreland basin. It is located within the Colorado Plateau, a broad tectonic province characterized by a thick sedimentary sequence that was segmented into smaller sub basins during the Late Cretaceous to Paleogene Laramide orogeny. The Hogback Monocline lies along the northwestern margin of the San Juan Basin and is considered a Laramide-age structure formed in response to compressional stress. In this study, we interpret surface and subsurface datasets to construct a structural geological model and evaluate its tectonic significance. Through seismic data, we identify key fault and fold geometries at depth. The seismic dataset used in this study was reprocessed in depth and constrained with well log velocity data to enhance seismic imaging quality. Additionally, we performed well log correlations to identify formation tops and assess variations in basin infill and thickness geometry. A series of structural cross-sections, constructed using seismic data and a high density of boreholes, are presented to evaluate geometric variations along the structure and its evolution during basin development. Furthermore, kinematic restoration and forward modeling analyses were conducted to validate our structural interpretation. This work suggests that the Hogback Monocline formed through fault-propagation folding and flexural slip affecting the pre-Laramide sedimentary sequence under compressional stresses associated with the Laramide orogeny. This structure is interpreted as a high-angle reverse fault that influenced the geometry of the late basin infill. Additionally, monocline bending along the structure may have been controlled by fault relay systems and, in some cases, influenced by strike-slip faulting.

Reyes, Martin [New Mexico Bureau o fGeology and Mi

Leveraging Natural Language Processing and Generative Models in Molecular Chemistry: Property Prediction and Novel Compound Generation

The accurate prediction of molecular properties is important for the rational design and the advancement of green chemistry and sustainable materials research. However, the predictive power of traditional computational chemistry methods is limited due to computational restrictions. Here, in this study, we examine an alternative approach to the accurate prediction of properties of organic compounds: natural language processing (NLP)-based molecular embedding. Using viscosity, partition coefficient (log P), and enthalpy of vaporization as test properties through a survey of comprehensive datasets comprising 5695 data points for viscosity, 25 870 data points for log P, and 2296 data points for enthalpy of vaporization. These are important properties for the design of greener, safer, and sustainable chemical processes. Models were trained using NLP methods such as Mol2vec and fine-tuned ChemBERTa, and results were compared with traditional input featurization techniques such as Morgan fingerprints and quantum chemistry derived sigma profiles and DFT features. Among the various machine learning models, Mol2vec demonstrated superior predictive capabilities, achieving the highest correlation coefficient (R 2 = 0.945) and lowest RMSE (0.106 mPa s) for viscosity, as well as high accuracy for log P and enthalpy of vaporization predictions. These findings establish the Mol2vec featurization technique, graph-convolutional neural networks (GCNN), and fine-tuned ChemBERTa model as powerful tools for predictive modeling of organic compounds properties, offering a significant improvement over previously used featurization techniques and opening up strategies for very-high-throughput computational screening. Finally, we integrated ML models with hybrid language-model-based generative adversarial networks (LM-GAN) to generate novel molecular sequences with desirable properties for different research applications. The ability to computationally design solvents with lower viscosity, lower log P, and lower enthalpy of vaporization offers a data-driven route to accelerating the discovery of sustainable alternatives to traditionally toxic solvents.

ChemBERTa

Optimizing Deep Geothermal Drilling for Energy Sustainability in the Appalachian Basin

This study investigates the geological and geomechanical characteristics of the MIP 1S geothermal well in the Appalachian Basin to optimize drilling and address the wellbore stability issues encountered. Data from well logs, sidewall core analysis, and injection tests were used to derive elastic and rock strength properties, as well as stress and pore pressure profiles. A robust 1D-geomechanical model was developed and validated, correlating strongly with wellbore instability observations. This revealed significant wellbore breakout, widening the diameter from 12 ¼ inches to over 16 inches. Advanced technologies like Cerebro Force™ In-Bit Sensing were used to monitor drilling performance with high accuracy. This technology tracks critical metrics such as bit acceleration, vibration in the x, y, and z directions, Gyro RPM, stick-slip indicators, and bending on the bit. Cerebro Force™ readings identified hole drag caused by poor hole conditions, including friction between the drill string and wellbore walls and the presence of cuttings or debris. This led to higher torque and weight on bit (WOB) readings at the surface compared to downhole measurements, affecting drilling efficiency and wellbore stability. Optimal drilling parameters for future deep geothermal wells were determined based on these findings.

Environmental Sciences & Ecology

Utah FORGE: Well 16B(78)-32 Drill Core Fracture Analysis Images and Data

This dataset contains drilling core data from well 16B(78)-32, including PDF documents with flattened core images annotated by feature type and core interval, as well as spreadsheets detailing feature morphologies by depth, planar feature measurements, and planar feature orientations rotated to in situ conditions. Core was recovered from three intervals, one per stimulation stage, in the crystalline rocks affected by the stimulation of well 16A(78)-32. Seven core runs were conducted, yielding 135.8 feet of recovered core. Features in the core were categorized into planar fractures, semi-planar fractures, unbroken mineralized fractures, rough fractures, curviplanar fractures, concave-convex surfaces, and planar compositional features such as mylonite or dike-like structures. Planar features were measured while the core was positioned horizontally, with the core axis aligned to a downhole azimuth of 42 degrees. Planar core measurements from stimulations 2 and 3 that could be confidently correlated with FMI data were rotated to in situ orientations. This was done by rotating the planes along vertical and horizontal axes to match the azimuth and inclination data recorded in the directional survey of well 16B(78)-32, as well as applying an axial rotation to resemble the fracture orientations observed in the FMI log at corresponding depths. Coherent sets of planar fracture measurements were made by aligning the core within each 3-foot section of the dissected core barrel, and between adjacent 3-foot sections within a core run by matching rock fabrics, saw cuts and/or tool marks. Where coherent fracture measurements could not be made within a core run, data sets are denoted by a subscript (i.e. 2-Ta and 2-Tb both come from tangent core run number 2).

15 GEOTHERMAL ENERGY

LTAU-FF: Loss Trajectory Analysis for Uncertainty in atomistic Force Fields

Model ensembles are effective tools for estimating prediction uncertainty in deep learning atomistic force fields. However, their widespread adoption is hindered by high computational costs and overconfident error estimates. In this work, we address these challenges by leveraging distributions of per-sample errors obtained during training and employing a distance-based similarity search in the model latent space. Our method, which we call LTAU (Loss Trajectory Analysis for Uncertainty), efficiently estimates the full probability distribution function of errors for any test point using the logged training errors, achieving speeds that are 2–3 orders of magnitudes faster than typical ensemble methods and allowing it to be used for tasks where training or evaluating multiple models would be infeasible. We apply LTAU towards estimating parametric uncertainty in atomistic force fields (LTAU-FF), demonstrating that it produces well-calibrated confidence intervals and predicts errors that correlate strongly with the true errors for data near the training domain. Furthermore, we show that the errors predicted by LTAU-FF can be used in practical applications for detecting out-of-domain data, tuning model performance, and predicting failure during simulations. We believe that LTAU will be a valuable tool for uncertainty quantification in atomistic force fields and is a promising method that should be further explored in other domains of machine learning.

97 MATHEMATICS AND COMPUTING

Heterogeneous Multi-Domain Dataset Synthesis to Facilitate Privacy and Risk Assessments in Smart City IoT

The emergence of the Smart Cities paradigm and the rapid expansion and integration of Internet of Things (IoT) technologies within this context have created unprecedented opportunities for high-resolution behavioral analytics, urban optimization, and context-aware services. However, this same proliferation intensifies privacy risks, particularly those arising from cross-modal data linkage across heterogeneous sensing platforms. To address these challenges, this paper introduces a comprehensive, statistically grounded framework for generating synthetic, multimodal IoT datasets tailored to Smart City research. The framework produces behaviorally plausible synthetic data suitable for preliminary privacy risk assessment and as a benchmark for future re-identification studies, as well as for evaluating algorithms in mobility modeling, urban informatics, and privacy-enhancing technologies. As part of our approach, we formalize probabilistic methods for synthesizing three heterogeneous and operationally relevant data streams—cellular mobility traces, payment terminal transaction logs, and Smart Retail nutrition records—capturing the behaviors of a large number of synthetically generated urban residents over a 12-week period. The framework integrates spatially explicit merchant selection using K-Dimensional (KD)-tree nearest-neighbor algorithms, temporally correlated anchor-based mobility simulation reflective of daily urban rhythms, and dietary-constraint filtering to preserve ecological validity in consumption patterns. In total, the system generates approximately 116 million mobility pings, 5.4 million transactions, and 1.9 million itemized purchases, yielding a reproducible benchmark for evaluating multimodal analytics, privacy-preserving computation, and secure IoT data-sharing protocols. To show the validity of this dataset, the underlying distributions of these residents were successfully validated against reported distributions in published research. We present preliminary uniqueness and cross-modal linkage indicators; comprehensive re-identification benchmarking against specific attack algorithms is planned as future work. This framework can be easily adapted to various scenarios of interest in Smart Cities and other IoT applications. By aligning methodological rigor with the operational needs of Smart City ecosystems, this work fills critical gaps in synthetic data generation for privacy-sensitive domains, including intelligent transportation systems, urban health informatics, and next-generation digital commerce infrastructures.

IoT

Modeling of the high column density systems in the Lyman-Alpha forest

The Lyman-α forests observed in the spectra of high-redshift quasars can be used as a tracer of the cosmological matter density to study baryon acoustic oscillations (BAO) and the Alcock-Paczynski effect. Extraction of cosmological information from these studies requires modeling of the forest correlations. While the models depend most importantly on the bias parameters of the intergalactic medium (IGM), they also depend on the numbers and characteristics of high-column-density systems (HCDs) ranging from Lyman-limit systems with column densities log N HI /1cm -2 > 17 to damped Lyman-α systems (DLAs) with log N HI /1cm -2 > 20.2. These HCDs introduce broad damped absorption characteristic of a Voigt profile. Consequently they imprint a component on the power spectrum whose modes in the radial direction are suppressed, leading to a scale-dependent bias. Using mock data sets of known HCD content, we test a model that describes this effect in terms of the distribution of column densities of HCDs, the Fourier transforms of their Voigt profiles and the bias of the halos containing the HCDs. Our results show that this physically well-motivated model describes the effects of HCDs with an accuracy comparable to that of the ad-hoc models used in published forest analyses. We also discuss the problems of applying the model to real data, where the HCD content and their bias is uncertain.

Lyman alpha forest

Development of the United States GReenhouse Gas and Air Pollutants Emissions System (GRA 2 PES)

In the U.S., emissions of greenhouse gases and air pollutants are often developed independently. Here, we describe the GReenhouse gas And Air Pollutants Emissions System (GRA 2 PES), which provides gridded emissions of fossil-fuel carbon dioxide (ffCO 2 ) and 93 air quality (AQ) species for 17 combustion and non-combustion sectors at 4 km × 4 km spatial resolution across the contiguous US. We find that the AQ emissions most spatially correlated with ffCO 2 are nitrogen oxides (NO x , ρ = 0.67), followed by sulfur dioxide (SO 2 , ρ = 0.51), carbon monoxide (CO, ρ = 0.44), and fine particulate matter (PM 2.5 , ρ = 0.38). We evaluate GRA 2 PES ffCO 2 emissions with an ensemble of publicly available regional and global inventories at national (Normalized Mean Bias (NMB) = +1.4%), state (NMB = +1.5%, R 2 = 0.98), and urban (NMB = +11.5%, R 2 = 0.97) scales. Nationally, the differences of publicly available inventories from the ensemble average range from −10.0% to +5.7%, and consistency diverges at state and urban scales. We simulate GRA 2 PES ffCO 2 in a particle dispersion model and compare to measurements of radiocarbon ( 14 C)-derived ffCO 2 collected in Los Angeles (August 2021), with results suggesting that GRA 2 PES ffCO 2 may be low by 19% for this city, but well within model-observation differences for other publicly available inventories (−43% to +94%). GRA 2 PES AQ/ffCO 2 ratios converted to concentration space generally agree with field observations (NMB = +4%, log R 2 = 0.90). Lastly, we present a method by which to utilize GRA 2 PES to derive AQ emission fluxes from ffCO 2 emissions.

Lyu, Congmeng [National Oceanic and Atmospheric Ad