Search NASA⌕ Search

SEARCH · Search NASA

Results for “Geostatistics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Accelerating geostatistical modeling using geostatistics-informed machine Learning

Ordinary Kriging (OK) is a popular geostatistical algorithm for spatial interpolation and estimation. The computational complexity of OK changes quadratically and cubically for memory and speed, respectively, given the number of data. Therefore, it is computationally intensive and also challenging to process a large set of data, especially in three-dimensional (3D) cases. This paper develops a geostatistics-informed machine learning (GIML) model to improve the efficiency of OK by reducing the number of points required to be estimated using OK. Specifically, only a very few of the unknown points are estimated by OK to get the weights and estimations, which are used as the training dataset. Moreover, the governing equations of OK are used to guide our proposed machine learning to better reproduce the spatial distributions. Our results show that the proposed GIML can reduce the computational time of OK by at least one order of magnitude. The effectiveness of the GIML is evaluated and compared using a 2D case. Furthermore, we demonstrate its efficiency and robustness by considering a different number of training samples on various 3D simulation grids.

58 GEOSCIENCES↗

Special Issue: Geostatistics and Machine Learning

Abstract Recent years have seen a steady growth in the number of papers that apply machine learning methods to problems in the earth sciences. Although they have different origins, machine learning and geostatistics share concepts and methods. For example, the kriging formalism can be cast in the machine learning framework of Gaussian process regression. Machine learning, with its focus on algorithms and ability to seek, identify, and exploit hidden structures in big data sets, is providing new tools for exploration and prediction in the earth sciences. Geostatistics, on the other hand, offers interpretable models of spatial (and spatiotemporal) dependence. This special issue on Geostatistics and Machine Learning aims to investigate applications of machine learning methods as well as hybrid approaches combining machine learning and geostatistics which advance our understanding and predictive ability of spatial processes.

58 GEOSCIENCES↗

Integration of Soft Data Into Geostatistical Simulation of Categorical Variables

Uncertain or indirect “soft” data, such as geologic interpretation, driller’s logs, geophysical logs or imaging, offer potential constraints or “soft conditioning” to stochastic models of discrete categorical subsurface variables in hydrogeology such as hydrofacies. Previous bivariate geostatistical simulation algorithms have not fully addressed the impact of data uncertainty in formulation of the (co) kriging equations and the objective function in simulated annealing (or quenching). This paper introduces the geostatistical simulation code tsim-s, which accounts for categorical data uncertainty through a data “hardness” parameter. In generating geostatistical realizations with tsim-s, the uncertainty inherent to soft conditioning is factored into both 1) the data declustering and spatial correlation functions in cokriging and 2) the acceptance probability for change of category in simulated quenching. The degree or sensitivity to which soft data conditions a realization as a function of hardness can be quantified by mapping category probabilities derived from multiple realizations. In addition to point or borehole data, arrays of data (e.g., as derived from a depth-dependency function, probability map, or “prior realization”) can be used as soft conditioning. The tsim-s algorithm provides a theoretically sound and general framework for integrating datasets of variable location, resolution, and uncertainty into geostatistical simulation of categorical variables. A practical example shows how tsim-s is capable of generating a large-scale three-dimensional simulation including curvilinear features.

54 ENVIRONMENTAL SCIENCES↗

Geostatistical interpolation of streambed hydrologic attributes with addition of left censored data and anisotropy

Spatial geostatistical interpolation of point measurements of streambed attributes in the hyporheic zone may be constrained by the streambed anisotropy, and data density and spatial distribution may significantly impact the results. Spatial clustering and low spatial data density can be caused by bedrock outcropping at the streambed limiting installation of in-stream piezometers. This study examines parameter error variability of the geostatistical interpolation using anisotropic interpolation methods and increasing the data density by adding left censored values (i.e., data below measurement limit) to locations where measurements were limited by exposed bedrock lining the streambed. The reduction in relative standard error of the interpolation was determined for the spatial distributions of streambed attributes including hydraulic conductivity, seepage flux, and mercury solute flux measured in two different years along a study reach in East Fork Poplar Creek, Tennessee, USA. Here, two methods to impute the left censored values were compared including the conventional half the detection limit substitution method, and the Stochastic Approximation of Expectation-Maximization (SAEM) algorithm, which both had comparable results. Imputing left censored data increased the data density to recommended ranges, reduced data clustering, increased the spatial dependence for some attributes, and reduced the standard error for each of the three attributes. For the reach considered herein, addition of the left censored values resulted in a larger error reduction than the consideration of anisotropy within the interpolation, which confirms the benefit of data addition to increase data density within data-limited river corridors.

58 GEOSCIENCES↗

Temperature uncertainty modelling with proxy structural data as geostatistical constraints for well siting: an example applied to Granite Springs Valley, NV, USA

Utilizing existing temperature and structural geology information around Granite Springs Valley, Nevada, we build 3D stochastic temperature models with the aims of evaluating the 3D uncertainty of temperature and choosing between candidate exploration well locations. The data used to support the modelling are measured temperatures and structural proxies from 3D geologic modelling (distance to fault, distance to fault intersections and terminations, Coulomb stress change and dilation tendency), the latter considered ‘secondary’ data. Two stochastic geostatistical techniques are explored for incorporating the structural proxies: cosimulation and local varying mean. With both the cosimulation and local varying mean methods, many equally-likely temperature models (i.e. realizations) are produced, from which temperature probability profiles are calculated at candidate well locations. To aid in choosing between the candidate locations, two quantities summarize the temperature probabilities: V prior and entropy. V prior quantifies the likelihood for economic temperatures at each candidate location, whereas entropy identifies where new information has the most potential to reduce uncertainty. In general, the cosimulation realizations have smoother spatial structure, and extrapolate high temperatures at candidate locations that are located along the direction of the longest spatial correlation, which are down dip from existing temperature logs. The smooth realizations result in tight temperature probability profiles that are easier to interpret, but they have unrealistic temperature reversals in some locations because of the dipping ellipsoid shape created and that the cosimulation technique does not enforce a conductive geothermal gradient as a baseline (i.e. linearly increasing temperature with depth). The local varying mean results produce realizations with more realistic geothermal gradients, with temperatures increasing downward since a depth-temperature relationship is included. However, because they have much noisier spatial nature compared to cosimulation, it is harder to interpret the temperature probability profiles. The different local varying mean results allow the geologist to determine which proxy (e.g. dilation v. distance to fault termination) should be used given the specific geothermal system. In general, V prior from local varying mean results identify locations that are close to high values for the structural proxies: areas with higher probabilities for higher temperatures. The entropy results identify where uncertainty is greatest and therefore new drilling information could be most useful. Though these techniques provide useful information, even when applied to areas of sparse data, our comparison of these two techniques demonstrates the need for new geothermal geostatistics techniques that combine the advantages of these two methods and that are tailored to the spatial uncertainty issues inherent in geothermal exploration.

15 GEOTHERMAL ENERGY↗

Temperature Uncertainty Modeling with Proxy Structural Data as Geostatistical Constraints for Well Siting: An Example Applied to Granite Springs Valley, NV, USA

Utilizing existing temperature and structural information around Granite Springs Valley, Nevada, we build 3D stochastic temperature models with the aim of evaluating the 3D uncertainty of temperature and choosing between candidate exploration well locations . The data used to support the modeling are measured temperatures and structural proxies from 3D geologic modeling, the latter considered "secondary" data. Two stochastic geostatistical techniques are explored for incorporating the structural proxies: cosimulation and local varying mean. With both the cosimulation and local varying mean methods, many equally likely temperature models (i.e., realizations) are produced, from which temperature probability profiles are calculated at candidate well locations. To aid in choosing between the candidate locations, two quantities summarize the temperature probabilities: Vprior and entropy. Vprior quantifies the likelihood for economic temperatures at each candidate location, whereas entropy identifies where new information has the most potential to reduce uncertainty. In general, the cosimulation realizations have smoother spatial structure, and extrapolate high temperatures at candidate locations that are located along the direction of the longest spatial correlation, which are down dip from existing temperature logs. The smooth realizations result in tight temperature probability profiles that are easier to interpret, but they have unrealistic temperature reversals in some locations because the cosimulation technique does not enforce a conductive geothermal gradient as a baseline (i.e., linearly increasing temperature with depth). The local varying mean results produce realizations with more realistic geothermal gradients, with temperatures increasing downward since a depth-temperature relationship is included. However, because they have much noisier spatial nature compared to cosimulation, it is harder to interpret the temperature probability profiles. The different local varying mean results allow the geologist to determine which proxy (e.g., dilation versus distance to fault termination) should be used given the specific geothermal system. In general, Vprior from local varying mean results identify locations that are close to high values for the structural proxies: areas with highe r probabilities for higher temperatures. The entropy results identify where uncertainty is greatest and therefore new drilling information could be most useful. Though these techniques provide useful information, even when applied to areas of sparse data, our comp arison of these two techniques demonstrates the need for new geothermal geostatistics techniques that combine the advantages of these two methods and that are tailored to the spatial uncertainty issues inherent in geothermal exploration.

3D temperature modeling↗

Variational encoder geostatistical analysis (VEGAS) with an application to large scale riverine bathymetry

Estimation of riverbed profiles, also known as bathymetry, plays a vital role in many applications, such as safe and efficient inland navigation, prediction of bank erosion, land subsidence, and flood risk management. The high cost and complex logistics of direct bathymetry surveys, i.e, depth imaging, have encouraged the use of indirect measurements such as surface flow velocities. However, estimating high-resolution bathymetry from indirect measurements is an inverse problem that can be computationally challenging. Here, we propose a reduced-order model (ROM) based approach that utilizes a variational autoencoder (VAE), a type of deep neural network with a narrow layer in the middle, to compress bathymetry and flow velocity information and accelerate bathymetry inverse problems from flow velocity measurements. In our application, the shallow-water equations (SWE) with appropriate boundary conditions (BCs), e.g., the discharge and/or the free surface elevation, constitute the forward problem, to predict flow velocity. Then, ROMs of the SWEs are constructed on a nonlinear manifold of low dimensionality through a variational encoder and the bathymetry inversion problem is derived on the low-dimensional latent space in a Hierarchical Bayesian setting. Further, the reformulation allows variational inference with a small number (e.g., $\mathscr{O}$ (100) of ROM runs and efficient uncertainty quantification. We have tested our inversion approach on a one-mile reach of the Savannah River, GA, USA. Once the neural network is trained (offline stage), the proposed technique can perform the inversion operation orders of magnitude faster than traditional inversion methods that are commonly based on linear projections, such as principal component analysis (PCA), or the principal component geostatistical approach (PCGA). Furthermore, tests show that the algorithm can estimate the bathymetry with good accuracy even with sparse flow velocity measurements.

54 ENVIRONMENTAL SCIENCES↗

Geostatistical Mapping of Salinity Conditioned on Borehole Logs, Montebello Oil Field, California

We present a geostatistics-based stochastic salinity estimation framework for the Montebello Oil Field that capitalizes on available total dissolved solids (TDS) data from groundwater samples as well as electrical resistivity (ER) data from borehole logging. Data from TDS samples (n = 4924) was coded into an indicator framework based on falling below four selected thresholds (500, 1000, 3000, and 10,000 mg/L). Collocated TDS-ER data from the surrounding groundwater basin were then employed to produce a kernel density estimator to establish conditional probabilities for ER data (n = 8 boreholes) falling below the selected TDS thresholds within the Montebello Oil Field area. Directional variograms were estimated from these indicator coded data, and 500 TDS realizations from conditional indicator simulation were generated for the subsurface region above the Montebello Oil Field reservoir. Simulations were summarized as 3D maps of median TDS, most likely salinity class, and probability for exceeding each of the specified TDS thresholds. Results suggested TDS was below 500 mg/L in most of the study area, with a trend toward higher values (500 to 1000 mg/L) to the southwest; consistent with the average regional groundwater flow direction. Discrete localized zones of TDS greater than 1000 mg/L were observed, with one of these zones in the greater than 10,000 mg/L range; however, these areas were not prevalent. The probabilistic approach used here is adaptable and is readily modified to include additional data and types and can be employed in time-lapse salinity modeling through Bayesian updating.

54 ENVIRONMENTAL SCIENCES↗

Overcoming Data Scarcity in Carbon Storage Assessment: Estimating Petrophysical Properties in Legacy Wells Integrating Sample Logs and Geostatistical Approaches

Geological CO2 sequestration feasibility studies in mature basins frequently face challenges associated with legacy well data, which often lack the comprehensive logging suites required for accurate reservoir characterization. This study evaluates the storage potential of Ordovician– Devonian formations in a filed in the Delaware Basin portion of the Permian Basin, using a dataset of 47 wells.

58 GEOSCIENCES↗

Mitigation of spatial nonstationarity with vision transformers

Spatial nonstationarity, the location variance of features’ statistical distributions, is ubiquitous in many natural settings. For example, in geological reservoirs rock matrix porosity varies vertically due to geomechanical compaction trends, in mineral deposits grades vary due to sedimentation and concentration processes, in hydrology rainfall varies due to the atmosphere and topography interactions, and in metallurgy crystalline structures vary due to differential cooling. Conventional geostatistical modeling workflows rely on the assumption of stationarity to be able to model spatial features for geostatistical inference. Nevertheless, this is often not a realistic assumption when dealing with nonstationary spatial data and this has motivated a variety of nonstationary spatial modeling workflows such as trend and residual decomposition, cosimulation with secondary features, and spatial segmentation and independent modeling over stationary subdomains. The advent of deep learning technologies has enabled new workflows for modeling spatial relationships. However, there is a paucity of demonstrated best practice and general guidance on mitigation of spatial nonstationarity with deep learning in the geospatial context. We demonstrate the impact of two common types of geostatistical spatial nonstationarity on deep learning model prediction performance and propose the mitigation of such impacts using self-attention (vision transformer) models. We demonstrate the utility of vision transformers for the mitigation of nonstationarity with relative errors as low as 10%, exceeding the performance of alternative deep learning methods such as convolutional neural networks. We establish best practice by demonstrating the ability of self-attention networks for modeling large-scale spatial relationships in the presence of commonly observed geospatial nonstationarity.

58 GEOSCIENCES↗

The influence of permeability anisotropy in the upper ocean crust on advective heat transport by a ridge-flank hydrothermal system

Here, in this study, we highlight the importance of permeability anisotropy on the hydrogeological regime of a ridge-flank hydrothermal system. Our study site, North Pond, is a marine sediment pond on ~8 Ma seafloor in the North Atlantic, and represents a low-temperature, end-member ridge-flank hydrothermal system. Previous simulations of North Pond elucidated long-standing hypotheses concerning hydrothermal fluid and heat transport in the upper volcanic crust but failed to fully explain observed patterns of seafloor heat flux in this area. Here we use variography, a geostatistical method, to quantify relations between seafloor heat-flux measurements, and coupled numerical simulations of fluid and heat flow to simulate the hydrogeologic regime. Directional variography shows that heat-flux observations are correlated along-strike of the regional crustal fabric. Three-dimensional simulations that include permeability anisotropy are able to replicate seafloor heat-flux patterns across North Pond. The simulations that result in the best match to thermal data incorporate permeability anisotropy in the horizontal plane. We find that the feedback between permeability anisotropy and the asymmetric geometry of North Pond combine to promote advective removal of heat and mass within the crustal aquifer. These findings suggest that permeability anisotropy in the oceanic crust may influence ridge-flank hydrothermal circulation more broadly.

58 GEOSCIENCES↗

Multi-site evaluation of stratified and balanced sampling of soil organic carbon stocks in agricultural fields

Estimating soil organic carbon (SOC) stocks in agricultural fields is essential for environmental and agronomic research, management, and policy. Stratified sampling is a classic strategy for estimating mean soil properties, and has recently been codified in SOC monitoring protocols. However, for the specific task of estimating the SOC stock of an agricultural field, concrete guidance is needed for which covariates to stratify on and how much stratification can improve estimation efficiency. It is also unknown how stratified sampling of SOC stocks compares to modern alternatives, notably doubly balanced sampling. To address these gaps, we collected high-density (average of 7 samples ha -1 ) and deep (average of 75 cm) measurements of SOC stocks at eight commercial fields under maize-soybean production in two US Midwestern states. We combined these measurements with a Bayesian geostatistical model to evaluate stratified and balanced sampling strategies that use a set of readily-available geographic, topographic, spectroscopic, and soil survey data. We examined the number of samples needed to achieve a given level of SOC stock estimation accuracy. While stratified sampling using these variables enables an average sample size reduction of 17% (95% CI, 11% to 23%) compared to simple random sampling, doubly balanced sampling is consistently more efficient, reducing sample sizes by 32% (95% CI, 25% to 37%). The data most important to these efficiency gains are a remotely-sensed SOC index, SSURGO estimates of SOC stocks, and the topographic wetness index. We conclude that in order to meet the urgent challenge of climate change, SOC stocks in agricultural fields could be more efficiently estimated by taking advantage of this readily-available data, especially with doubly balanced sampling.

54 ENVIRONMENTAL SCIENCES↗

A hierarchical stochastic modeling approach for representing point bar geometries and petrophysical property variations

The flow of fluids in point bars is affected by the existence of heterogeneities like shale drapes that are found on the surfaces of inclined heterolithic stratifications. In fact, these shale drapes can act as fluid flow baffles; therefore, developing a framework for modeling point bars and their associated heterogeneities is vital. In this study, a stochastic process-based modeling approach is presented for capturing the main point bar heterogeneities: accretion surfaces (i.e., the aerial heterogeneity) and inclined heterolithic stratifications (i.e., the vertical heterogeneity). The former was modeled using a sine-generation function and the latter, with a sigmoidal function after which they were combined into a 3D point bar model. To ensure proper modeling of petrophysical properties, we developed a more representative gridding scheme which generates curvilinear grids representative of the point bar geometry. This grid was then transformed into a rectilinear grid to allow for geostatistical simulation after which all petrophysical properties were mapped back into the original curvilinear grid. An essential element of this modeling approach is the stochastic representation of shale drapes at the interface between successive accretion surfaces. The workflow was tested using a real field dataset for the Cranfield field, Mississippi. The constructed model was then subjected to a flow simulation study mimicking a CO 2 storage scenario. Various sensitivities were simulated to evaluate the effect of heterogeneities on CO 2 flow within the point-bar. Results demonstrate the importance of representing point-bar related heterogeneity and the spatial distribution of shale drapes on CO 2 plume migration and storage.

58 GEOSCIENCES↗

Quantifying the hierarchy of structural and mechanical length scales in granular systems

Continuum modeling of granular media is made possible by the existence of a length scale at and above which grain-resolved properties can be meaningfully homogenized. Progress has been made in identifying such length scales relevant to local structural properties such as porosity. However, a systematic analysis of scales above which different mechanical properties can be homogenized has yet to emerge. Here, X-ray tomography and 3D X-ray diffraction data are examined to identify such length scales. The data was obtained in-situ in compressed granular materials with rigid and flexible confinement. The experimental data are supplemented with validated discrete element simulations which examine different system sizes and different boundary conditions. Overall, our study reveals a hierarchy in the length scales of granular solids, with lengths governing structural variables being the shortest, lengths of stress variables being intermediate, and lengths of energy dissipation being the longest. All structural and mechanical length scales obey a power law based on the theory of Geostatistics, implying that the length scales can be found by analyzing samples significantly smaller than the length scales themselves. The length scales are also found to be sensitive to boundary conditions, implying that they are extrinsic features of granular media.

36 MATERIALS SCIENCE↗

Simulating water dynamics related to pedogenesis across space and time: Implications for four-dimensional digital soil mapping

Digital soil mapping (DSM) relies on machine-learning and geostatistics to represent soil property observations across space. DSM techniques are powerful but often empirical, being limited to the quality and density of point samples. Water dynamics are closely related to soil variability, and the physics that govern water movement are well known. Hydrological properties can hence be simulated by physical models through space and time, unveiling key characteristics about soils. We propose the use of hydrologic models to map soils across the surface (2D), depth (1D), and time (1D)–which provides a 4D approach to digital soil mapping (4DSM). The Distributed Hydrology Soil Vegetation Model (DHSVM) was applied to a watershed currently under pasture. Moisture sensors and wells were installed at different depths in the watershed on summit, sideslope and toeslope positions to validate the model. DHSVM simulations of soil moisture distribution and depth to saturation were performed during the hydrological year (October 2008-September 2009). Clusters of similar pixels based on soil moisture values were determined using Dynamic Time Warping (DTW) to align temporal data and K-means. Clustering was performed both seasonally and for the entire year. Temporal patterns simulated by DHSVM matched measurements given by moisture sensors and wells. Seasonal clusters differed from the annual cluster. Distinct clusters were observed for each season and with depth, showing that spatiotemporal soil variability is lost when statically assessing soils. Spatiotemporal clusters corroborated field observations of fragipan occurrence not explicitly spatially mapped by Soil Survey Geographic Database (SSURGO). If a connection can be made between water and soils, static and dynamic soil variability can be predicted using physically based hydrologic models. Hydrologic models can benefit soil mapping by enabling reliable 4D simulation of water dynamics, which are fundamental to soil variability and soil classification and directly relate to biological, physical and chemical soil processes not captured by typical soil sampling protocols.

54 ENVIRONMENTAL SCIENCES↗