Search NASA⌕ Search

SEARCH · Search NASA

Results for “data driven upscaling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Machine learning models inaccurately predict current and future high-latitude C balances

The high-latitude carbon (C) cycle is a key feedback to the global climate system, yet because of system complexity and data limitations, there is currently disagreement over whether the region is a source or sink of C. Recent advances in big data analytics and computing power have popularized the use of machine learning (ML) algorithms to upscale site measurements of ecosystem processes, and in some cases forecast the response of these processes to climate change. Due to data limitations, however, ML model predictions of these processes are almost never validated with independent datasets. To better understand and characterize the limitations of these methods, we develop an approach to independently evaluate ML upscaling and forecasting. We mimic data-driven upscaling and forecasting efforts by applying ML algorithms to different subsets of regional process-model simulation gridcells, and then test ML performance using the remaining gridcells. In this study, we simulate C fluxes and environmental data across Alaska using ecosys, a process-rich terrestrial ecosystem model, and then apply boosted regression tree ML algorithms to training data configurations that mirror and expand upon existing AmeriFLUX eddy-covariance data availability. We first show that a ML model trained using ecosys outputs from currently-available Alaska AmeriFLUX sites incorrectly predicts that Alaska is presently a modeled net C source. Increased spatial coverage of the training dataset improves ML predictions, halving the bias when 240 modeled sites are used instead of 15. However, even this more accurate ML model incorrectly predicts Alaska C fluxes under 21st century climate change because of changes in atmospheric CO 2 , litter inputs, and vegetation composition that have impacts on C fluxes which cannot be inferred from the training data. Our results provide key insights to future C flux upscaling efforts and expose the potential for inaccurate ML upscaling and forecasting of high-latitude C cycle dynamics.

54 ENVIRONMENTAL SCIENCES↗

Upscaling Wetland Methane Emissions From the FLUXNET–CH4 Eddy Covariance Network (UpCH4 v1.0): Model Development, Network Assessment, and Budget Comparison

Wetlands are responsible for 20%–31% of global methane (CH 4 ) emissions and account for a large source of uncertainty in the global CH 4 budget. Data-driven upscaling of CH 4 fluxes from eddy covariance measurements can provide new and independent bottom-up estimates of wetland CH 4 emissions. Here, we develop a six-predictor random forest upscaling model (UpCH4), trained on 119 site-years of eddy covariance CH 4 flux data from 43 freshwater wetland sites in the FLUXNET-CH4 Community Product. Network patterns in site-level annual means and mean seasonal cycles of CH 4 fluxes were reproduced accurately in tundra, boreal, and temperate regions (Nash-Sutcliffe Efficiency ~0.52–0.63 and 0.53). UpCH4 estimated annual global wetland CH 4 emissions of 146 ± 43 TgCH 4 y –1 for 2001–2018 which agrees closely with current bottom-up land surface models (102–181 TgCH 4 y –1 ) and overlaps with top-down atmospheric inversion models (155–200 TgCH 4 y –1 ). However, UpCH4 diverged from both types of models in the spatial pattern and seasonal dynamics of tropical wetland emissions. We conclude that upscaling of eddy covariance CH 4 fluxes has the potential to produce realistic extra-tropical wetland CH 4 emissions estimates which will improve with more flux data. To reduce uncertainty in upscaled estimates, researchers could prioritize new wetland flux sites along humid-to-arid tropical climate gradients, from major rainforest basins (Congo, Amazon, and SE Asia), into monsoon (Bangladesh and India) and savannah regions (African Sahel) and be paired with improved knowledge of wetland extent seasonal dynamics in these regions.

54 ENVIRONMENTAL SCIENCES↗

Permafrost Region Greenhouse Gas Budgets Suggest a Weak CO 2 Sink and CH 4 and N 2 O Sources, But Magnitudes Differ Between Top-Down and Bottom-Up Methods

Large stocks of soil carbon (C) and nitrogen (N) in northern permafrost soils are vulnerable to remobilization under climate change. However, there are large uncertainties in present-day greenhouse gas (GHG) budgets. We compare bottom-up (data-driven upscaling and process-based models) and top-down (atmospheric inversion models) budgets of carbon dioxide (CO 2 ), methane (CH 4 ) and nitrous oxide (N 2 O) as well as lateral fluxes of C and N across the region over 2000–2020. Bottom-up approaches estimate higher land-to-atmosphere fluxes for all GHGs. Both bottom-up and top-down approaches show a sink of CO 2 in natural ecosystems (bottom-up: -29 (-709, 455), top-down: -587 (-862, -312) Tg CO 2 -C yr -1 ) and sources of CH 4 (bottom-up: 38 (22, 53), top-down: 15 (11, 18) Tg CH 4 -C y -1 ) and N 2 O (bottom-up: 0.7 (0.1, 1.3), top-down: 0.09 (-0.19, 0.37) Tg N 2 O-N yr -1 ). The combined global warming potential of all three gases (GWP-100) cannot be distinguished from neutral. Over shorter timescales (GWP-20), the region is a net GHG source because CH 4 dominates the total forcing. The net CO 2 sink in Boreal forests and wetlands is largely offset by fires and inland water CO 2 emissions as well as CH 4 emissions from wetlands and inland waters, with a smaller contribution from N 2 O emissions. Priorities for future research include the representation of inland waters in process-based models and the compilation of process-model ensembles for CH 4 and N 2 O. Discrepancies between bottom-up and top-down methods call for analyses of how prior flux ensembles impact inversion budgets, more and well-distributed in situ GHG measurements and improved resolution in upscaling techniques.

54 ENVIRONMENTAL SCIENCES↗

A data-driven peridynamic continuum model for upscaling molecular dynamics

Nonlocal models, including peridynamics, often use integral operators that embed lengthscales in their definition. However, the integrands in these operators are difficult to define from the data that are typically available for a given physical system, such as laboratory mechanical property tests. In contrast, molecular dynamics (MD) does not require these integrands, but it suffers from computational limitations in the length and time scales it can address. To combine the strengths of both methods and to obtain a coarse-grained, homogenized continuum model that efficiently and accurately captures materials’ behavior, we propose a learning framework to extract, from MD data, an optimal Linear Peridynamic Solid (LPS) model as a surrogate for MD displacements. To maximize the accuracy of the learnt model we allow the peridynamic influence function to be partially negative, while preserving the well-posedness of the resulting model. To achieve this, we provide sufficient well-posedness conditions for discretized LPS models with sign-changing influence functions and develop a constrained optimization algorithm that minimizes the equation residual while enforcing such solvability conditions. This framework guarantees that the resulting model is mathematically well-posed, physically consistent, and that it generalizes well to settings that are different from the ones used during training. We illustrate the efficacy of the proposed approach with several numerical tests for single layer graphene. Our two-dimensional tests show the robustness of the proposed algorithm on validation data sets that include thermal noise, different domain shapes and external loadings, and discretizations substantially different from the ones used for training.

homogenization↗

Mesoscale informed parameter estimation through machine learning: A case-study in fracture modeling

Scale bridging is a critical need in computational sciences, where the modeling community has developed accurate physics models from first principles, of processes at lower length and time scales that influence the behavior at the higher scales of interest. However, it is not computationally feasible to incorporate all of the lower length scale physics directly into upscaled models. This is an area where machine learning has shown promise in building emulators of the lower length scale models, which incur a mere fraction of the computational cost of the original higher fidelity models. We demonstrate the use of machine learning using an example in materials science estimating continuum scale parameters by emulating, with uncertainties, complicated mesoscale physics. Additionally, we describe a new framework to emulate the fine scale physics, especially in the presence of microstructures, using machine learning, and showcase its usefulness by providing an example from modeling fracture propagation. Our approach can be thought of as a data-driven dimension reduction technique that yields probabilistic emulators. Our results show well-calibrated predictions for the quantities of interests in a low-strain simulation of fracture propagation at the mesoscale level. Furthermore, on average, we achieve ~10% relative errors on time-varying quantities like total damage and maximum stresses. Successfully replicating mesoscale scale physics within the continuum models is a crucial step towards predictive capability in multi-scale problems.

36 MATERIALS SCIENCE↗

Explicit physics-informed neural networks for nonlinear closure: The case of transport in tissues

In upscaling methods, closures for nonlinear problems present a well-known challenge. While a number of theoretical methods have been proposed for handling such closures, nonlinearities still remain a significant obstacle for many problems. In this work, we use a combination of formal upscaling and data-driven machine learning for explicitly closing a nonlinear transport and reaction process in multiscale tissues. The classical effectiveness factor model is used to formulate the macroscale reaction kinetics. We train a multilayer perceptron network using training data generated by direct numerical simulations over microscale examples. Once trained, the network is used in an algorithm for numerically solving the upscaled (coarse-grained) differential equation describing mass transport and reaction in two example tissues. The network is described as being explicit in the sense that the network is trained using macroscale concentrations and gradients of concentration as components of the feature space rather than incorporating them as part of a constraint in the optimization process. Network training and solutions to the macroscale transport equations were computed for two different tissues. The two tissue types (brain and liver) exhibit markedly different geometrical complexity and spatial scale (cell size and sample size). The upscaled solutions for the average concentration are compared with numerical solutions derived from the microscale concentration fields by a posteriori averaging. There are three outcomes of this work of particular note. 1) Our overall approach results in an upscaled nonlinear PDE. The PDE is closed using a neural network, and our approach results in the definition of the classical effectiveness factor for effecting closure. 2) We identify particular source terms for the closure problem that are important for representing the structure of the closure. These source terms involve macroscale concentrations and their gradients. We adopt these source terms to use as explicit features in the learning algorithm. We find the trained networks that include the macroscale source terms generate models that are able to predict the correction factor with increased fidelity over those that do not. 3) We find that the trained network exhibits good generalizability, and it is able to predict the effectiveness factor with high fidelity for realistically-structured tissues despite the significantly different scale and geometrical complexity of the two example tissue types. This latter result emphasizes our purposeful connection between conventional averaging methods with the use of machine learning for closure; this contrasts with some machine learning methods for upscaling where the exact form of the macroscale equation remains unknown.

97 MATHEMATICS AND COMPUTING↗

Network of networks: Time series clustering of AmeriFlux sites

Environmental observation networks, such as AmeriFlux, are foundational for monitoring ecosystem response to climate change, management practices, and natural disturbances; however, their effectiveness depends on their representativeness for the regions or continents. We proposed an empirical, time series approach to quantify the similarity of ecosystem fluxes across AmeriFlux sites. We extracted the diel and seasonal characteristics (i.e., amplitudes, phases) from carbon dioxide, water vapor, energy, and momentum fluxes, which reflect the effects of climate, plant phenology, and ecophysiology on the observations, and explored the potential aggregations of AmeriFlux sites through hierarchical clustering. While net radiation and temperature showed latitudinal clustering as expected, flux variables revealed a more uneven clustering with many small (number of sites < 5), unique groups and a few large (> 100) to intermediate (15–70) groups, highlighting the significant ecological regulations of ecosystem fluxes. Many identified unique groups were from under-sampled ecoregions and biome types of the International Geosphere-Biosphere Programme (IGBP), with distinct flux dynamics compared to the rest of the network. At the finer spatial scale, local topography, disturbance, management, edaphic, and hydrological regimes further enlarge the difference in flux dynamics within the groups. Nonetheless, our clustering approach is a data-driven method to interpret the AmeriFlux network, informing future cross-site syntheses, upscaling, and model-data benchmarking research. Finally, we highlighted the unique and underrepresented sites in the AmeriFlux network, which were found mainly in Hawaii and Latin America, mountains, and at under-sampled IGBP types (e.g., urban, open water), motivating the incorporation of new/unregistered sites from these groups.

54 ENVIRONMENTAL SCIENCES↗

X-BASE: the first terrestrial carbon and water flux products from an extended data-driven scaling framework, FLUXCOM-X

Mapping in situ eddy covariance measurements of terrestrial land–atmosphere fluxes to the globe is a key method for diagnosing the Earth system from a data-driven perspective. We describe the first global products (called X-BASE) from a newly implemented upscaling framework, FLUXCOM-X, representing an advancement from the previous generation of FLUXCOM products in terms of flexibility and technical capabilities. The X-BASE products are comprised of estimates of CO 2 net ecosystem exchange (NEE), gross primary productivity (GPP), evapotranspiration (ET), and for the first time a novel, fully data-driven global transpiration product (ETT), at high spatial (0.05°) and temporal (hourly) resolution. X-BASE estimates the global NEE at −5.75 ± 0.33 Pg C yr −1 for the period 2001–2020, showing a much higher consistency with independent atmospheric carbon cycle constraints compared to the previous versions of FLUXCOM. The improvement of global NEE was likely only possible thanks to the international effort to increase the precision and consistency of eddy covariance collection and processing pipelines, as well as to the extension of the measurements to more site years resulting in a wider coverage of bioclimatic conditions. However, X-BASE global net ecosystem exchange shows a very low interannual variability, which is common to state-of-the-art data-driven flux products and remains a scientific challenge. With 125 ± 2.1 Pg C yr −1 for the same period, X-BASE GPP is slightly higher than previous FLUXCOM estimates, mostly in temperate and boreal areas. X-BASE evapotranspiration amounts to 74.7×10 3 ± 0.9×10 3 km 3 globally for the years 2001–2020 but exceeds precipitation in many dry areas, likely indicating overestimation in these regions. On average 57 % of evapotranspiration is estimated to be transpiration, in good agreement with isotope-based approaches, but higher than estimates from many land surface models. Despite considerable improvements to the previous upscaling products, many further opportunities for development exist. Pathways of exploration include methodological choices in the selection and processing of eddy covariance and satellite observations, their ingestion into the framework, and the configuration of machine learning methods. For this, the new FLUXCOM-X framework was specifically designed to have the necessary flexibility to experiment, diagnose, and converge to more accurate global flux estimates.

Nelson, Jacob A.↗

Causality guided machine learning model on wetland CH 4 emissions across global wetlands

Wetland CH 4 emissions are among the most uncertain components of the global CH 4 budget. The complex nature of wetland CH 4 processes makes it challenging to identify causal relationships for improving our understanding and predictability of CH 4 emissions. In this study, we used the flux measurements of CH 4 from eddy covariance towers (30 sites from 4 wetlands types: bog, fen, marsh, and wet tundra) to construct a causality-constrained machine learning (ML) framework to explain the regulative factors and to capture CH 4 emissions at sub-seasonal scale. We found that soil temperature is the dominant factor for CH 4 emissions in all studied wetland types. Ecosystem respiration (CO 2 ) and gross primary productivity exert controls at bog, fen, and marsh sites with lagged responses of days to weeks. Integrating these asynchronous environmental and biological causal relationships in predictive models significantly improved model performance. More importantly, modeled CH 4 emissions differed by up to a factor of 4 under a +1°C warming scenario when causality constraints were considered. These results highlight the significant role of causality in modeling wetland CH 4 emissions especially under future warming conditions, while traditional data-driven ML models may reproduce observations for the wrong reasons. Our proposed causality-guided model could benefit predictive modeling, large-scale upscaling, data gap-filling, and surrogate modeling of wetland CH 4 emissions within earth system land models.

54 ENVIRONMENTAL SCIENCES↗

Improving Simulations of Vegetation Dynamics over the Tibetan Plateau: Role of Atmospheric Forcing Data and Spatial Resolution

The efficacy of vegetation dynamics simulations in offline land surface models (LSMs) largely depends on the quality and spatial resolution of meteorological forcing data. In this study, the Princeton Global Meteorological Forcing Data (PMFD) and the high spatial resolution and upscaled China Meteorological Forcing Data (CMFD) were used to drive the Simplified Simple Biosphere model version 4/Top-down Representation of Interactive Foliage and Flora Including Dynamics (SSiB4/TRIFFID) and investigate how meteorological forcing datasets with different spatial resolutions affect simulations over the Tibetan Plateau (TP), a region with complex topography and sparse observations. By comparing the monthly Leaf Area Index (LAI) and Gross Primary Production (GPP) against observations, we found that SSiB4/TRIFFID driven by upscaled CMFD improved the performance in simulating the spatial distributions of LAI and GPP over the TP, reducing RMSEs by 24.3% and 20.5%, respectively. The multi-year averaged GPP decreased from 364.68 gC m –2 yr –1 to 241.21 gC m –2 yr –1 with the percentage bias dropping from 50.2% to –1.7%. When using the high spatial resolution CMFD, the RMSEs of the spatial distributions of LAI and GPP simulations were further reduced by 7.5% and 9.5%, respectively. This study highlights the importance of more realistic and high-resolution forcing data in simulating vegetation growth and carbon exchange between the atmosphere and biosphere over the TP.

54 ENVIRONMENTAL SCIENCES↗

Downscaled hyper-resolution (400 m) gridded datasets of daily precipitation and temperature (2008–2019) for the East–Taylor subbasin (western United States)

Abstract. High-resolution gridded datasets of meteorological variables are needed in order to resolve fine-scale hydrological gradients in complex mountainous terrain. Across the United States, the highest available spatial resolution of gridded datasets of daily meteorological records is approximately 800 m. This work presents gridded datasets of daily precipitation and mean temperature for the East–Taylor subbasin (in the western United States) covering a 12-year period (2008–2019) at a high spatial resolution (400 m). The datasets are generated using a downscaling framework that uses data-driven models to learn relationships between climate variables and topography. We observe that downscaled datasets of precipitation and mean temperature exhibit smoother spatial gradients (while preserving the spatial variability) when compared to their coarser counterparts. Additionally, we also observe that when downscaled datasets are upscaled to the original resolution (800 m), the mean residual error is almost zero, ensuring no bias when compared with the original data. Furthermore, the downscaled datasets are observed to be linearly related to elevation, which is consistent with the methodology underlying the original 800 m product. Finally, we validate the spatial patterns exhibited by downscaled datasets via an example use case that models lidar-derived estimates of snowpack. The presented dataset constitutes a valuable resource to resolve fine-scale hydrological gradients in the mountainous terrain of the East–Taylor subbasin, which is an important study area in the context of water security for the southwestern United States and Mexico. The dataset is publicly available at https://doi.org/10.15485/1822259 (Mital et al., 2021).

54 ENVIRONMENTAL SCIENCES↗

Gaining Perspective on Unconventional Well Design Choices through Play-level Application of Machine Learning Modeling

The recent development of unconventional oil and gas (O&G) reservoirs has led to an abundant hydrocarbon supply, both domestically and globally. However, there is a continued push to develop new and innovative approaches to improve exploration and extraction efficiencies and overall well productivity moving forward. Substantial improvements in unconventional O&G development are expected through optimized well completion and stimulation strategies aimed at maximizing well productivity. Optimizing well designs will require tailoring to the distinctive geologic conditions present for any newly placed well. To better evaluate the impact of well design attributes and their associated interactions on productivity in a major unconventional play, multivariate machine learning-based models that use empirical datasets were developed. A gradient boosted regression tree (GBRT) algorithm was applied. GBRT has been narrowly investigated for O&G applications but enables straightforward parametric importance and influence evaluation, as well as assessment of parameter interaction effects. Models were trained on well design and locational parameters that serve as a proxy for variable geologic conditions to estimate two types of productivity indicator response variables strongly correlated to estimated ultimate recovery (EUR). The dataset utilized consists of over 7,000 well observations that cover the majority of the productive region of the Marcellus Shale. Model performance was evaluated and algorithm parameters tuned by analyzing the goodness-of-fit for simulated results against observed data in a cross-validation approach. Models were found capable of 73–79 percent prediction accuracy on held out testing data of gas equivalent production and can be used to inform future well design and placement decisions for increasing EUR per well and improving overall field-level recovery. Study results indicate that Marcellus well performance improves most with upscaling perforated interval lengths and water and proppant volumes per foot; but relative productivity improvements are spatially dependent across the play. Finally, optimal combinations of water and proppant on well performance were found to vary depending on well location, emphasizing the utility of data-driven models capable of broad application across a play of interest for informing tailored well design approaches prior to their field deployment.

04 OIL SHALES AND TAR SANDS↗

The Seasonal Cycle of Satellite Chlorophyll Fluorescence Observations and its Relationship to Vegetation Phenology and Ecosystem Atmosphere Carbon Exchange

Mapping of terrestrial chlorophyll uorescence from space has shown potentialfor providing global measurements related to gross primary productivity(GPP). In particular, space-based fluorescence may provide information onthe length of the carbon uptake period that can be of use for global carboncycle modeling. Here, we examine the seasonal cycle of photosynthesis asestimated from satellite fluorescence retrievals at wavelengths surroundingthe 740nm emission feature. These retrievals are from the Global OzoneMonitoring Experiment 2 (GOME-2) flying on the MetOp A satellite. Wecompare the fluorescence seasonal cycle with that of GPP as estimated froma diverse set of North American tower gas exchange measurements. Because the GOME-2 has a large ground footprint (40 x 80km2) as compared with that of the flux towers and requires averaging to reduce random errors, we additionally compare with seasonal cycles of upscaled GPP in the satellite averaging area surrounding the tower locations estimated from the Max Planck Institute for Biogeochemistry (MPI-BGC) machine learning algorithm. We also examine the seasonality of absorbed photosynthetically-active radiation(APAR) derived with reflectances from the MODerate-resolution Imaging Spectroradiometer (MODIS). Finally, we examine seasonal cycles of GPP as produced from an ensemble of vegetation models. Several of the data-driven models rely on satellite reflectance-based vegetation parameters to derive estimates of APAR that are used to compute GPP. For forested sites(particularly deciduous broadleaf and mixed forests), the GOME-2 fluorescence captures the spring onset and autumn shutoff of photosynthesis as delineated by the tower-based GPP estimates. In contrast, the reflectance-based indicators and many of the models tend to overestimate the length of the photosynthetically-active period for these and other biomes as has been noted previously in the literature. Satellite fluorescence measurements therefore show potential for improving model GPP estimates.

Joiner, J.↗

Simulation of Radon-222 with the GEOS-Chem Global Model: Emissions, Seasonality, and Convective Transport

Radon-222 (Rn-222) is a short-lived radioactive gas naturally emitted from land surfaces and has long been used to assess convective transport in atmospheric models. In this study, we simulate Rn-222 using the GEOS-Chem chemical transport model to improve our understanding of Rn-222 emissions and surface concentration seasonality and characterize convective transport associated with two Goddard Earth Observing System (GEOS) meteorological products, the Modern-Era Retrospective analysis for Research and Applications (MERRA) and GEOS Forward Processing (GEOS-FP). We evaluate four global Rn-222 emission scenarios by comparing model results with observations at 51 surface sites. The default emission scenario in GEOS-Chem yields a moderate agreement with surface observations globally (68.9 % of data within a factor of 2) and a large underestimate of winter surface Rn-222 concentrations at Northern Hemisphere midlatitudes and high latitudes due to an oversimplified formulation of Rn-222 emission fluxes (1 atom cm−2 s−1 over land with a reduction by a factor of 3 under freezing conditions). We compose a new global Rn-222 emission scenario based on Zhang et al. (2011) and demonstrate its potential to improve simulated surface Rn-222 concentrations and seasonality. The regional components of this scenario include spatially and temporally varying emission fluxes derived from previous measurements of soil radium content and soil exhalation models, which are key factors in determining Rn-222 emission flux rates. However, large model underestimates of surface Rn-222 concentrations still exist in Asia, suggesting unusually high regional Rn-222 emissions. We therefore propose a conservative upscaling factor of 1.2 for Rn-222 emission fluxes in China, which was also constrained by observed deposition fluxes of 210Pb (a progeny of Rn-222). With this modification, the model shows better agreement with observations in Europe and North America (> 80 % of data within a factor of 2) and reasonable agreement in Asia (close to 70 %). Further constraints on Rn-222 emissions would require additional concentration and emission flux observations in the central United States, Canada, Africa, and Asia. We also compare and assess convective transport in model simulations driven by MERRA and GEOS-FP using observed Rn-222 vertical profiles in northern midlatitude summer and from three short-term airborne campaigns. While simulations with both GEOS products are able to capture the observed vertical gradient of Rn-222 concentrations in the lower troposphere (0–4 km), neither correctly represents the level of convective detrainment, resulting in biases in the middle and upper troposphere. Compared with GEOS-FP, MERRA leads to stronger convective transport of Rn-222, which is partially compensated for by its weaker large-scale vertical advection, resulting in similar global vertical distributions of Rn-222 concentrations between the two simulations. This has important implications for using chemical transport models to interpret the transport of other trace species when these GEOS products are used as driving meteorology.

Bo Zhang↗

Modeling of thermal pressurization in tight claystone using sequential THM coupling: Benchmarking and validation against in-situ heating experiments in COx claystone

We apply thermoporoelasticity and a sequentially coupling technique for modeling thermally-driven coupled Thermo-Hydro-Mechanical (THM) processes in tight claystone. A THM benchmark case with a corresponding analytic solution for thermoporoelasticity under a constant heat loading verifies the model. Thereafter, two in situ heating experiments are simulated for model validation: a smaller-scale heating experiment (TED experiment) and a larger-scale experiment (ALC experiment) in Callovo-Oxfordian (COx) claystone at the Meuse/Haute-Marne underground research laboratory in France. The model exhibits good performance to match the observed temperature and pore pressure evolution for the smaller-scale TED experiment. For the larger-scale ALC experiment, general trends of thermal-pressurization are captured in the modeling, but pressure is underestimated at some monitoring points during cool-down. This indicates that the THM response in the field may be affected by the variability of rock's properties or irreversible or time-dependent mechanical processes that are not included in the current thermoporoelastic model. The main contributions of this work are as follows: (1) we verify and validate the numerical simulator, TOUGH-FLAC, to be a valuable coupled THM modeling tool; (2) prove that the laboratory determined material parameters can be used as reference values for upscaling experiments. However, to better identify and quantify THM processes with modeling of in situ tests, more emphasize should be dedicated to obtaining high-quality mechanical deformation data.

58 GEOSCIENCES↗