Search NASA⌕ Search

SEARCH · Search NASA

Results for “model-data integration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Machine learning-enabled model-data integration for predicting subsurface water storage

Subsurface water storage (SWS) is a key variable of the climate system and a storage component for precipitation and radiation anomalies, inducing persistence in the climate system. It plays a critical role in climate-change projections and can mitigate the impacts of climate change on ecosystems. However, because of the difficult accessibility of the underground, hydrologic properties and dynamics of SWS are poorly known. Direct observations of SWS are limited, and accurate incorporation of SWS dynamics into Earth system land models remains challenging. We propose a machine learning-enabled model-data integration framework to improve the SWS prediction at local to conus scales in a changing climate by leveraging all the available observation and simulation resources, as well as to inform the model development and guide the observation collection. The accurate prediction will enable an optimal decision of water management and land use and improve the ecosystem's resilience to the climate change.

Lu, Dan↗

Using machine learning and artificial intelligence to improve model-data integrated earth system model predictions of water and carbon cycle extremes

The research proposed here focuses on improving the predictive power of the land component of earth system models (ESMs) using (1) model-data fusion enabled by machine learning (ML) and artificial intelligence (AI), (2) predictive modeling through the combination of ML, AI, and big-data (comprising both model output and observations), and (3) insight of ESM structure and process mechanisms gleaned from complex data using ML and AI.

54 ENVIRONMENTAL SCIENCES↗

Automated Cloud Based Long Short-Term Memory Neural Network Based SWE Prediction

Snow derived water is a critical component of the US water supply. Measurements of the Snow Water Equivalent (SWE) and associated predictions of peak SWE and snowmelt onset are essential inputs for water management efforts. This paper aims to develop an integrated framework for real-time data ingestion, estimation, prediction and visualization of SWE based on daily snow datasets. In particular, we develop a data-driven approach for estimating and predicting SWE dynamics using the Long Short-Term Memory neural network (LSTM) method. Our approach uses historical datasets (precipitation, air temperature, SWE, and snow thickness) collected at NRCS Snow Telemetry (SNOTEL) stations to train the LSTM network and current year data to predict SWE behavior. The performance of our prediction was compared for different prediction dates and prediction training datasets. Our results suggest that the proposed LSTM network can be an efficient tool for forecasting the SWE timeseries, as well as Peak SWE and snowmelt timing. Results showed that the window size impacts the model performance (where the Nash Sutcliffe efficiency (NSE) ranged from 0.96 to 0.85 and the Rooted Mean Square Error (RMSE) ranged from 0.038 to 0.07) with an optimum number that should be calibrated for different stations and climate conditions. In addition, by implementing the LSTM prediction capability in a cloud based site-monitoring platform, we automate model-data integration. By making the data accessible through a graphical web interface and an underlying API which exposes both training and prediction capabilities. The associated results can be made easily accessible to a broad range of stakeholders.

54 ENVIRONMENTAL SCIENCES↗

Perspectives on AI Architectures and Codesign for Earth System Predictability

Abstract Recently, the U.S. Department of Energy (DOE), Office of Science, Biological and Environmental Research (BER), and Advanced Scientific Computing Research (ASCR) programs organized and held the Artificial Intelligence for Earth System Predictability (AI4ESP) workshop series. From this workshop, a critical conclusion that the DOE BER and ASCR community came to is the requirement to develop a new paradigm for Earth system predictability focused on enabling artificial intelligence (AI) across the field, laboratory, modeling, and analysis activities, called model experimentation (ModEx). BER’s ModEx is an iterative approach that enables process models to generate hypotheses. The developed hypotheses inform field and laboratory efforts to collect measurement and observation data, which are subsequently used to parameterize, drive, and test model (e.g., process based) predictions. A total of 17 technical sessions were held in this AI4ESP workshop series. This paper discusses the topic of the AI Architectures and Codesign session and associated outcomes. The AI Architectures and Codesign session included two invited talks, two plenary discussion panels, and three breakout rooms that covered specific topics, including 1) DOE high-performance computing (HPC) systems, 2) cloud HPC systems, and 3) edge computing and Internet of Things (IoT). We also provide forward-looking ideas and perspectives on potential research in this codesign area that can be achieved by synergies with the other 16 session topics. These ideas include topics such as 1) reimagining codesign, 2) data acquisition to distribution, 3) heterogeneous HPC solutions for integration of AI/ML and other data analytics like uncertainty quantification with Earth system modeling and simulation, and 4) AI-enabled sensor integration into Earth system measurements and observations. Such perspectives are a distinguishing aspect of this paper. Significance Statement This study aims to provide perspectives on AI architectures and codesign approaches for Earth system predictability. Such visionary perspectives are essential because AI-enabled model-data integration has shown promise in improving predictions associated with climate change, perturbations, and extreme events. Our forward-looking ideas guide what is next in codesign to enhance Earth system models, observations, and theory using state-of-the-art and futuristic computational infrastructure.

54 ENVIRONMENTAL SCIENCES↗

WHONDRS-GUI: a web application for global survey of surface water metabolites

Background The Worldwide Hydrobiogeochemistry Observation Network for Dynamic River Systems (WHONDRS) is a consortium that aims to understand complex hydrologic, biogeochemical, and microbial connections within river corridors experiencing perturbations such as dam operations, floods, and droughts. For one ongoing WHONDRS sampling campaign, surface water metabolite and microbiome samples are collected through a global survey to generate knowledge across diverse river corridors. Metabolomics analysis and a suite of geochemical analyses have been performed for collected samples through the Environmental Molecular Sciences Laboratory (EMSL). The obtained knowledge and data package inform mechanistic and data-driven models to enhance predictions of outcomes of hydrologic perturbations and watershed function, one of the most critical components in model-data integration. To support efforts of the multi-domain integration and make the ever-growing data package more accessible for researchers across the world, a Shiny/R Graphical User Interface (GUI) called WHONDRS-GUI was created. Results The web application can be run on any modern web browser without any programming or operational system requirements, thus providing an open, well-structured, discoverable dataset for WHONDRS. Together with a context-aware dynamic user interface, the WHONDRS-GUI has functionality for searching, compiling, integrating, visualizing and exporting different data types that can easily be used by the community. The web application and data package are available at https://data.ess-dive.lbl.gov/view/doi:10.15485/1484811 , which enables users to simultaneously obtain access to the data and code and to subsequently run the web app locally. The WHONDRS-GUI is also available for online use at Shiny Server ( https://xmlin.shinyapps.io/whondrs/ ).

59 BASIC BIOLOGICAL SCIENCES↗

A research agenda for nonvascular photoautotrophs under climate change

Nonvascular photoautotrophs (NVP), including bryophytes, lichens, terrestrial algae, and cyanobacteria, are increasingly recognized as being essential to ecosystem functioning in many regions of the world. Current research suggests that climate change may pose a substantial threat to NVP, but the extent to which this will affect the associated ecosystem functions and services is highly uncertain. Here, we propose a research agenda to address this urgent question, focusing on physiological and ecological processes that link NVP to ecosystem functions while also taking into account the substantial taxonomic diversity across multiple ecosystem types. Accordingly, we developed a new categorization scheme, based on microclimatic gradients, which simplifies the high physiological and morphological diversity of NVP and world-wide distribution with respect to several broad habitat types. We found that habitat-specific ecosystem functions of NVP will likely be substantially affected by climate change, and more quantitative process understanding is required on (1) potential for acclimation, (2) response to elevated CO 2 , (3) role of the microbiome, and (4) feedback to (micro)climate. We suggest an integrative approach of innovative, multimethod laboratory and field experiments and ecophysiological modelling, for which sustained scientific collaboration on NVP research will be essential.

54 ENVIRONMENTAL SCIENCES↗

Mass Spectrometry Sample Submission Portal

Each step in the scientific process generates contextual information about the data that is important to consider when performing data integration, developing models of biological process, or training AI models. We will develop a flexible, template-driven tool that will log biological samples, capture metadata about those samples, and track the type(s) of analysis being performed by researchers providing samples for analysis by mass spectrometry.

97 MATHEMATICS AND COMPUTING↗

A scalable framework for quantifying field-level agricultural carbon outcomes

Agriculture contributes nearly a quarter of global greenhouse gas (GHG) emissions, which is motivating interest in adopting certain farming practices that have the potential to reduce GHG emissions or sequester carbon in soil. The related GHG emission (including N 2 O and CH 4 ) and changes in soil carbon stock are defined here as “agricultural carbon outcomes”. Accurate quantification of agricultural carbon outcomes is the basis for achieving emission reductions for agriculture, but existing approaches for measuring carbon outcomes (including direct measurements, emission factors, and process-based modeling) fall short of achieving the required accuracy and scalability necessary to support credible, verifiable, and cost-effective measurement and improvement of these carbon outcomes. Here we propose a foundational and scalable framework to quantify field-level carbon outcomes for farmland, which is based on the holistic carbon balance of the agroecosystem: Agroecosystem Carbon Outcomes = Environment (E) × Management (M) × Crop (C). Following a comprehensive review of the scientific challenges associated with existing approaches, as well as their tradeoffs between cost and accuracy, we propose that the most viable path for the quantification of field-level carbon outcomes in agricultural land is through an effective integration of various approaches (e.g. diverse observations, sensor/in-situ data, and modeling), defined as the “System-of-Systems” solution. Such a “System-of-Systems” solution should simultaneously comprise the following components: (1) scalable collection of ground truth data and cross-scale sensing of environment variables (E), management practices (M), and crop conditions (C) at the local field level; (2) advanced modeling with necessary processes to support the quantification of carbon outcomes; (3) systematic Model-Data Fusion (MDF), i.e. robust and efficient methods to integrate sensing data and models at each local farmland level; (4) high computation efficiency and artificial intelligence (AI) to scale to millions of individual fields with low cost; and (5) robust and multi-tier validation systems and infrastructures to ensure solution fidelity and true scalability, i.e. the ability of a solution to perform robustly with accepted accuracy on all targeted fields. In this regard, we provide here the detailed scientific rationale, current progress, and future research and development (R&D) priorities to achieve different components of the “System-of-Systems” solution, thus accomplishing the Environment×Management×Crop framework to quantify field-level agricultural carbon outcomes.

54 ENVIRONMENTAL SCIENCES↗

Modeling the Effects of Artificial Drainage on Agriculture-dominated Watersheds using a Fully Distributed Integrated Hydrology Model: Datasets, scripts, model files

This model-data archive supports the research paper that demonstrates the integration of agricultural drainage features—specifically, narrow engineered ditches and tile drains—into a fully distributed, basin-scale integrated surface-subsurface hydrology model (ISSHM), Amanzi-ATS. The model employs innovative computational meshes aligned with agricultural ditches and incorporates the physically based Hooghoudt's drainage equation to simulate tile drainage, offering a novel strategy that enhances the accuracy of hydrological simulations.The archived dataset includes input parameters, model configurations, and select simulation outputs for the Amanzi-ATS model that successfully captured the streamflow patterns in the Portage River Watershed as validated by USGS gauge readings. Jupyter notebook for the preparation of model inputs and post-processing of outputs are also included. The model's predictive performance achieved a normalized Kling-Gupta Efficiency (KGE) of 0.81, surpassing SWAT without the necessity for site-specific calibration.The Amanzi-ATS model presented in this modeL-data archive allows for numerical experiments to explore the shifts in the flow structure under different drainage scenarios. As a tool for advancing the understanding of distributed hydrological responses and nutrient cycling, this archived model provides valuable insights for researchers, modelers, and decision-makers involved in watershed management and environmental modeling.The Watershed Workflow package is implemented in Python3. The Jupyter notebooks can be executed through multiple open-source tools, for example, Anaconda Jupyter Lab, VS Studio Code, etc. Other data files include CSV and HDF5 files, which can be read through Python scripts. The input files for the ATS model, open-source integrated hydrology, and transport model, are in XML format and can be edited in any commonly used text editors.

54 ENVIRONMENTAL SCIENCES↗

Information theory optimization of signals from small-angle scattering measurements

Small-angle X-ray scattering (SAXS) of particles in solution informs on the conformational states and assemblies of biological macromolecules (bioSAXS) outside of cryo- and solid-state conditions. In bioSAXS, the SAXS measurement under dilute conditions is resolution limited, and through an inverse Fourier transform, the measured SAXS intensities directly relate to the physical space occupied by the particles via the P (r)-distribution. Yet, this inverse transform of SAXS data has been historically cast as an ill-posed, ill-conditioned problem requiring an indirect approach. Here, we show that through the applications of matrix and information theories, the inverse transform of SAXS intensity data is a well-conditioned problem. The so-called ill-conditioning of the inverse problem is directly related to the Shannon number. By exploiting the oversampling enabled by modern detectors, a direct inverse Fourier transform of the SAXS data is possible, provided the recovered information does not exceed the Shannon number. The Shannon limit corresponds to the maximum number of significant singular values that can be recovered in a SAXS experiment, suggesting this relationship is a fundamental property of band-limited inverse integral transform problems. This correspondence reduces the complexity of the inverse problem to the Shannon limit and maximum dimension. We propose a hybrid scoring function using an information theory framework that assesses both the quality of the model-data fit as well as the quality of the recovered P (r)-distribution. The hybrid score utilizes the Akaike information criteria and Durbin-Watson statistic that considers parameter-model complexity, i.e., degrees of freedom, and the randomness of the model-data residuals. The described tests and findings extend the boundaries for bioSAXS by completing the information theory formalism initiated by Peter B. Moore to enable a quantitative measure of resolution in SAXS, robustly determine maximum dimension, and more precisely define the best parameter model appropriately representing the observed scattering data.

Rambo, Robert P. [Science and Technology Facilitie↗

Bayesian model-data comparison incorporating theoretical uncertainties

Accurate comparisons between theoretical models and experimental data are critical for scientific progress. However, inferred physical model parameters can vary significantly with the chosen physics model, highlighting the importance of properly accounting for theoretical uncertainties. In this Letter, we present a Bayesian framework that explicitly quantifies these uncertainties by statistically modeling theory errors, guided by qualitative knowledge of a theory’s varying reliability across the input domain. We demonstrate the effectiveness of this approach using two systems: a simple ball drop experiment and multi-stage heavy-ion simulations. In both cases incorporating model discrepancy leads to improved parameter estimates, with systematic improvements observed as additional experimental observables are integrated.

Bayesian methods↗

Combining Observations and Models: A Review of the CARDAMOM Framework for Data‐Constrained Terrestrial Ecosystem Modeling

The rapid increase in the volume and variety of terrestrial biosphere observations (i.e., remote sensing data and in situ measurements) offers a unique opportunity to derive ecological insights, refine process‐based models, and improve forecasting for decision support. However, despite their potential, ecological observations have primarily been used to benchmark process‐based models, as many past and current models lack the capability to directly integrate observations and their associated uncertainties for parameterization. In contrast, data assimilation frameworks such as the CARbon DAta MOdel fraMework (CARDAMOM) and its suite of process‐based models, known as the Data Assimilation Linked Ecosystem Carbon Model (DALEC), are specifically designed for model‐data fusion. This review, motivated by a recent CARDAMOM community workshop, examines the development and applications of CARDAMOM, with an emphasis on its role in advancing ecosystem process understanding. CARDAMOM employs a Bayesian approach, using a Markov Chain Monte Carlo algorithm to enable data‐driven calibration of DALEC parameters and initial states (i.e., carbon pool sizes) through observation operators. CARDAMOM's unique ability to retrieve localized model process parameters from diverse datasets—ranging from in situ measurements to global satellite observations—makes it a highly flexible tool for analyzing spatially variable ecosystem responses to environmental change. However, assimilating these data also presents challenges, including data quality issues that propagate into model skill, as well as trade‐offs between model complexity, parameter equifinality, and predictive performance. We discuss potential solutions to these challenges, such as reducing parameter equifinality by incorporating new observations. This review also offers community recommendations for incorporating emerging datasets, integrating machine learning techniques, strengthening collaboration with remote sensing, field, and modeling communities, and expanding CARDAMOM's relevance for localized ecosystem monitoring and decision‐making. CARDAMOM enables a deep, mechanistic understanding of terrestrial ecosystem dynamics that cannot be achieved through empirical analyses of observational datasets or weakly constrained models alone.

Bayesian inference↗

The future of Earth system prediction: Advances in model-data fusion

Predictions of the Earth system, such as weather forecasts and climate projections, require models informed by observations at many levels. Some methods for integrating models and observations are very systematic and comprehensive (e.g., data assimilation), and some are single purpose and customized (e.g., for model validation). We review current methods and best practices for integrating models and observations. We highlight how future developments can enable advanced heterogeneous observation networks and models to improve predictions of the Earth system (including atmosphere, land surface, oceans, cryosphere, and chemistry) across scales from weather to climate. As the community pushes to develop the next generation of models and data systems, there is a need to take a more holistic, integrated, and coordinated approach to models, observations, and their uncertainties to maximize the benefit for Earth system prediction and impacts on society.

54 ENVIRONMENTAL SCIENCES↗

Carbon Organisms Rhizosphere and Protection in Soil Environment model script and input data for soil moisture-respiration responses in tropical forests

Objectives: Climatic drying is predicted for many tropical forests, yet models remain poorly parameterized for tropical forests, hampering predictions of forest-climate feedbacks. We applied an integrated model–experiment approach, parameterizing an ecosystem model Carbon Organisms Rhizosphere and Protection in the Soil Environment (CORPSE) with tropical forest observational data, and comparing model predictions with a field drying manipulation. We hypothesized that drying would suppress soil CO2 fluxes (i.e., respiration) in already-drier tropical forests, but increases CO2 fluxes in wetter tropical forests by alleviating anaerobiosis. We measured soil CO2 fluxes, soil moisture, soil temperature, and forest floor biomass during wet-dry cycles (2015 – 2022) in four Panamanian forests that vary in rainfall and soil fertility. We used the field data to parameterize and run tests in the model.Results: Measured CO2 fluxes declined in the dry season and peaked in the early wet season ahead of peak soil moisture, resulting in a lower soil moisture optimum for respiration than previously modeled. We used this data to parameterize the model, which then predicted increased soil CO2 fluxes in wetter and fertile forests with drying, and decreased fluxes in drier, infertile forests. In contrast to model predictions, a chronic throughfall exclusion experiment in the forests initially suppressed soil CO2 fluxes across forests, with sustained suppression after four years in the wettest forest only (-28 ± 4% during the dry season), but elevated soil CO2 fluxes in a fertile forest after four years (+75 ± 28% during the late wet season), as predicted by the model. The unexpected negative drying effect in the wettest, most infertile forest could have resulted from reduced vertical flushing of nutrients into soils. Including hydro-nutrient interactions in ecosystem models could improve predictions of tropical forest-climate feedbacks (results presented in Cusack et al. 2023). Datasets included: Code files:CORPSE_array.py: Defines the equations of the CORPSE modelCORPSE_solvers: Functions for running the CORPSE model using either iterative or ordinary differential equation (ODE) solversrun_Panama_sims.py: Read in datasets and run the model simulations for this studyInput data:PanamaGradientEcosystemChem_BT_CPools_20152016CO2_DC_20190615.xlsx: Plot characteristics used in running model simulationsLiCor compiled surface flux only to 2020_03 DC_20200825.xlsx: Surface gas exchange fluxes used in model-data comparisonsPARCHED litterfall data for Ben Sulman LD 20200902.xlsx: Litterfall data used to drive model simulationsInitialization data:state_500y_20190823.csv: Initial state of model pools based on previous spinup runsOutput data:Outputs/prev_moisture_response.csv: Simulations of multiple sites using original model moisture response function.Outputs/updated_moisture_response.csv: Simulations of multiple sites using updated model moisture response function.Outputs/dry15_prev_moisture_response.csv: Simulations with soil moisture reduced by 15%, using original moisture response function.Outputs/dry15_updated_moisture_response.csv: Simulations with soil moisture reduced by 15%, using updated moisture response function.Outputs/dry30_prev_moisture_response.csv: Simulations with soil moisture reduced by 30%, using original moisture response function.Outputs/dry30_updated_moisture_response.csv: Simulations with soil moisture reduced by 30%, using updated moisture response function.Outputs/latestart_prev_moisture_response.csv: Simulations with extended dry season, using original moisture response function.Outputs/latestart_updated_moisture_response.csv: Simulations with extended dry season, using updated moisture response function.Outputs/[site name]_oneyear.csv: One-year simulation for each site in expanded site list using original moisture response function.Outputs/[site name]_oneyear_dried.csv: One-year simulation for each site in expanded site list using original moisture response function, with soil moisture reduced by 25%.Outputs/[site name]_oneyear_updated_moisture_response.csv: One-year simulation for each site in expanded site list using updated moisture response function.Outputs/[site name]_oneyear_updated_moisture_response_dried.csv: One-year simulation for each site in expanded site list using updated moisture response function, with soil moisture reduced by 25%.Field plot location data:There is also a .kml file that includes coordinates for all 32 plots included in the study of four forests (n = 4 throughfall reduction and n = 4 control plots per site).

54 ENVIRONMENTAL SCIENCES↗

Modeling the Effects of Artificial Drainage on Agriculture-dominated Watersheds using a Fully Distributed Integrated Hydrology Model: Datasets, code, models files

This model-data archive is for a modeling study focused on the effects of artificial drainage on agriculture-dominated watersheds using a fully distributed integrated hydrology model. This study introduced a novel strategy to represent the effects of tile drains and agricultural ditches in a fully distributed basin-scale integrated hydrology model. The numerical experiments in this study reveal the effects of surface and subsurface drainage on various hydrological states and fluxes. The details about the datasets can be found in the "readme" document attached with the dataset.

Agricultural Watersheds↗