Search NASASearch

SEARCH · Search NASA

Results for “Retrieval Augmented Generation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

78 records · Page 5

Results from an Aeromagnetic Survey to Detect Steel-Cased Wells at a Marcellus Shale Well Site in Washington County, Pennsylvania

Pennsylvania has a 150-year history of oil and gas production—the longest of any state—and this enduring activity has resulted in the drilling of more than 300,000 recorded wells. However, unknown wells likely exist because innumerable wells were drilled during Pennsylvania’s intense early oil and gas history when incomplete records were kept of well locations. There is concern that early wells are likely to be ineffectively sealed because there were no laws that required plugging when the wells were abandoned. Today, many undocumented and unplugged wells are thought to be in areas of emerging shale gas and shale oil development where open wellbores can provide a pathway for undesired upward migration of fluids and gas from hydraulically fractured reservoirs. Due to this concern, Pennsylvania regulators have asked operators to locate orphaned and abandoned wells within a 1,000-ft buffer of proposed new wells. The objective of this report is to demonstrate that high-resolution aeromagnetic surveys, historic air photos, and Light Detection and Ranging (LiDAR) imagery can be rapid and effective methods to reconnoiter large, forested areas of moderate terrain for the presence of abandoned wells. These well-finding methods were evaluated at a proposed Marcellus Shale gas drilling site in Washington County, Pennsylvania, where the methods collectively located 18 confirmed wells: 15 wells were identified from aeromagnetic surveys, two wells were identified from inspection of historical air photos, and one well was identified by evaluation of state-wide LiDAR imagery. Only six wells were previously known, and their locations, as recorded in Pennsylvania’s statewide oil and gas wells database (PA/IRIS/WIS), were often too inaccurate for the wells to be found in the dense underbrush. Twelve wells identified in this study were abandoned, unmarked, and undocumented. Aeromagnetic surveys locate wells by detecting the unique magnetic signature of vertical, steel well casing, which is depicted on magnetic maps as a “bull’s eye” type anomaly that is centered directly over the well. However, when wells were drilled and found to be sub-economic, their casing was sometimes pulled and salvaged for reuse. Such wellbores provide no magnetic response and go undetected if all casing was removed. Oftentimes attempts to retrieve well casing were not 100% successful. For example, historical records for one well in the study area indicate that the well was completed in 1902 as a dry hole and that, to the extent possible, the casing was pulled for reuse. However, a section of 10-in. diameter steel casing was not recovered and remains at an unknown depth in the wellbore. This well was easily detected by the aeromagnetic survey although only deep casing remained in the well. To mitigate for the likelihood that wellbores exist where most or all casing has been removed, this study augmented aeromagnetic data with historic air photos and digital terrain models generated from LiDAR datasets—both databases are publicly available at no cost for areas within Pennsylvania. These complementary methods located three wells where the aeromagnetic anomaly, although present, was subtle and overlooked. Together, these methods determined accurate locations for six known wells within the study area and located 12 previously unknown wells. Although it is not certain that these methods successfully located all wells in the study area, the application of these methods does represent a significant improvement over relying on existing databases for well locations. For the Appendix to the report, see: https://www.netl.doe.gov/energy-analysis/details?id=b46c417a-7c9e-4d25-b810-e6248b0217f4</p>

04 OIL SHALES AND TAR SANDS

GIScience in the era of Artificial Intelligence: a research agenda towards Autonomous GIS

The advent of generative AI exemplified by large language models (LLMs) opens new ways to represent and compute geographic information and transcends the process of geographic knowledge production, driving geographic information systems (GIS) towards autonomous GIS. Leveraging LLMs as the decision core, autonomous GIS can independently generate and execute geoprocessing workflows to perform spatial analysis. In this vision paper, we further elaborate on the concept of autonomous GIS and present a conceptual framework that defines its five autonomous goals, five levels of autonomy, five core functions, and three operational scales. We demonstrate how autonomous GIS could perform geospatial data retrieval, spatial analysis, and map making with four proof-of-concept GIS agents. We conclude by identifying critical challenges and future research directions, including fine-tuning and self-growing decision-cores, autonomous modelling, and examining the societal and practical implications of autonomous GIS. By establishing the groundwork for a paradigm shift in GIScience, this paper envisions a future where GIS moves beyond traditional workflows to autonomously reason, derive, innovate, and advance geospatial solutions to pressing global challenges. Meanwhile, we emphasize that as we design and deploy increasingly intelligent geospatial systems, we carry a responsibility to ensure they are developed in socially responsible ways, serve the public good, and support the continued value of human geographic insight in an AI-augmented future.

Autonomous GI

Mechanisms Regulating Deep Moist Convection and Sea-Surface Temperatures of the Tropics

Despite numerous previous studies, two relationships between deep convection and the sea-surface temperature (SST) of the tropics remain unclear. The first is the cause for the sudden emergence of deep convection at about 28 deg SST, and the second is its proximity to the highest observed SST of about 30 C. Our analysis provides a rational explanation for both by utilizing the Improved Meteorological (IMET) buoy data together with radar rainfall retrievals and atmospheric soundings provided by the Tropical Ocean Global Atmosphere Coupled Ocean-Atmosphere Response Experiment (TOGA-COARE). The explanation relies on the basic principles of moist convection as enunciated in the Arakawa-Schubert cumulus parameterization. Our analysis shows that an SST range of 28-29 C is necessary for "charging" the atmospheric boundary layer with sufficient moist static energy that can enable the towering convection to reach up to the 200 hPa level. In the IMET buoy data, the changes in surface energy fluxes associated with different rainfall amounts show that the deep convection not only reduces the solar flux into the ocean with a thick cloud cover, but it also generates downdrafts which bring significantly cooler and drier air into the boundary-layer thereby augmenting oceanic cooling by increased sensible and latent heat fluxes. In this way, the ocean seasaws between a net energy absorber for non-raining and a net energy supplier for deep-convective raining conditions. These processes produce a thermostat-like control of the SST. The data also shows that convection over the warm pool is modulated by dynamical influences of large-scale circulation embodying tropical easterly waves (with a 5-day period) and MJOs (with 40-day period); however, the quasi-permanent feature of the vertical profile of moist static energy, which is primarily maintained by the large-scale circulation and thermodynamical forcings, is vital for both the 28 C SST for deep convection and its upper limit at about 30 C.

Sud, Y. C.

An ML-based terrestrial data fusion and augmentation framework to enable advanced understanding of the terrestrial carbon and water interactions

Soil moisture is essential to the terrestrial carbon and water cycles and land–atmosphere interactions. There are various types of soil moisture data, and each type has the distinct spatiotemporal strengths and limitations, depending on the diverse applications and retrieval methodologies of different data types (Li et al., in review; The PNNL-82151 FY23 Report). However, the limitations of different soil moisture data in terms of accuracy and spatiotemporal coverage hinder our ability to further understand the soil moisture dynamics across scales. To have a gap free soil moisture data product with a fine spatiotemporal coverage and vertical profiles, we train extreme gradient boosting (XGBoost) models by using (1) in-situ soil moisture measurements from the International Soil Moisture Network (ISMN), (2) soil moisture from the ECMWF reanalysis (ERA) at the 9 km and sub-daily spatiotemporal resolution, (3) the Daymet meteorological fields, and (4) data products that characterize surface conditions, including soil texture, organic content, topography, vegetation type, and rooting depth. We use the trained XGBoost models that have consistent performance across seven soil layers, i.e., 0–5 cm, 5–10 cm, 10–20 cm, 20–40 cm, 40–60 cm, 60–100 cm, and 100–200 cm, and the gridded model predictors to generate a soil moisture data at the 1 km and daily spatiotemporal resolution for the Continental United States (CONUS) from 2001–2020. This dataset can be broadly used for Earth system model benchmark, monitoring extreme weathers, making informed decisions regarding agriculture, water resource management, climate change mitigation, and ecosystem preservation.

58 GEOSCIENCES

Knowledge Graph of RB-Tnseq Data from Fitness Browser (KP-DP1)

Motivation: Predicting microbial gene fitness across environmental conditions remains a central challenge for predictive phenomics and autonomous experimentation. Fitness assays generate large volumes of genotype–phenotype measurements difficult to integrate with experimental metadata and biological function in a form that supports mechanistic reasoning. Knowledge graphs offer a semantic framework for unifying modalities and enabling context-aware inference. Results: We build GIMME (Graph Inference for Microbial Metabolism Exploration), a semantically grounded knowledge graph that unifies gene fitness measurements spanning 10 Pseudomonas species with experimental metadata and biological context. Media are decomposed into chemical components and experiments carry structured links to natural-language descriptions. The resulting graph supports two inference modes: (1) symbolic graph traversal to surface candidate gene–environment and gene–chemical associations, and (2) learned inference using heterogeneous graph neural networks that propagate information across neighborhoods. We formulate link regression over (gene, media, experiment) triplets, combining learned gene embeddings with pretrained LLM sourced text embeddings of node descriptions to predict gene fitness. We then augment a baseline MLP with an auxiliary message-passing encoder (GraphSAGE/GAT) that propagates information over gene–protein–function and media–chemical subgraphs, and fuse the two pathways with a gated residual connection. This approach produces strong agreement with held-out fitness measurements (GraphSAGE Pearson r 0.74) while also highlighting inference challenges in extreme-fitness regimes. We aggregate GAT edge-attention weights by relation type and layer to estimate which biological and environmental relations most influence fitness predictions. Conclusion: This work explores using knowledge graphs as “context graphs” for microbial phenotype prediction. They provide a rich substrate which enables explainable retrieval of supporting evidence, and provides a natural bridge to autonomous workflows that prioritize the next experiment.

59 BASIC BIOLOGICAL SCIENCES

Understanding and Utilizing PBL Height Data from Multiple Observing Systems in the GEOS System

The accuracy of PBL height simulation is a key issue in many applications including forecasting near surface meteorology and air quality, however, it is a very challenging problem due to the lack of not only comprehensive, global Planetary Boundary Layer (PBL) observations but also a strategy and infrastructure to utilize PBL height data from a variety of sensors. Following the designation of PBL as an incubation class observable in the 2017 Decadal Survey, the PBL Incubation Study Team Report [14] made clear that “a future global PBL observing system requires modeling and data assimilation as essential components.” There is an urgent need for global modeling development in order to utilize Program of Record (POR) observations, assess their impacts, and identify gaps to be filled by future PBL missions. Our overall objective is to develop PBL data assimilation capabilities in the NASA Global Earth Observing System (GEOS), focusing on PBL height from multiple observing systems, to support the assessment and use of future PBL observations. The NASA GEOS system is composed of the GEOS global atmospheric general circulation model (AGCM) and the atmospheric data assimilation system (ADAS). The PBL parameterizations include the “Lock” K-profile scheme driven by surface and cloud-top buoyancy fluxes ([4]), and the “Louis” local scheme for stable conditions based on the Richardson number ([5]). Above the mixed layer defined by the Lock surface plume, shallow cumulus convection is represented by the mass flux scheme of [9]. Additional parameterizations are summarized in [1]. The ADAS employs the hybrid 4D Ensemble- Variational (EnVar) configuration ([15]), with the ensemble providing flow-dependent background error covariance information. The resultant analysis increments are fed back to the forecast model through the 4D incremental analysis update (IAU) approach ([11]). In this study, PBL height data are being or have been generated from radiosondes, GNSS RO, satellite (CATS, CALIPSO and ICESat-2) and ground-based (MPLNET) lidars, and wind profiler. Investigations have been conducted to specify quality marks for PBL height retrievals for the data assimilation purpose. These PBL height data have different strengths and weaknesses ([2], [3], [6], [7], [8], [10]), and the satellite PBL height data provide better global coverage and complement in-situ PBL height data. Radiosondes offer high accuracy and in situ measurement of temperature and humidity profiles, but with poor spatio-temporal sampling. The in-situ observing systems like MPLNET and wind profiler provide long history of PBL height records at each station. The GNSS RO based PBL height is retrieved based on the sharp gradients in refractivity profile that represent the fine vertical structure of temperature and moisture changes above the PBL. However, not all RO refractivity profiles reach the surface depending on location and regime, and RO refractivity retrievals can be negatively biased below 2km. The PBL height data from satellite lidars provide high resolution along track PBL height retrievals, but over land they are affected by previous day convective PBL aerosol and strongly associated with mixing layer and retrievals cannot be made below thick, attenuating clouds. A successful assimilation of PBL height data requires a thorough understanding of the observing method and the retrieval algorithm for each observing system in order to use the PBL height data from multiple observing systems properly. Due to the sensitivity of PBL height data to the observing method and choice of algorithm, it is important to use a model definition appropriate for each observation type to compute differences between PBL height data and model PBL height (OmFs). The GEOS model currently includes two PBL height definitions suitable for direct comparison with observed PBL height, and additional definitions are being added in this study. Evaluation of different model PBL height definitions is underway. Meanwhile, efforts have been made in the GEOS data assimilation system to develop PBL height data assimilation capability. PBL height data can be assimilated using two different approaches. The traditional approach is to construct an observation operator and its tangent linear and adjoint, which link control variables to PBL height data from each observing system. This observation operator can be very complicated, e.g., the lidar-based PBL height observation operator includes the backscatter lidar forward observation operator, the algorithm to derive PBL height from attenuated total backscatter, interpolation, and calculations handling the mismatch between observed and model scales. The other approach is to augment PBL height to the control variable vector, and it is adopted in this study. The latter approach was also used in previous studies, e.g., the assimilation of PBL height data from radiosonde and aircraft in the Real Time Mesoscale Analysis (RTMA) system for a dispersion modelling study ([13]); the PBL height assimilation study using lidar PBL height data at Greensburg, Kansas for a field campaign ([12]). The PBL height assimilation from multiple observing systems in this study allows us to take advantage of the diverse PBL height data that provide much better global coverage collectively under different meteorological conditions and with different temporal and spatial scales. As all the PBL heights are tightly coupled with the PBL thermodynamic variables, the strong correlations, which are provided by the 4D ensemble forecast, enable PBL height data from various sources to interact and combine coherently and provide additional information for PBL temperature and moisture fields. The results of comparisons among PBL height data from different sources and the evaluation of the model PBL height definitions with the PBL height data will be presented, and the PBL height data synergy strategies and preliminary results will also be discussed at the conference.

Y. Zhu