Search NASASearch

SEARCH · Search NASA

Results for “Retrieval Augmented Generation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

ARCS: Agentic Retrieval-Augmented Code Synthesis with Iterative Refinement

Agentic Retrieval-Augmented Code Synthesis with Iterative RefinementIn supercomputing, efficient and optimized code generation is essential to leverage high-performance systems effectively. We have developed Agentic Retrieval-Augmented Code Synthesis (ARCS), an advanced framework for accurate, robust, and efficient code generation, completion, and translation. ARCS integrates Retrieval-Augmented Generation (RAG) with Chain-of-Thought (CoT) reasoning to systematically break down and iteratively refine complex programming tasks. An agent-based RAG mechanism retrieves relevant code snippets, while real-time execution feedback drives the synthesis of candidate solutions. This process is formalized as a state-action search tree optimization, balancing code correctness with editing efficiency. Evaluations on the Geeks4Geeks and HumanEval benchmarks demonstrate that ARCS significantly outperforms traditional prompting methods in translation and generation quality. By enabling scalable and precise code synthesis, ARCS offers transformative potential for automating and optimizing code development in supercomputing applications, enhancing computational resource utilization

Bhattarai, Manish [Los Alamos National Labs]

Artificial Intelligence (AI) Methods for Augmenting the IMPACT Tool Evidence Library

Development of the Evidence Library for use with the IMPACT probability risk assessment tool took several years and involved a staggering amount of effort from a multi-disciplinary team. A very significant amount of the labor effort to collect, assess and finalize the Clinical Finding Form (CliFF) for each of the 119 medical conditions was provided by physician subject matter experts from the Exploration Medical Capability (ExMC) Element Clinical and Science Team. Many AI tools such as ChatGPT are excellent at summarizing large amounts of information and the current project was initiated to determine how such tools might streamline laborious processes, e.g., review and summarization of many scientific research publications, to execute key steps more efficiently in the process of developing CliFFs. The process for collecting the evidence which is found in the CliFFs is well documented in the Evidence Library Methods document (ELM; HRP-48036*). Using ELM and the CliFF development instructions as a guideline, a team of developers is leveraging Microsoft Azure AI tools and services along with open-source frameworks, to construct an AI-assisted automated pipeline. This pipeline is designed to search, retrieve, and process the necessary data sources, and ultimately help generate the final version of a CliFF. Currently, the large language model evaluates the relevance of each source material to spaceflights, either as direct evidence or as an analog. Additionally, the model assists in extracting keywords and generating brief summaries to enhance augmented retrieval and search processes in later stages of CliFF development. Once the data is ready, the model can perform semantic search and retrieval, generating and extracting valuable information for the CliFF. For instance, it can handle epidemiological statistical data, such as incidence rates and the likelihood of best or worst-case scenarios. The steps that required reading and summarizing articles were viewed as providing the greatest return on investment since large language models are very efficient and accurate in summarizing large amounts of text. Since labor effort to complete the original CliFF was not recorded with sufficient granularity, comparisons with an AI tool-generated CliFF will provide merely an approximation of time saved. Upon completion of the process, the CliFF for the medical condition “appendicitis” generated with the support of AI-based methods will serve as a proof-of-concept and will be compared to the original appendicitis CliFF to determine if use of the tools resulted in content and conclusory similarity. Based upon the results from face validation of the two CliFFs, modifications to the process will be made if necessary and additional condition CliFFs will be evaluated. Ultimately, CliFFs for the entire set of medical conditions will be created with the assistance of AI tools. Depending on the cost savings realized, CliFFs for additional medical conditions can be created to expand the Evidence Library. Future direction includes specifying the characteristics of the reviewer (prompting the AI tools to generate output assuming the reviewer is a sub-specialist physician, or nurse or EMT/medic) to determine if the effects on AI-generated output are different based on knowledge, skills and abilities. *Exploration Medical Capability Evidence Library Methods, HRP-48036 Rev A, July 2022.

Ali Al

Augmented Reality Data Generation for Training Deep Learning Neural Network

One of the major challenges in deep learning is retrieving sufficiently large labeled training datasets, which can become expensive and time consuming to collect. A unique approach to training segmentation is to use Deep Neural Network (DNN) models with a minimal amount of initial labeled training samples. The procedure involves creating synthetic data and using image registration to calculate affine transformations to apply to the synthetic data. The method takes a small dataset and generates a highquality augmented reality synthetic dataset with strong variance while maintaining consistency with real cases. Results illustrate segmentation improvements in various target features and increased average target confidence.

Torres, Gil

Opportunities for retrieval and tool augmented large language models in scientific facilities

Upgrades to advanced scientific user facilities such as next-generation x-ray light sources, nanoscience centers, and neutron facilities are revolutionizing our understanding of materials across the spectrum of the physical sciences, from life sciences to microelectronics. However, these facility and instrument upgrades come with a significant increase in complexity. Driven by more exacting scientific needs, instruments and experiments become more intricate each year. This increased operational complexity makes it ever more challenging for domain scientists to design experiments that effectively leverage the capabilities of and operate on these advanced instruments. Large language models (LLMs) can perform complex information retrieval, assist in knowledge-intensive tasks across applications, and provide guidance on tool usage. Using x-ray light sources, leadership computing, and nanoscience centers as representative examples, we describe preliminary experiments with a Context-Aware Language Model for Science (CALMS) to assist scientists with instrument operations and complex experimentation. With the ability to retrieve relevant information from facility documentation, CALMS can answer simple questions on scientific capabilities and other operational procedures. With the ability to interface with software tools and experimental hardware, CALMS can conversationally operate scientific instruments. By making information more accessible and acting on user needs, LLMs could expand and diversify scientific facilities’ users and accelerate scientific output.

97 MATHEMATICS AND COMPUTING

A hypertext system that learns from user feedback

Retrieving specific information from large amounts of documentation is not an easy task. It could be facilitated if information relevant in the current problem solving context could be automatically supplied to the user. As a first step towards this goal, we have developed an intelligent hypertext system called CID (Computer Integrated Documentation). Besides providing an hypertext interface for browsing large documents, the CID system automatically acquires and reuses the context in which previous searches were appropriate. This mechanism utilizes on-line user information requirements and relevance feedback either to reinforce current indexing in case of success or to generate new knowledge in case of failure. Thus, the user continually augments and refines the intelligence of the retrieval system. This allows the CID system to provide helpful responses, based on previous usage of the documentation, and to improve its performance over time. We successfully tested the CID system with users of the Space Station Freedom requirements documents. We are currently extending CID to other application domains (Space Shuttle operations documents, airplane maintenance manuals, and on-line training). We are also exploring the potential commercialization of this technique.

Mathe, Nathalie

Microwave Remote Sensing of Snow: Advances over Ice Sheet, Land, and Sea Ice

Satellite microwave radiometers have enabled us to observe the cryosphere and its changes. A variety of algorithms have been developed since the late 1970s. These convert observed microwave radiation into geophysical, glaciological properties relevant to study ice sheets’ snow accumulation and melt, terrestrial snow water equivalent and freeze/thaw state, and various sea ice cover characteristics like extent, concentration, and thickness. In spite of the continuous availability of satellite microwave radiometers for the past 40 years, potential for advancing our understanding of the relationship between microwave radiation and snow/ice properties still exists. Original empirical algorithms have matured and are becoming more physically based. This presentation offers insights into some advances made during the past 10 years in monitoring ice sheet, terrestrial snow, and sea ice. These advances made it possible to provide new, more reliable climate-related variables to the community for the satellite era using the typical 18-37 GHz frequency range (e.g., retrievals of grain size profiles in Antarctica, snow cover stratification, and therefore accumulation with applications to climate studies). Recent NASA instruments have recorded low microwave frequencies (at 1.4 GHz). These observations have a large penetration depth, they emanate from deep into the ice, where properties including temperature are very stable. Nonetheless, it has been found at both Dome C, Antarctica and Summit, Greenland that changes in surface snow properties significantly influence these observations. Compared to ice sheets, terrestrial snow and sea ice present higher spatial heterogeneities. Over land, presence of canopy and lakes, though with known locations, add ambiguities in the retrievals of snow properties. Over sea ice, ridges, leads, changes in salinity, and sea ice drift augment further the level of difficulty in obtaining robust geophysical properties. Assessing the quality of satellite retrievals, often requiring field activities, is a necessary step in designing the next generation of microwave algorithms to monitor changes in the cryosphere.

Brucker, Ludovic

GEONEX: Land Monitoring From a New Generation of Geostationary Satellite Sensors

The latest generation of geostationary satellites carry sensors such as ABI (Advanced Baseline Imager on GOES-16) and the AHI (Advanced Himawari Imager on Himawari) that closely mimic the spatial and spectral characteristics of Earth Observing System flagship MODIS for monitoring land surface conditions. More importantly they provide observations at 5-15 minute intervals. Such high frequency data offer exciting possibilities for producing robust estimates of land surface conditions by overcoming cloud cover, enabling studies of diurnally varying local-to-regional biosphere-atmosphere interactions, and operational decision-making in agriculture, forestry and disaster management. But the data come with challenges that need special attention. For instance, geostationary data feature changing sun angle at constant view for each pixel, which is reciprocal to sun-synchronous observations, and thus require careful adaptation of EOS algorithms. Our goal is to produce a set of land surface products from geostationary sensors by leveraging NASA's investments in EOS algorithms and in the data/compute facility NEX. The land surface variables of interest include atmospherically corrected surface reflectances, snow cover, vegetation indices and leaf area index (LAI)/fraction of photosynthetically absorbed radiation (FPAR), as well as land surface temperature and fires. In order to get ready to produce operational products over the US from GOES-16 starting 2018, we have utilized 18 months of data from Himawari AHI over Australia to test the production pipeline and the performance of various algorithms for our initial tests. The end-to-end processing pipeline consists of a suite of modules to (a) perform calibration and automatic georeference correction of the AHI L1b data, (b) adopt the Multi-Angle Implementation of Atmospheric Correction (MAIAC) algorithm to produce surface spectral reflectances along with compositing schemes and QA, and (c) modify relevant EOS retrieval algorithms (e.g., LAI and FPAR, GPP, etc.) for subsequent science product generation. Initial evaluation of Himawari AHI products against standard MODIS products indicate general agreement, suggesting that data from geostationary sensors can augment low earth orbit (LEO) satellite observations.

geostationary

Management of the orbital environment

Data regarding orbital debris are presented to shed light on the requirements of environmental management in space, and strategies are given for active intervention and operational strategies. Debris are generated by inadvertent explosions of upper stages, intentional military explosions, and collisional breakups. Design and operation practices are set forth for minimizing debris generation and removing useless debris from orbit in the low-earth and geosynchronous orbits. Self-disposal options include propulsive maneuvers, drag-augmentation devices, and tether systems, and the drag devices are described as simple and passive. Active retrieval and disposition are considered, and the difficulty is examined of removing small debris. Active intervention techniques are required since pollution prevention is more effective than remediation for the problems of both earth and space.

Loftus, Joseph P., Jr.

Global Assimilation of Multi-Sensor Snow Observations for Improved Characterization of Snow Processes

Snow conditions on the land surface are recognized to be key components of the global hydrological cycle as they play a critical role in the determination of local and regional climate. In many mid-latitude and high-latitude regions, the seasonal water storage and associated spring snowmelt dominate the local hydrology. The contribution to the runoff and moisture conditions from snow is vital in supporting agriculture and in determining water resources management practices. Consequently, accurate characterization of snow properties becomes important for both end-use applications and weather and climate research. Recently a joint effort between the u.S. Air Force and NASA has enabled a blended, multi-sensor snow product known as the AFWA NASA Snow Algorithm (ANSA). This global snow dataset has been generated by utilizing the Earth Observation System (EOS) Moderate Resolution Imaging Spectroradiometer (MODIS) and Advanced Microwave Scanning Radiometer for EOS (AMSR-E) datasets. ANSA product includes estimates of snow cover extent, snow water equivalent (SWE) and SWE-derived snow depth fields. The MODIS-based products enable snow cover mappings under cloud-free conditions whereas the passive microwave data from AMSR-E provides measurements under cloudy conditions. These remotely-sensed snow observations are further augmented with the information from ground-based snow measurements through data fusion techniques. The resulting ANSA products are employed in the NASA Land Information System (LIS) data assimilation framework, which provides a comprehensive environment for integrating community land surface models, ground and satellite-based observations, and ensemble-based data assimilation tools. LIS incorporates the multisensor ANSA snow retrievals with the land surface model estimates to generate spatially and temporally continuous estimates of snow states, through data assimilation. A suite of experiments to assimilate ANSA snow cover, SWE and snow depth estimates with different land surface models in LIS are conducted and the resulting estimates of snow conditions are evaluated against a number of in-situ observational datasets, over several regions of the world. These evaluations are used to compare and contrast the advantages and disadvantages of these multi-sensor snow observations.

Kumar, Sujay

Generating Land Surface Reflectance for the New Generation of Geostationary Satellite Sensors with the MAIAC Algorithm

The latest generation of geostationary satellite sensors, including the GOES-16/ABI and the Himawari 8/AHI, provide exciting capability to monitor land surface at very high temporal resolutions (5-15 minute intervals) and with spatial and spectral characteristics that mimic the Earth Observing System flagship MODIS. However, geostationary data feature changing sun angles at constant view geometry, which is almost reciprocal to sun-synchronous observations. Such a challenge needs to be carefully addressed before one can exploit the full potential of the new sources of data. Here we take on this challenge with Multi-Angle Implementation of Atmospheric Correction (MAIAC) algorithm, recently developed for accurate and globally robust applications like the MODIS Collection 6 re-processing. MAIAC first grids the top-of- atmosphere measurements to a fixed grid so that the spectral and physical signatures of each grid cell are stacked (“remembered”) over time and used to dramatically improve cloud/shadow/snow detection, which is by far the dominant error source in the remote sensing. It also exploits the changing sun-view geometry of the geostationary sensor to characterize surface BRDF with augmented angular resolution for accurate aerosol retrievals and atmospheric correction. The high temporal resolutions of the geostationary data indeed make the BRDF retrieval much simpler and more robust as compared with sun-synchronous sensors such as MODIS. As a prototype test for the geostationary-data processing pipeline on NASA Earth Exchange (GEONEX), we apply MAIAC to process 18 months of data from Himawari 8/AHI over Australia. We generate a suite of test results, including the input TOA reflectance and the output cloud mask, aerosol optical depth (AOD), and the atmospherically-corrected surface reflectance for a variety of geographic locations, terrain, and land cover types. Comparison with MODIS data indicates a general agreement between the retrieved surface reflectance products. Furthermore, the geostationary results satisfactorily capture the movement of clouds and variations in atmospheric dust/aerosol concentrations, suggesting that high quality land surface and vegetation datasets from the advanced geostationary sensors can help complement and improve the corresponding EOS products.

geostationary satellite sensors

Real-Time Multimission Event Notification System for Mars Relay

As the Mars Relay Network is in constant flux (missions and teams going through their daily workflow), it is imperative that users are aware of such state changes. For example, a change by an orbiter team can affect operations on a lander team. This software provides an ambient view of the real-time status of the Mars network. The Mars Relay Operations Service (MaROS) comprises a number of tools to coordinate, plan, and visualize various aspects of the Mars Relay Network. As part of MaROS, a feature set was developed that operates on several levels of the software architecture. These levels include a Web-based user interface, a back-end "ReSTlet" built in Java, and databases that store the data as it is received from the network. The result is a real-time event notification and management system, so mission teams can track and act upon events on a moment-by-moment basis. This software retrieves events from MaROS and displays them to the end user. Updates happen in real time, i.e., messages are pushed to the user while logged into the system, and queued when the user is not online for later viewing. The software does not do away with the email notifications, but augments them with in-line notifications. Further, this software expands the events that can generate a notification, and allows user-generated notifications. Existing software sends a smaller subset of mission-generated notifications via email. A common complaint of users was that the system-generated e-mails often "get lost" with other e-mail that comes in. This software allows for an expanded set (including user-generated) of notifications displayed in-line of the program. By separating notifications, this can improve a user's workflow.

Wallick, Michael N.

Delaware Basin Ecological Forecasting: Identifying Vegetation Trends and Atmospheric Stressors in the Guadalupe Mountains and Carlsbad Caverns National Parks

The Guadalupe Mountains and Carlsbad Caverns National Parks, located in the Delaware Basin in the southwestern United States, observed both a decrease in precipitation and an increase in temperature over the last decade. Furthermore, activity from local oil fields generated nitrogen dioxide (NO2) plumes that spread over the parks and augmented the effects of the drought. NO2 is a precursor for tropospheric ozone (O3) which is known to have adverse effects on vegetation and ecosystems at large. These new climate dynamics prompted the National Park Service (NPS) to collaborate with NASA DEVELOP to assess the impact on vegetation within the parks. We used NASA Earth observations including Landsat 5 Thematic Mapper (TM), Landsat 7 Enhanced Thematic Mapper (ETM+), Landsat 8 Operational Land Imager (OLI), and Global Precipitation Measurement Integrated Multi-Satellite Retrievals (GPM IMERG) to assess vegetation health, water stress, and precipitation in the affected parks. After creating a homogenous reference area in the Sierra Diablo Mountains (SDM), the team visualized vegetation health through a Normalized Difference Vegetation Index (NDVI) time series map from 2010-2021. This did not show strong evidence that the NO2 plume is causing vegetation decline. Following this, we created a water stress map with a Normalized Difference Moisture Index (NDMI) time series map from 2010-2021, which revealed a pattern of increasing water stress. We also confirmed that precipitation in the region decreased over the span of 2010-2021. These observations and findings will allow the NPS Intermountain Region to more effectively plan for the preservation and maintenance of vegetation health within the parks.

Jack Mezger

Delaware Basin Ecological Forecasting: Identifying Vegetation Trends and Atmospheric Stressors in the Guadalupe Mountains and Carlsbad Caverns National Parks

The Guadalupe Mountains and Carlsbad Caverns National Parks, located in the Delaware Basin in the southwestern United States, observed both a decrease in precipitation and an increase in temperature over the last decade. Furthermore, activity from local oil fields generated nitrogen dioxide (NO2) plumes that spread over the parks and augmented the effects of the drought. NO2 is a precursor for tropospheric ozone (O3) which is known to have adverse effects on vegetation and ecosystems at large. These new climate dynamics prompted the National Park Service (NPS) to collaborate with NASA DEVELOP to assess the impact on vegetation within the parks. We used NASA Earth observations including Landsat 5 Thematic Mapper (TM), Landsat 7 Enhanced Thematic Mapper (ETM+), Landsat 8 Operational Land Imager (OLI), and Global Precipitation Measurement Integrated Multi-Satellite Retrievals (GPM IMERG) to assess vegetation health, water stress, and precipitation in the affected parks. After creating a homogeneous reference area in the Sierra Diablo Mountains, the team visualized vegetation health through a Normalized Difference Vegetation Index (NDVI) time series map from 2010-2021. This did not show strong evidence that the NO2 plume is causing vegetation decline. Following this, we created a water stress map with a Normalized Difference Moisture Index (NDMI) time series map from 2010-2021, which revealed a pattern of increasing water stress. We also confirmed that precipitation in the region decreased over the span of 2010-2021. These observations and findings will allow the NPS Intermountain Region to more effectively plan for the preservation and maintenance of vegetation health within the parks.

Jack Mezger

Biogeochemical Response to Mesoscale Physical Forcing in the California Current System

In the first part of the project, we investigated the local response of the coastal ocean ecosystems (changes in chlorophyll, concentration and chlorophyll, fluorescence quantum yield) to physical forcing by developing and deploying Autonomous Drifting Ocean Stations (ADOS) within several mesoscale features along the U.S. west coast. Also, we compared the temporal and spatial variability registered by sensors mounted in the drifters to that registered by the sensors mounted in the satellites in order to assess the scales of variability that are not resolved by the ocean color satellite. The second part of the project used the existing WOCE SVP Surface Lagrangian drifters to track individual water parcels through time. The individual drifter tracks were used to generate multivariate time series by interpolating/extracting the biological and physical data fields retrieved by remote sensors (ocean color, SST, wind speed and direction, wind stress curl, and sea level topography). The individual time series of the physical data (AVHRR, TOPEX, NCEP) were analyzed against the ocean color (SeaWiFS) time-series to determine the time scale of biological response to the physical forcing. The results from this part of the research is being used to compare the decorrelation scales of chlorophyll from a Lagrangian and Eulerian framework. The results from both parts of this research augmented the necessary time series data needed to investigate the interactions between the ocean mesoscale features, wind, and the biogeochemical processes. Using the historical Lagrangian data sets, we have completed a comparison of the decorrelation scales in both the Eulerian and Lagrangian reference frame for the SeaWiFS data set. We are continuing to investigate how these results might be used in objective mapping efforts.

Niiler, Pearn P.

NASA Pilot-Engaged Expert Response Using IBM Watson Technology: Prototype Evaluation of Knowledge Retrieval System

NASA Langley Research Center and IBM have been investigating the use of IBM Watson technology in aerospace research and development. One application of Watson technology is the Pilot-Engaged Expert Response (PEER) use case. The PEER system is envisioned as an in-cockpit advisor that will act as a source of situationally-relevant information for pilots and other flight crew members to assist in decision making about real-time events and situations that arise in the course of aircraft operations. PEER will make available vast stores of knowledge and information quickly and directly, putting important informational resources where they are needed most. IBM has worked with NASA to develop an architecture and articulate a roadmap for the development of the PEER system. That vision is built around Watson Discovery Advisor (WDA) software solution, derived from IBM's Jeopardy!-winning automatic question answering system. PEER makes use of WDA's sophisticated question-answering capabilities as its core, adding important User Interface components and other customizations for the cockpit environment, including communication with flight systems and other external data sources. The development plan for PEER includes four development stages, with the current project constituting the first phase. In this project, a prototype instance of PEER was successfully adapted to the aviation domain, enabling users to ask questions about aviation topics and receive useful and accurate answers to these questions. Major tasks accomplished include the development of procedures for domain adaptation through automatic lexicon extraction from domain glossaries; generation of question-answer training data which was used to train the system; and assessment of the effectiveness of domain adaptation, which showed a dramatic improvement in the ability of the PEER system to answer domain-relevant questions. In addition, the vision for the PEER system was pushed forward by the articulation of a plan for the automatic enhancement of question-answering with contextual information. This initial phase focused on two main goals: 1) the targeted domain adaptation of the underlying WDA system to the aviation domain; and, 2) the design of the software systems needed to leverage flight-contextual data. Domain adaptation of the WDA system proceeds via three main activities: Domain data ingestion, lexical customization and model training. A textual corpus consisting of 1,147 individual documents with more than 7.5 million words of text was ingested into the system and this served as the basis of all further development. A domain lexicon of over 3,500 aviation-domain terms was semi-automatically generated from domain documents and used to train the system. In addition, a set of over 500 question-answer (QA) pairs relevant to the PEER use case was developed; these were used to train and assess the system. These important first steps established the basis for the PEER system. In addition, steps were taken towards the integration of the PEER system into the cockpit environment with the development of a functional design for the Contextual Data Augmentation (CDA) subsystem. This subsystem brings to bear contextual data to improve system responses. It has three main submodules: the Contextual Data Collection module, the Contextual Data Selection module, and the Contextual QA Augmentation module. These modules form a processing pipeline that addresses the problems associated with automatically integrating information from external resources into the knowledge-retrieval mechanism.

Machine learning

Results from an Aeromagnetic Survey to Detect Steel-Cased Wells at a Marcellus Shale Well Site in Washington County, Pennsylvania

Pennsylvania has a 150-year history of oil and gas production—the longest of any state—and this enduring activity has resulted in the drilling of more than 300,000 recorded wells. However, unknown wells likely exist because innumerable wells were drilled during Pennsylvania’s intense early oil and gas history when incomplete records were kept of well locations. There is concern that early wells are likely to be ineffectively sealed because there were no laws that required plugging when the wells were abandoned. Today, many undocumented and unplugged wells are thought to be in areas of emerging shale gas and shale oil development where open wellbores can provide a pathway for undesired upward migration of fluids and gas from hydraulically fractured reservoirs. Due to this concern, Pennsylvania regulators have asked operators to locate orphaned and abandoned wells within a 1,000-ft buffer of proposed new wells. The objective of this report is to demonstrate that high-resolution aeromagnetic surveys, historic air photos, and Light Detection and Ranging (LiDAR) imagery can be rapid and effective methods to reconnoiter large, forested areas of moderate terrain for the presence of abandoned wells. These well-finding methods were evaluated at a proposed Marcellus Shale gas drilling site in Washington County, Pennsylvania, where the methods collectively located 18 confirmed wells: 15 wells were identified from aeromagnetic surveys, two wells were identified from inspection of historical air photos, and one well was identified by evaluation of state-wide LiDAR imagery. Only six wells were previously known, and their locations, as recorded in Pennsylvania’s statewide oil and gas wells database (PA/IRIS/WIS), were often too inaccurate for the wells to be found in the dense underbrush. Twelve wells identified in this study were abandoned, unmarked, and undocumented. Aeromagnetic surveys locate wells by detecting the unique magnetic signature of vertical, steel well casing, which is depicted on magnetic maps as a “bull’s eye” type anomaly that is centered directly over the well. However, when wells were drilled and found to be sub-economic, their casing was sometimes pulled and salvaged for reuse. Such wellbores provide no magnetic response and go undetected if all casing was removed. Oftentimes attempts to retrieve well casing were not 100% successful. For example, historical records for one well in the study area indicate that the well was completed in 1902 as a dry hole and that, to the extent possible, the casing was pulled for reuse. However, a section of 10-in. diameter steel casing was not recovered and remains at an unknown depth in the wellbore. This well was easily detected by the aeromagnetic survey although only deep casing remained in the well. To mitigate for the likelihood that wellbores exist where most or all casing has been removed, this study augmented aeromagnetic data with historic air photos and digital terrain models generated from LiDAR datasets—both databases are publicly available at no cost for areas within Pennsylvania. These complementary methods located three wells where the aeromagnetic anomaly, although present, was subtle and overlooked. Together, these methods determined accurate locations for six known wells within the study area and located 12 previously unknown wells. Although it is not certain that these methods successfully located all wells in the study area, the application of these methods does represent a significant improvement over relying on existing databases for well locations. For the Appendix to the report, see: https://www.netl.doe.gov/energy-analysis/details?id=b46c417a-7c9e-4d25-b810-e6248b0217f4</p>

04 OIL SHALES AND TAR SANDS

GIScience in the era of Artificial Intelligence: a research agenda towards Autonomous GIS

The advent of generative AI exemplified by large language models (LLMs) opens new ways to represent and compute geographic information and transcends the process of geographic knowledge production, driving geographic information systems (GIS) towards autonomous GIS. Leveraging LLMs as the decision core, autonomous GIS can independently generate and execute geoprocessing workflows to perform spatial analysis. In this vision paper, we further elaborate on the concept of autonomous GIS and present a conceptual framework that defines its five autonomous goals, five levels of autonomy, five core functions, and three operational scales. We demonstrate how autonomous GIS could perform geospatial data retrieval, spatial analysis, and map making with four proof-of-concept GIS agents. We conclude by identifying critical challenges and future research directions, including fine-tuning and self-growing decision-cores, autonomous modelling, and examining the societal and practical implications of autonomous GIS. By establishing the groundwork for a paradigm shift in GIScience, this paper envisions a future where GIS moves beyond traditional workflows to autonomously reason, derive, innovate, and advance geospatial solutions to pressing global challenges. Meanwhile, we emphasize that as we design and deploy increasingly intelligent geospatial systems, we carry a responsibility to ensure they are developed in socially responsible ways, serve the public good, and support the continued value of human geographic insight in an AI-augmented future.

Autonomous GI