Search NASASearch

SEARCH · Search NASA

Results for “knowledge engineering”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Plant Bioengineering Atlas: A Knowledge Graph of Genes, DNA Constructs, and Plant Traits.

Plant bioengineering has generated tens of thousands of genotype-to-phenotype relationships, but this knowledge remains fragmented across narrative literature and difficult to use computationally. Inconsistent descriptions of DNA constructs, host species, and traits, including variable species names, omitted regulatory elements, and inconsistent gene symbols, impede data reuse, comparative analysis, and design-build-test-learn cycles. Here, we present the Plant Bioengineering Atlas, a literature-mined, ontology-grounded knowledge base assembled using an artificial intelligence (AI)-aided extraction pipeline. A large language model parsed open-access primary research articles to generate structured, provenance-anchored records of engineered genes, modification types, promoter-gene-terminator constructs, host species, target traits, and reported phenotypes, with every record traceable to its source. The current release contains 14,358 curated records encompassing 6,998 distinct genes across 436 plant species from 6,452 papers published between 2000 and 2026. Corpus analysis reveals that experiments are concentrated in a small group of model and crop species, disease and pathogen resistance is the most frequently engineered trait class, and constitutive regulatory parts (particularly the CaMV 35S promoter and NOS terminator) remain pervasive. Two in five records omit one or both flanking regulatory elements (i.e., promoter and terminator), while only 23.4% describe cassettes in which both elements resolve to named part classes, exposing a systematic reproducibility gap. We organize these data into a knowledge graph linking genes, constructs, species, and traits; provide access through an interactive web portal; and propose an AI-compatible documentation standard for AI-ready reporting. The Plant Bioengineering Atlas provides a foundation for data-driven hypothesis generation and AI-aided plant biodesign.

, Genes, DNA Constructs

Inaugural Molten Salt Technologies Workshop Powering the Future

The inaugural “Molten Salt Technologies – Powering the Future” Workshop marked a significant convergence of minds from diverse industries, each harnessing molten-salt technologies in innovative ways. Participants from sectors such as solar, geothermal, and advanced nuclear energy, as well as those involved in cutting-edge applications like thermal transport and rare-earth metals extraction, gathered to discuss their shared technical challenges and opportunities. Industry leaders in metal extraction and recycling, alongside experts in high-temperature sensor technology and advanced material manufacturing, also brought their unique perspectives to the table. This workshop served as a crucial platform for these varied industries to delve into the engineering intricacies that molten salt technologies entail. Common challenges such as understanding the thermophysical properties of salts, tackling corrosion mechanisms in harsh environments, and enhancing material resilience under extreme conditions were at the forefront of discussions. These technical sessions highlighted the critical need for cross-industry collaboration to address issues like salt life-cycle process engineering, impurity mitigation, and the development of durable, high-performance materials. By bringing together academia, industry, and representatives from national laboratories and the U.S. Department of Energy (DOE), the workshop facilitated a rich exchange of knowledge and experiences. This interaction not only fostered new partnerships but also strengthened the network among existing collaborators, setting the stage for joint solutions to the complex problems faced by all sectors using molten salt technologies. The event underscored the importance of collaborative efforts in overcoming common engineering challenges and advancing the application of molten-salt technologies across various industries. The workshop not only provided an essential forum for networking and idea exchange but also highlighted the collective drive towards innovative solutions that could benefit multiple fields.

Department of Energy

Inaugural Molten Salt Technologies Workshop Powering the Future

The inaugural “Molten Salt Technologies – Powering the Future” Workshop marked a significant convergence of minds from diverse industries, each harnessing molten-salt technologies in innovative ways. Participants from sectors such as solar, geothermal, and advanced nuclear energy, as well as those involved in cutting-edge applications like thermal transport and rare-earth metals extraction, gathered to discuss their shared technical challenges and opportunities. Industry leaders in metal extraction and recycling, alongside experts in high-temperature sensor technology and advanced material manufacturing, also brought their unique perspectives to the table. This workshop served as a crucial platform for these varied industries to delve into the engineering intricacies that molten salt technologies entail. Common challenges such as understanding the thermophysical properties of salts, tackling corrosion mechanisms in harsh environments, and enhancing material resilience under extreme conditions were at the forefront of discussions. These technical sessions highlighted the critical need for cross-industry collaboration to address issues like salt life-cycle process engineering, impurity mitigation, and the development of durable, high-performance materials. By bringing together academia, industry, and representatives from national laboratories and the U.S. Department of Energy (DOE), the workshop facilitated a rich exchange of knowledge and experiences. This interaction not only fostered new partnerships but also strengthened the network among existing collaborators, setting the stage for joint solutions to the complex problems faced by all sectors using molten salt technologies. The event underscored the importance of collaborative efforts in overcoming common engineering challenges and advancing the application of molten-salt technologies across various industries. The workshop not only provided an essential forum for networking and idea exchange but also highlighted the collective drive towards innovative solutions that could benefit multiple fields.

Department of Energy

A roadmap to understanding and anticipating microbial gene transfer in soil communities

Engineered microbes are being programmed using synthetic DNA for applications in soil to overcome global challenges related to climate change, energy, food security, and pollution. However, we cannot yet predict gene transfer processes in soil to assess the frequency of unintentional transfer of engineered DNA to environmental microbes when applying synthetic biology technologies at scale. This challenge exists because of the complex and heterogeneous characteristics of soils, which contribute to the fitness and transport of cells and the exchange of genetic material within communities. Here, we describe knowledge gaps about gene transfer across soil microbiomes. Here, we propose strategies to improve our understanding of gene transfer across soil communities, highlight the need to benchmark the performance of biocontainment measures in situ, and discuss responsibly engaging community stakeholders. We highlight opportunities to address knowledge gaps, such as creating a set of soil standards for studying gene transfer across diverse soil types and measuring gene transfer host range across microbiomes using emerging technologies. By comparing gene transfer rates, host range, and persistence of engineered microbes across different soils, we posit that community-scale, environment-specific models can be built that anticipate biotechnology risks. Such studies will enable the design of safer biotechnologies that allow us to realize the benefits of synthetic biology and mitigate risks associated with the release of such technologies.

bioccontainment

ARM Lead Mentor Selection Process

The Atmospheric Radiation Measurement (ARM) Program was created in 1989 with funding from the U.S. Department of Energy (DOE) to develop several highly instrumented ground stations to study cloud-formation processes and their influence on radiative transfer. This scientific infrastructure provides for fixed sites, mobile facilities, an aerial facility, and a data archive available for use by scientists worldwide through the ARM Climate Research Facility—a scientific user facility. The ARM Climate Research Facility currently operates more than 300 instrument systems that provide ground-based observations of the atmospheric column. To keep ARM at the forefront of climate observations, the ARM infrastructure depends heavily on instrument scientists and engineers, known as Mentors. Mentors must have an excellent understanding of instrumentation theory and operation for their instrument areas and have comprehensive knowledge of critical scale-dependent atmospheric processes. They must also possess the technical and analytical skills to develop new data retrievals that provide innovative approaches for creating research-quality data sets. The ARM Facility seeks the best overall qualified candidate, or team when appropriate, that can fulfill Mentor requirements in a timely manner. The roles and responsibilities of the ARM Instrument Operations Manager are provided in Appendix A. The key role and responsibilities and detailed responsibilities of ARM Lead Mentors are provided in Appendix B and Appendix C, respectively.

47 OTHER INSTRUMENTATION

Grid Resilience to Extreme Events (ResiliEX 2.0)

The Grid Resilience to Extreme Events (ResiliEX) 2.0 workshop, co-hosted by Pacific Northwest National Laboratory and Seattle City Light, was held at the Seattle City Hall April 23–25, 2024. This is the second workshop of its kind, with the first one having occurred in Seattle in November 2022. Participants from the second workshop hailed from research organizations, utilities, professional associations, consultants, government organizations, and communities. The purpose of the workshop was to Connect scientists, energy professionals, and policy experts to build knowledge and partnerships Advance the understanding of the science of extreme events and application to the energy system Promote grid planning and engineering that addresses the increasingly complex interdependencies as society combats the climate crisis Understand the role of different decision-makers and policymakers in increasing and accelerating grid resilience Identify new approaches, processes, and structures that should be pursued to increase grid resilience to extreme events.

24 POWER TRANSMISSION AND DISTRIBUTION

PUFFIn Software Modeling for Quality Management

PUFFIn (PENELOPE User Friendly Fast Interface) was designed as a fast and simple Monte Carlo simulation tool for the transport of photons and electrons, with a primary purpose as a learning and education tool for a broad range of static configurations in the radiation processing industry. Development of the PUFFIn software is funded by the Office of Radiological Security (ORS) within the United States National Nuclear Security Administration (NNSA). PUFFIn helps fill the education and knowledge gaps in the industry, as identified in reports by Fermilab (2017) and the IAEA (2020). PUFFIn uses the PENELOPE (NEA-2023) physics engine to perform simulations on static configurations. PUFFIn has support for multiple geometry types from simple, single material simulations to full 3D configurations created from CAD input files or images from X-Ray Tomography scans. PUFFin was designed to be easy for the novice user, it will generate the input and geometry files required by PENLOPE and will display the output plots within the PUFFIn interface. PUFFin is distributed for free but requires a free workshop so users can be adequately trained in its use. Workshops have been presented in the past at Texas A&M university, the Aerial-CRT facility in Strasbourg France and Jakarta Indonesia. PUFFIn simulations have been validated by 10 MeV ebeam experiments done at Aerial-CRT in France (Radiation Physics and Chemistry 222 (2024) 111774). Further user experimental comparisons were made at the medical product hands on workshop at Texas A&M in October 2024.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

Grid Architecture Mapping to Understand Transformation (GAMUT): Methods and Framework Architecture

Grid architecture (GA) is a concept that was developed to address the need for a comprehensive view of power grid challenges. GA can be viewed as a relatively consistent and fixed high-level approach; however, for any instantiation of grid structures, a combinatorial explosion results from each lower-layer expansion. This constitutes the main challenge with GA—it is a grid architect’s view of the system, which might not be very informative at the implementation level. Grid Architecture Mapping to Understand Transformation (GAMUT project) seeks to bridge that gap by integrating subject matter expertise across GA structures, providing users who lack expertise in GA approaches with valuable insights and informational materials. GAMUT seeks to answer feasibility questions for the approach. System-level expectations are that a GA baseline needs to be established in order for GA to be the common framework to which any lower layer approach is tied. This report explores a potential information ingestion and documentation framework to support GAMUT. The main concepts that enable the solution domain of GAMUT are discussed, and examples are provided. The solution domain leverages already-existing technology and concepts related to GA, knowledge management, and other relevant areas. To assess GAMUT building blocks and the overall approach, a feasibility assessment is proposed, rooted in systems engineering and GA architecture evaluation concepts.

24 POWER TRANSMISSION AND DISTRIBUTION

Evaluation of a Reduced-Order Model for IBR Fault Response Representation via OEM Blackbox Models: Preprint

Driven by the need to capture the electromagnetic transients of transmission lines, inverter switching behavior, and detailed control systems, electromagnetic transient (EMT) studies have become increasingly important in industry, such as IBR interconnection study and fault study. However, original equipment manufacturer (OEM) inverter models typically include extensive parameters and proprietary settings that are unavailable to protection engineers. This paper introduces a data-driven, reduced-order model (ROM) developed as a PSCAD library component for use in EMT-based fault studies. The ROM replicates key OEM model behaviors without requiring detailed knowledge of control design or parameterization. The accompanying Python automation scripts streamline data generation, parameter fitting, and validation. The ROM's performance is demonstrated through comparison with both IEEE 2800-compliant and non-compliant OEM models in a real-world power system. Relay responses show nearly identical results, while simulation runtime is reduced by an average of 32.8\%, highlighting the ROM's practicality for protection engineers.

14 SOLAR ENERGY

Stakeholder Questionnaire on Fish Passage Facilities at U.S. Hydropower Developments

To mitigate the environmental impacts of hydropower dams on fish populations in rivers, fish passage infrastructure may be required by regulatory authorities or recommended by state or federal agencies to allow migratory fish to pass hydropower barriers and complete their life histories. However, information on the location, types, and characteristics of fish passage infrastructure is incomplete at the national scale. This lack of information limits understanding of current fish passage capabilities and hinders efforts to anticipate future passage requirements during project licensing or relicensing. To address this knowledge gap, researchers at Oak Ridge National Laboratory partnered with federal agency and industry stakeholders to create the first national-scale database of fish passage infrastructure at U.S. hydropower developments. As part of this effort, a stakeholder questionnaire was created and deployed to solicit information on fish passage infrastructure, including engineering characteristics, operational schedules, targeted species, and costs. This online questionnaire was designed to allow respondents to provide highly detailed information regarding fish passage capabilities at U.S. hydropower developments as efficiently as possible. The questionnaire improved database coverage and accuracy by providing valuable development- or passage-specific stakeholder knowledge that is not otherwise publicly available. However, questionnaire response rates were low, and stakeholders showed a preference for providing data via the spreadsheet template. This is likely because most stakeholders had high-level knowledge about fish passage existence and type at many features, and few stakeholders had highly detailed, feature-specific knowledge of fish passage infrastructure, cost, operational scheduling, and capabilities. Nonetheless, the development and deployment of the stakeholder questionnaire advanced the project in several fundamental ways, and a similar questionnaire may be valuable for future database expansion via targeted dissemination to stakeholders with confirmed knowledge of detailed fish passage information at specific hydropower projects.

13 HYDRO ENERGY

Transportability of exogenous microbial community correlates with interwell connectivity in deep aquifers

Subsurface resource engineering operations often utilize continuous injection of externally-sourced water into geological reservoirs for formation pressure maintenance, resource recovery or energy/waste storage. Such injected water generally contains naturally occurring microbes. Little is known, however, about how the injectate microbes transport through geological media as a community, how such transportability is affected by injector-producer connectivity, and whether such knowledge can be utilized for flowpath characterization. In this study, we analyzed daily-to-weekly timeseries microbial community data from the injected- and produced-fluids of a ten-month flow test at a deep, well-characterized engineered aquifer. We found that the injectate microbial community was distinct from the indigenous community at the amplicon sequence variant (ASV) level, and that the transportability of injectate community towards a given producer, quantified by an “nASV-Overlap” metric we propose, had strong and significant positive correlation with known injector-producer connectivities at our site. This suggests that the better the connectivity, the higher the probability for more injectate species to flow through the interwell region and arrive at a producer. Because interwell connectivity is an important yet usually unknown parameter in subsurface resource engineering, such correlation in turn points to nASV-Overlap as a useful indicator of interwell connectivity for aquifer characterization and long-term monitoring. Based on our findings, an nASV-Overlap-based microbial tracing approach was developed for characterizing and monitoring the relative connectivities across multiple producers with a given injector. A side-by-side comparison between the new nASV-Overlap approach and traditional artificial tracer methods is presented, and their respective strengths and limitations are discussed.

Deep biosphere

FlowDash Geothermal Energy Enhancer: Where is Next Geothermal Resource? Machine Learning + Multiple Datasets => Geothermal Exploration Indication?

This is the presentation delivered at the 2025 GEODE Datathon competition. GEODE is a consortium of experts that addresses technology and knowledge gaps in geothermal energy, leveraging technology and best practices from the oil and gas industry. NETL team was awarded the 1st place in the engineering track. 2025 GEODE Datathon had a total of 42 teams from top universities and several major industrial companies. This awarded work is founded on a robust idea and innovative approach that uses machine learning coupled to multiple datasets to visualize geothermal “sweet” spots/indications in Great Basin based on the data provided from the GEODE Datathon. The use case also leveraged other datasets and demonstrated insightful and valuable indications for geothermal exploration.

Geothermal energy, Machine learning, Multiple Data

Science of Scale-Up: Accelerating chemical manufacturing technology development workshop report

The Science of Scale-Up: Accelerating chemical manufacturing technology development workshop report outlines key insights and actionable recommendations for accelerating the scale-up of disruptive chemical manufacturing technologies. Convened in October 2024, the workshop brought together approximately fifty experts from academia, industry, national laboratories, and government agencies to address the barriers and solutions for maturing technologies from proof-of-concept to commercialization. The report identifies seven critical themes for enabling faster scale-up. These themes were explored through general discussions and breakout sessions focused on three specific chemical manufacturing technologies—electrochemical, thermochemical, and biological conversion processes. The findings emphasize the importance of interdisciplinary collaboration, robust funding mechanisms, and shared resources to overcome technical barriers and accelerate technology deployment. The report also highlights technology-specific challenges and opportunities, including the need for advanced materials, scalable manufacturing processes, and integrated testing environments. For electrochemical manufacturing processes, durability and material optimization are key priorities, while thermochemical processes require novel reactor designs and better supply chain integration. Biological conversion processes face hurdles in strain engineering, reactor design, and process integration. Across all technologies, the workshop emphasized the importance of leveraging computational tools, standardized protocols, and collaborative networks to address knowledge gaps and technical barriers. By acting on these insights, stakeholders can reduce the timeline for scaling up critical chemical manufacturing technologies, ensuring their timely impact on manufacturing competitiveness, and environmental sustainability.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Quantifying dispersity in size and shape of nanoparticles from small-angle scattering data using machine learning based CREASE

Here, we use machine learning (ML) enhanced computational reverse engineering analysis of scattering experiments (CREASE) to interpret small-angle X-ray scattering (SAXS) data obtained from a system of nanoparticles without a priori knowledge of their exact shapes (e.g. spheres or ellipsoids), sizes (0.5–50 nm) and distributions. The SAXS measurements yielded three categories of scattering profiles exhibiting 'strong', 'weak' and 'no' features. Diminishing features (e.g. broadening or disappearing peaks) in scattering profiles have always been attributed to the presence of significant dispersity in the system. Such featureless SAXS data are not suitable for traditional analysis using analytical models. If one were to fit a relevant analytical model (e.g. the lmfit analytical model for polydisperse spheres) to these 'weak' and 'no' SAXS profiles from our nanoparticle systems, one would obtain non-unique interpretations of the data. Relying on electron microscopy to identify the distributions of nanoparticle shapes and sizes is also unfeasible, especially in high-throughput synthesis and characterization loops. In such situations, to identify the distributions of particle sizes and shapes that could be present in the sample, one must rely on methods like ML-CREASE to interpret the data quickly and output all relevant interpretations about the structure present in the system. The ML-CREASE optimization loop takes the experimental scattering profile as input and outputs multiple candidate solutions whose computed scattering profiles match the SAXS profile input. The ML-CREASE method outputs distributions of relevant structural features, such as the volume fraction of the nanoparticles in the system and the mean and standard deviation of the particle size and aspect ratio, assuming a type of distribution (e.g. normal, log-normal) for size and aspect ratio. We find that, for the SAXS profiles analyzed here, accounting for the shape dispersity along with size dispersity of the nanoparticles using ML-CREASE improved the match between the computed scattering profiles and input experimental profiles.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND

Advancing Building Energy Modeling with Large Language Models: Exploration and Case Studies

The rapid progression in artificial intelligence has facilitated the emergence of large language models like ChatGPT, offering potential applications extending into specialized engineering modeling, especially physics-based building energy modeling. This paper investigates the innovative integration of large language models with building energy modeling software, focusing specifically on the fusion of ChatGPT with EnergyPlus. A literature review is first conducted to reveal a growing trend of incorporating large language models in engineering modeling, albeit limited research on their application in building energy modeling. We underscore the potential of large language models in addressing building energy modeling challenges and outline potential applications including simulation input generation, simulation output analysis and visualization, conducting error analysis, co-simulation, simulation knowledge extraction and training, and simulation optimization. Three case studies reveal the transformative potential of large language models in automating and optimizing building energy modeling tasks, underscoring the pivotal role of artificial intelligence in advancing sustainable building practices and energy efficiency. The case studies demonstrate that selecting the right large language model techniques is essential to enhance performance and reduce engineering efforts. The findings advocate a multidisciplinary approach in future artificial intelligence research, with implications extending beyond building energy modeling to other specialized engineering modeling.

building energy modeling

Workflow for Developing and Operating Subsurface Hydrogen Storage Facilities in Porous Reservoirs

Long-duration (seasonal) storage of natural gas (NG), which primarily consists of methane (CH 4 ), has been practiced for more than a hundred years at underground gas storage (UGS) facilities that use depleted hydrocarbon reservoirs, saline aquifers, and salt caverns. To enable hydrogen (H 2 ) to be used as a long-duration, energy-storage medium, similar facilities are envisioned for underground H 2 storage (UHS) of either H 2 or H 2 /NG mixtures. Experience with UGS can be used to guide recommended practices for developing and operating UHS facilities in porous reservoirs. The most important factors (formation/fluid properties and engineering choices) that influence the performance of UHS reservoirs have been identified and quantified in previous studies. These factors and choices influence phenomena that determine the sweep efficiency of the stored working gas. These phenomena include viscous fingering, hysteretic capillary trapping, and gravity override of the working gas, as well as the upconing of nonproductive fluid that determine the sweep efficiency of the stored working gas. This report describes initial recommended-practices and a project-development workflow for UHS facilities that utilize porous reservoirs, based on the current state-of-knowledge about H 2 behavior in the subsurface. The workflow sequentially addresses all aspects of UHS project development, including the identification of H 2 sources and users, site ranking and down-selection, geologic and reservoir-engineering characterization, reservoir design, testing, risk management, commissioning, operations, and monitoring for a UHS facility. The goal is to enable UHS facilities to be developed in an efficient and timely manner, while carefully managing project risks. This workflow is similar to that which has been developed for UGS facilities (see Figure 1 of API, 2022), with the addition of tasks and subtasks specific to H 2 and UHS. The project-development workflow is broken down into three major stages: (1) define the H 2 use case; (2) rank, down-select, and characterize potential, candidate UHS sites; and (3) reservoir design, integrity testing, risk assessment, commissioning, operations, and monitoring for selected UHS sites. Each major stage is further broken down into tasks and subtasks, which are described at a high level. This report also provides more detailed descriptions of all tasks and subtasks that involve reservoir analysis and testing.

08 HYDROGEN

Leveraging large language models to address data scarcity in machine learning for graphene synthesis

Machine learning in experimental materials science faces significant challenges due to the scarcity of data, which are costly and time-consuming to generate, particularly when relying on in-house experiments. Literature data mining offers a potential solution but introduces issues like mixed data quality, inconsistent formats, and non-uniform reporting of synthesis parameters, resulting in partially missing and heterogeneous features across the dataset. Here, we propose data imputation and feature engineering methods that employ pre-trained large language models (LLMs) to enhance machine learning performance on scarce, heterogeneous datasets, demonstrated on graphene CVD synthesis data and the ML-HydPARK hydrogen storage dataset. GPT models perform data imputation via tailored prompting and semantic normalization of inconsistently reported features through embeddings, for example, to harmonize the complex nomenclature of CVD substrates. Beyond yielding more diverse and richer feature representations than traditional methods such as K-nearest neighbors (KNN) and Multivariate Imputation by Chained Equations (MICE), LLM-based data imputation is evaluated against dataset characteristics and prompting strategies. We vary the level of autonomy granted to the LLM, from generic prompting that leverages pre-trained knowledge for autonomous data generation to data-informed prompting that constrains outputs using target-specific information, and demonstrate which level of autonomy yields superior imputation performance across datasets and feature types. The proposed data engineering methods markedly improve downstream performance; for example, in graphene layer number classification using a support vector machine (SVM), binary accuracy increases from 39% to 65% and ternary accuracy from 52% to 72%. Fine-tuning experiments on both datasets show that combining our proposed LLM-based data imputation and feature encoding methods with numerical machine learning predictors outperforms standalone fine-tuned LLM predictors in data-scarce settings. The proposed strategies emphasize data enhancement techniques rather than refining learning architectures or regularizing loss functions, offering a broadly applicable framework for improving machine learning performance on scarce, inhomogeneous datasets.

Chemical vapor deposition

Geospatial Data Platform for All

Spatiotemporal data has evolved in scale due to augmented use in cross-domain applications. Simultaneously, there is substantial growth in the availability of Geographic Information Systems (GIS) data provided by the United States Geological Survey (USGS) along with other federal, state, county, or local agencies through open-data portals and public access APIs. However, data availability does not equate with accessibility. Large-scale analyses and applications require robust, performant data management with co-location of data storage and computing. The insufficiency of data management infrastructure compels researchers to adopt ad hoc project- specific GIS data storage solutions (e.g., copying data to High-Performance computer file systems). As an ad hoc storage strategy does not scale, it hampers cross-domain analyses causing difficulty in data reuse and utilizing existing code bases. Furthermore, GIS data is complex and requires expertise to analyze and manipulate due to its intricate data structures and data-specific projection transformations. Despite the challenges, we recognize that derived GIS data products, e.g., satellite or LIDAR-based images, can be used in downstream applications such as AI by domain, but non-GIS experts. To address the data needs and overcome the challenges, we are working towards a GIS Data Platform focused on efficient data storage, data discovery and access, and an API to enable common workflows. We propose a knowledge-graph (KG) approach for data discovery, whereby datasets are semantically linked to higher- level constructs such as projects and research areas. The semantic data links enable researchers to explore datasets in a top-down approach by specifying relevant and meaningful terms (assists in finding hidden data). An advantage is that the nodes and edges in a knowledge graph create built-in semantic documentation. Deeper spatiotemporal connections between data sources can be encoded via Graph Neural Networks (GNN) (Zhang et al., 2021). The KG approach can be extended to integrate the data itself in a Virtual KG (VKG). Our work will derive inspiration from large-scale VKG efforts that have been undertaken or are currently underway as part of the OpenStreetMap project (Ding et al., 2021). For DOE Data Days, we share the proposed geospatial data platform hybrid (cloud/on-prem) architecture, our work-to-date on storing, retrieving, and transforming LiDAR and raster data relevant to two important NREL use-cases, including the Renewable Energy Potential (reV) Model, and present our proposal for a KG based data discovery engine.

data platform