Search NASASearch

SEARCH · Search NASA

Results for “data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Trust Not Verify? The Critical Need for Data Curation Standards in Materials Informatics

The importance of data curation has been recognized in multiple areas of research; however, the discussion of this important issue is only beginning to emerge in materials science. In this Perspective, we highlight the benefits of using the standardized data curation protocols in materials science and discuss current gaps in accurate and reproducible data reporting using case studies drawn from high-impact materials science papers and well-known databases such as the Crystallography Open Database (COD) and the Cambridge Structural Database (CSD). We argue that both experimental and computational materials scientists need to embrace a culture of rigorous data curation as part of modern research data management. We propose a sample data curation pipeline for materials chemistry and illustrate its use by creating two new materials chemistry databases. Here, we hope that this perspective will serve to catalyze further discussion and promote the continuous development of rigorous data curation practices within the materials science research community. We posit that adherence to best practices of data curation will promote and enhance the reliability, reproducibility, and integrity of materials research and enable the development of reliable AI and machine learning models that critically depend on the use of quality data.

Chemical structure

Fine-Root Ecology Database (FRED): A Global Collection of Root Trait Data with Coincident Site, Vegetation, Edaphic, and Climatic Data, Version 4.

To address the need for a centralized root trait database, we compiled the Fine-Root Ecology Database (FRED) from published and unpublished data sources. We have continued to add to the FRED database since the release of FRED 1.0 in 2017, followed by 2.0 in 2018, and 3.0 in 2021. This new release of FRED 4.0 now has 213,941 observations of 238 root traits, for a combined total of roughly 3.4 million data fields for root traits and ancillary data together. FRED 4.0 has 39.8% more root trait observations than FRED 3.0 and a 34.4% increase in unique data sources. This release of FRED 4.0 also includes significant increases in geographic regions that have long been underrepresented in global datasets, notably in the tropical low latitudes. Ancillary data on associated site, vegetation, edaphic, and climatic conditions from across the globe have also increased concurrently with root trait observations. FRED is focused on fine roots (traditionally defined as roots less than 2 mm in diameter), as coarse roots are studied using different methodology, often at very different scales, and have different traits and trait interpretations. Despite this fine-root focus, FRED accepts data collected from roots of all sizes and contains observations of many root classes including coarse roots. Data collection will continue for the foreseeable future. The FRED4_Entire_Database_2026.csv file is the flat csv data file for FRED 4.0, and the FRED4_dd.csv file is the data dictionary of all columns available in FRED, including column IDs, column names, definitions, and unit (where applicable).

54 ENVIRONMENTAL SCIENCES

Transportation Secure Data Center: Frequently Asked Questions for Data Owners/Contributors

The Transportation Secure Data Center is a centralized repository for detailed transportation data from travel and transit surveys and studies conducted across the nation. It makes vital transportation data broadly available to users while preserving the privacy of survey participants. Hundreds of datasets from surveys and studies of household travel and transit passenger travel are archived in the TSDC, including surveys and studies conducted by state departments of transportation, metropolitan planning organizations, transit agencies, cities, and other public agencies. Detailed data from travel surveys and studies are extremely valuable for research purposes. However, the fine-grained information they contain could potentially be misused to identify individual travelers, so access to these data should only be granted with safeguards in place to protect participant privacy. The TSDC was created to address this challenge and to relieve public agencies from the burden of archiving their data and responding to data requests.

33 ADVANCED PROPULSION SYSTEMS

Data for A Hybrid Biophysical-Machine Learning Framework for Diurnal Surface Energy Flux Estimation Using Proximal Sensing

Thermal infrared-based remote sensing of surface energy fluxes has traditionally relied on high spatial resolution satellite data with revisit frequencies on the order of weeks. In this study, we evaluate a biophysics-based analytical surface energy balance model for predicting latent energy (LE) and sensible heat (H) fluxes using proximal sensing observations. The Surface Temperature Initiated Closure (STIC1.2) model has been extensively validated across a wide range of spatial and temporal scales using various satellite-derived thermal infrared data sets. Here we extend this validation by applying STIC at sub-hourly temporal resolution over multiple growing seasons for four distinct agricultural systems. We further develop and evaluate novel STIC variants that incorporate machine learning (ML) techniques to eliminate the need for surface energy balance observations, specifically net radiation and soil heat flux, thereby enhancing model applicability in data-sparse settings. The integration of a ML component to estimate surface available energy is shown to have strong predictive performance for both LE (R2 = 0.81–0.94) and H (R2 = 0.46–0.72) across all agricultural systems examined here, demonstrating the potential of hybrid biophysical-machine learning approaches for surface energy balance modeling with minimal data requirements. This study concludes with a novel application of explainable machine learning (exML) to diagnose sources of model error. This exML framework attributes residual prediction errors to both model input variables and environmental drivers not explicitly included in the simulation experiments. This approach provides a new pathway for improving model design and integrating previously overlooked yet influential variables into future model iterations.

AI/ML

Evaluating the risk of data loss due to particle radiation damage in a DNA data storage system

DNA data storage is a potential alternative to magnetic tape for archival storage purposes, promising substantial gains in information density. Critical to the success of DNA as a storage media is an understanding of the role of environmental factors on the longevity of the stored information. In this paper, we evaluate the effect of exposure to ionizing particle radiation, a cause of data loss in traditional magnetic media, on the longevity of data in DNA data storage pools. We develop a mass action kinetics model to estimate the rate of damage accumulation in DNA strands due to neutron interactions with both nucleotides and residual water molecules, then utilize the model to evaluate the effect several design parameters of a typical DNA data storage scheme have on expected data longevity. Finally, we experimentally validate our model by exposing dried DNA samples to different levels of neutron irradiation and analyzing the resulting error profile. Our results show that particle radiation is not a significant contributor to data loss in DNA data storage pools under typical storage conditions.

97 MATHEMATICS AND COMPUTING

Identifying genomic data use with the Data Citation Explorer

Increases in sequencing capacity, combined with rapid accumulation of publications and associated data resources, have increased the complexity of maintaining associations between literature and genomic data. As the volume of literature and data have exceeded the capacity of manual curation, automated approaches to maintaining and confirming associations among these resources have become necessary. Here we present the Data Citation Explorer (DCE), which discovers literature incorporating genomic data that was not formally cited. This service provides advantages over manual curation methods including consistent resource coverage, metadata enrichment, documentation of new use cases, and identification of conflicting metadata. The service reduces labor costs associated with manual review, improves the quality of genome metadata maintained by the U.S. Department of Energy Joint Genome Institute (JGI), and increases the number of known publications that incorporate its data products. The DCE facilitates an understanding of JGI impact, improves credit attribution for data generators, and can encourage data sharing by allowing scientists to see how reuse amplifies the impact of their original studies.

59 BASIC BIOLOGICAL SCIENCES

Data Quality Assessment Process for Real-Time Data-Driven Traffic Microsimulation of Smart Corridor

Smart corridor digital twins are often created for the development and evaluation of emerging intelligent transportation systems and Connected and Autonomous Vehicle (CAV) technologies. However, limited guidance exists for data quality assessment for digital twin development. To address this, this paper discusses the data quality assessment utilized to develop data-driven real-time microscopic simulation models, i.e., digital twins, for two separate smart corridors: the North Avenue Smart Corridor in Atlanta, GA, and the Martin Luther King Smart Corridor in Chattanooga, Tennessee. This paper provides a summary of the author’s investigations of data requirements and data characteristics for the given smart corridor digital twin development efforts. With a focus on data, this summary includes a description of the data investigation process, key data issues observed, and strategies to address observed issues. Discussion is provided to help expand the lessons from these studies to other digital twin development efforts.

Saroj, Abhilasha [ORNL] (ORCID:0000000191178063)

DESI 2024: Constraints on physics-focused aspects of dark energy using DESI DR1 BAO data

Baryon acoustic oscillation data from the first year of the Dark Energy Spectroscopic Instrument (DESI) provide near percent-level precision of cosmic distances in seven bins over the redshift range z=0.1–4.2. Here, this paper is the follow-up to the original DESI BAO cosmology paper [A. G. Adame et al. (DESI Collaboration), arXiv:2404.03002], which considered the conventional w 0 w a cold dark matter (CDM) model. We use the novel DESI data, together with other cosmic probes, to constrain the background expansion history using some well-motivated physical classes of dark energy. In particular, we explore three physics-focused behaviors of dark energy from the equation of state and energy density perspectives: the thawing class (matching many simple quintessence potentials), emergent class (where dark energy comes into being recently, as in phase transition models), and mirage class [where phenomenologically the distance to cosmic microwave background (CMB) last scattering is close to that from a cosmological constant Λ despite dark energy dynamics]. All three classes fit the data at least as well as Λ ⁢CDM, and indeed can improve on it by Δ⁢χ 2 ≈ –5 to –17 for the combination of DESI BAO with CMB and supernova data while having one more parameter. The mirage class does essentially as well as w 0 ⁢w a CDM and exhibits moderate to strong Bayesian evidence preference with respect to Λ⁢ CDM. These classes of dynamical behaviors highlight worthwhile avenues for further exploration into the nature of dark energy.

79 ASTRONOMY AND ASTROPHYSICS

Data and Scripts Associated with "Modeling Ecohydrological Responses of Vegetation to Urban Microclimates Using the E3SM Land Model"

This dataset supports the study of vegetation ecohydrological responses to urban microclimates using the land component of the Energy Exascale Earth System Model (ELM) at four urban sites in Knoxville, Tennessee, USA. It includes the model inputs, simulation outputs, and associated scripts for running ELM simulations and analyzing the resulting data. The Model_Inputs folder includes static surface data, satellite-derived phenology (i.e., leaf area index), and atmospheric forcing data used to drive ELM simulations. Detailed descriptions of these datasets are provided in Section 2.3.2 of the associated manuscript. The Model_Outputs folder contains simulation results for the baseline, treatment, and ensemble experiments. Outputs from the baseline and treatment simulations are provided as raw ELM NetCDF files. Because the raw outputs from the 4,000-member ensemble are prohibitively large, the ensemble results are provided as summarized CSV files, which also serve as the source data for Figure 5 of the associated manuscript. The Scripts folder contains three components: E3SM, the core codebase of the Energy Exascale Earth System Model (E3SM); elm-olmt, the Offline Land Model Testbed (OLMT) used to perform the simulations; and knoxville_elm, which contains the analysis scripts used to process model outputs and generate the figures and results presented in the associated manuscript. Additional information is provided in Scripts_readme.txt within the Scripts directory.

Lu, Xiaoman [ORNL] (ORCID:0000000306698780)

Package Data for CERF-Data Centers

This dataset contains sample input 100m resolution raster files for running the CERF-DC python package (see https://github.com/IMMM-SFA/cerf_data_centers) at the state level across the CONUS. Due to data availability constraints, some of the items included in this dataset are proxies or assumptions for siting factors used in the model. These are individually noted in the item descriptions and can be exchanged with more detailed information upon availability. Data Descriptions The following raster files are included in the data download: state_siting_region.tif — State areas identified by state FIPS code composite_siting_suitability.tif — Value of 1 indicates suitable siting location, 0 otherwise. The following areas are excluded from siting: Areas within 300m of a federal airport runway Waterbodies Areas with slope >16% Areas susceptible to sinkholes High coastal or inland flood risk areas Local, state, and federal parks, leisure areas, and cemeteries Areas >2 km away from electric substations Areas >5 km away from a municipal water supplier service area Areas >2 km away from high-speed fiber provider service territory Protected Areas Database of the United States (PAD-US) areas Railroads, major roadways, and minor roadways Military areas and training grounds Developed lands Areas >0.8 km (0.5 miles) from developed lands land_value_dollar_per_sqft.tif — USD per square foot (sqft) derived from USDA $/acre land cost personal_property_tax_rate.tif — Personal property tax rate by state. Uses an assumed 0.0125 personal property tax rate for states with personal property tax, 0 for states without personal property tax. real_property_tax_rate.tif — Real property tax rate. Based on county level residential real estate property tax rates. sales_tax_rate.tif — Sales tax rate by state. mechanical_cooling_fraction.tif — Fraction of year (values between 0 and 1, inclusive) that the data center would be cooled through mechanical processes based on local water stress and humidity levels. water_cooling_fraction.tif — Fraction of year (values between 0 and 1, inclusive) that the data center would be cooled through evaporative (water cooled) processes based on local water stress and humidity levels. distance_to_substation.tif — Distance to nearest substation in hundreds of meters (i.e., value of 1 equals a distance of 100m). Offshore areas have a value of 0. industrial_electricity_rates_dollar_per_kwh.tif — USD/kWh industrial electricity rates. Represents the average industrial rate across all utilities that operate within a given county. Values are derived from the US Utility Rate Database. commercial_electricity_rates_dollar_per_kwh.tif — USD/kWh commercial electricity rates. Represents the average commercial rate across all utilities that operate within a given county. Values are derived from the US Utility Rate Database. data_center_market_locations.tif — Grid cells with positive values represent the centroid of existing data center market clusters. The value of non-zero grid cells represents the number of data centers in the market cluster. All other grid cells have a value of 0. Geospatial Metadata CRS: Albers Equal Area Conic (ESRI:102003) Extent: -2415585.0000000023283064,-1441981.2605773280374706 : 2384414.9999999976716936,1708018.7394226719625294 Dimensions: X: 48000 Y: 31500 Bands: 1 Origin: -2415585.0000000023283064,1708018.7394226719625294 Pixel Size: 100,-100 Acknowledgment IM3 is a multi-institutional effort led by Pacific Northwest National Laboratory and supported by the U.S. Department of Energy's Office of Science as part of research in MultiSector Dynamics, Earth and Environmental Systems Modeling Program. License This data is made available under a CCBY4.0 License Disclaimer This material was prepared as an account of work sponsored by an agency of the United States Government. Neither the United States Government nor the United States Department of Energy, nor the Contractor, nor any or their employees, nor any jurisdiction or organization that has cooperated in the development of these materials, makes any warranty, express or implied, or assumes any legal liability or responsibility for the accuracy, completeness, or usefulness or any information, apparatus, product, software, or process disclosed, or represents that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by trade name, trademark, manufacturer, or otherwise does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States Government or any agency thereof, or Battelle Memorial Institute. The views and opinions of authors expressed herein do not necessarily state or reflect those of the United States Government or any agency thereof. PACIFIC NORTHWEST NATIONAL LABORATORYoperated byBATTELLEfor theUNITED STATES DEPARTMENT OF ENERGYunder Contract DE-AC05-76RL01830

Mongird, Kendall

Package Data for CERF-Data Centers

This dataset contains sample input 100m resolution raster files for running the CERF-DC python package (see https://github.com/IMMM-SFA/cerf_data_centers) at the state level across the CONUS. Due to data availability constraints, some of the items included in this dataset are proxies or assumptions for siting factors used in the model. These are individually noted in the item descriptions and can be exchanged with more detailed information upon availability. Data Descriptions The following raster files are included in the data download: state_siting_region.tif — State areas identified by state FIPS code composite_siting_suitability.tif — Value of 1 indicates suitable siting location, 0 otherwise. The following areas are excluded from siting: Areas within 300 m of a federal airport runway or within an airport area boundary Waterbodies Areas with slope >16% Areas susceptible to sinkholes High coastal or inland flood risk areas Local, state, and federal parks, leisure areas, and cemeteries Areas >2 km away from electric substations Areas >5 km away from a municipal water supplier service area Areas >2 km away from high-speed fiber provider service territory USGS Protected Areas Database of the United States (PAD-US) GAP status 1, 2, or 3 areas US National Parks Wetlands USFWS critical habitats BIA land areas Railroads, major roadways, and minor roadways Military areas and training grounds NLCD developed lands Areas >0.8 km (0.5 miles) from NLCD developed lands land_value_dollar_per_sqft.tif — USD per square foot (sqft) derived from USDA $/acre land cost personal_property_tax_rate.tif — Personal property tax rate by state. Uses an assumed 0.0125 personal property tax rate for states with personal property tax, 0 for states without personal property tax. real_property_tax_rate.tif — Real property tax rate. Based on county level residential real estate property tax rates. sales_tax_rate.tif — Sales tax rate by state. mechanical_cooling_fraction.tif — Fraction of year (values between 0 and 1, inclusive) that the data center would be cooled through mechanical processes based on local water stress and humidity levels. water_cooling_fraction.tif — Fraction of year (values between 0 and 1, inclusive) that the data center would be cooled through evaporative (water cooled) processes based on local water stress and humidity levels. distance_to_substation.tif — Distance to nearest substation in hundreds of meters (i.e., value of 1 equals a distance of 100m). Offshore areas have a value of 0. industrial_electricity_rates_dollar_per_kwh.tif — USD/kWh industrial electricity rates. Represents the average industrial rate across all utilities that operate within a given county. Values are derived from the US Utility Rate Database. commercial_electricity_rates_dollar_per_kwh.tif — USD/kWh commercial electricity rates. Represents the average commercial rate across all utilities that operate within a given county. Values are derived from the US Utility Rate Database. data_center_market_locations.tif — Grid cells with positive values represent the centroid of existing data center market clusters. The value of non-zero grid cells represents the number of data centers in the market cluster. All other grid cells have a value of 0. Geospatial Metadata CRS: Albers Equal Area Conic (ESRI:102003) Extent: -2415585.0000000023283064,-1441981.2605773280374706 : 2384414.9999999976716936,1708018.7394226719625294 Dimensions: X: 48000 Y: 31500 Bands: 1 Origin: -2415585.0000000023283064,1708018.7394226719625294 Pixel Size: 100,-100 Acknowledgment IM3 is a multi-institutional effort led by Pacific Northwest National Laboratory and supported by the U.S. Department of Energy's Office of Science as part of research in MultiSector Dynamics, Earth and Environmental Systems Modeling Program. License This data is made available under a CCBY4.0 License Disclaimer This material was prepared as an account of work sponsored by an agency of the United States Government. Neither the United States Government nor the United States Department of Energy, nor the Contractor, nor any or their employees, nor any jurisdiction or organization that has cooperated in the development of these materials, makes any warranty, express or implied, or assumes any legal liability or responsibility for the accuracy, completeness, or usefulness or any information, apparatus, product, software, or process disclosed, or represents that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by trade name, trademark, manufacturer, or otherwise does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States Government or any agency thereof, or Battelle Memorial Institute. The views and opinions of authors expressed herein do not necessarily state or reflect those of the United States Government or any agency thereof. PACIFIC NORTHWEST NATIONAL LABORATORYoperated byBATTELLEfor theUNITED STATES DEPARTMENT OF ENERGYunder Contract DE-AC05-76RL01830

Mongird, Kendall

Data about data – when, why and how metadata can support the digital plant

A structured approach for recording data quality and contextual information about how and why a signal exists – i.e. metadata – is central to interpret and use sensor data correctly. This is becoming increasingly important with the global trend with data-driven applications such as digital twins and AI-models. But a structured metadata collection and organization of sensor data is not routine in most plants, which can result in lost information and missed opportunities to make use of the investments made in the data collection. Therefore, the IWA task group on Metadata Collection and Organization in wastewater resource recovery systems (MetaCO) was initiated in 2020 and recently delivered the IWA scientific and technical report number 31. The report gives and in-depth description about metadata in water resources recovery facilities (WRRFs) and is available as open access at IWA publishing. The report is the outcome of the collaboration between more than 80 water professionals with the intention to serve WRRF data users with a guide on how to structure and make use of metadata throughout the data pipeline in order to maximize the value of sensor data.

Alferes, Janelcy [VITO, Belgium]

TSDC: Transportation Secure Data Center: Real-World Data for Planning, Modeling, and Analysis

The Transportation Secure Data Center is a centralized repository for high-resolution transportation data from hundreds of travel and transit surveys and studies. It makes vital transportation data broadly available to users while preserving the privacy of survey participants. It houses surveys and studies conducted by state departments of transportation, metropolitan planning organizations, transit agencies, cities, and other public agencies. Meanwhile, the Livewire Data Platform empowers research, industry, and academic partners to easily and securely preserve, maintain, share, discover, and gain access to transportation and mobility data. Livewire accommodates a range of datasets, including behavioral, experimental, model, analytical, and raw data at the vehicle, traveler, and system levels. Datasets support mobility research and planning spanning urban science, connected and automated vehicles, fueling and charging infrastructure, mobility decision science, multimodal transportation, vehicle efficiency, and more.

33 ADVANCED PROPULSION SYSTEMS

Leveraging Large Language Models for Real-World Data Evidence: A Framework for Automated Treatment Extraction and Data Harmonization

Background: The ability to comprehensively collect treatment information from cancer patient medical records would enable studies to evaluate real-world benefits and risks tied to specific treatments. Currently, it is difficult to system- atically collect high-quality treatment information because it is often stored in unstructured text. Manually extracting and standardizing drug and regimen data is time-intensive. Recent advances in large language models (LLMs) offer a potential solution for automated extraction of structured treatment information from clinical text. Objective: This study systematically evaluates the utility of four LLMs from the Llama family for automated extraction of oncology treatment information from clinical text. This information can guide researchers using cancer registry data to provide insights into cancer care and outcomes beyond clinical trials. Methods: Four instruction-tuned Llama models with varying parameter counts (1B, 3B, 8B, and 70B) were evaluated for their ability to extract treatment information from clinical documents. A unified oncology knowledge base integrating seven major public data sources was developed to standardize and normalize extracted entities—a critical step for harmonizing data from diverse sources. Extracted treatment data were compared against expert-annotated ground truth. Model performance was assessed using accuracy metrics (Precision, Recall, F1-Score) and opera- tional feasibility metrics, including processing speed and structural compliance of the output. Results: A strong positive correlation was observed between model size and extraction accuracy. F1-score improved from 0.609 for the 1B model to 0.710 (3B), 0.807 (8B), and 0.828 (70B). While larger models demonstrated superior accuracy and compliance, they incurred higher computational costs. The modest performance difference between 8B and 70B suggests diminishing returns with increasing model size. Conclusions: LLMs represent a viable technology for automating oncology treatment extraction. The 8B-parameter model emerged as a highly effective option, balancing high accuracy and computational efficiency. Selecting an appropriate LLM for deployment in cancer registries involves a trade-off between desired accuracy and available operational resources. Harmonizing extracted entities with the oncology knowledge base facilitates standardized integration into common data models, enhancing data quality for real-world evidence analyses.

artificial intelligence

GenAI-Based Digital Twins Aided Data Augmentation Increases Accuracy in Real-Time Cokurtosis-Based Anomaly Detection of Wearable Data

Early detection of potential infectious disease outbreaks is crucial for developing effective interventions. In this study, we introduce advanced anomaly detection methods tailored for health datasets collected from wearables, offering insights at both individual and population levels. Leveraging real-world physiological data from wearables, including heart rate and activity, we developed a framework for the early detection of infection in individuals. Despite the availability of data from recent pandemics, substantial gaps remain in data collection, hindering method development. To bridge this gap, we utilized Wasserstein Generative Adversarial Networks (WGANs) to generate realistic synthetic wearable data, augmenting our dataset for training. Subsequently, we use these augmented datasets to implement a cokurtosis-based technique for anomaly detection in multivariate time-series data. Our approach includes a comprehensive assessment of uncertainties in synthetic data compared to the actual data upon which it was modeled, as well as the uncertainty associated with fine-tuning anomaly detection thresholds in physiological measurements. Through our work, we present an enhanced method for early anomaly detection in multivariate datasets, with promising applications in healthcare and beyond. This framework could revolutionize early detection strategies and significantly impact public health response efforts in future pandemics.

Data-Driven Digital Twins

Asi Nuclear Energy Sensors Data Portal Chatbot And Data Structuring Tool

The Idaho National Laboratory (INL) is advancing the development of an AI-powered chatbot and data structuring tool specifically designed to accelerate data mining processes for sensor-related information and seamlessly integrate the results into the ASI Sensors Data Portal (https://nes.energy.gov/). By doing so, the software aims to enhance the accessibility, usability, and organization of sensor data for nuclear energy applications. The software initial phase focuses on retrieving comprehensive datasets, prioritizing the past five years of publicly available information from the Office of Scientific and Technical Information (OSTI). These datasets will be meticulously processed to ensure compatibility, employing cleaning and preprocessing steps to eliminate irrelevant, incomplete, or corrupted information, thus establishing a robust foundation for subsequent AI use. The data will serve as the backbone for training an AI model and chatbot, which will act as an interactive tool enabling users to ask complex, context-specific questions and receive accurate, validated answers derived from constrained literature. In parallel, the project incorporates a data structuring process supported by AI to organize sensor information from multiple sources into a standardized format. This structured data will include detailed sensor specifications, such as measurement range, applications, accuracy, and operating conditions, generated and documented with AI. These specifications will be systematically integrated into the sensor portal. To maintain the highest levels of accuracy and relevance, all AI-generated outputs will be reviewed and validated by subject matter experts (SMEs), with additional fields or parameters added as needed. Future stages of the project aim to expand the dataset beyond OSTI to include other sources and potentially incorporate unclassified controlled information (UCI) with restricted access protocols to address security and confidentiality requirements.

Mapes, NormanJ. [Idaho National Laboratory (INL),

Raw Lidar and Camera Data Synchronized with Precipitation and Present Weather Data

As part of the sensor characterization task of the SMART 2.0 project, this dataset includes raw data from three spinning lidars ([Ouster OS2-128](https://ouster.com/products/scanning-lidar/os2-sensor/), [Velodyne Puck (VLP-16)](https://velodynelidar.com/products/puck/), and [Velodyne Ultra Puck (VLP-32)](https://velodynelidar.com/products/ultra-puck/)), one camera ([Mako G-319](https://www.alliedvision.com/en/camera-selector/detail/mako/g-319/)), and one present weather sensor ([Vaisala FD-70](https://www.vaisala.com/en/products/weather-environmental-sensors/forward-scatter-fd70)). All data were synchronized, with the log start time indicated in the file name (HHMMSS). The data can be filtered by date, log time (HHMMSS), sensor, frame ID, and weather classification. These data were gathered statically at the Argonne Testbed for Multiscale Observational Science (ATMOS). Two target stop signs were placed in view of the sensors to contribute a target for comparing sensor data under different conditions. The weather data for each day are stored in netCDF “.nc” files. The lidar data contain the X, Y, Z, intensity, reflectivity, and ring from Ouster OS2-128 rev6, Velodyne VLP-16, and Velodyne VLP-32 lidars. ![raw lidar image](LiDAR_pointcloud_ATMOS.png)

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Where are the Data? Automating a Workflow for Carbon Storage Data Gap Analyses

This presentation demonstrates a spatial analysis workflow to assess data availability for the many components of geologic carbon storage technical viability. The workflow relies upon a knowledge-data framework that links the different components of GCS technical viability to the data types needed for evaluation. Using this contextual information, a combination of data science methods (e.g., natural language processing) and spatial analyses are applied to identify areas where sufficient data exists for a given component. The results are aggregated into maps illustrating data density and spatial gaps across all technical viability factors and data categories, as well as the individual component and category level for a more nuanced understanding. Presented at the FECM NETL Carbon Management Program Review Meeting 2024.

Creason, Christopher