Search NASASearch

SEARCH · Search NASA

Results for “data augmentation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

2012 California Household Travel Survey Supplement

# 2012 California Household Travel Survey Supplement The 2012 California Household Travel Survey Supplement focused on gathering specific travel information from residents for the development of next-generation, activity-based models. Called the "Augment Survey," it supplemented the [2010–2012 California Household Travel Survey](https://www.nrel.gov/transportation/secure-transportation-data/tsdc-california-travel-survey). ## Data Collection Agency The Southern California Association of Governments (SCAG) hired Abt-SRBI, Inc. to conduct the survey. ## Methodology Travel data were collected from households via in-vehicle (625 vehicles) and wearable (244 participants) global positioning system (GPS) devices. ## Drive Cycle Processing and Filtering NREL has developed a GPS data filtration routine to filter erroneous data points in individual drive cycles sourced from GPS devices mounted in vehicles. Second-by-second drive cycle data collected from GPS-instrumented vehicles during this survey have passed through NREL's drive cycle processing and filtering routines. ## Survey Records Study records include 473 households. ## Transportation Data The SCAG data set contains data from 473 households that participated in one or more areas of study. Of these, 141 completed the wearable GPS portion of the study and 332 completed the vehicle GPS portion. There was no overlap between households participating in the two study areas (wearable and vehicle GPS). For details on available travel survey data and variable definitions, see the [data dictionary](https://www.nrel.gov/media/docs/libraries/tsdc/caltrans_scag_data_dictionary.pdf?sfvrsn=6ec36d7a_1). NREL-generated drive cycle data are also available for this survey. For details on available data and variable definitions, see the [drive cycle data dictionary](https://www.nrel.gov/media/docs/libraries/tsdc/drive_cycles_data_dictionary.pdf?sfvrsn=7de7e888_1). Transportation data are available as zipped files. [Download Winzip](http://www.winzip.com/downwz.htm).

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

2012 California Household Travel Survey Supplement

# 2012 California Household Travel Survey Supplement The 2012 California Household Travel Survey Supplement focused on gathering specific travel information from residents for the development of next-generation, activity-based models. Called the "Augment Survey," it supplemented the [2010–2012 California Household Travel Survey](https://www.nrel.gov/transportation/secure-transportation-data/tsdc-california-travel-survey). ## Data Collection Agency The Southern California Association of Governments (SCAG) hired Abt-SRBI, Inc. to conduct the survey. ## Methodology Travel data were collected from households via in-vehicle (625 vehicles) and wearable (244 participants) global positioning system (GPS) devices. ## Drive Cycle Processing and Filtering NREL has developed a GPS data filtration routine to filter erroneous data points in individual drive cycles sourced from GPS devices mounted in vehicles. Second-by-second drive cycle data collected from GPS-instrumented vehicles during this survey have passed through NREL's drive cycle processing and filtering routines. ## Survey Records Study records include 473 households. ## Transportation Data The SCAG data set contains data from 473 households that participated in one or more areas of study. Of these, 141 completed the wearable GPS portion of the study and 332 completed the vehicle GPS portion. There was no overlap between households participating in the two study areas (wearable and vehicle GPS). For details on available travel survey data and variable definitions, see the [data dictionary](https://www.nrel.gov/media/docs/libraries/tsdc/caltrans_scag_data_dictionary.pdf?sfvrsn=6ec36d7a_1). NREL-generated drive cycle data are also available for this survey. For details on available data and variable definitions, see the [drive cycle data dictionary](https://www.nrel.gov/media/docs/libraries/tsdc/drive_cycles_data_dictionary.pdf?sfvrsn=7de7e888_1). Transportation data are available as zipped files. [Download Winzip](http://www.winzip.com/downwz.htm).

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

2012 California Household Travel Survey Supplement

# 2012 California Household Travel Survey Supplement The 2012 California Household Travel Survey Supplement focused on gathering specific travel information from residents for the development of next-generation, activity-based models. Called the "Augment Survey," it supplemented the [2010–2012 California Household Travel Survey](https://www.nrel.gov/transportation/secure-transportation-data/tsdc-california-travel-survey). ## Data Collection Agency The Southern California Association of Governments (SCAG) hired Abt-SRBI, Inc. to conduct the survey. ## Methodology Travel data were collected from households via in-vehicle (625 vehicles) and wearable (244 participants) global positioning system (GPS) devices. ## Drive Cycle Processing and Filtering NREL has developed a GPS data filtration routine to filter erroneous data points in individual drive cycles sourced from GPS devices mounted in vehicles. Second-by-second drive cycle data collected from GPS-instrumented vehicles during this survey have passed through NREL's drive cycle processing and filtering routines. ## Survey Records Study records include 473 households. ## Transportation Data The SCAG data set contains data from 473 households that participated in one or more areas of study. Of these, 141 completed the wearable GPS portion of the study and 332 completed the vehicle GPS portion. There was no overlap between households participating in the two study areas (wearable and vehicle GPS). For details on available travel survey data and variable definitions, see the [data dictionary](https://www.nrel.gov/media/docs/libraries/tsdc/caltrans_scag_data_dictionary.pdf?sfvrsn=6ec36d7a_1). NREL-generated drive cycle data are also available for this survey. For details on available data and variable definitions, see the [drive cycle data dictionary](https://www.nrel.gov/media/docs/libraries/tsdc/drive_cycles_data_dictionary.pdf?sfvrsn=7de7e888_1). Transportation data are available as zipped files. [Download Winzip](http://www.winzip.com/downwz.htm).

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

2012 California Household Travel Survey Supplement

# 2012 California Household Travel Survey Supplement The 2012 California Household Travel Survey Supplement focused on gathering specific travel information from residents for the development of next-generation, activity-based models. Called the "Augment Survey," it supplemented the [2010–2012 California Household Travel Survey](https://www.nrel.gov/transportation/secure-transportation-data/tsdc-california-travel-survey). ## Data Collection Agency The Southern California Association of Governments (SCAG) hired Abt-SRBI, Inc. to conduct the survey. ## Methodology Travel data were collected from households via in-vehicle (625 vehicles) and wearable (244 participants) global positioning system (GPS) devices. ## Drive Cycle Processing and Filtering NREL has developed a GPS data filtration routine to filter erroneous data points in individual drive cycles sourced from GPS devices mounted in vehicles. Second-by-second drive cycle data collected from GPS-instrumented vehicles during this survey have passed through NREL's drive cycle processing and filtering routines. ## Survey Records Study records include 473 households. ## Transportation Data The SCAG data set contains data from 473 households that participated in one or more areas of study. Of these, 141 completed the wearable GPS portion of the study and 332 completed the vehicle GPS portion. There was no overlap between households participating in the two study areas (wearable and vehicle GPS). For details on available travel survey data and variable definitions, see the [data dictionary](https://www.nrel.gov/media/docs/libraries/tsdc/caltrans_scag_data_dictionary.pdf?sfvrsn=6ec36d7a_1). NREL-generated drive cycle data are also available for this survey. For details on available data and variable definitions, see the [drive cycle data dictionary](https://www.nrel.gov/media/docs/libraries/tsdc/drive_cycles_data_dictionary.pdf?sfvrsn=7de7e888_1). Transportation data are available as zipped files. [Download Winzip](http://www.winzip.com/downwz.htm).

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

2012 California Household Travel Survey Supplement

# 2012 California Household Travel Survey Supplement The 2012 California Household Travel Survey Supplement focused on gathering specific travel information from residents for the development of next-generation, activity-based models. Called the "Augment Survey," it supplemented the [2010–2012 California Household Travel Survey](https://www.nrel.gov/transportation/secure-transportation-data/tsdc-california-travel-survey). ## Data Collection Agency The Southern California Association of Governments (SCAG) hired Abt-SRBI, Inc. to conduct the survey. ## Methodology Travel data were collected from households via in-vehicle (625 vehicles) and wearable (244 participants) global positioning system (GPS) devices. ## Drive Cycle Processing and Filtering NREL has developed a GPS data filtration routine to filter erroneous data points in individual drive cycles sourced from GPS devices mounted in vehicles. Second-by-second drive cycle data collected from GPS-instrumented vehicles during this survey have passed through NREL's drive cycle processing and filtering routines. ## Survey Records Study records include 473 households. ## Transportation Data The SCAG data set contains data from 473 households that participated in one or more areas of study. Of these, 141 completed the wearable GPS portion of the study and 332 completed the vehicle GPS portion. There was no overlap between households participating in the two study areas (wearable and vehicle GPS). For details on available travel survey data and variable definitions, see the [data dictionary](https://www.nrel.gov/media/docs/libraries/tsdc/caltrans_scag_data_dictionary.pdf?sfvrsn=6ec36d7a_1). NREL-generated drive cycle data are also available for this survey. For details on available data and variable definitions, see the [drive cycle data dictionary](https://www.nrel.gov/media/docs/libraries/tsdc/drive_cycles_data_dictionary.pdf?sfvrsn=7de7e888_1). Transportation data are available as zipped files. [Download Winzip](http://www.winzip.com/downwz.htm).

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Distributed Augmentation, Hypersweeps, and Branch Decomposition of Contour Trees for Scientific Exploration

Contour trees describe the topology of level sets in scalar fields and are widely used in topological data analysis and visualization. A main challenge of utilizing contour trees for large-scale scientific data is their computation at scale using highperformance computing. To address this challenge, recent work has introduced distributed hierarchical contour trees for distributed computation and storage of contour trees. However, effective use of these distributed structures in analysis and visualization requires subsequent computation of geometric properties and branch decomposition to support contour extraction and exploration. In this work, we introduce distributed algorithms for augmentation, hypersweeps, and branch decomposition that enable parallel computation of geometric properties, and support the use of distributed contour trees as query structures for scientific exploration. Finally, we evaluate the parallel performance of these algorithms and apply them to identify and extract important contours for scientific visualization.

97 MATHEMATICS AND COMPUTING

Computer vision-based rock bolt detection in orthomosaic imagery obtained in the Waste Isolation Pilot Plant underground facility

Assessing structural integrity of large underground tunnel facilities is often a time consuming and human-labor intensive task. Thus, research using various modes of sensing and automated detection of key structural components in mines is posed to aid in safety assessments and establishing overall structural health. We propose an approach utilizing off-the-shelf camera and lidar technology fixed to a custom sensing platform, image stitching techniques, fine-tuned object detection models, and specialized model-inference methods to automatically detect, count, and map roof bolts for assessment of structural safety in man-made underground tunnels. Results show a novel workflow for effective object counting in orthomosaic tunnel ceiling images generated from collections in GPS-denied mining environments. Additionally, we demonstrate effective fine-tuning of EfficientDet object detectors utilizing state-of-the-art image augmentation techniques known as the mosaic and mixup transformations. Our work is demonstrated on sensed data and imagery collected from the Department of Energy (DOE) Waste Isolation Pilot Plant (WIPP) where miles of tunnel ceiling must be assessed for structural integrity.

42 ENGINEERING

Selection function of clusters in Dark Energy Survey year 3 data from cross-matching with South Pole Telescope detections

Context. Galaxy clusters selected based on overdensities of galaxies in photometric surveys provide the largest cluster samples. However, modeling the selection function of such samples is complicated by noncluster members projected along the line of sight (projection effects) and the potential detection of unvirialized objects (contamination). Aims. We empirically constrained the magnitude of these effects by cross-matching galaxy clusters selected in the Dark Energy Survey data with the redMaPPer algorithm with significant detections in three South Pole Telescope surveys (SZ, pol-ECS, pol-500d). Methods. For matched clusters, we augmented the redMaPPer catalog with the SPT detection significance. For unmatched objects we used the SPT detection threshold as an upper limit on the SZe signature. Using a Bayesian population model applied to the collected multiwavelength data, we explored various physically motivated models to describe the relationship between observed richness and halo mass. Results. Our analysis reveals a clear preference for models with an additional skewed scatter component associated with projection effects over a purely log-normal scatter model. We rule out significant contamination by unvirialized objects at the high-richness end of the sample. While dedicated simulations offer a well-fitting calibration of projection effects, our findings suggest the presence of redshift-dependent trends that these simulations may not have captured. Our findings highlight that modeling the selection function of optically detected clusters remains a complicated challenge that requires a combination of simulation and data-driven approaches.

79 ASTRONOMY AND ASTROPHYSICS

Polyconvex neural network models of thermoelasticity

Machine-learning function representations such as neural networks have proven to be excellent constructs for constitutive modeling due to their flexibility to represent highly nonlinear data and their ability to incorporate constitutive constraints, which also allows them to generalize well to unseen data. Here, in this work, we extend a polyconvex hyperelastic neural network framework to (isotropic) thermo-hyperelasticity by specifying the thermodynamic and material theoretic requirements for an expansion of the Helmholtz free energy expressed in terms of deformation invariants and temperature. Different formulations which a priori ensure polyconvexity with respect to deformation and concavity with respect to temperature are proposed and discussed. The physics-augmented neural networks are furthermore calibrated with a recently proposed sparsification algorithm that not only aims to fit the training data but also penalizes the number of active parameters, which prevents overfitting in the low data regime and promotes generalization. The performance of the proposed framework is demonstrated on synthetic data, which illustrate the expected thermomechanical phenomena, and existing temperature-dependent uniaxial tension and tension-torsion experimental datasets.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC

CAFE AU LAIT: Compute-Aware Federated Augmented Low-Rank AI Training

Federated finetuning is crucial for unlocking the knowledge embedded in pretrained Large Language Models (LLMs) when data are geographically distributed across clients. Unlike finetuning with data from a single institution, federated finetuning allows collaboration across multiple institutions, enabling the utilization of diverse and decentralized datasets while preserving data privacy. Given the high computing costs of LLM training and the emphasis on energy efficiency in Federated Learning (FL), Low-Rank Adaptation (LoRA) has emerged as a widely adopted algorithm due to its significantly reduced number of trainable parameters. However, this assumes that all data silos have the necessary computing resources to compute local updates of LLMs. Nevertheless, in practice, the computing resources across clients are highly heterogeneous: while some may have access to hundreds of GPUs, others might have limited or no GPU access. Recently, federated finetuning using synthetic data has been proposed, allowing clients to participate in a collaborative training run without training LLMs locally. However, our experimental results reveal a performance gap between models trained using synthetic data and those trained using local updates. Motivated by the observed heterogeneity in computing resources and the performance gap, we propose a novel two-stage algorithm that leverages the storage and computing capabilities of a strong server. In the first stage, under the coordination of the strong server, clients with limited computing resources collaborate to generate synthetic data, which is transferred to and stored on the strong server. In the second stage, the strong server uses this synthetic data on behalf of the resource-constrained clients to perform federated LoRA finetuning alongside clients with sufficient computing resources. This approach ensures that all clients can participate in the finetuning process. Experimental results demonstrate that incorporating local updates from even a small fraction of clients improves performance compared to using synthetic data for all clients. Furthermore, we incorporate the Gaussian mechanism in both stages to guarantee client-level differential privacy.

Wang, Jiayi [ORNL]

Integration of 5G and Time Sensitive Networks in Fossil Energy Generation Systems: A Case Study

Precise timing and data transmission within stringent time constraints are critical for numerous applications, such as robotics, virtual and augmented reality, industrial automation, energy and medical business, and various other sectors. Time-sensitive networks (TSN) and fifth-generation wireless communications (5G) are crucial for industrial communications, enabling convergent communication for various services using a common network core. Applications that are time-sensitive and require deterministic communications with low latency fall into this category, such as generation plant systems. This paper presents a simulated model of 5G-TSN for a fossil power plant based on the wired network parameters implemented in situ. Metrics analysis, comparison and future work are presented. © 2024 IEEE.

01 COAL, LIGNITE, AND PEAT

HERO CarbonSAFE Phase 2 Project in the Columbia River Basalt Group

The Hermiston, Oregon Basalt CarbonSAFE Phase II project (HERO CarbonSAFE) seeks to accelerate the deployment of commercial carbon dioxide (CO2) storage projects in basaltic rocks. Basalt CO2 storage has several advantages to conventional saline storage reservoirs including 1. The potential for rapid mineralization of CO2, 2. Associated decreases in pressure and CO2 migration risks, 3. Reduced long-term monitoring requirements with respect to plume tracking, 4. Widespread geographic distribution and, 5. Large storage potential due to thickness, porosity, and CO2 interactions with basalt. And for locations such as the Pacific Northwest, Hawaii, Iceland, India and Japan, basalts may offer the only economically feasible option for local CO2 storage. However, there are limited field-scale assessments of CO2 storage in basalt, and current carbon capture utilization and storage (CCUS) permitting and regulatory frameworks were developed for conventional saline reservoirs. HERO CarbonSAFE is designed to address research gaps and uncertainties associated with basalt storage. Specifically, the project will assess the feasibility of CO2 injection in the deep layered basalts, long-term storage (mineralization), practical approaches for large-scale implementation (50+ million metric tons of CO2 over 30 years), lithology-specific risks, and the technoeconomic potential for CO2 storage in basalts. The HERO CarbonSAFE project will assess feasibility of developing a commercial-scale (50+ million metric tons of CO2) geological storage complex within the Columbia River Basalt Group (CRBG), a layered continental flood basalt complex that underlies Calpine’s natural gas-fired Hermiston Power Project (HPP) in Hermiston, OR (Figure 1). Under this 2-year CarbonSAFE Phase II project, the HERO team will conduct a data acquisition campaign that includes drilling a stratigraphic well to a total depth of ~1,500 m into the thick layered basalts proximal to HPP. A comprehensive well logging and hydrologic testing program will be augmented with new core collected from flow zones and sealing units, and comprehensive laboratory testing to help refine the kinetic rates of mineralization. The newly acquired information will be integrated with existing data from regional wells to correlate basalt injection zone properties to develop storage hub/commercial-scale models. Using these models, the project team will evaluate injection scenarios to define the technical and economic potential for storing a minimum of 50 million metric tons of CO2 over a 30-year period, along with a robust sensitivity analysis on key parameters governing reservoir viability for sustainable injection over a commercial project lifetime. Specific technical objectives of HERO are: (1) assessing the reservoir response of a series of stacked layered reservoir flowtop sequences occurring in this area of the CRBG to commercial-scale injection volumes; (2) extending prior efforts by the project team to characterize the deep layered basalts encountered in regional studies, to leverage prior investments by U.S. Department of Energy’s (DOE) Carbon Storage program; (3) leveraging DOE’s mineralization characterization efforts to advance model parametrization for commercial scale injection of CO2 in basalts; (4) conducting risk assessments associated with scaling up to commercial storage hub injection goals, while validating DOE’s National Risk Assessment Partnership (NRAP) tools, to identify potential constraints that would prevent the CRBG from serving as a commercial-scale storage complex; (5) developing mitigation plans to address identified risks; (6) developing a commercial-scale injection and monitoring, verification and accounting (MVA) strategy; (7) utilizing computational models to define and minimize, if possible, the Area of Review (AoR) under Class VI regulations; and (8) developing a robust CO2 management strategy for CRBG that also considers a regional source/sink approach that is responsive to stakeholder needs and industrial demand. Specific institutional objectives are: (1) identifying and developing plans to mitigate the nontechnical challenges associated with the build-out of a commercial-scale storage complex within the CRBG with integrated CO2 sources; (2) implementing the community outreach plan; (3) conducting regulatory research, including a survey of issues related to pore space ownership, MVA and long-term assurance of mineralization-based storage, to support an eventual application for a UIC Class VI permit; (4) advancing the project’s plan for CO2 liability management; and (5) continuing to refine and update the project’s economic model. The final objective is the preparation of a comprehensive Site Characterization Plan that draws upon the technical and institutional feasibility assessments to prepare the project for future commercialization efforts.

58 GEOSCIENCES

MPEX AI Digital Twins

All magnetically confined plasma fusion power plant concepts (Tokamak, Spherical Tokamak, Stellarator, Mirror, ...) must exhaust the heat and plasma from the core confinement region to the material walls. The primary channel for this exhaust is through a plasma divertor which directs plasma along open magnetic field lines to a material target. The Material Plasma Exposure eXperiment (MPEX) illustrated in Figure 1, is a high-power, steady-state linear plasma device designed to produce the plasma material interaction (PMI) conditions of the divertor of future magnetic confinement fusion power plants: energy flux 20MW/m 2 , ion fluence 1031/m 2 , pulse duration 106 sec. These goals of plasma exposure in MPEX are well beyond those achieved in magnetic fusion experimental devices. Successfully achieving these high power steady state conditions for long pulses requires operational control of the heating and particle sources and the plasma flux to the walls and target. The MPEX AI Hot Spot Controller, proposed in this project, will help achieve the operational milestones of MPEX. The MPEX device will begin commissioning at the end of FY26. A smaller proto-MPEX was operated for 14,666 plasma discharges and will resume operation in September of 2025 as proto-MPEX-lite, with reduced capability, to test a new window for the Helicon plasma source. The proto-MPEX data has undergone surrogate modeling with machine learning methods (R. Archibald, 2022 IEEE International Conference on Big Data). This proto-MPEX data will be used to begin development of the AI digital twins described in this white paper. The scientific mission of MPEX is to qualify materials of different composition for use in the high energy and plasma flux conditions of a fusion power plant. The materials exposed in MPEX will in some cases be exposed to high neutron fluxes at other ORNL facilities to measure the changes to their PMI properties. The targets exposed in MPEX will be transported under vacuum to a Surface Analysis Station (SAS). The SAS will be equipped with the following diagnostics: Focused Ion Beam (FIB) for trench milling, 100-400 angstrom resolution scanning electron microscope (SEM), surface mapping x-ray spectrometer, high resolution camera, and a future upgrade to a laser induced breakdown spectroscopy quadruple mass spectrometer (LIBS-QMS). The MPEX experiments will generate diverse pre- and post-exposure measurement data of detailed material properties down to the crystal grain level in 3D for post-exposure assessment of PMI damage (e.g. cracking, melting, erosion and redeposition of the material). Physics models for the PMI, and how the material composition and manufacturing impact its performance under high energy plasma exposure, need to be validated with MPEX data to guide the selection of new candidate materials. Our vision for the MPEX AI Digital Twins project is to supply experimental and physics model simulation data to train Artificial Intelligence (AI) models for data processing, analysis, operational control, PMI and materials simulation to maximize the scientific output of the MPEX device. Ultimately, an AI digital twin of MPEX material assessment metrics for tested and synthetic material types with simulated PMI will be trained by the AI Modeling Teams on the experimental and physics simulation data submitted to the American Science Cloud by this project. A purely empirical search for the best material is inefficient given the finite number of samples that can be tested on MPEX. In order to expand the material properties database for training the MPEX Material Assessment AI Digital Twin, and to gain physics understanding of the PMI processes, physics models of the material properties and PMI processes are required. The physics simulations provide detailed simulation data, like impact angles for plasma ions, sputtering yields, transport of the ionized sputtered target material in the plasma, and redeposition locations. This simulation data expands the measurement data for deeper physics understanding. The experimental data is essential to validate the PMI and material structure simulation models. The validated models can then be used to generate new simulation data of MPEX material assessments for synthetic material compositions that have not been exposed in MPEX. These predictive simulations, plus the whole experimental dataset, will be used to train the MPEX Material Assessment AI Digital Twin allowing a rapid generative AI search for new materials with reduced PMI damage by interpolating the domain of the training set. These new optimum materials can be simulated with the physics codes and/or tested in MPEX. The ability of AI neural networks to interpolate multi-dimensional parameter spaces and generate virtual data is exploited for a more efficient search for optimum materials. The advent of the Transformational AI Models Consortium (TAIMC) is an opportunity to engage with state of the art private and public AI developers to achieve the goals of the AI digital twins and AI accelerated physics models proposed in this project. Our partners at ORNL from the Advance Scientific Computing Research (ASCR) organization will collaborate in accelerating the integrated plasma material interaction simulation framework. This simulation framework will provide a platform for generating simulation data across a range of physical fidelities, including hybrid methods that produce multi-fidelity results. This data will be leveraged for AI model development, both for generation of surrogates and the automation of simulation campaigns. A part of the research below will include collaborative efforts with the TAIMC to (i) adapt data storage approaches to ensure AI-readiness, (ii) provide a protypical exemplar to inform and exercise constructed workflows, and (iii) generate and share data, using the TAIMC unified AI data standard, for foundational models that will be trained from multiple sources across the DOE complex. We will also collaborate with the TAIMC, as well as the planned AI modeling teams, to develop approaches for reducing the cost of data generation. These include tailored multi-fidelity approaches as well as fine-tuning strategies to augment general, large-scale foundational models.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

Site and endmember spectra of terrestrial vegetation and soils for the Colorado Headwaters Ecological Spectroscopy Study, June-July 2025

This dataset provides site and endmember spectra collected during the 2025 Colorado Headwaters Ecological Spectroscopy Study (CHESS) campaign. The site spectra were collected to help validate airborne hyperspectral data acquired by the National Ecological Observatory Network's aerial observation platform (NEON AOP). Endmember spectra were collected to augment existing spectral libraries with additional samples of bare surfaces and non-photosynthetic vegetation. All measurements were acquired with an Analytical Spectral Devices (ASD) FieldSpec4 Hi-Res NG (Next Generation) spectroradiometer, which records radiance at 1nm (nanometer) intervals from the ultraviolet to the short-wave infrared (350-2500 nm). The dataset includes spectra measured at meadow sites where the CHESS team also collected vegetation samples for trait analyses. The site spectra were collected with the ASD FieldSpec4 palm grip attachment using an 8° field-of-view foreoptic. Site spectra are integrated measurements of the entire surface within the foreoptic’s field of view. For site-level spectra, the sun is the illumination source. A Spectralon panel mounted on a tripod was used for instrument optimization and white reference measurements for all site spectra. Site spectra were acquired within two hours of solar noon and within 48 hours of a NEON AOP overflight. Site spectra are labeled by date, sampling area, and site number according to the naming conventions of the CHESS campaign’s data management plan. The dataset also contains endmember spectra in the following categories: photosynthetic vegetation (PV), non-photosynthetic vegetation (NPV), bare (soil/rock), and flowers. Endmember measurements were acquired using either the contact probe or the leaf clip attachments of the ASD FieldSpec4. In these configurations, the bulb inside the spectrometer provides the light source for the measurements. The spectrometer was optimized and white reference measurements were recorded using the circular white pucks attached to the contact probe and leaf clip. Because they do not rely on solar illumination, contact probe and leaf clip measurements were collected during a broader time frame than the palm grip site spectra. Some endmembers were measured at CHESS meadow sites, while others were collected within the larger sampling area or in nearby locations (e.g. Gothic Townsite) with similar characteristics. Radiance, reflectance, and metadata files are split into three subfolders according to measurement type: proximal/palm grip (prx), contact probe (cp), and leaf clip (lc). Radiance spectra are provided in ASD file format (.asd file extension). All ASD files can be opened using the provided scripts. Metadata is provided in two formats: CSV file format (no geolocation) and GEOJSON file format (includes geolocation for each spectra). The dataset includes a set of pre-processed reflectance spectra as CSV files (yyyymmdd_rfl.csv). The python scripts and jupyter notebook used to calculate reflectance spectra from the ASD radiance data is included here and was previously published at: https://doi.org/10.3334/ORNLDAAC/2446. There is also a folder of JPEG photographs corresponding to selected spectra. We include a protocol document with detailed steps for ASD FieldSpec4 assembly and operations. This data additionally contains a file level metadata (flmd.csv) and data dictionary (dd.csv) file. Geospatial information: Geospatial data for mapping measurement site locations are in the files CHESS_polygons_lai_UTM.geojson, CHESS_polygons_shrub_UTM.geojson, and CHESS_polygons_meadow_UTM.geojson in the companion geospatial package for the 2025 CHESS campaign, ‘CHESS 2025: Location data for field observations and sampling’ (Henderson et al., 2026). CHESS Project Description: The Colorado Headwaters Ecological Spectroscopy Study (CHESS) comprised a multi-week airborne remote sensing and field observation campaign in the Upper Gunnison Basin, Colorado, conducted in June and July of 2025. Airborne remote sensing was conducted by the National Ecological Observatory Network Airborne Observation Platform (NEON AOP), concurrent with a field campaign run by the Rocky Mountain Biological Laboratory (RMBL), the Lawrence Berkeley National Laboratory (LBNL) and SLAC National Accelerator Laboratory Watershed Function Science Focus Area (SFA), and NASA-JPL (Jet Propulsion Laboratory) Earth Surface Mineral Dust Source Investigation (EMIT) program. Between June 10 and July 18, 2025, the NEON AOP flight team collected high-resolution aerial imaging spectroscopy and Light Detection and Ranging (LiDAR) data over three domains: the Upper East River (CRBU), Almont Triangle (ALMO), and the Upper Taylor Basin (UPTA). In coordination with the flights, a field campaign acquired ground-truth observations, including observations of vegetation composition, foliar traits, forest demography, and subsurface properties in 18 core sampling areas within the domains. Additional surface water observations were taken at over 380 point locations. All CHESS campaign datasets can be found within the CHESS ESS-DIVE data portal: https://data.ess-dive.lbl.gov/portals/chess. Funding Acknowledgment: This research was carried out at the Jet Propulsion Laboratory, California Institute of Technology, under a contract with the National Aeronautics and Space Administration (80NM0018D0004) and was funded by EMIT Extended Mission Phase E Science.

2018 NEON and 2025 CHESS Campaigns

Lessons Learned from AskGDR: Usage and Impact Analysis of the Geothermal Data Repository's AI Research Assistant: Preprint

In October of 2024, the Department of Energy's (DOE) Geothermal Data Repository (GDR) team officially launched AskGDR, an AI research assistant resulting from the integration of a Large Language Model (LLM) with the metadata and supporting documents associated with GDR datasets. AskGDR allows GDR users to ask deeper questions about the origin of datasets, the methods used to collect them, and the findings they help support. Using Retrieval Augmented Generation (RAG), AskGDR can be used to summarize findings spread across dozens of papers and technical reports or to extract relevant information describing a single data field. However, generative AI is experimental. The National Renewable Energy Laboratory (NREL) has been collecting metrics on AskGDR and documenting lessons learned during its deployment. This paper will outline the efficacy and impact of AskGDR through analysis of its use, operating costs, number and types of questions asked, and the quality of answers provided.

15 GEOTHERMAL ENERGY

CAHS: Context-Aware Homology Search

Protein homology search is foundational to bioinformatics: it supports annotation transfer, structure/function inference, and evolutionary analysis over rapidly expanding sequence repositories (e.g., UniProtKB). Profile hidden Markov models (pHMMs), as implemented in HMMER, remain the most widely trusted approach because they provide statistically calibrated E-values; however, their gap behavior is fixed once a profile is trained, despite biological evidence that insertion/deletion tolerance varies across flexible loops and intrinsically disordered regions. We present CAHS (Context-Aware Homology Search), a lightweight query-time adapter for pHMM search that incorporates learned and biologically motivated signals without changing HMMER's downstream search pipeline or its calibrated E-value reporting. Given a query sequence, CAHS computes per-residue representations from a protein language model and a disorder predictor, maps these to profile coordinates, and modulates only match-state transition rows (gap-open and gap-extension probabilities) while preserving Plan7 constraints. We comprehensively evaluate CAHS across six structurally diverse protein families and multi-domain architectures against a 570k-sequence target corpus. CAHS expands detection capability, retrieving thousands of additional remote homologs at relaxed thresholds by maintaining alignment quality through flexible regions. For multi-domain proteins, context-aware modulation resolves 94% of fragmented alignments. Crucially, CAHS preserves hit-set invariance at stringent operating points (E<10-10), demonstrating increased statistical confidence without inflating false positives. Furthermore, sharper statistical distinction between homologs and background noise during early filter stages yields up to a 3.87× acceleration in end-to-end wall-clock time on high-performance computing clusters. Overall, CAHS illustrates a practical AI-for-science design pattern: augmenting a trusted probabilistic model with query-specific learned signals to improve interpretable, reproducible inference in data-rich biology.

Bhattaram, Swethasree [Georgia Institute of Techno

Antarctic lake viromes reveal potential virus associated influences on nutrient cycling in ice-covered lakes

The McMurdo Dry Valleys (MDVs) of Antarctica are a mosaic of extreme habitats which are dominated by microbial life. The MDVs include glacial melt holes, streams, lakes, and soils, which are interconnected through the transfer of energy and flux of inorganic and organic material via wind and hydrology. For the first time, we provide new data on the viral community structure and function in the MDVs through metagenomics of the planktonic and benthic mat communities of Lakes Bonney and Fryxell. Viral taxonomic diversity was compared across lakes and ecological function was investigated by characterizing auxiliary metabolic genes (AMGs) and predicting viral hosts. Our data suggest that viral communities differed between the lakes and among sites: these differences were connected to microbial host communities. AMGs were associated with the potential augmentation of multiple biogeochemical processes in host, most notably with phosphorus acquisition, organic nitrogen acquisition, sulfur oxidation, and photosynthesis. Viral genome abundances containing AMGs differed between the lakes and microbial mats, indicating site specialization. Using procrustes analysis, we also identified significant coupling between viral and bacterial communities (p = 0.001). Finally, host predictions indicate viral host preference among the assembled viromes. Collectively, our data show that: (i) viruses are uniquely distributed through the McMurdo Dry Valley lakes, (ii) their AMGs can contribute to overcoming host nutrient limitation and, (iii) viral and bacterial MDV communities are tightly coupled.

Microbiology

Large Language Model Integration for Knowledge Retrieval and Interaction for the DUNE Experiment

The Deep Underground Neutrino Experiment (DUNE) is a next-generation neutrino experiment that will generate an unprecedented volume of heterogeneous information-from documentation and technical notes to experimental data and reconstruction pipelines. Efficient knowledge retrieval and contextual understanding are increasingly critical for collaboration-wide productivity and onboarding. In this work, we present DUNE-GPT, a prototype framework that leverages large language models (LLMs) and retrieval-augmented generation (RAG) to enable natural-language querying of DUNE's internal documentation and technical resources. The system provides an intelligent interface for DUNE collaborators to interact with experiment-specific knowledge while maintaining data privacy and infrastructure compliance within Fermilab computing resources.

Rafique, A. [Argonne (main)]