Search NASASearch

SEARCH · Search NASA

Results for “DATA RETRIEVAL”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Error-controlled Progressive Retrieval of Scientific Data under Derivable Quantities of Interest

The unprecedented amount of scientific data has introduced heavy pressure on the current data storage and transmission systems. Progressive compression has been proposed to mitigate this problem, which offers data access with on-demand precision. However, existing approaches only consider precision control on primary data, leaving uncertainties on the quantities of interest (QoIs) derived from it. In this work, we present a progressive data retrieval framework with guaranteed error control on derivable QoIs. Our contributions are three-fold. (1) We carefully derive the theories to strictly control QoI errors during progressive retrieval. Our theory is generic and can be applied to any QoIs that can be composited by the basis of derivable QoIs proved in the paper. (2) We design and develop a generic progressive retrieval framework based on the proposed theories, and optimize it by exploring feasible progressive representations. (3) We evaluate our framework using five real-world datasets with a diverse set of QoIs. Experiments demonstrate that our framework can faithfully respect any user-specified QoI error bounds in the evaluated applications. This leads to over 2.02× performance gain in data transfer tasks compared to transferring the primary data while guaranteeing a QoI error that is less than 1E-5.

Wu, Xuan

Satellite Retrievals

Data from the SENTINEL-1 satellite constellation (SENTINEL-1a and SENTINEL-1b), from PNNL.

17 WIND ENERGY

QProR: An Efficient Framework for Quantity-of-Interest Based Progressive Retrieval with Guaranteed Error Control

Scientific applications generate an unprecedented volume of data, overwhelming the network and file systems’ bandwidth and posing challenges for efficient and scalable data retrieval and analysis. Progressive data compression offers a promising solution by enabling on-demand retrieval at reduced size. However, existing progressive methods either fail to bound the errors in essential quantities of interest (QoIs) derived from raw data or suffer from suboptimal retrieval efficiency. In this work, we propose QProR, an efficient QoI-based progressive framework that optimizes progressive retrieval for target QoIs. Our key contributions include: (1) a systematic framework that integrates error-controlled lossy compressors with bitplane encoding while decoupling the two processes for high flexibility and adaptability; (2) a novel weighted bitplane encoding method which incorperates QoI knowledge into data refactoring to enhance retrieval efficiency; (3) an optimized retrieval strategy that accounts for the varying impacts of different variables on multivariate QoIs; (4) comprehensive evaluations using six real-world datasets from multiple scientific applications and thorough comparisons against state of the arts. Experimental results demonstrate that QProR achieves up to 80.38% reduction in the retrieval size under the same requested QoI error tolerance, when compared with the best-performing existing methods. When transferring 384 GB of scientific data to remote sites, QProR delivers up to 1.68 × speedup in the end-to-end data transfer performance.

Li, Wenbo [University of Kentucky]

Sensos Smart Label Performance Summary as Observed by Oak Ridge National Laboratory

The Oak Ridge National Laboratory (ORNL) team performed an evaluation of the Sensos Smart Label Gen 2.0, as shown in Figure 1, for package tracking. A long-distance round-trip shipment between Oak Ridge, Tennessee, and Seattle, Washington, was completed to assess the device’s performance in location tracking, environment sensing capabilities, alerting features, threshold options, battery life, and real-time and historical data retrieval from the “Sync” data dashboard provided by Sensos. The evaluation was conducted to gain a general understanding of the capabilities of the device. Furthermore, due to time and resource constraints, ORNL did not conduct exhaustive testing to confirm reliability, availability, or effectiveness of alerting and tracking features. On equipment arrangement, Sensos (sensos.ai) graciously agreed to provide a Sensos Smart Label Gen 2.0 device to ORNL, at no cost, for testing and evaluation purposes. ORNL conducted assessments along with other commercial off-the-shelf (COTS) tracking devices. As a courtesy, ORNL will provide Sensos with this report summarizing the observations and findings specific to the Sensos label based on the tests performed.

42 ENGINEERING

HP-MDR: High-performance and Portable Data Refactoring and Progressive Retrieval with Advanced GPUs

Scientific applications produce vast amounts of data, posing grand challenges in the underlying data management and analytic tasks. Progressive compression is a promising way to address this problem, as it allows for on-demand data retrieval with significantly reduced data movement cost. However, most existing progressive methods are designed for CPUs, leaving a gap for them to unleash the power of today’s heterogeneous computing systems with GPUs.In this work, we propose HP-MDR, a high-performance and portable data refactoring and progressive retrieval framework for GPUs. Our contributions are four-fold: (1) We carefully optimize the bitplane encoding and lossless encoding, two key stages in progressive methods, to achieve high performance on GPUs; (2) We propose pipeline optimization and incorporate it with data refactoring and progressive retrieval workflows to further enhance the performance for large data process; (3) We leverage our framework to enable high-performance data retrieval with guaranteed error control for common Quantities of Interest; (4) We evaluate HP-MDR and compare it with state of the arts using five real-world datasets. Experimental results demonstrate that HP-MDR delivers an average 13.68 × and 6.31 × throughput in data refactoring and progressive retrieval tasks, respectively. It also leads to 11.22 × throughput for recomposing required data representations under Quantity-of-Interest error control and 6.04 × performance for the corresponding end-to-end data retrieval, when compared with state-of-the-art solutions.

Li, Yanliang [University of Oregon]

Radio Frequency Spectrum Audit to Inventory Private Cellular Base Station Infrastructure

The ever-changing cellular communication landscape makes it difficult to identify, map, and localize cellular base stations. Localizing cellular base stations provides various advantages, including information security, cybersecurity, spectrum management, and interference detection. For example, the MITRE ATT&CK® (Adversarial Tactics, Techniques, and Common Knowledge architecture) [1] and Common Attack Pattern Enumeration and Classification [2] emphasize the importance of being able to minimize the cyber security threat presented by unregulated private cellular base stations (PCBS). The majority of published research looks at the malicious use of PCBSs and focuses on using data retrieved from user equipment (UE), data obtained from an application on the UE, or data shared between the UE and a mobile network to locate it. This innovative strategy, however, focuses on the passively discovered uniqueness of radio frequency (RF) transmissions from commercial cellular infrastructure received in a designated monitoring position (DMP).

42 ENGINEERING

Remotely Sensed High‐Resolution Soil Moisture and Evapotranspiration: Bridging the Gap Between Science and Society

This paper reviews the current state of high‐resolution remotely sensed soil moisture (SM) and evapotranspiration (ET) products and modeling, and the coupling relationship between SM and ET. SM downscaling approaches for satellite passive microwave products leverage advances in artificial intelligence and high‐resolution remote sensing using visible, near‐infrared, thermal‐infrared, and synthetic aperture radar sensors. Remotely sensed ET continues to advance in spatiotemporal resolutions from MODIS to ECOSTRESS to Hydrosat and beyond. These advances enable a new understanding of bio‐geo‐physical controls and coupled feedback mechanisms between SM and ET reflecting the land cover and land use at field scale (3–30 m, daily). Still, the state‐of‐the‐science products have their challenges and limitations, which we detail across data, retrieval algorithms, and applications. We describe the roles of these data in advancing 10 application areas: drought assessment, food security, precision agriculture, soil salinization, wildfire modeling, dust monitoring, flood forecasting, urban water, energy, and ecosystem management, ecohydrology, and biodiversity conservation. We discuss that future scientific advancement should focus on developing open‐access, high‐resolution (3–30 m), sub‐daily SM and ET products, enabling the evaluation of hydrological processes at finer scales and revolutionizing the societal applications in data‐limited regions of the world, especially the Global South for socio‐economic development.

54 ENVIRONMENTAL SCIENCES

ARM Lead Mentor Selection Process

The Atmospheric Radiation Measurement (ARM) Program was created in 1989 with funding from the U.S. Department of Energy (DOE) to develop several highly instrumented ground stations to study cloud-formation processes and their influence on radiative transfer. This scientific infrastructure provides for fixed sites, mobile facilities, an aerial facility, and a data archive available for use by scientists worldwide through the ARM Climate Research Facility—a scientific user facility. The ARM Climate Research Facility currently operates more than 300 instrument systems that provide ground-based observations of the atmospheric column. To keep ARM at the forefront of climate observations, the ARM infrastructure depends heavily on instrument scientists and engineers, known as Mentors. Mentors must have an excellent understanding of instrumentation theory and operation for their instrument areas and have comprehensive knowledge of critical scale-dependent atmospheric processes. They must also possess the technical and analytical skills to develop new data retrievals that provide innovative approaches for creating research-quality data sets. The ARM Facility seeks the best overall qualified candidate, or team when appropriate, that can fulfill Mentor requirements in a timely manner. The roles and responsibilities of the ARM Instrument Operations Manager are provided in Appendix A. The key role and responsibilities and detailed responsibilities of ARM Lead Mentors are provided in Appendix B and Appendix C, respectively.

47 OTHER INSTRUMENTATION

Heat load measurements for the PIP-II pHB650 cryomodule

This study presents a brief overview of the 1st and 2nd phases and an in-depth analysis of the 3rd phase heat load testing performed on the pHB650 (prototype High Beta 650 MHz) cryomodule at PIP2IT (PIP-II Injector Test Facility), with a focus on both the results and the methodological advancements that have improved testing efficiency and accuracy. A key challenge identified in the testing campaign is the higher-than-expected heat loads observed in the first PIP-II (Proton Improvement Plan II) prototype cryomodules (pSSR1 and pHB650) tested at PIP2IT. Elevated heat loads are concerning given the fixed capacity of the PIP-II cryoplant that is currently being installed at Fermilab. However, understanding the sources of these elevated heat loads offers a critical opportunity to implement effective heat load mitigations on upcoming PIP-II cryomodules to stay within the available capacity of the PIP-II cryoplant. The study includes a summary of test results, descriptions of measurement procedures, and key observations on parameters directly and indirectly related to heat load measurements. Direct observations include measured heat loads and the effectiveness of JT heat exchanger under varying conditions, while indirect observation analyze factors such as the temperature distribution on the two-phase pipe and relief piping under varying conditions. Thermal acoustic oscillations (TAO) were identified during testing, which was mitigated by replacing the original G10 stem with a stainless steel stem equipped with wipers for the cryomodule cooldown valve. A major innovation during pHB650 Phase 3 testing was the development of an automated Python script to streamline data acquisition, analysis, and reporting of heat load results. This script automatically retrieved data from ACNET (Accelerator Control Network), performed heat load calculations, and generated detailed reports featuring plots and tables. This advancement significantly reduced manual labor and enhanced the thoroughness of data analysis compared to earlier campaigns. The heat load test reports were promptly uploaded to the electronic logbook shortly after each test, enabling rapid feedback and collaboration between the SRF and cryogenic teams. The heat load measurements included various components: HTTS (high-temperature thermal shield), LTTS (low-temperature thermal shield), 2K isothermal and non-isothermal heat loads. Results were recorded both within the cryomodule and between the bayonet can supply and return. Measurements were conducted under different operating conditions such as "standard", "linac", and "simulated dynamic". Additionally, HTTS and LTTS heat loads were calculated in real time, allowing for the tracking of thermal stability and identification of changes during testing, both in steady-state and transient conditions. The results of this testing campaign not only provide valuable insights into the performance of the pHB650 cryomodule but also highlight best practices and lessons learned that will inform future cryomodule testing at PIP2IT. These include adopting automated tools for data analysis, refining real-time measurement capabilities, and emphasizing detailed pre-test planning. The framework established in this campaign aims to set an improved standard for cryomodule testing and heat load reporting in future cryomodule test campaigns.

Porwisiak, D. [Fermilab; Wroclaw Tech. U.]

Performance and Reliability Assessment of the U.S. Department of Energy Atmospheric Radiation Measurement (ARM) Data Advisor (ADA)

The Atmospheric Radiation Measurement (ARM) User Facility provides one of the world's largest openly accessible repositories of atmospheric observations through the ARM Data Discovery platform. Although the repository contains more than three decades of measurements collected from permanent observatories, mobile facilities, aircraft campaigns, and field experiments, identifying appropriate datasets can be challenging, particularly for new users unfamiliar with ARM instrumentation and datastream organization. To improve data accessibility, the ARM Data Center developed the ARM Data Advisor (ADA), an artificial intelligence-powered assistant designed to facilitate scientific data discovery, dataset interpretation, and user guidance. This report evaluates ADA's performance as a domain-specific scientific assistant using realistic atmospheric science workflows. The evaluation examines five key capabilities: data retrieval and curation efficiency, hallucination resistance, scientific reasoning, response to ambiguous queries, and content retention and session continuity. Representative prompts were developed to simulate typical interactions between researchers and the ARM Data Discovery platform, and ADA's responses were assessed for retrieval completeness, scientific accuracy, consistency, and practical usefulness. In these representative tests, ADA reduced the complexity of discovering and accessing ARM datasets by recommending appropriate datastreams, explaining instrumentation, interpreting metadata, and assisting with data processing workflows. ADA also exhibits strong domain knowledge of atmospheric science terminology and generally resists hallucination by acknowledging unavailable datasets and requesting clarification when appropriate. Overall, the results indicate that ADA represents a promising advancement in scientific data discovery within the ARM User Facility and has considerable potential to improve researcher productivity, particularly for new users and interdisciplinary scientists seeking efficient access to ARM observations.

Salvador, Christian [ORNL] (ORCID:0000000283287777

Compressed sensing methods with applications to advanced air sampling

Environmental sampling methods developed by the Savannah River National Laboratory (SRNL) employ collectors with sorbent media tubes set at various locations to collect airborne emissions. Laboratory analyses of these tubes results in one-dimensional signals regarding what chemicals are being released and transported within the atmosphere. The analysis process is time consuming especially when analyzing a full year’s worth of tubes (hourly sample collection results in nearly 9,000 tubes per year). Using a signal processing method such as compressed sensing allows for recreation of the full signal while greatly reducing the number of analyzed samples required. Due to the sparsity of data retrieved from the air tubes, it is possible to use measurements a fraction of the size of the original data to gain much of the same information. This would improve the overall time and cost of analysis when modeling one-dimensional sampling signals.

54 ENVIRONMENTAL SCIENCES

Compressed Sensing Methods with Applications to Advanced Air Sampling [Poster]

Environmental sampling methods developed by the Savannah River National Laboratory (SRNL) employ collectors with sorbent media tubes set at various locations to collect airborne emissions. Laboratory analyses of these tubes results in one-dimensional signals regarding what chemicals are being released and transported within the atmosphere. The analysis process is time consuming especially when analyzing a full year’s worth of tubes (hourly sample collection results in nearly 9,000 tubes per year). Using a signal processing method such as compressed sensing allows for recreation of the full signal while greatly reducing the number of analyzed samples required. Due to the sparsity of data retrieved from the air tubes, it is possible to use measurements a fraction of the size of the original data to gain much of the same information. This would improve the overall time and cost of analysis when modeling one-dimensional sampling signals.

Campbell, Cassidy [Savannah River National Laborat

JHTDB-wind: a web-accessible large-eddy simulation database of a wind farm with virtual sensor querying

This paper introduces JHTDB-wind (https://turbulence.idies.jhu.edu/datasets/windfarms, last access: 11 November 2025), a publicly accessible database containing large-eddy simulation (LES) data from wind farms. Building on the framework of the Johns Hopkins Turbulence Database (JHTDB), which hosts direct numerical simulation (DNS) and some LES datasets of canonical turbulent flows, JHTDB-wind stores the 4D space–time history of the flow and provides users the ability to access and query the data via a web-based virtual sensor interface. The initial dataset comprises LES results from a large wind farm with 10×6 turbines, modeled using a filtered actuator line method, under conventionally neutral atmospheric conditions. These data comprise 1 h (hour) of flow field data (velocity, pressure, potential temperature deviation, subgrid-scale (SGS) eddy viscosity, and turbine forces, approximately 15 TB (terabytes) and wind turbine data – including both turbine-level operational quantities and blade-level aerodynamic quantities (approximately 1.3 TB) – stored in Zarr and Parquet formats, respectively. Data retrieval is facilitated by the giverny Python package, allowing remote users to query the database in Python or MATLAB (C and Fortran support are available for flow field data). This paper details the simulation setup and demonstrates data access through examples that analyze wind farm flow structures and turbine performance. The framework is extensible to future datasets, including the JHTDB-wind diurnal cycle simulation analyzed in Xiao et al. (2025).

17 WIND ENERGY

Comparison of horizontal wind speed and direction measurements from dual-Doppler radar and profiling lidars

Dual-Doppler radar is a relatively new technology in the wind energy community and thus not yet studied vastly. This paper aims to compare horizontal wind speed and direction data retrieved from dual-Doppler radar and profiling lidars within the American WAKE experimeNt (AWAKEN) to investigate the influence of measurement height, wind direction and speed on the comparison. The 10-min averaged data show a better agreement of the measurements for higher altitudes, especially at faster wind speeds. For the wind direction, two sectors of larger differences in the measurements were detected: around 270° transient winds occur with a higher frequency than in other sectors. To explain the different measurement values in the wind direction sector around 90°, further studies, e.g. on the influence of atmospheric stability, are necessary.

17 WIND ENERGY

ATcT — Active Thermochemical Tables Python Interface

SF-25-140 atct is a lightweight, Python client for the ATcT v1 API that enables programmatic access to high-accuracy thermochemical data and turnkey reaction-enthalpy analysis. The package implements full v1 endpoint coverage (species lookup by ATcT ID, name, formula, SMILES, InChI, CAS RN; covariance queries; health checks) with robust error handling, retries, and environment-based configuration for local/production endpoints. Beyond data retrieval, atct provides rigorously implemented reaction calculators that propagate uncertainties via either (i) a conventional independent-errors method (0 K or 298.15 K) or (ii) covariance-aware propagation using provided covariances at 298.15 K. Typed data classes ensure transparent, reproducible data structures and carry ATcT Thermochemical Network (TN) version identifiers for provenance. Dual import paths and comprehensive examples facilitate integration into research pipelines, enabling reproducible thermochemical calculations, automated validation, and downstream method development.

Bross, DavidHamilton [Argonne National Laboratory

Digitizing and Enhancing Accessibility of the Fusion Safety Archives

This project focuses on the digitization and public accessibility to the Fusion Safety Archives at the Idaho National Laboratory. The first phase involves a thorough review of each document in the physical archives to determine its online availability. For documents that are available online, PDF copies and unique identifiers are collected for database integration. Documents not available online are delivered to Red Inc. for digitization. Additionally, defunct storage devices such as diskettes are sent to INL’s archival department for data retrieval where possible. The second phase of the project involves the creation of a comprehensive database to house the digital copies of the archives. The database will facilitate easy access and management of the digitized documents. Following the database creation, we plan to train a Retrieval-Augmented Generation (RAG) based AI on publicly available documents. The trained AI will be integrated into a front-facing application, allowing the public to easily access information from the Fusion Safety Archives. This project aims to preserve valuable historical data, improve accessibility, and promote transparency in fusion safety research.

70 - PLASMA PHYSICS AND FUSION TECHNOLOGY

Towards a RAG-based summarization for the Electron Ion Collider

Abstract The complexity and sheer volume of information — encompassing documents, papers, data, and other resources — from large-scale experiments demand significant time and effort to navigate, making the task of accessing and utilizing these varied forms of information daunting, particularly for new collaborators and early-career scientists.To tackle this issue, a Retrieval Augmented Generation (RAG)-based Summarization AI for EIC (RAGS4EIC) is under development. This AI-Agent not only condenses information but also effectively references relevant responses, offering substantial advantages for collaborators. Our project involves a two-step approach: first, querying a comprehensive vector database containing all pertinent experiment information; second, utilizing a Large Language Model (LLM) to generate concise summaries enriched with citations based on user queries and retrieved data. We describe the evaluation methods that use RAG assessments (RAGAs) scoring mechanisms to assess the effectiveness of responses. Furthermore, we describe the concept of prompt template based instruction-tuning which provides flexibility and accuracy in summarization. Importantly, the implementation relies on LangChain [1], which serves as the foundation of our entire workflow. This integration ensures efficiency and scalability, facilitating smooth deployment and accessibility for various user groups within the Electron Ion Collider (EIC) community. This innovative AI-driven framework not only simplifies the understanding of vast datasets but also encourages collaborative participation, thereby empowering researchers. As a demonstration, a web application has been developed to explain each stage of the RAG Agent development in detail. The application can be accessed athttps://rags4eic-ai4eic.streamlit.app.[A tagged version of the source code can be found inhttps://github.com/ai4eic/EIC-RAG-Project/releases/tag/AI4EIC2023_PROCEEDING.]

Instruments & Instrumentation