Search NASA⌕ Search

SEARCH · Search NASA

Results for “data retrieval”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Hanford 200 West Area Flowsheet Data to Support Waste Treatment and Disposal Request for Proposal

This report documents the 200 West Area (200W) flowsheet supplemental data needed to support the Request for Proposal (RFP) to procure onsite and/or offsite treatment and disposal capabilities for West Area pretreated1 tank waste (PTW). The data provided in this document is based on the results of a 200W flowsheet model run evaluating single-shell tank (SST) retrievals for all S, SX, and U Farms, except for Tank S-112, which was retrieved in March of 2007 (HNF-EP-0182, Waste Tank Summary Report for Month End August 31, 2024). The evaluation also includes the waste inventory in double-shell tanks (DSTs) in SY Farm.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Custom surface reflectance, shade mask, and equivalent water thickness maps for the Colorado Headwaters Ecological Spectroscopy Study (2025)

This dataset contains land surface reflectance estimates and additional derived products generated from NEON Imaging Spectrometer (NIS) data collected in the Upper Gunnison river basin during June and July of 2025. Data was collected over three domains: the Upper East River (CRBU), Almont Triangle (ALMO), and the Upper Taylor Basin (UPTA). These products were derived from radiance and LiDAR data collected by the NEON Airborne Observation Platform (AOP) campaign funded by the Colorado Headwaters Ecological Spectroscopy Study (CHESS) (doi:10.15485/3017965). Products include per-pixel surface reflectance (rfl) and reflectance uncertainty (rfl_unc), observational data (obs), canopy equivalent water thickness (ewt), and shade masks. Atmospheric correction was performed per flightline using the ISOFIT (Imaging Spectrometer Optimal FITting) optimal estimation framework to estimate surface reflectance and the associated per-band reflectance uncertainty. Reflectance retrievals achieved a mean absolute error of 1.5% across diverse validation surfaces (see validation report.pdf). Equivalent water thickness was calculated from surface reflectance using the Beer–Lambert absorption of liquid water. Shade masks were generated based on the geometry between the sun angle, ground surface, and sensor at the time of flight. Data products are provided per-flightline and as mosaics for each domain. Flightline data products are provided as ENVI-formatted binary files (rfl, rfl_unc, ewt) and GeoTIFFs (shade). Reflectance and uncertainty mosaics are provided as tiled NetCDFs, while all other mosaicked products are provided as cloud-optimized GeoTIFFs. These formats are supported by common geospatial software (e.g., QGIS, ArcGIS, ENVI) and programmatic libraries in Python (e.g., rasterio, xarray, spectral, netCDF4) and R (e.g., terra, ncdf4). Processing workflows were designed to be equivalent to those used to generate the 2018 CHESS campaign airborne imaging spectroscopy data products (doi:10.15485/3013527). All outputs were co-registered to a common spatial grid to support time series analyses. CHESS Project Description: The Colorado Headwaters Ecological Spectroscopy Study (CHESS) comprised a multi-week airborne remote sensing and field observation campaign in the Upper Gunnison Basin, Colorado, conducted in June and July of 2025. Airborne remote sensing was conducted by the National Ecological Observatory Network Airborne Observation Platform (NEON AOP), concurrent with a field campaign run by the Rocky Mountain Biological Laboratory (RMBL), the Lawrence Berkeley National Laboratory (LBNL) and SLAC National Accelerator Laboratory Watershed Function Science Focus Area (SFA), and NASA-JPL (Jet Propulsion Laboratory) Earth Surface Mineral Dust Source Investigation (EMIT) program. Between June 10 and July 18, 2025, the NEON AOP flight team collected high-resolution aerial imaging spectroscopy and Light Detection and Ranging (LiDAR) data over three domains: the Upper East River (CRBU), Almont Triangle (ALMO), and the Upper Taylor Basin (UPTA). In coordination with the flights, a field campaign acquired ground-truth observations, including observations of vegetation composition, foliar traits, forest demography, and subsurface properties in 18 core sampling areas within the domains. Additional surface water observations were taken at over 380 point locations. All CHESS campaign datasets can be found within the CHESS ESS-DIVE data portal: https://data.ess-dive.lbl.gov/portals/chess. Funding Acknowledgment: Data acquisition was performed under a grant from the National Aeronautics and Space Administration (80NSSC24K1005). Computational research was carried out at the Jet Propulsion Laboratory, California Institute of Technology, under a contract with the National Aeronautics and Space Administration (80NM0018D0004) and was funded by EMIT Extended Mission Phase E Science.

2018 NEON and 2025 CHESS Campaigns↗

NMF-Based Anomaly Detection in CMS 2D Tracking Occupancy Histograms

The CMS experiment relies on Data Quality Monitoring (DQM) to ensure that recorded collision data are suitable for physics analysis. During LHC Run 3, each run contains many lumisections and tracking monitoring elements, making offline inspection challenging, especially for localized detector effects that may appear only for short periods of time. This poster presents an unsupervised machine-learning approach to identify anomalous lumisections in CMS tracking occupancy histograms using Non-Negative Matrix Factorization (NMF). The workflow uses offline CMS DQMIO tracking histograms retrieved with the CMS DIALS API and organized as two-dimensional occupancy maps for each lumisection. After selecting stable lumisections, the occupancy maps are normalized and arranged into a non-negative data matrix. The NMF model learns a compact set of basis patterns describing normal tracking occupancy. Each lumisection is then reconstructed from these learned components, and the reconstruction error is used as an anomaly score. Large residuals indicate occupancy patterns that deviate from normal detector behavior and are flagged for further inspection. This NMF-based approach provides a fast and interpretable way to flag lumisections whose tracking occupancy patterns differ from normal detector behavior. Preliminary studies show sensitivity to known tracking anomalies, and ongoing work is focused on validating the method across additional Run 3 Pixel and Strip detector issues.

Rodríguez Ramos, Iliomar [Puerto Rico U., Mayaguez↗

FATHOMS-RAG: A Framework for the Assessment of Thinking and Observation in Multimodal Systems that use Retrieval Augmented Generation

Retrieval-augmented generation (RAG) has emerged as a promising paradigm for improving factual accuracy in large language models (LLMs). We introduce a benchmark designed to evaluate RAG pipelines as a whole, evaluating a pipelines ability to ingest several modalities of information. We present (1) a curated dataset of 93 questions designed to evaluate a pipeline's ability to ingest textual data, tables, images, multimodal data, and cross-document multimodal data; (2) a phrase-level recall metric for correctness; (3) a nearest-neighbor embedding classifier in an attempt to classify pipeline hallucinations; (4) a comparative evaluation of 2 pipelines built with open-source retrieval mechanisms and 4 closed-source foundational models; and (5) a third-party human evaluation of the alignment of our correctness and hallucination metrics. We find that closed-source pipelines significantly outperform open-source pipelines in both the correctness and halucination metrics, with a wider performance gap in questions relying on multimodal and cross-document information. We also find after a human evaluation of our correctness and hallucination metric compared with our questions and pipeline responses, average agreement was 4.62 for correctness 4.53 for hallucination detection on a 1-5 Likert scale with 5 being strongly agree with our determination.

Hildebrand, Samuel [ORNL] (ORCID:0009000465963104)↗

X-ray scattering based scanning tomography for imaging and structural characterization of cellulose in plants

X-ray and neutron scattering have long been used for structural characterization of cellulose in plants. Due to averaging over the illuminated sample volume, these measurements traditionally overlooked the compositional and morphological heterogeneity within the sample. Here, a scanning tomographic imaging method is described, using contrast derived from the X-ray scattering intensity, for virtually sectioning the sample to reveal its internal structure at a resolution of a few micrometres. This method provides a means for retrieving the local scattering signal that corresponds to any voxel within the virtual section, enabling characterization of the local structure using traditional data-analysis methods. This is accomplished through tomographic reconstruction of the spatial distribution of a handful of mathematical components identified by non-negative matrix factorization from the large dataset of X-ray scattering intensity. Joint analysis of multiple datasets, to find similarity between voxels by clustering of the decomposed data, could help elucidate systematic differences between samples, such as those expected from genetic modifications, chemical treatments or fungal decay. The spatial distribution of the microfibril angle can also be analyzed, based on the tomographically reconstructed scattering intensity as a function of the azimuthal angle.

36 MATERIALS SCIENCE↗

Parallax-corrected VISST-derived pixel-level products from satellite GOES-16

The NASA Langley group led by William Smith produced GOES-16 satellite cloud retrievals over an approximate 10 by 10 degree region over the CACTI field campaign location. These retrievals are described here: https://www.arm.gov/capabilities/vaps/visst and are available for download here . They use algorithms historically called VISST that are now referred to as SatCORPS. More information can be found in Trepte et al. (2019), Minnis et al. (2021), and Yost et al. (2021). If using this dataset, please cite these references, the CACTI VISST dataset DOI found at the download link above, and this dataset’s DOI. The CACTI VISST pixel-level retrievals are on a 2 km spatial grid and available every 15 minutes (every 10 minutes late in the campaign), producing 21,765 files for the entire field campaign between October 2018 and April 2019. They are not corrected for parallax error, which is an offset in the actual geographical location of a cloud above the surface due to the satellite viewing the cloud partly from the side off nadir. This dataset applies a correction for parallax using the location relative to the satellite and the retrieved cloud top height above the surface, which allows the dataset to be geo-located with surface-based observations. The parallax correction for each location depends on the longitude, latitude and cloud top height above ground level (AGL) for that longitude and latitude in the original VISST files. The cloud top height AGL requires first computing the surface elevation at each VISST grid point. Data from the Advanced Spaceborne Thermal Emission and Reflection (ASTER) Global Digital Elevation Map Version 3 at 30-m resolution is projected onto the VISST grid using conservative coarsening (conserving surface elevation) in the xESMF Python package. The surface elevation is then subtracted from the VISST-retrieved cloud top height above mean sea level. These cloud top heights AGL are then combined with longitude and latitude to estimate the latitude and longitude corrections. Due to variability in cloud top height, the parallax shifts produce an irregular grid of values since higher cloud tops are shifted further than lower cloud tops. A ball tree-based neighbor search with Haversine distance is performed using the Python-based scikit-learn library to find the nearest VISST grid point to each parallax correction-shifted point. The data value of the shifted point is then assigned to that VISST grid point. In this manner, the irregular geographical shifts to correct for parallax are projected back to the rectilinear VISST grid. Because relatively higher clouds should obscure lower clouds, the variable values for the highest cloud top are preferentially chosen if two or more values are assigned to a grid point. The parallax correction should be viewed as an improved but still imperfect estimation of the cloud top locations, largely because the cloud top height is an imperfect retrieval. Please see the attached README document for further information. Users are encouraged to contact the authors with any additional questions.

54 ENVIRONMENTAL SCIENCES↗

Lost in OCR Translation? Vision-Based Approaches to Robust Document Retrieval

Code for ‘Lost in OCR Translation?’: robust document retrieval under degradation. Compares OCR-based, vision-only, and hybrid pipelines; includes SambaNova LLaMA Vision OCR, Nougat, and ViDoRe baselines. Provides QA data generation, RAG evaluation, and metrics (Levenshtein, nDCG@k, Recall@k, EM/F1) with reproducible scripts. Includes dataset guides

Bhattarai, Manish [Los Alamos National Labs]↗

AI-Ready Semantic Infrastructure for CEBAF: From CED to PALS Knowledge Graphs

JLab and PNNL are jointly developing an AI-ready data ecosystem that exposes the Continuous Electron Beam Acceleration Facility’s (CEBAF’s) operational configuration, lattice description, and control-system channels to agentic optimization frameworks through a standards-based semantic layer. The effort integrates the existing facility-specific CEBAF Element Database (CED) with extensions of the emerging facility-agnostic Particle Accelerator Lattice Standard (PALS) to produce a knowledge graph (KG) containing coherent, machine-interpretable views of devices, signals, and regions. With this KG, CEBAF’s setpoints, readbacks, and device hierarchies become queryable using a uniform declarative graph query language (e.g., Neo4j Cypher), providing intents and inspectable semantics suitable for agentic control. The resulting graph-backed interfaces will allow autonomous agents to retrieve authoritative machine configurations, reason over device- and signal-level relationships, and execute tuning and diagnostic workflows without bespoke CEBAF-specific logic, thereby delivering a scalable pathway from operational data to trustworthy agentic accelerator tuning frameworks.

Zhang, He [Thomas Jefferson National Accelerator F↗

The NASA ACTIVATE Mission

The NASA Aerosol Cloud Meteorology Interactions over the Western Atlantic Experiment (ACTIVATE) conducted 162 joint flights with two aircraft over the northwest Atlantic to study aerosol–cloud interactions (ACIs), which represent the largest uncertainty in estimating total anthropogenic radiative forcing. The combination of a high-flying King Air and low-flying HU-25 Falcon, equipped with remote sensing and in situ instruments, characterized trace gases, aerosol particles, clouds, and meteorological variables with data collected nearly simultaneously below, within, and above marine boundary layer (MBL) clouds. Flights spanning warm and cold seasons across 3 years (2020–22) provided a broad range of conditions associated with aerosol particles, cloud properties (including particle size and phase), and meteorology, ideally suited for robust ACI calculations and assessing how well models simulate a wide range of MBL clouds from stratiform to cumulus. ACTIVATE data suggest that drivers of cloud droplet number concentration N d , including aerosol particles and MBL dynamics, vary between winter and summer months with a stronger potential to convert aerosol particles into cloud droplets in winter. Models of varying complexity not only highlight some skills in simulating winter and summer cloud types but also identify challenges that still need to be addressed such as treatment of turbulence, wet scavenging, and mesoscale organization. Remote sensing advances range from new retrieval methods for N d , cloud phase classification, vertically resolved aerosol and cloud condensation nuclei number concentration, and ocean surface wind speed. This work describes these scientific and technological advances along with efforts in outreach and open data science.

aerosol indirect effect↗

SEED: Semantic Energy Exploration and Discovery

The Bioenergy Knowledge Discovery Framework (KDF) hosts a vast repository of specialized data, yet traditional keyword-based search methods often struggle to provide direct answers, requiring significant domain expertise and manual effort to filter through raw documents. To overcome these barriers, this software introduces a semantic search engine that enables both specialists and non-specialists to query the KDF using natural language. By shifting from rigid keyword matching to intent-based retrieval, the tool automatically identifies and ranks the most relevant sources within the database. The system functions by processing natural language queries to extract the most pertinent information, delivering an AI-generated plain-language summary alongside exact supporting quotes from retrieved documents. This integrated approach provides users with immediate, evidence-based answers while eliminating the need for exhaustive manual review. By surfacing direct insights and contextual evidence, the software enhances the usability of existing KDF resources and democratizes access to complex bioenergy data. Ultimately, this semantic search solution accelerates the discovery process and supports faster, more informed decision-making across the bioenergy sector.

Pan, Meiyu (Melrose) [Oak Ridge National Laborator↗

1998/99 Thurston County Household Travel Study

The survey was conducted under the auspices of the Thurston Regional Planning Council, and it was funded through a state grant awarded to Intercity Transit of Olympia, Washington. Data collection was from September 1998 through March 1999. The purpose of the study was to provide data for the continuing development and refinement of the Regional Travel Demand Forecasting Model, as well as to provide a better understanding of travel behavior in the southern Puget Sound region of Washington. The resultant data set will be used to fulfill the model's functions of estimating trip generation and distribution, mode choice, and assignments. Participating households were assigned specific “travel days” to record their travel over a 48-hour period. A total of 2,465 households were recruited to participate in the study. Of these, 1,537 households completed travel diaries, and the information was retrieved from 3,653 household members regardless of age. Households member made 25,278 total trips during their 48-hour diary period.

1Hz data↗

GPM IMERG V07B and V06B: Evaluation Using Ground-Based Radar Observations and Application in Global Mesoscale Convective System Tracking

This study evaluates the latest Global Precipitation Measurement (GPM) Integrated Multi-satellitE Retrievals for GPM (IMERG V07B) against its predecessor V06B, for studying mesoscale convective systems (MCSs). Both versions are compared using ground-based radar and rain gauge data from five meteorologically diverse regions: the contiguous United States (including eastern coastlines), Amazon rainforest, central Argentina mountains, equatorial Indian Ocean, and northern Australia across multiple temporal (0.5–6 hours) and spatial scales (0.1°–0.25°). An updated global MCS tracking dataset is developed by integrating satellite-observed infrared brightness temperature with IMERG V07B. Comparation of IMERG against radar observations reveals that IMERG demonstrates better performance in capturing the probability distribution and quantitative contributions of rainfall (from no-rain to intense-rain conditions) over tropical oceans than over land, with marked improvements in IMERG V07B for heavy-to-intense rain (> 10 mm h-1). Over land, systematic biases persist: IMERG tends to overestimate light-to-moderate rain (1–10 mm h-1) while underestimating heavy-to-intense rain. Additionally, aggregating IMERG to coarser resolutions (3-hourly or 0.25°) improves consistency with radar observations, outperforming the 1-hourly/0.1° resolution. The new IMERG V07B-based global MCS dataset exhibits consistent statistical characteristics with the V06B-based dataset, despite lower mean rain rates and reduced heavy precipitation contributions. These findings offer valuable insights for utilizing IMERG V07B in global precipitation studies, MCS characterization, and model evaluation.

Zhang, Sihan↗