Search NASA⌕ Search

SEARCH · Search NASA

Results for “Retrieval”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

High-Resolution LiDAR Observations for Coupled Effects of Dry-Air Entrainment and Haze-Cloud Interactions on Cloud Vertical Structure

This study demonstrates the high-resolution profiling of cloud microphysics in a laboratory chamber using Time-Correlated Single Photon Counting (TCSPC) LiDAR. We present a novel retrieval method to derive vertical extinction (𝜎) profiles, constrained by in situ measurements, to diagnose responses to dry-air entrainment. In clean clouds, the LiDAR signals and retrieved 𝜎 remain relatively uniform, with entrainment effects confined to the upper layer. In contrast, polluted clouds exhibit strong vertical variability and a transition to a water-vapor-limited state. Entrainment significantly enhances the haze number concentration (𝑁ℎ), particularly near the bottom, creating highly height-dependent extinction profiles for polluted clouds. Our results highlight the capability of high-resolution LiDAR in capturing fine-scale vertical inhomogeneities. This approach provides a robust framework for quantifying how aerosol loading modulates entrainment sensitivity, offering new insights into the transition between buffered and water-vapor-limited regimes.

54 ENVIRONMENTAL SCIENCES↗

AI-enabled Lorentz microscopy for quantitative imaging of nanoscale magnetic spin textures

The manipulation and control of nanoscale magnetic spin textures are of rising interest as they are potential foundational units in next-generation computing paradigms. Achieving this requires a quantitative understanding of the spin texture behavior under external stimuli using in situ experiments. Lorentz transmission electron microscopy (LTEM) enables real-space imaging of spin textures at the nanoscale, but quantitative characterization of in situ data is extremely challenging. Here, we present an AI-enabled phase-retrieval method based on integrating a generative deep image prior with an image formation forward model for LTEM. Our approach uses a single out-of-focus image for phase retrieval and achieves significantly higher accuracy and robustness to noise compared to existing methods. Furthermore, our method is capable of isolating sample heterogeneities from magnetic contrast, as shown by application to simulated and experimental data. This approach allows quantitative phase reconstruction of in situ data and can also enable near real-time quantitative magnetic imaging.

36 MATERIALS SCIENCE↗

Modeling performance of data collection systems for high-energy physics

Exponential increases in scientific experimental data are outpacing silicon technology progress, necessitating heterogeneous computing systems—particularly those utilizing machine learning (ML)—to meet future scientific computing demands. The growing importance and complexity of heterogeneous computing systems require systematic modeling to understand and predict the effective roles for ML. We present a model that addresses this need by framing the key aspects of data collection pipelines and constraints and combining them with the important vectors of technology that shape alternatives, computing metrics that allow complex alternatives to be compared. For instance, a data collection pipeline may be characterized by parameters such as sensor sampling rates and the overall relevancy of retrieved samples. Alternatives to this pipeline are enabled by development vectors including ML, parallelization, advancing CMOS, and neuromorphic computing. By calculating metrics for each alternative such as overall F1 score, power, hardware cost, and energy expended per relevant sample, our model allows alternative data collection systems to be rigorously compared. We apply this model to the Compact Muon Solenoid experiment and its planned high luminosity-large hadron collider upgrade, evaluating novel technologies for the data acquisition system (DAQ), including ML-based filtering and parallelized software. The results demonstrate that improvements to early DAQ stages significantly reduce resources required later, with a power reduction of 60% and increased relevant data retrieval per unit power (from 0.065 to 0.31 samples/kJ). However, we predict that further advances will be required in order to meet overall power and cost constraints for the DAQ.

Olin-Ammentorp, Wilkie (ORCID:0000000224729862)↗

Towards a RAG-based summarization for the Electron Ion Collider

Abstract The complexity and sheer volume of information — encompassing documents, papers, data, and other resources — from large-scale experiments demand significant time and effort to navigate, making the task of accessing and utilizing these varied forms of information daunting, particularly for new collaborators and early-career scientists.To tackle this issue, a Retrieval Augmented Generation (RAG)-based Summarization AI for EIC (RAGS4EIC) is under development. This AI-Agent not only condenses information but also effectively references relevant responses, offering substantial advantages for collaborators. Our project involves a two-step approach: first, querying a comprehensive vector database containing all pertinent experiment information; second, utilizing a Large Language Model (LLM) to generate concise summaries enriched with citations based on user queries and retrieved data. We describe the evaluation methods that use RAG assessments (RAGAs) scoring mechanisms to assess the effectiveness of responses. Furthermore, we describe the concept of prompt template based instruction-tuning which provides flexibility and accuracy in summarization. Importantly, the implementation relies on LangChain [1], which serves as the foundation of our entire workflow. This integration ensures efficiency and scalability, facilitating smooth deployment and accessibility for various user groups within the Electron Ion Collider (EIC) community. This innovative AI-driven framework not only simplifies the understanding of vast datasets but also encourages collaborative participation, thereby empowering researchers. As a demonstration, a web application has been developed to explain each stage of the RAG Agent development in detail. The application can be accessed athttps://rags4eic-ai4eic.streamlit.app.[A tagged version of the source code can be found inhttps://github.com/ai4eic/EIC-RAG-Project/releases/tag/AI4EIC2023_PROCEEDING.]

Instruments & Instrumentation↗

Enhancement of Rydberg Blockade via Microwave Dressing

Experimental control over the strength and angular dependence of interactions between atoms is a key capability for advancing quantum technologies. Here, in this work, we use microwave dressing to manipulate and enhance Rydberg-Rydberg interactions in an atomic ensemble. By varying the cloud length relative to the blockade radius and measuring the statistics of the light retrieved from the ensemble, we demonstrate a clear enhancement of the interaction strength due to microwave dressing. These results are successfully captured by a theoretical model that accounts for the excitation dynamics, atomic density distribution, and phase-matched retrieval efficiency. Our approach offers a versatile platform for further engineering interactions by exploiting additional features of the microwave fields, such as polarization and detuning, opening pathways for new quantum control strategies.

collective effects in quantum optics↗

Mnemosyne

SAND2024-08570O Mnemosyne is an interactive tool for finding, retrieving, and exploring information about U.S. nuclear tests documented in the National Nuclear Security Administration’s NV-209 report. It also acts as an information architecture and codebase for integrating additional information and computational tools related to these tests at the unclassified and classified levels. Users can search by any number of nuclear test attributes—name, yield range, altitude ranges, purpose of a test—and find all matching tests. Users can also retrieve specific test information published in NV-209. The tool displays geospatial and topological data about test location, and it provides an information architecture for storing additional contextual material, such as photographs. Mnemosyne provides a capability of interfacing with HYCHEM, Sandia's nuclear detonation optical waveform tool. The software is designed for use by government, academia, military, and research institutions. The software will likely be advanced to integrate seismic data and the nuclear detonation optical signal simulation code radCTH. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Fisher, Dustin↗

SEED: Semantic Energy Exploration and Discovery

The Bioenergy Knowledge Discovery Framework (KDF) hosts a vast repository of specialized data, yet traditional keyword-based search methods often struggle to provide direct answers, requiring significant domain expertise and manual effort to filter through raw documents. To overcome these barriers, this software introduces a semantic search engine that enables both specialists and non-specialists to query the KDF using natural language. By shifting from rigid keyword matching to intent-based retrieval, the tool automatically identifies and ranks the most relevant sources within the database. The system functions by processing natural language queries to extract the most pertinent information, delivering an AI-generated plain-language summary alongside exact supporting quotes from retrieved documents. This integrated approach provides users with immediate, evidence-based answers while eliminating the need for exhaustive manual review. By surfacing direct insights and contextual evidence, the software enhances the usability of existing KDF resources and democratizes access to complex bioenergy data. Ultimately, this semantic search solution accelerates the discovery process and supports faster, more informed decision-making across the bioenergy sector.

Pan, Meiyu (Melrose) [Oak Ridge National Laborator↗

An Observational Evaluation of RKW Theory over the U.S. Southern Great Plains

The theory of Rotunno et al. (“RKW” theory) addresses the behavior of squall-line cold pools in vertically sheared flows. It predicts that, within a given thermodynamic environment, a balance between baroclinic vorticity generation by the cold pool and low-level environmental vertical wind shear induces an upright updraft along the gust front that maximizes the initiation of new convective cells. Although this theory has been evaluated numerically, its applicability to observed systems remains unclear and is limited by a lack of critical measurements, including high-frequency thermodynamic and wind profiles across the gust front. Herein, observations from the Atmospheric Radiation Measurement Southern Great Plains (ARM-SGP) observatory near Lamont, Oklahoma, are used to evaluate RKW theory for 10 well-observed squall lines over a 11-yr period. For this evaluation, RKW parameters including cold-pool intensity (c), low-level ambient, line-normal vertical shear (ΔV n ), subcloud and cloud-layer updraft tilts, and multiple measures of system intensity are estimated. Furthermore, the c estimates rely on thermodynamic retrievals from the Atmosphere Emitted Radiance Interferometer (AERI), which are uncertain but verify reasonably well against independent observations. As predicted by the theory, for c/ΔV n ≥ 1, c/ΔV n correlates positively with updraft tilt and negatively with system intensity, but these results are not always statistically significant and are also sensitive to the method by which ΔV n is evaluated. Specifically, ΔV n evaluations that extend above the cold-pool top yield greater consistency with RKW predictions. Also, some measures of intensity correlate more strongly with standard moist instability metrics than with RKW parameters.

Cold pools↗

Temperature Inversions below 1 km from a V-Band Scanning Radiometer at the North Slope of Alaska

A single-channel (56.7 GHz) scanning radiometer was deployed in August 2022 at the Atmospheric Radiation Measurement (ARM) North Slope of Alaska site near Utqiaġvik. The radiometer is designed to provide temperature profiles between 0 and 1 km every 5 min. Averaging kernels show that this single-channel radiometer, taking observations at 10 discrete elevation angles, yields approximately the same information as a seven-channel V-band radiometer scanning three elevation angles. The instrument is able to reproduce the occurrence of temperature inversions between the surface and 1 km and their strength showing a correlation of 0.85, bias of −0.6 K, and slope of 1.04 with respect to radiosondes. Uncertainty in the inversion base height varies from 50 m near the surface to ∼300 m above 0.4 km when compared with radiosondes. Conversely, the inversion top height is overestimated and has higher uncertainty due to the degrading effects of the averaging kernels on the vertical resolution of the retrievals. Thanks to the high temporal resolution of the retrievals, the diurnal cycle of boundary layer temperature was evaluated showing that the radiometer can capture some aspects of the boundary layer thermal structure. The present analysis provides an overview of the capabilities of this simple observing configuration for selected applications.

Atmospheric profilers↗

Spatial and Temporal Variability of Vertical Velocity under Shallow Cumulus

Vertical velocity distribution below cloud is one of the key determinants of cloud life cycle, but observations of this variable are extremely sparse in space. Doppler lidar retrievals and large-eddy simulations at the U.S. Department of Energy’s Atmospheric Radiation Measurement User Facility Southern Great Plains site are used to determine whether vertical velocity statistics from temporally dense profiles at a single location can be substituted for spatial vertical velocity statistics. We show that even a small number (five) of widely distributed [ O (1°) latitude/longitude spacing] lidars is sufficient sampling to reconstruct domainwide spatial vertical velocity variance, but not higher moments of the vertical velocity distribution. Spatial and temporal vertical velocity variances in the Doppler lidar observations are nearly interchangeable as long as the spatial variance is temporally averaged and the temporal variance is averaged across lidars. This is true even though the dominant spatial scales of vertical velocity variability are ≲ 3 km, more than an order of magnitude smaller than the spacing between the lidars. Further, in the limit where the temporal variance does not vary across a spatial domain (e.g., if the meteorological and surface forcing of the atmospheric turbulence is homogeneous across the domain) and the domain-mean vertical velocity is zero, the commonly available retrieval of temporal vertical velocity variance at one site is equivalent to the spatial variance over the domain. We use an updraft parcel model to show that substituting temporal for spatial vertical velocity statistics will have a relatively minor effect on cloud droplet number concentrations.

54 ENVIRONMENTAL SCIENCES↗

MOSAIC-CONUS: A Multimodal, Multi-Temporally Paired Dataset for Earth Sciences

Earth embeddings—vector representations of geographic locations indexed in space and time—are emerging as a unifying interface for geospatial AI. However, their quality depends not only on model design, but on how multimodal Earth observation (EO) data are spatially indexed, temporally aligned, and cross-modally associated during pretraining. We introduce MOSAIC-CONUS (Multimodal Observations with Spatially Aligned Imagery, Urban Points of Interest, In-Situ Measurements and Text Captions), a large-scale EO dataset over the contiguous United States, organized around 250,000 stratified point indices that serve as stable spatial keys across seven modalities: active radar, passive optical imagery, lidar-derived elevation, land cover, functional context, hydrometeorological measurements, and textual summaries. Unlike existing EO datasets, MOSAIC-CONUS introduces four contributions not jointly addressed in prior work: 1. an open-source, large-scale multimodal EO corpus structured around point-indexed data designed to support Earth embedding learning; 2. explicit radar-optical pairing tables spanning twelve temporal alignment regimes, formalizing cross-sensor alignment as a controllable variable for analyzing how temporal mismatch across modalities influences learned embeddings quality; 3. a benchmark suite spanning cross-modal retrieval, annual nightlights regression, and basin-held-out streamflow prediction, positioning MOSAIC-CONUS as a benchmark-ready resource for multimodal AI systems; and 4. a language-based embedding layer through co-registered textual summaries, enabling Earth embeddings to function as a queryable interface for agentic AI systems. The dataset and pairing protocols are publicly released.

54 ENVIRONMENTAL SCIENCES↗

2002 St. Louis Region Travel Survey

The 2002 St. Louis Household Travel Survey entailed the collection of weekday travel behavior characteristics of households residing in each of the eight counties that comprise the St. Louis region. In addition to collecting basic demographic and socioeconomic information about each household and its members, the survey documented specific characteristics of activities and trips, including the number and purpose of trips, trip duration, time of day, mode of transportation, and specifics of school- and work-related travel. The survey instruments contained three components: 1) the recruitment questionnaire, 2) the travel log, and 3) the retrieval questionnaire. In total, 7,046 households were recruited to participate in the study via telephone interview. Of these, 5,094 completed travel logs during a specific 24-hour period, and the information was retrieved from all household members, regardless of age. Demographic information for this study includes age, gender, education level, employment status, and household income.

1Hz data↗

SPRUCE FT-ICR MS, Bulk Chemistry, and Mass Loss from Litter Decomposition Study in Experimental Plots, Marcell Experimental Forest, Minnesota, 2015-2017

This dataset contains molecular, bulk chemical, and mass loss measurements from a litter decomposition study at the Spruce and Peatland Responses Under Changing Environments (SPRUCE) experimental site within the Marcell Experimental Forest in northern Minnesota, USA. This site is in a Sphagnum spp. ombrotrophic bog forest. Litterbags were deployed into the peat in September 2015 across three warming levels (+0, +4.5, and +9°C) under ambient and elevated carbon dioxide (CO₂ - +500 ppm) and retrieved after roughly 0.5, 1, and 2 years of field incubation (2015-09-23 to 2017-08-02). Litterbags containing six peatland litter types: black spruce needles (Picea mariana - SPL), spruce fine roots (SPR), Sphagnum angustifolium (ANG), Sphagnum magellanicum (MAG), Labrador tea leaves (Rhododendron groenlandicum - LTL), and Labrador tea roots (LTR). Molecular composition of water-soluble organic matter extracts was characterized using Fourier Transform Ion Cyclotron Resonance Mass Spectrometry (FT-ICR MS) at 9.4 Tesla, operated in negative ion mode with electrospray ionization, providing molecular formula assignments and compound-class distributions across the decomposition time series. Bulk chemical characterization included elemental analysis (percent carbon, nitrogen, and phosphorus) and Fourier Transform Infrared Spectroscopy (FTIR) to quantify functional group composition. Litter mass loss was tracked gravimetrically at each retrieval interval, expressed as percent mass remaining relative to initial dry mass for each litter type and treatment combination. These data are valuable for understanding how vegetation shifts driven by increased atmospheric CO2 and temperature in peatlands alter litter inputs and organic matter stabilization trajectories, with implications for projecting and modeling peatland carbon cycling. This dataset contains two data files in comma-separated value (.csv) format. Additional metadata are provided: two data dictionaries and a file-level metadata file in comma separate (.csv) format and a user guide in PDF (*.pdf) format.

decomposition↗

Towards Unlocking Insights from Logbooks Using AI

Electronic logbooks contain valuable information about activities and events concerning their associated particle accelerator facilities. However, the highly technical nature of logbook entries can hinder their usability and automation. As natural language processing (NLP) continues advancing, it offers opportunities to address various challenges that logbooks present. This work explores jointly testing a tailored Retrieval Augmented Generation (RAG) model for enhancing the usability of particle accelerator logbooks at institutes like DESY, BESSY, Fermilab, BNL, SLAC, LBNL and CERN. The RAG model uses a corpus built on logbook contributions and aims to unlock insights from these logbooks by leveraging retrieval over facility datasets, including discussion about potential multimodal sources. Our goals are to increase the FAIR-ness (findability, accessibility, interoperability, and reusability) of logbooks by exploiting their information content to streamline everyday use, to enable macro-analysis for root cause analysis, and to facilitate problem-solving automation.

43 PARTICLE ACCELERATORS↗

AIACHNE's contribution for Nuclear Energy Agency Working Party on International Nuclear Data Evaluation Co-operation Subgroup 50

The AIACHNE (AI/ML Informed cAlifornium CHi Nuclear data Experiment) project aims at designing an experiment for the 252 Cf Prompt Fission Neutron Spectrum (PFNS) that explores systematic biases in an experimental database retrieved from the EXFOR databases. To that end, machine learning (ML) methods were applied to pint-point measurement features likely related to bias. From that information, we selected a feature that should be explored by the AIACHNE experiment. Measurement features are metadata encapsulating all pertinent information about the physical measurement and analysis techniques. Examples are, for instance, what neutron and fission detectors were used for the physical metadata, and what background reduction techniques were employed for analysis techniques. Such metadata were retrieved both from EXFOR entries as well as the literature of data sets described in detail in Reference 2 (at the end of the article).

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

AIACHNE's contribution for Nuclear Energy Agency Working Party on International Nuclear Data Evaluation Co-operation Subgroup 50

The AIACHNE (AI/ML Informed cAlifornium CHi Nuclear data Experiment) project aims at designing an experiment for the 252 Cf Prompt Fission Neutron Spectrum (PFNS) that explores systematic biases in an experimental database retrieved from the EXFOR databases. To that end, machine learning (ML) methods were applied to pint-point measurement features likely related to bia. From that information, we selected a feature that should be explored by the AIACHNE experiment. Measurement features are metadata encapsulating all pertinent information about the physical measurement and analysis techniques. Examples are, for instance, what neutron and fission detectors were used for the physical metadata, and what background reduction techniques were employed for analysis techniques. Such metadata were retrieved both from EXFOR entries as well as the literature of data sets described in detail in Ref. [2]. The prerequisite for applying machine learning techniques is casting the metadata into a format that can be parsed by the algorithm. This step might seem trivial but requires to find a unique language where metadata that carry the same physics meaning across several experiments must have the same identifier. One example is, for instance, the neutron detector. As seen in Figure 1, the machine learning code identified the use of 6 Li detectors as being related to bias in some datasets of the AIACHNE 252 Cf PFNS experimental database. In fact, here are several experiments that used neutron detectors containing 6Li in the database, for instance for the example below. EXFOR format has a unique keywords describing detectors such as “SCIN” or “GLASD”. One may think that these keywords are already sufficient descriptors for ML to uniquely find an issue. However, “SCIN” (used for [3, 4]) and “GLASD” (used for [5]) fail to inform the algorithm what is the active material in the detector. And, the key common issue leading to bias in 252 Cf related to neutron detectors is not whether it is a glass detector or a scintillator. No, the issue is that 6 Li was within both detector types and that even small mistakes in the detector response functions around approximately 200 keV are amplified by the 6 Li(n,α) resonance there leading to bias in data as highlighted in Fig. 1 and Ref. [1]. Hence, the features describing the neutron detector must call out the active material in the detector, rather than the existing EXFOR detector keyword, that the ML algorithm can find physically meaningful features related to bias. The AIACHNE team used a precursor of the WPEC (Working Party on International Nuclear Data Evaluation Co-operation) SG(Subgroup)-50 format to store the metadata for the ML analysis.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Vitrification of Hanford Tank 241-AN-107 Waste and Equivalent Simulant

Hanford Site nuclear waste is to be vitrified at the Waste Treatment and Immobilization Plant (WTP), which is a part of the safe and efficient retrieval, treatment, and disposal mission of the U.S. Department of Energy Office of River Protection. Hanford tank 241-AN-107 (referred to herein as AN-107) is one of the initial Hanford radioactive tank wastes planned to be processed and vitrified. A portion of AN-107 waste was retrieved by Washington River Protection Solutions, LLC (WRPS) and transferred to Pacific Northwest National Laboratory (PNNL). Compared to previously received and vitrified wastes (AP-107, AP-105, and AP-105), the concentration of organics in AN-107 was greater by an order of magnitude, while the activity of radionuclides was multiple orders of magnitude greater.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗