Search NASA⌕ Search

SEARCH · Search NASA

Results for “Retrieval”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Wake-Resolving Acoustic Tomography: Advances through Numerical Covariance Methods

Acoustic tomography offers path-integrated measurements of atmospheric velocity and temperature fluctuations with high spatial resolution. Classical implementations of time-dependent stochastic inversion rely on homogeneous, isotropic covariance models that are poorly suited to the anisotropic structure of wind turbine wakes. By directly estimating heterogeneous covariances from large-eddy simulations (LESs) into the time-dependent stochastic inversion operator, we relax implicit assumptions in the analytical models used historically. Retrievals using these LES-informed models improve agreement with true fields in variance, turbulent kinetic energy, and spectral content compared to analytical and precursor-based covariance models. The results indicate that LES-informed covariance models can enhance the accuracy of acoustic tomography retrievals in complex, anisotropic flows such as wind turbine wakes in some cases and highlight instances where analytical models still offer competitive performance, despite their simplifying assumptions.

17 WIND ENERGY↗

PSU-BSEC Doppler Lidar Processed Scans

This data repository includes the processed vertically staring measurements (stare files) and the angled scans (profile files) necessary to calculate horizontal winds retrieved by the Pennsylvania State University (PSU) Doppler Lidar as a part of the Baltimore Social-Environmental Collaborative Urban Integrated Field Laboratory (BSEC UIFL). The PSU Doppler Lidar measures aerosol backscatter intensity (m-1 sr-1), signal-to-noise ratio, and radial velocity (i.e., vertical velocity in the case of stare files, units: m/s) at approximately 1 Hz temporal resolution and 30 m spatial resolution. Within the "stare_data" subdirectory there exist example figures of all retrieved data. For further information, please email Nicholas Prince, nec5299@psu.edu.

Air Quality↗

Projected Urban Morphology of the Los Angeles Area by the Year 2100

This dataset provides projections of urban building morphologies for the Los Angeles urban area at 30-meter spatial resolution. It contains 192 raster files that detail two primary building attributes: building footprint fractions (ranging from 0 to 1) and average building heights (ranging from 0 to 75 meters). The projections account for a wide range of future pathways, covering two Shared Socioeconomic Pathway (SSP) scenarios (SSP3 and SSP5), two population scenarios, two developed land intensification scenarios, and four distinct levels of intensification. The dataset was created using dual Generative Adversarial Networks (GANs) trained on 2015 land cover and building properties from the National Land Cover Database (NLCD) and Model America datasets. Supporting information on the dataset has been described in the LAUrbanAreaMorphologyProjections2100_README.txt file.

Pandey, Bhartendu↗

Towards Unlocking Insights from Logbooks Using AI

Electronic logbooks contain valuable information about activities and events concerning their associated particle accelerator facilities. However, the highly technical nature of logbook entries can hinder their usability and automation. As natural language processing (NLP) continues advancing, it offers opportunities to address various challenges that logbooks present. This work explores jointly testing a tailored Retrieval Augmented Generation (RAG) model for enhancing the usability of particle accelerator logbooks at institutes like DESY, BESSY, Fermilab, BNL, SLAC, LBNL, and CERN. The RAG model uses a corpus built on logbook contributions and aims to unlock insights from these logbooks by leveraging retrieval over facility datasets, including discussion about potential multimodal sources. Our goals are to increase the FAIR-ness (findability, accessibility, interoperability, and reusability) of logbooks by exploiting their information content to streamline everyday use, enable macro-analysis for root cause analysis, and facilitate problem-solving automation.

43 PARTICLE ACCELERATORS↗

Automating and Evaluating Large Language Models for Accurate Text Summarization Under Zero-Shot Conditions

Automated text summarization (ATS) is crucial for collecting specialized, domain-specific information. Zero-shot learning (ZSL) allows large language models (LLMs) to respond to prompts on information not included in their training, playing a vital role in this process. This study evaluates LLMs' effectiveness in generating accurate summaries under ZSL conditions and explores using retrieval augmented generation (RAG) and prompt engineering to enhance factual accuracy and understanding. We combined LLMs with summarization modeling, prompt engineering, and RAG, evaluating the summaries using the METEOR metric and keyword frequencies through word clouds. Results indicate that LLMs are generally well-suited for ATS tasks, demonstrating an ability to handle specialized information under ZSL conditions with RAG. However, web scraping limitations hinder a single generalized retrieval mechanism. While LLMs show promise for ATS under ZSL conditions with RAG, challenges like goal misgeneralization and web scraping limitations need addressing. Future research should focus on solutions to these issues.

Priebe Mendes Rocha, Maria Eduarda [ORNL]↗

Automation of Vulnerability and Patch Management: Information Extraction, Association, and Optimization

Vulnerability and patch management is an integral part of a robust cybersecurity program, yet it grows increasingly complex due to the sheer amount of data that must be analyzed. Particularly in Operational Technology (OT) environments, analysis must be done manually because of the lack of automated solutions. Additionally, there are many steps in this process, from the initial discovery of the vulnerability to the implementation of its remediation, and each step in the process requires different data in order to be performed effectively. In this work, we provide approaches and strategies to assist operators in industrial or OT environments throughout the vulnerability management cycle. Security advisories provide key information about mitigation strategies, or actions that can be taken when a patch is unavailable or cannot be installed. Details of these strategies are not shared in public vulnerability databases and must be found manually. We approach this problem by designing a solution to automatically identify that information within vendor security advisories and retrieve it for operator use. We start with an approach that requires domain-specific knowledge of certain frequently-seen reference websites. Next, an approach that can work on an arbitrary website but relies on certain keywords. Finally, an approach that uses Natural Language Processing (NLP) methods and does not require specific knowledge or keywords. Each of these approaches is more general than its predecessor; we demonstrate high accuracy for all approaches Advisories also often contain details of affected products in non-standard or natural language formats. While this information can be easily understood when read by an operator, the non-standard format acts as a barrier to effective automation. We provide an approach for the first step in this process: identifying vendors in security advisories and mapping them to a standard framework for representing digital assets and software products. We evaluate five established string similarity algorithms, plus one of our own design that combines string similarity and information theory, on the task of mapping vendors to their corresponding entries in the Common Platform Enumeration (CPE) repository. Our results show that our proposed metric outperforms all others. Due to the constraints on time, finances, and personnel for organizations, Large Language Models (LLMs) may seem like attractive opportunities for security operators to speed up information gathering; however, it is still not clear whether LLMs can handle vulnerability management tasks well. To answer this question, we perform an empirical study of LLMs’ ability to provide consistent, accurate information about vulnerabilities in order to guide organizations in their adoption of LLMs. We observe poor performance for all models tested, suggesting that these models are not well-suited to the consistent retrieval of accurate vulnerability information. Finally, once vulnerabilities have been identified and any additional information has been obtained, operators must decide which remediation actions to implement based on their available resources. This already-complex problem becomes even more so when we consider that a vulnerability may have multiple avenues for remediation. We formulate this scenario as two knapsack problems and provide solutions, which we then compare against several existing strategies for vulnerability prioritization seen in real operational environments.

McClanahan, Kylie↗

Digitizing and Enhancing Accessibility of the Fusion Safety Archives

This project focuses on the digitization and public accessibility to the Fusion Safety Archives at the Idaho National Laboratory. The first phase involves a thorough review of each document in the physical archives to determine its online availability. For documents that are available online, PDF copies and unique identifiers are collected for database integration. Documents not available online are delivered to Red Inc. for digitization. Additionally, defunct storage devices such as diskettes are sent to INL’s archival department for data retrieval where possible. The second phase of the project involves the creation of a comprehensive database to house the digital copies of the archives. The database will facilitate easy access and management of the digitized documents. Following the database creation, we plan to train a Retrieval-Augmented Generation (RAG) based AI on publicly available documents. The trained AI will be integrated into a front-facing application, allowing the public to easily access information from the Fusion Safety Archives. This project aims to preserve valuable historical data, improve accessibility, and promote transparency in fusion safety research.

70 - PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Archi: Agentic Operations at the CMS Experiment

We present Archi, an open-source, end-to-end framework for scientific collaborations that combines the systematic ingestion and organization of heterogeneous data sources with the deployment of configurable, private, and extensible agents that retrieve and reason over them. An instance of Archi has been deployed for the Computing Operations team of the CMS experiment at CERN's LHC since February 2026 as a support agent for technical operators, offering retrieval and analysis capabilities by combining documentation, historical data, and live monitoring systems. We evaluate the system on operator feedback and a question set collected from production usage, graded by human and automated panels. The system proves effective at operational tasks, resolving real-world queries posed by CMS operators. We also observe that locally-hosted, open-weight models perform competitively, enabling fully private management of sensitive data.

Lugato, Pietro [MIT; CERN]↗

Hydrogen density mapping in biomolecular crystals through dynamic nuclear polarization

Many fundamental biological processes, including those in photosynthetic reaction centers and enzyme active sites, involve charge and energy transfer, bond cleavage, protonation and hydrogen bonding. Because H atoms play such central roles in these reactions, accurately determining their positions is essential. Yet, conventional X-ray crystallography primarily resolves the heavy atoms in biological structures and provides limited insight into hydrogen, even at atomic resolution. Neutron macromolecular crystallography (NMC) overcomes this limitation by offering exceptional sensitivity to hydrogen and deuterium. Here, we present a theoretical framework for the development of dynamic nuclear polarization NMC (DNP-NMC) techniques, which exploit the alignment of neutron and proton nuclear spins to enhance and tune the hydrogen signal contribution. The DNP-NMC approach advances the resolution of H atoms within biomolecular crystals, whether bound to protein residues or present in solvent. The method establishes key relationships for the coherent structure factor of polarized neutron scattering from hydrogenous matter. It theoretically achieves full accuracy in phase reconstruction and offers a path to improve neutron structure determination, achieving accuracies exceeding ≳80% by incorporating titration states. Using a variant of the hybrid input/output phase-retrieval algorithm, it allows recovery of the hydrogen density with ≳90% phase accuracy. In conclusion, we further discuss sources of experimental uncertainty for the upcoming DNP-enabled, quasi-Laue IMAGINE-X experiment at Oak Ridge National Laboratory's High Flux Isotope Reactor.

dynamic nuclear polarization↗

Reducing AI RAG Hallucination by Optimizing Routing Techniques

Large Language Models (LLMs), such as ChatGPT, tend to “hallucinate”, meaning they confidently generate false information. Retrieval Augmented Generation (RAG) attempts to diminish hallucination by providing context to the LLM from data stores (indexes) containing relevant information. The LLM uses this context to formulate its response. RAG systems can still suffer from hallucination because of bad embeddings or ineffective routing. For example, a router will often return context from an irrelevant index, resulting in a hallucinated answer. In this study, we aim to minimize the frequency of routing hallucinations by optimizing Index Summary Routing.

97 MATHEMATICS AND COMPUTING↗

The Central Role of Oxo Clusters in Zirconium‐Based Esterification Catalysis

Oxo clusters are a unique link between oxide nanocrystals and Metal‐Organic Frameworks (MOFs), representing the limit of downscaling each of the respective crystals. Herein, the superior catalytic activity of clusters, compared to zirconium MOF UiO‐66 and nanocrystals is shown. Focus is on esterification reactions given their general importance in consumer products and the challenge of converting large substrates. Oxo clusters have a higher surface‐to‐volume ratio than nanocrystals, rendering them more active. For large substrates, for example, oleic acid, MOF UiO‐66 has negligible catalytic activity while clusters provide almost quantitative conversion, a fact we ascribe to limited diffusion of large substrates through the MOF pores. Clusters do not suffer from limited mass transfer and we also obtain high conversion in solvent‐free reactions with sterically hindered alcohols (hexanol, 2‐ethyl hexanol, benzyl alcohol, and neopentyl alcohol). The cluster catalyst can be recovered and shows identical activity when reused. The structural integrity of the cluster is confirmed using X‐ray total scattering and pair distribution function analysis. Moreover, when homogeneous zirconium alkoxides are used as catalysts, the same oxo cluster is retrieved, showing that oxo clusters are the active catalytic species, even in previously assumed homogeneously catalyzed reactions.

Pulparayil Mathew, Jikson↗

Observations of the Marine Atmospheric Boundary Layer’s Response to a Solar Eclipse

The atmospheric response to the solar eclipse of 8 April 2024 in North America is investigated with a specific focus on the marine atmospheric boundary layer (MABL). We leverage measurements collected during the Third Wind Forecast Improvement Project (WFIP3), including Doppler lidars, sonic anemometers, and thermodynamic profiler data to investigate the atmospheric response across sites that experienced partial eclipse conditions with nearly 90% obscuration. Using these measurements, we examine eclipse-induced changes in key meteorological parameters, such as temperature, wind speed, and turbulent fluxes. Most previous eclipse studies have been conducted over land, whereas this study provides new observations for both coastal and marine environments, offering additional insight into eclipse-driven variability in the MABL. The findings confirm a notable decrease in downwelling shortwave radiation during the eclipse, which results in rapid cooling of surface air. The temperature reduction ranges from $1.2^\circ \text {C}$ to $1.4^\circ \text {C}$ in coastal regions and from $0.3^\circ \text {C}$ to $0.5^\circ \text {C}$ over the ocean. This analysis suggests that the MABL’s higher thermal inertia compared to coastal regions moderates the temperature decrease during the eclipse. Wind speed exhibits a more complex behavior, as it is influenced by both the MABL and preexisting synoptic conditions. Although a reduction in wind speed is observable up to approximately 140 m above ground level (AGL) at more inland sites, at other locations closer to the coast, this reduction is constrained to the lowest 100 m AGL. Turbulence parameters retrieved from sonic anemometers, such as turbulence kinetic energy, turbulent heat flux, and friction velocity, decrease during the eclipse at coastal sites, accompanied by a brief transition of atmospheric stability from unstable to neutral or weakly stable conditions. For the open-ocean sites, the variability in turbulence statistics and atmospheric stability is minimal during the occurrence of the eclipse.

16 TIDAL AND WAVE POWER↗

Agentic traffic intelligence: Augmented human-in-the-loop scenario generation for microscopic traffic simulation

Traditional microscopic traffic simulation generation often relies on static datasets and manual design, limiting its ability to simulate complex conditions easily. This paper presents a novel framework, Agentic Traffic Intelligence, which combines human approval large language models (LLMs), the Real-Twin tool, and multi-agent systems to perform realistic microscopic traffic simulation scenario generation. The proposed framework incorporates human-in-the-loop (HIL) control, retrieval-augmented generation (RAG), and multi-agent control mechanisms. HIL mechanisms are used to guide multiple LLMs focused on attributes for microscopic simulation generation and to improve the interpretability and transparency of LLM execution for users. RAG enhances context extraction by dynamically integrating external knowledge sources for traffic scenario generation foundations. A multi-agent architecture with supervisory control coordinates the interaction of simulation components, including traffic simulators, control logic, and calibration tools. This enables the synthesis of simulation-ready scenarios that reflect dynamic demand profiles and behavior controls. Furthermore, the framework fuses multisource traffic data with unstructured context and supports iterative refinement through interactive user feedback. Validated through microscopic simulation using Simulation of Urban Mobility, the generated scenarios demonstrate high-fidelity network generation with inflow and turn movement and behavioral calibration, offering a robust and efficient tool for stress-testing and optimizing urban mobility systems.

Hierarchical multi-agent control↗

An iterative bidirectional gradient boosting approach for CVR baseline estimation

Here this paper presents a novel Iterative Bidirectional Gradient Boosting Model (IBi-GBM) for estimating the baseline of Conservation Voltage Reduction (CVR) programs. In contrast to many existing methods, we treat CVR baseline estimation as a missing data retrieval problem. The approach involves dividing the load and its corresponding temperature profiles into three periods: pre-CVR, CVR, and post-CVR. To restore the missing load profile during the CVR period, the method employs a three-step process. First, a forward-pass GBM is executed using data from the pre-CVR period as inputs. Subsequently, a backward-pass GBM is applied using data from the post-CVR period. The two restored load profiles are reconciled, considering pre-calculated weights derived from forecasting accuracy, and only the leftmost and rightmost points are retained. The newly restored points are then included as inputs for the subsequent iteration. This iterative procedure continues until the original load data in the CVR period is fully restored. We develop IBi-GBM using actual smart meter and Supervisory Control and Data Acquisition (SCADA) data. Our results demonstrate that IBi-GBM exhibits robust performance across various data resolutions and in different seasons and outperforms existing methods by achieving a 1-2% reduction in normalized Root Mean Square Error (nRMSE).

42 ENGINEERING↗

Unraveling Hydrogen Induced Geochemical Reaction Mechanisms through Coupled Geochemical Modeling and Machine Learning

Underground hydrogen storage (UHS) provides a promising large-scale, long-term energy storage solution. A reasonable recovery of stored hydrogen is critical for a successful storage scheme. However, in subsurface reservoirs hydrogen is subject to active geochemical reactions that might result in hydrogen loss. In this study, we implemented a geochemical modeling approach coupled with an unsupervised machine learning technique called non-negative matrix factorization (NMF) to unravel the complex brine-rock-H 2 geochemical processes responsible for hydrogen losses, with particular focus on sulfate reduction reactions. NMF is applied to modeled mineral evolution and fluid component profiles to retrieve profiles that can be interpreted to more easily assess competing processes. NMF decouples simulated competing equilibrium reactions. This facilitates separation of overlapping reaction profiles from redox processes, dissolution fronts, and secondary precipitation while considering the effects of simulation parameters such as salinity, temperature, and total H 2 pressure. NMF successfully discriminates these competing effects in nonlinear ways, allowing robust interpretation. In addition, NMF reveals subtle coupled mineral associations and reaction fronts that are invisible to conventional model analysis. This integrated approach strengthens the conceptual understanding of complex nonlinear hydrogen-brine-rock interactions and advances geochemical research on UHS systems to resolve complexities in modeled geochemical systems without the need for direct experiments or prior knowledge. Furthermore, this study highlights the efficacy of combining geochemical modeling with machine learning techniques to enhance the interpretability of the intricate geochemical simulation output through deciphering the overlapping reaction path that cannot be achieved only using conventional analysis of geochemical models alone.

08 HYDROGEN↗

Mining waste-driven carbon capture via ocean alkalinity enhancement

Ocean alkalinity enhancement (OAE) has emerged as a promising strategy to mitigate ocean acidification and reduce global warming. Traditional metal (e.g., critical materials) mining industries release alkaline waste via mining tailings with high concentrations (99.2%) of calcium and magnesium oxide (CaO, MgO). Incorporating mining waste into OAE processes is less energy intensive than processes relying on calcination of limestone for CaO production. The solubility limit of simulated mining waste in American Society for Testing and Materials (ASTM) seawater is 75 mg⋅L -1 , which can sequester 118 mg of carbon dioxide (CO 2 ). The solubility in seawater retrieved from Sunset Beach, FL was 25 mg⋅L -1 . Changes in pH, total alkalinity, and total inorganic carbon were analyzed to confirm the successful addition of simulated alkaline mining waste without the formation of secondary precipitation. This study proposes a new OAE strategy where a facility is developed nearby ocean waters that mixes alkaline waste with seawater. Subsequently, the seawater is met with previously captured, pure CO 2 to bring the pH back to 8.2 and eliminate the risks of pH shock and secondary precipitation. Technoeconomic analysis estimated an energy requirement of 1.4 GJ per ton of CO 2 stored that resulted in a processing cost of $\$$266 per ton of CO 2 sequestered (4.2 GJ per ton, $\$$807 per ton of CO 2 for real seawater). Results from this study underscore the potential for utilizing mining waste in OAE processes and provide a pathway for practical deployment.

Absorption↗

A dynamic 2D Borehole Thermal Energy Storage (BTES) model for enhanced computational efficiency

Progressing toward a future increasingly reliant on renewable energy sources, the development of effective, durable energy storage solutions becomes essential to balance supply and demand fluctuations. Borehole Thermal Energy Storage (BTES) is a long-duration thermal energy storage technology that captures excess heat generated from renewable energy sources and stores it underground for later use, enabling the efficient utilization of sustainable energy. This approach is particularly valuable in district energy networks when integrated with Ground Source Heat Pumps (GSHP) to provide stable heating and cooling. However, traditional three-dimensional (3D) numerical models of BTES systems demand extensive computational resources, limiting their practicality for real-time and large-scale applications. This study introduces a novel two-dimensional (2D) modeling approach that reduces computational costs while maintaining high accuracy. By employing a radial ring-based discretization method, the model simulates heat injection, retention, and retrieval dynamics over seasonal cycles. A new thermal-mass weighted-average temperature parameter is introduced to evaluate the performance of BTES systems. Model validation against FEFLOW simulations demonstrates a 17-fold improvement in computational speed compared to traditional Computational Fluid Dynamics (CFD) models while achieving a mean absolute percentage error (MAPE) of 2 % during charging and 4 % during discharging. Additionally, a trade-off analysis between computational efficiency and accuracy is conducted, ensuring the model's applicability for real-world scenarios. The findings of this research contribute to the development of computationally efficient BTES models, facilitating better optimization, control, and integration into renewable energy systems. This work provides a foundation for further studies in techno-economic analysis, multi-year performance evaluation, and real-time operational strategies for BTES applications, supporting a more sustainable energy future.

2D modeling↗

Laboratory evaluation of cyclic underground hydrogen storage in the Temblor sandstone of the San Joaquin Basin, California

Underground Hydrogen Storage (UHS) in depleted oil and gas reservoirs could provide a cost-effective solution to balance seasonal fluctuations in renewable energy generation. However, data and knowledge on UHS at subsurface conditions are limited so it is difficult to estimate how effective this type of storage could be. In this study, we perform high pressure experiment to measure the effectiveness of cyclic hydrogen (H 2 ) storage in a specimen of Temblor sandstone retrieved from the San Joaquin Basin of California. Our experiment mimics reservoir pressure conditions to measure H 2 -brine relative permeability and fluid-rock interactions over the course of ten charging and discharging cycles. Initial gas breakthrough occurred at 15 % to 25 % H2 saturation in the specimen with 3 % NaCl brine as the resident fluid. Continuing injecting to 4 pore volumes (PV) of H 2 yielded an asymptotic H 2 saturation of 38 % to 41 %, a level often referred to as the irreducible gas saturation based on two-phase flow. The boundary condition in this study mimics the near wellbore region, which experiences bi-directional H 2 flow. This bi-directional flow led to evaporative drying of the specimen resulting in 94 % H 2 saturation at the end of 10th cycle. This indicates that cyclic flow and evaporative drying can lead to more efficient reservoir storage where a larger fraction of the reservoir porosity is usable to store H 2 . The produced gas stream consisted of H 2 mixed with 8 % to 22 % H 2 O, indicating formation dry-out by evaporation. Meanwhile, produced water chemistry indicated calcite and silicate dissolution, with calcite sourced from fossil fragments. This led to a loss of cementation and weakened the rock sample. Combined, our results indicate dry-out, compaction, increased H 2 saturation, rock weakening, and permeability loss during cyclic UHS. Overall, we anticipate that the combined effects should lead to higher than anticipated UHS storage efficiency per volume of sandstone reservoir rock.

08 HYDROGEN↗