Search NASA⌕ Search

SEARCH · Search NASA

Results for “Retrieval methodology”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

An ML-based terrestrial data fusion and augmentation framework to enable advanced understanding of the terrestrial carbon and water interactions

Soil moisture is essential to the terrestrial carbon and water cycles and land–atmosphere interactions. There are various types of soil moisture data, and each type has the distinct spatiotemporal strengths and limitations, depending on the diverse applications and retrieval methodologies of different data types (Li et al., in review; The PNNL-82151 FY23 Report). However, the limitations of different soil moisture data in terms of accuracy and spatiotemporal coverage hinder our ability to further understand the soil moisture dynamics across scales. To have a gap free soil moisture data product with a fine spatiotemporal coverage and vertical profiles, we train extreme gradient boosting (XGBoost) models by using (1) in-situ soil moisture measurements from the International Soil Moisture Network (ISMN), (2) soil moisture from the ECMWF reanalysis (ERA) at the 9 km and sub-daily spatiotemporal resolution, (3) the Daymet meteorological fields, and (4) data products that characterize surface conditions, including soil texture, organic content, topography, vegetation type, and rooting depth. We use the trained XGBoost models that have consistent performance across seven soil layers, i.e., 0–5 cm, 5–10 cm, 10–20 cm, 20–40 cm, 40–60 cm, 60–100 cm, and 100–200 cm, and the gridded model predictors to generate a soil moisture data at the 1 km and daily spatiotemporal resolution for the Continental United States (CONUS) from 2001–2020. This dataset can be broadly used for Earth system model benchmark, monitoring extreme weathers, making informed decisions regarding agriculture, water resource management, climate change mitigation, and ecosystem preservation.

58 GEOSCIENCES↗

eLog analysis for accelerators: status and future outlook

This work demonstrates electronic logbook (eLog) systems leveraging modern AI-driven information retrieval capabilities at the accelerator facilities of Fermilab, Jefferson Lab, Lawrence Berkeley National Laboratory (LBNL), SLAC National Accelerator Laboratory. We evaluate contemporary tools and methodologies for information retrieval with Retrieval Augmented Generation (RAGs), focusing on operational insights and integration with existing accelerator control systems. The study addresses challenges and proposes solutions for state-of-the-art eLog analysis through practical implementations, demonstrating applications and limitations. We present a framework for enhancing accelerator facility operations through improved information accessibility and knowledge management, which could potentially lead to more efficient operations.

Accelerator Physics↗

eLog Analysis for Accelerators: Status and Future Outlook

This work demonstrates electronic logbook (eLog) systems leveraging modern AI-driven information retrieval capabilities at the accelerator facilities of Fermilab, Jefferson Lab, Lawrence Berkeley National Laboratory (LBNL), SLAC National Accelerator Laboratory. We evaluate contemporary tools and methodologies for information retrieval with Retrieval Augmented Generation (RAGs), focusing on operational insights and integration with existing accelerator control systems. The study addresses challenges and proposes solutions for state-of-the-art eLog analysis through practical implementations, demonstrating applications and limitations. We present a framework for enhancing accelerator facility operations through improved information accessibility and knowledge management, which could potentially lead to more efficient operations.

Hellert, Thorsten [LBNL, ALS]↗

eLog analysis for accelerators: status and future outlook

This work demonstrates electronic logbook (eLog) systems leveraging modern AI-driven information retrieval capabilities at the accelerator facilities of Fermilab, Jefferson Lab, Lawrence Berkeley National Laboratory (LBNL), SLAC National Accelerator Laboratory. We evaluate contemporary tools and methodologies for information retrieval with Retrieval Augmented Generation (RAGs), focusing on operational insights and integration with existing accelerator control systems. The study addresses challenges and proposes solutions for state-of-the-art eLog analysis through practical implementations, demonstrating applications and limitations. We present a framework for enhancing accelerator facility operations through improved information accessibility and knowledge management, which could potentially lead to more efficient operations.

Sulc, A. [LBL, Berkeley]↗

Working Fluid Characterization and Performance Assessment of Subcritical Organic Rankine Cycles Based on the Lee–Kesler Approach for Energy Recovery

Here, a generalized model using the Lee–Kesler approach based on the corresponding states principle is developed to assess the performance of subcritical Organic Rankine Cycles operating with different working fluids. Each fluid is characterized by five parameters: the acentric factor, critical temperature, critical pressure, molar mass, and the ideal-gas ratio of specific heats at the critical temperature. The model was developed using the compressibility factor modified version of the Benedict–Webb–Rubin equation proposed by Lee and Kesler and the enthalpy and entropy functions to calculate thermodynamic state properties. The model was validated by comparing the results calculated with the model and working fluid thermodynamic properties obtained with the CoolProp database. This comparison was conducted for 91 working fluids, obtaining a relative error below 5% for 88 out of the 91 fluids (∼97%). A generalized parametric study was conducted to determine the influence of the pinch point and each fluid parameter on the performance of Organic Rankine Cycle (ORC) systems. It was found that efficiency increases with critical temperature, ideal-gas ratio of specific heats at the critical temperature, and acentric factor, reaching up to 13%. The developed model enables the evaluation of ORC system performance for existing working fluids. It also allows the formulation and evaluation of new fluids to enhance the performance of the ORC while retrieving energy from any kind of source; and likewise, the methodology can be applied to other power generation cycles.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Refining Planetary Boundary Layer Height Retrievals From Micropulse‐Lidar at Multiple ARM Sites Around the World

Abstract Knowledge of the planetary boundary layer height (PBLH) is crucial for various applications in atmospheric and environmental sciences. Lidar measurements are frequently used to monitor the evolution of the PBLH, providing more frequent observations than traditional radiosonde‐based methods. However, lidar‐derived PBLH estimates have substantial uncertainties, contingent upon the retrieval algorithm used. In addressing this, we applied the Different Thermo‐Dynamic Stabilities (DTDS) algorithm to establish a PBLH data set at five separate Department of Energy's Atmospheric Radiation Measurement sites across the globe. Both the PBLH methodology and the products are subject to rigorous assessments in terms of their uncertainties and constraints, juxtaposing them with other products. The DTDS‐derived product consistently aligns with radiosonde PBLH estimates, with correlation coefficients exceeding 0.77 across all sites. This study delves into a detailed examination of the strengths and limitations of PBLH data sets with respect to both radiosonde‐derived and other lidar‐based estimates of the PBLH by exploring their respective errors and uncertainties. It is found that varying techniques and definitions can lead to diverse PBLH retrievals due to the inherent intricacy and variability of the boundary layer. Our DTDS‐derived PBLH data set outperforms existing products derived from ceilometer data, offering a more precise representation of the PBLH. This extensive data set paves the way for advanced studies and an improved understanding of boundary‐layer dynamics, with valuable applications in weather forecasting, climate modeling, and environmental studies.

54 ENVIRONMENTAL SCIENCES↗

Can We Rely on Satellite Visible/Infrared Microphysical Retrievals of Boundary Layer Clouds in Partially Cloudy Scenes? Implications for Climate Research

This study addresses the longstanding question of the reliability of gridded visible/infrared satellite cloud properties in partially cloudy scenes. By using in-situ cloud probes and airborne Research Scanning Polarimeter (RSP) observations, we analyze bias changes in satellite retrievals from the Spinning Enhanced Visible Infra-Red Imager (SEVIRI) geostationary sensor during the ORACLES campaign. Biases in cloud optical depth (τ) and droplet effective radius (r e ) modestly change for cloud area fraction greater than 35%. The agreement between SEVIRI and RSP r e substantially improves when the retrievals are averaged after removing pixels with τ < 3.0, yielding biases indistinguishable from overcast scenes. In addition, satellite and RSP show an excellent agreement for closed- and open-cell stratocumulus clouds, showing that the satellite retrievals capture spatial changes of r e , and confirming that satellites can faithfully reproduce real physical features for optically thick and partially cloudy scenes. We demonstrate that a simple methodology can minimize uncertainties in satellite-based climate studies.

Painemal, David [NASA Langley Research Center, Ham↗

Draft Feasibility Assessment for Use of AI in Preparing Transportation Safety Analysis Reports

Preparing transportation safety analysis reports for microreactors is time and labor intensive, requiring extensive cross referencing to Federal regulations, previously approved documents, and expert review comments across structural, thermal, criticality, shielding, containment, and security. These burdens are magnified by the novelty of microreactor technologies and the evolving regulatory landscape, as well as current workforce constraints. Generative AI and supporting machine learning tools present an opportunity to accelerate drafting timelines, lift generalized writing burdens, and systematically enforce regulatory adherence through retrieval augmented generation and other knowledge retrieval and mapping methods. This draft report presents a preliminary feasibility assessment of the use of AI to expedite the preparation of microreactor transportation safety analysis reports and proposes an initial methodology for doing so.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Four Years of Atmospheric Boundary Layer Height Retrievals Using COSMIC-2 Satellite Data

This work aimed to study the atmospheric boundary layer height (ABLH) from COSMIC-2 refractivity data, endeavoring to refine existing ABLH detection algorithms and scrutinize the resulting spatial and seasonal distributions. Through validation analyses involving different ground-based methodologies (involving data from lidar, ceilometer, microwave radiometers, and radiosondes), the optimal ABLH determination relied on identifying the lowest refractivity gradient negative peak with a magnitude at least $τ$% times the minimum refractivity gradient magnitude, where $τ$ is a fitting parameter representing the minimum peak strength relative to the absolute minimum refractivity gradient. Different $τ$ values were derived accounting for the moment of the day (daytime, nighttime, or sunrise/sunset) and the underlying surface (land or sea). Results show discernible relations between ABLH and various features, notably, the land cover and latitude. On average, ABLH is higher over oceans (≈1.5 km), but extreme values (maximums > 2.5 km, and minimums < 1 km) are reached over intertropical lands. Variability is generally subtle over oceans, whereas seasonality and daily evolution are pronounced over continents, with higher ABLHs during daytime and local wintertime (summertime) in intertropical (middle) latitudes.

54 ENVIRONMENTAL SCIENCES↗

From Rules to Reasoning: A Survey of Large Language Model-Based Approaches to Scientific Hypothesis and Idea Generation

Scientific hypothesis generation represents a fundamental challenge in contemporary research due to exponentially expanding literature volumes and increasing disciplinary specialization. Large language models (LLMs) have emerged as transformative tools for automated scientific discovery, moving beyond traditional rule-based and literature-mining approaches. Four paradigmatic approaches define current LLM-driven hypothesis generation: direct prompting and fine-tuning methods, knowledge-enhanced frameworks integrating retrieval-augmented generation (RAG), multi-agent collaborative systems simulating research teams, and reasoning-focused approaches implementing cognitive architectures. Domain-specific applications demonstrate statistical equivalence to human expert performance in social psychology, experimental validation in biomedical research, and near-expert quality in astronomy. Evaluation methodologies encompass human expert assessment, LLM-as-judge frameworks, and comprehensive benchmarking systems. Technical challenges include hallucination management, knowledge integration limitations, and balancing novelty with feasibility. Future directions emphasize hybrid neural-symbolic architectures and sophisticated human-AI collaboration models for responsible scientific discovery acceleration.

AI-driven discovery↗

Projected Urban Morphology of the Los Angeles Area by the Year 2100

This dataset provides projections of urban building morphologies for the Los Angeles urban area at 30-meter spatial resolution. It contains 192 raster files that detail two primary building attributes: building footprint fractions (ranging from 0 to 1) and average building heights (ranging from 0 to 75 meters). The projections account for a wide range of future pathways, covering two Shared Socioeconomic Pathway (SSP) scenarios (SSP3 and SSP5), two population scenarios, two developed land intensification scenarios, and four distinct levels of intensification. The dataset was created using dual Generative Adversarial Networks (GANs) trained on 2015 land cover and building properties from the National Land Cover Database (NLCD) and Model America datasets. Supporting information on the dataset has been described in the LAUrbanAreaMorphologyProjections2100_README.txt file.

Pandey, Bhartendu↗

Roadmap on data-centric materials science

Science is and always has been based on data, but the terms ‘data-centric’ and the ‘4th paradigm’ of materials research indicate a radical change in how information is retrieved, handled and research is performed. It signifies a transformative shift towards managing vast data collections, digital repositories, and innovative data analytics methods. The integration of artificial intelligence and its subset machine learning, has become pivotal in addressing all these challenges. This Roadmap on Data-Centric Materials Science explores fundamental concepts and methodologies, illustrating diverse applications in electronic-structure theory, soft matter theory, microstructure research, and experimental techniques like photoemission, atom probe tomography, and electron microscopy. While the roadmap delves into specific areas within the broad interdisciplinary field of materials science, the provided examples elucidate key concepts applicable to a wider range of topics. The discussed instances offer insights into addressing the multifaceted challenges encountered in contemporary materials research.

36 MATERIALS SCIENCE↗

Cloud micro- and macrophysical properties from ground-based remote sensing during the MOSAiC drift experiment

In the framework of the Multidisciplinary drifting Observatory for the Study of Arctic Climate Polarstern expedition, the Leibniz Institute for Tropospheric Research, Leipzig, Germany, operated the shipborne OCEANET-Atmosphere facility for cloud and aerosol observations throughout the whole year. OCEANET-Atmosphere comprises, amongst others, a multiwavelength Raman lidar, a microwave radiometer, and an optical disdrometer. A cloud radar was operated aboard Polarstern by the US Atmospheric Radiation Measurement program. These measurements were processed by applying the so-called Cloudnet methodology to derive cloud properties. To gain a comprehensive view of the clouds, lidar and cloud radar capabilities for low- and high-altitude observations were combined. Cloudnet offers a variety of products with a spatiotemporal resolution of 30 s and 30 m, such as the target classification, and liquid and ice microphysical properties. Additionally, a lidar-based low-level stratus retrieval was applied for cloud detection below the lowest range gate of the cloud radar. Based on the presented dataset, e.g., studies on cloud formation processes and their radiative impact, and model evaluation studies can be conducted.

54 ENVIRONMENTAL SCIENCES↗

Generative large language models for predictive maintenance planning

Maintenance planning and the generation of necessary components for tasks can prove time-consuming and complex. Automating the creation of recurring or similar tasks by leveraging previous planning packages and data, while uncovering insights to automate planning package generation, presents an opportunity to conserve valuable time and resources. This work aims to harness the textual and probabilistic capabilities of large language models (LLMs) to automate the generation of planning packages. Utilizing diverse data sources ranging from raw data to handwritten text, both singular and collaborative LLMs are trained and tested. Results demonstrate their capability to generate essential planning package components, effectively replicating the statistical patterns in the data. This demonstrates the use of these tools inside a digital asset for automated planning. This work outlines a methodology for constructing datasets, a training suite, and evaluation methods for LLM-based textual and conversational planning tools utilized in an asset digital twin. Results indicate that the fine-tuned models generate estimated planning information within the statistical ranges observed in real maintenance data. The models achieve high accuracy (>90%) in document question-answering and instruction generation tasks. Furthermore, the conversational retrieval-augmented generation (RAG) assistant system achieves 100% document retrieval accuracy, while conversational information capture exceeds 98% across the majority of work-package assistant modules.

97 MATHEMATICS AND COMPUTING↗

Marine Boundary Layer Cloud Boundaries and Phase Estimation Using Airborne Radar and In Situ Measurements During the SOCRATES Campaign over Southern Ocean

The Southern Ocean Clouds, Radiation, Aerosol Transport Experimental Study (SOCRATES) was an aircraft-based campaign (15 January–26 February 2018) that deployed in situ probes and remote sensors to investigate low-level clouds over the Southern Ocean (SO). A novel methodology was developed to identify cloud boundaries and classify cloud phases in single-layer, low-level marine boundary layer (MBL) clouds below 3 km using the HIAPER Cloud Radar (HCR) and in situ measurements. The cloud base and top heights derived from HCR reflectivity, Doppler velocity, and spectrum width measurements agreed well with corresponding lidar-based and in situ estimates of cloud boundaries, with mean differences below 100 m. A liquid water content–reflectivity (LWC-Z) relationship, LWC = 0.70Z0.29, was derived to retrieve the LWC and liquid water path (LWP) from HCR profiles. The cloud phase was classified using HCR measurements, temperature, and LWP, yielding 40.6% liquid, 18.3% mixed-phase, and 5.1% ice samples, along with drizzle (29.1%), rain (3.2%), and snow (3.7%) for drizzling cloud cases. The classification algorithm demonstrates good consistency with established methods. This study provides a framework for the boundary and phase detection of MBL clouds, offering insights into SO cloud microphysics and supporting future efforts in satellite retrievals and climate model evaluation.

MBL clouds over Southern Ocean↗

Sea spray aerosol production flux retrieval based on Doppler lidar measurements

Supermicron sea spray aerosols (SSA) play a crucial role in stratocumulus cloud microphysics and drizzle formation. However, our understanding of SSA's spatiotemporal distribution, production mechanisms, and fluxes in the marine boundary layer remains limited. This study introduces a novel approach to determine supermicron SSA production flux and number concentration at various heights above the ocean surface. Our method leverages Doppler lidar data, specifically attenuated backscatter and vertical velocity, collected at a 105 m range gate during the Eastern Pacific Clouds and Aerosols for Precipitation Experiment (EPCAPE) at Scripps Pier, La Jolla, California. To focus on nascent SSA, we analyzed data from periods with wind directions between 225° and 315° and wind speeds exceeding 4 m s −1 . Cloud and precipitation events were excluded using relative humidity, backscatter thresholds, and 2D Video Disdrometer observations. Results show that calculated supermicron SSA number concentrations at 105 m ranged from 0.26 to 2.93 cm −3 for surface wind speeds between 4 and 6 ms −1 . When exceeding the method's Limit of Detection, estimated production fluxes ranged from 0.20 to 1.53 cm −2 s −1 . This methodology has potential for application in various oceanic locations, likely enhancing our understanding of SSA's role in atmospheric processes and climate interactions.

54 ENVIRONMENTAL SCIENCES↗

Benchmark Tracking System for Performance Monitoring

Benchmarking is essential for high-performance software development, particularly for monitoring performance across code iterations. This project focused on enhancing the benchmarking process for Lamellar, an asynchronous runtime for High-Performance Computing (HPC) systems developed at Pacific Northwest National Laboratory. Prior to this work, benchmark results were difficult to track and compare across code versions, presenting significant challenges in identifying performance regressions and long-term trends. The primary objective was to establish a systematic, reproducible approach for measuring performance and detecting regressions following code commits. Our methodology involved three key components: standardizing benchmark outputs, implementing data versioning, and developing analysis tools. We standardized the benchmark output format to JSON Line records containing specific fields (execution time, hardware specifications, and environmental variables). To address data management challenges, we evaluated several options and eventually chose a git repository dedicated to benchmark data. We developed a suite of Python tools that processed benchmark results, enriched them with metadata, and facilitated search in the repository. The resulting system enables more efficient filtering and comparison of performance metrics across commit histories, hardware configurations, and benchmark variants through a unified query interface. Our implementation reduces computational overhead by first checking for existing results through configuration matching before initiating new benchmark runs, thereby conserving resources. The system has been validated by Lamellar developers. It organizes results by benchmark type and build configurations for efficient retrieval. Future developments include a planned Large Language Model interface for predicting benchmark performance, incorporating the criterion package for statistical analysis, which will enable automated detection of statistically significant performance changes, and integration with continuous integration pipelines. Despite these enhancements being reserved for future work, this project has successfully provided the Lamellar development team with a framework for maintaining consistent performance standards and identifying optimization opportunities across workloads and hardware environments.

97 MATHEMATICS AND COMPUTING↗

Reservoir Sediment Management and Monitoring Database

Overview This dataset compiles dam sediment management and monitoring information from surveys, case studies, and journal articles. Additionally, features described by the National Inventory of Dams (i.e., presence of sluice gates) are included to indicate known infrastructure features that may address sediment releases. The location and description of records from downstream monitoring gages are catalogued in order to help with tracking conditions over time (e.g., before and after management actions, as operations change, etc.). The data help address national scale understanding of challenges and solutions related to the accumulation of sediment behind a dam as well as downstream passage. Sediment trapping causes problems as it reduces storage capacity, disrupts dam and reservoir function, impedes access for recreation, alters water quality/habitat conditions, and contributes to riverbank and coastal erosion within the reservoir. Data compilation from a variety of sources is a first step towards assessing system-wide efficacy of management solutions. This dataset was developed under the Water Power Technologies Office funded effort which began as a Seedling on Reservoir Sedimentation Data, and was supported by the Reservoir Sedimentation Modeling Framework and Data Analysis project. These projects have addressed challenges in describing sediment transport, trapping, and management at dams throughout the US. Methodology An outer join on dams/reservoirs with surveys and survey reports (documented in the RESSED database, USBR or USACE databases, project websites, etc.) with the National Inventory of Dams, based on the NIDID to determine dams with documented management and/or sluice gates. Additional dams with documented management activity were identified through review of technical articles from the past 25 years in Journal of Hydrology, Journal of Water Resources Planning and Management, Geomorphology, Journal of Hydraulic Engineering, Water, Journal of Cleaner Production, International Journal of Sediment Research, Nature Scientific Reports, Earth Surface Processes and Landforms, and Environmental Science and Pollution Research. Individual records were created for each survey or management activity documented. To evaluate downstream sediment monitoring records, the nhdPlusTools and dataRetrieval packages in R were used to find gages within 10km of each dam in the management database. Length of record and location of matched gages were retrieved for those parameters relevant to sediment concentration or total sediment discharge.

Hansen, Carly [ORNL] (ORCID:0000000193280838)↗