Search NASA⌕ Search

SEARCH · Search NASA

Results for “data retrieval”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Observational Data for Next-Generation Climate Model Evaluation: Requirements, Considerations, and Best Practices

Climate model simulations are an important source of information about our planet’s climate system and also enable informed decision-making under different future scenarios. As a new archive of results from the next generation of climate models is anticipated to become available with the Coupled Model Intercomparison Project phase 7 (CMIP7), the need to develop efficient and robust methods to evaluate models is paramount. Observations are an integral part of model evaluation, providing a means to quantify and understand the degree to which climate models can faithfully reproduce Earth system processes. Such analysis is critical for constraining climate projections, identifying areas of focus for model development, and assisting analysts in deciphering the utility of models for specific applications. Observations of Earth system come from a diversity of sources, span different space–time domains, and are produced by different communities, and each dataset features different data structures and formats, metadata standards, and its own unique uncertainties. Uncertainties in an observational dataset may stem from gaps in temporal and spatial coverage, instrumentation errors, or assumptions in retrieval and processing methods. How then does one ensure that observational data are ready for use and utilized in the most appropriate way for robust, rapid, and routine climate model evaluation? The CMIP7 Model Benchmarking Task Team with input from the broader climate modeling, model evaluation, and observational data communities present a vision and considerations for best practices toward the optimal and appropriate use of observational data to support next-generation climate model evaluation.

Climate models↗

Evaluation of Physical Microphysical Property Retrieval Algorithms During the 2020 IMPACTS Field Campaign

The NASA Investigation of Microphysics and Precipitation for Atlantic Coast Threatening Snowstorms (IMPACTS) field campaign provides high-quality, high-altitude aircraft lidar (532 nm), radar (W-band) and in-cloud microphysical aircraft data taken during wintertime storm events impacting the United States. This study evaluates two mass-dimensional relationships (Brown and Francis (1995, BF95); Heymsfield (2014, H14) and two lidar-radar microphysical retrieval algorithms (Cloudsat and CALIPSO Ice Cloud Property Product (2C-ICE); VarPy (a variational method derived from the satellite lidar-radar data community)) to estimate aircraft-retrieved volume extinction coefficient (σ), ice water content (IWC), and effective radius (r e ) during the 2020 IMPACTS deployment. BF95 and H14 have a close 1:1 correlation (R 2 = 0.98) with in-situ observations of σ. However, only BF95 displays a linear, consistent, and almost temperature-independent low bias for IWC and r e , which likely arises from the environmental conditions used to determine each. Unlike the field-campaign-derived BF95 and H14 relationships, VarPy and 2C-ICE directly ingest the aircraft-based lidar and radar data to simulate σ, IWC, and r e . For all three microphysical parameters, VarPy and 2C-ICE retrieval errors became notably more pronounced around the dendritic growth zone (-15°C to -10°C) and near freezing (≥-5°C), which suggests that both algorithms experience difficulty addressing riming and aggregation processes and with larger particles (dendrites and plates) due in part to their simplified ice particle assumptions. However, the mean-melt diameter ice-particle assumption did yield more accurate IWC estimates, which led to slightly better overall results for VarPy.

54 ENVIRONMENTAL SCIENCES↗

Physics informed neural network can retrieve rate and state friction parameters from acoustic monitoring of laboratory stick-slip experiments

Various machine learning (ML) and deep learning (DL) techniques have been recently applied to the forecasting of laboratory earthquakes from friction experiments. The magnitude and timing of shear failures in stick-slip cycles are predicted using features extracted from the recorded ultrasonic or acoustic emission (AE) signals. In addition, the Rate and State Friction (RSF) constitutive laws are extensively used to model the frictional behavior of faults. In this work, we use data from shear experiments coupled with passive acoustic (variance, kurtosis, and AE rate) interleaved with active source ultrasonic monitoring (transmitted wave amplitude) to develop physics-informed neural network (PINN) models incorporating the RSF law and AE rate generation equation with wave amplitude serving as a proxy for friction state variable. This PINN framework allows learning RSF parameters from stick-slip experiments rather than measuring them through a series of velocity step experiments. We observe that when the stick-slip cycles are irregular, the PINN models outperform the data-driven DL models. Transfer learning (TL) PINN models are also developed by pre-training on data collected at one normal stress level followed by forecasting shear failures and retrieving RSF parameters at other stress levels (i.e., with different recurrence intervals) after retraining on a limited amount of new data. Our findings suggest that TL models perform better compared to standalone models. Both standalone and TL PINN-estimated RSF parameters and their ground truth values show excellent agreements thus demonstrating that RSF parameters can be retrieved from laboratory stick-slip experiments using the corresponding acoustic data and that the transmitted wave amplitude provides a good representation of the evolving frictional state during stick-slips.

58 GEOSCIENCES↗

NWTC Site 4.0 - NREL ASSIST (SN10) / Thermodynamic retrievals TROPoe

This dataset contains daily files with thermodynamic profiles retrieved with the optimal estimation physical retrieval TROPoe v0.12 (Turner and Löhnert 2014; Turner and Blumberg 2019; Turner and Löhnert 2021). The profiles are retrieved every 10 minutes from instantaneous observations from the NREL ASSIST-II (SN 10) infrared spectrometer. Observations are noise-filtered but not averaged in time to minimize errors due to non-uniform clouds. Additional input data in TROPoe are cloud base height from a Vaisala CL51 ceilometer. The full pipeline for running the retrieval is available at https://github.com/StefanoWind/TROPoe_processor. Met data were not ingested. In addition to these temporally resolved input data, TROPoe requires an a priori dataset (prior) that provides mean climatological estimates of thermodynamic profiles and specifies how temperature and humidity covary with height as an input (for details see, e.g., Djalalova et al. 2022). The prior is a key component of the retrieval and provides a constraint on the ill-posed inversion problem. A monthly prior was computed from operational radiosonde launches at Denver, CO.

17 WIND ENERGY↗

NWTC Site 3.2 - NREL ASSIST (SN12) / Thermodynamic retrievals TROPoe

This dataset contains daily files with thermodynamic profiles retrieved with the optimal estimation physical retrieval TROPoe v0.12 (Turner and Löhnert 2014; Turner and Blumberg 2019; Turner and Löhnert 2021). The profiles are retrieved every 10 minutes from instantaneous observations from the NREL ASSIST-II (SN 12) infrared spectrometer. Observations are noise-filtered but not averaged in time to minimize errors due to non-uniform clouds. Additional input data in TROPoe are cloud base height from a Vaisala CL51 ceilometer. The full pipeline for running the retrieval is available at https://github.com/StefanoWind/TROPoe_processor. Met data were not ingested. In addition to these temporally resolved input data, TROPoe requires an a priori dataset (prior) that provides mean climatological estimates of thermodynamic profiles and specifies how temperature and humidity covary with height as an input (for details see, e.g., Djalalova et al. 2022). The prior is a key component of the retrieval and provides a constraint on the ill-posed inversion problem. A monthly prior was computed from operational radiosonde launches at Denver, CO.

17 WIND ENERGY↗

Title NWTC Site 3.2 - NREL ASSIST (SN11) / Thermodynamic retrievals TROPoe

This dataset contains daily files with thermodynamic profiles retrieved with the optimal estimation physical retrieval TROPoe v0.12 (Turner and Löhnert 2014; Turner and Blumberg 2019; Turner and Löhnert 2021). The profiles are retrieved every 10 minutes from instantaneous observations from the NREL ASSIST-II (SN 11) infrared spectrometer. Observations are noise-filtered but not averaged in time to minimize errors due to non-uniform clouds. Additional input data in TROPoe are cloud base height from a Vaisala CL51 ceilometer. The full pipeline for running the retrieval is available at https://github.com/StefanoWind/TROPoe_processor. Met data were not ingested. In addition to these temporally resolved input data, TROPoe requires an a priori dataset (prior) that provides mean climatological estimates of thermodynamic profiles and specifies how temperature and humidity covary with height as an input (for details see, e.g., Djalalova et al. 2022). The prior is a key component of the retrieval and provides a constraint on the ill-posed inversion problem. A monthly prior was computed from operational radiosonde launches at Denver, CO.

17 WIND ENERGY↗

FC Site 4.0 - NLR Thermodynamic profiler (ASSIST II-11) Thermodynamic Retrievals TROPoe

This dataset contains daily files with thermodynamic profiles retrieved with the optimal estimation physical retrieval TROPoe v0.19 (Turner and Löhnert 2014; Turner and Blumberg 2019; Turner and Löhnert 2021). The profiles are retrieved every 10 minutes from instantaneous observations from the NLR ASSIST II infrared spectrometer. Observations are noise-filtered but not averaged in time to minimize errors due to non-uniform clouds. Additional input data in TROPoe are cloud base height (CBH) from co-located scanning lidar. The full pipeline for running the retrieval is available at https://github.com/StefanoWind/TROPoe_processor. Met data was not ingested. In addition to these temporally resolved input data, TROPoe requires an a priori dataset (prior) that provides mean climatological estimates of thermodynamic profiles and specifies how temperature and humidity covary with height as an input (for details see, e.g., Djalalova et al. 2022). The prior is a key component of the retrieval and provides a constraint on the ill-posed inversion problem. A monthly prior was computed from operational radiosonde launches in Denver, CO.

17 WIND ENERGY↗

The influence of cloud cover on the reliability of satellite-based solar resource data

Satellite-based solar resource data are often developed and validated by using binary cloudiness categories: clear sky or overcast cloudy sky. To investigate the reliability of solar resource data in partially cloudy conditions, we estimate cloud fraction using two distinct algorithms: a physical retrieval model using surface observed global horizontal irradiance (GHI) and direct normal irradiance (DNI) and a temporal average of cloud mask data estimated by the observed DNI. Our analysis reveals a significant presence of scattered clouds, broken clouds, and mismatches between satellite- and surface-based cloud data at 17 surface sites across the contiguous United States, though confidently clear and cloudy conditions collectively account for more than 70 % of the data. Solar radiation is computed using the National Solar Radiation Database (NSRDB) algorithm and validated using surface observations. Here, our findings suggest that, in the presence of scattered clouds, NSRDB data for clear-sky conditions can be subject to significant overestimation. In cloudy-sky conditions classified by satellite data, DNI computed by the Fast All-sky Radiation Model for Solar applications with DNI (FARMS-DNI) can be underestimated when limited clouds are detected by surface observations. The bias observed in several cloudiness categories indicates that the NSRDB is exceptionally accurate in confidently clear conditions. However, clear-sky conditions with scattered clouds and mismatched cloud data contribute significantly to the overall uncertainties in the NSRDB. Therefore, future improvements in solar resource data should involve development and implementation of satellite-derived cloud fraction and should consider a novel radiative transfer model accounting for amplified cloud reflection. The evaluation within cloudiness categories also provides a physical rationale for the superior performance of FARMS-DNI compared to the Direct Insolation Simulation Code (DISC) in both cloudy-sky and all-sky conditions.

14 SOLAR ENERGY↗

Statistically Resolved Planetary Boundary Layer Height Diurnal Variability Using Spaceborne Lidar Data

The Planetary Boundary Layer Height (PBLH) significantly impacts weather, climate, and air quality. Understanding the global diurnal variation of the PBLH is particularly challenging due to the necessity of extensive observations and suitable retrieval algorithms that can adapt to diverse thermodynamic and dynamic conditions. This study utilized data from the Cloud-Aerosol Transport System (CATS) to analyze the diurnal variation of PBLH in both continental and marine regions. By leveraging CATS data and a modified version of the Different Thermo-Dynamics Stability (DTDS) algorithm, along with machine learning denoising, the study determined the diurnal variation of the PBLH in continental mid-latitude and marine regions. The CATS DTDS-PBLH closely matches ground-based lidar and radiosonde measurements at the continental sites, with correlation coefficients above 0.6 and well-aligned diurnal variability, although slightly overestimated at nighttime. In contrast, PBLH at the marine site was consistently overestimated due to the viewing geometry of CATS and complex cloud structures. The study emphasizes the importance of integrating meteorological data with lidar signals for accurate and robust PBLH estimations, which are essential for effective boundary layer assessment from satellite observations.

54 ENVIRONMENTAL SCIENCES↗

Multi-Doppler radar analysis from CSAPR, CHIVO, COW, and RMA-1 radars during the CACTI/RELAMPAGO experiments in Argentina in 2018

This data set contains multi-Doppler radar analysis from CSAPR-2, CSU-CHIVO, COW, and RMA-1 radars. These radars were collecting dual-polarization data during the CACTI/RELAMPAGO experiments in Argentina in 2018. Doppler analysis is systematically conducted for 31 days with convection. Three-dimensional wind fields are retrieved using the PyDDA algorithm. Dual-polarization information is also included in the data set.

3D cartesian gridded corrected mean Doppler veloci↗

Applying Gaussian Process Machine Learning and Modern Probabilistic Programming to Satellite Data to Infer CO 2 Emissions

Satellite data provides essential insights into the spatiotemporal distribution of CO 2 concentrations. However, many atmospheric inverse models fail to adequately incorporate the spatial and temporal correlations inherent in satellite observations and often lack rigorous methods for estimating parameters like spatial length scales. We introduce an inference model that processes the spatiotemporal covariance in satellite data and estimates hyperparameters such as covariance length scales. Our approach uses the Gaussian process (GP) machine learning (ML) and modern probabilistic programming languages (PPLs) to perform atmospheric inversions of emissions from satellite data. We develop a GP ML inversion system based on modern PPLs and the GEOS-Chem chemical transport model, simulating atmospheric CO 2 concentrations corresponding to the Orbiting Carbon Observatory-2/3 (OCO-2/3) data for July 2020. In our supervised learning framework, we treat the GEOS-Chem simulated data set as the target, with predictors derived by scaling the target with sector-specific factors hidden from the GP machine. Our results show that the GP model, combined with GPU-enabled PPLs, effectively retrieves true emission scaling factors and infers noise levels concealed within the data. This suggests that our method could be applied over larger areas with more complex covariance structures, enabling comprehensive analysis of the spatiotemporal patterns observed in OCO-2/3 and similar satellite data sets.

54 ENVIRONMENTAL SCIENCES↗

Advanced Instrumentation for Metal Additive Manufacturing

Laser powder bed fusion (LPBF) is the most widely used process for metal additive manufacturing (AM), particularly where complex geometries provide performance advantages unattainable with traditional manufacturing techniques. However, LPBF is highly sensitive to innate variability in both the powder spreading and fusion steps, often leading to defects such as pores that are difficult to detect yet significantly impair component mechanical properties and fatigue life. This thesis presents a range of novel instruments enabling both precise assessment of powder layer characteristics and in-situ thermal metrology of metal AM to advance the quality control of LPBF. First, leveraging a custom X-ray microscope and a radiation-transport model developed through this work, transmission X-ray imaging is used to study spreading of thin metal powder layers. Effective layer depth is directly mapped at a process-relevant size scale, surpassing optical techniques that can only estimate local deposition from layer surface topography. Layer packing density and quality are shown to be influenced by powder flowability and particle size relative to nominal powder layer thickness. Layer quality is additionally connected to the geometry of the spreading implement and its velocity. This technique and its presented findings enable pairing feedstocks with spreading strategies that create layers with consistent packing density and uniformity. Second, a twofold approach is employed to optically interrogate the laser fusion step of LPBF for observing signatures of defect formation. Aperture division multiplexing is conceptualized, providing for simultaneous laser delivery and high-fidelity infrared (IR) process monitoring through a common optic. In-situ microscopy at 50 μm spatial resolution and at mid-wave IR wavelengths is proven readily achievable with the first purpose-built optic of this type. Next, a bespoke imaging spectrometer, along with a temperature-emissivity separation technique, is used to retrieve accurate process temperatures over a 1000 K range. Data from these instruments are correlated to porosity as fine as 4.3 μm in two LPBF test artifacts, as verified using computed tomography (CT), establishing the viability of robust optically-based component qualification.

Penny, Ryan↗

Asi Nuclear Energy Sensors Data Portal Chatbot And Data Structuring Tool

The Idaho National Laboratory (INL) is advancing the development of an AI-powered chatbot and data structuring tool specifically designed to accelerate data mining processes for sensor-related information and seamlessly integrate the results into the ASI Sensors Data Portal (https://nes.energy.gov/). By doing so, the software aims to enhance the accessibility, usability, and organization of sensor data for nuclear energy applications. The software initial phase focuses on retrieving comprehensive datasets, prioritizing the past five years of publicly available information from the Office of Scientific and Technical Information (OSTI). These datasets will be meticulously processed to ensure compatibility, employing cleaning and preprocessing steps to eliminate irrelevant, incomplete, or corrupted information, thus establishing a robust foundation for subsequent AI use. The data will serve as the backbone for training an AI model and chatbot, which will act as an interactive tool enabling users to ask complex, context-specific questions and receive accurate, validated answers derived from constrained literature. In parallel, the project incorporates a data structuring process supported by AI to organize sensor information from multiple sources into a standardized format. This structured data will include detailed sensor specifications, such as measurement range, applications, accuracy, and operating conditions, generated and documented with AI. These specifications will be systematically integrated into the sensor portal. To maintain the highest levels of accuracy and relevance, all AI-generated outputs will be reviewed and validated by subject matter experts (SMEs), with additional fields or parameters added as needed. Future stages of the project aim to expand the dataset beyond OSTI to include other sources and potentially incorporate unclassified controlled information (UCI) with restricted access protocols to address security and confidentiality requirements.

Mapes, NormanJ. [Idaho National Laboratory (INL), ↗

AERO-MAP: a data compilation and modeling approach to understand spatial variability in fine- and coarse-mode aerosol composition

Abstract. Aerosol particles are an important part of the Earth climate system, and their concentrations are spatially and temporally heterogeneous, as well as being variable in size and composition. Particles can interact with incoming solar radiation and outgoing longwave radiation, change cloud properties, affect photochemistry, impact surface air quality, change the albedo of snow and ice, and modulate carbon dioxide uptake by the land and ocean. High particulate matter concentrations at the surface represent an important public health hazard. There are substantial data sets describing aerosol particles in the literature or in public health databases, but they have not been compiled for easy use by the climate and air quality modeling community. Here, we present a new compilation of PM2.5 and PM10 surface observations, including measurements of aerosol composition, focusing on the spatial variability across different observational stations. Climate modelers are constantly looking for multiple independent lines of evidence to verify their models, and in situ surface concentration measurements, taken at the level of human settlement, present a valuable source of information about aerosols and their human impacts complementarily to the column averages or integrals often retrieved from satellites. We demonstrate a method for comparing the data sets to outputs from global climate models that are the basis for projections of future climate and large-scale aerosol transport patterns that influence local air quality. Annual trends and seasonal cycles are discussed briefly and are included in the compilation. Overall, most of the planet or even the land fraction does not have sufficient observations of surface concentrations – and, especially, particle composition – to characterize and understand the current distribution of particles. Climate models without ammonium nitrate aerosols omit ∼ 10 % of the globally averaged surface concentration of aerosol particles in both PM2.5 and PM10 size fractions, with up to 50 % of the surface concentrations not being included in some regions. In these regions, climate model aerosol forcing projections are likely to be incorrect as they do not include important trends in short-lived climate forcers.

Mahowald, Natalie M. (ORCID:000000022873997X)↗

Extraction and Analysis of Time Series Data from Building Automation Systems Using Large Language Models

Semantic schemas like Haystack 4, Brick and ASHRAE standard 223 enable the structured, standardized, and machine-readable representation of building data, facilitating interoperability, data integration, and advanced analytics. However, extracting information from these models requires specialized expertise in SPARQL and other programming languages, skills that are not commonly found among building professionals. Recent advancements in Large Language Models (LLMs), such as ChatGPT, enable the construction of queries using natural language, making it easier for individuals to interact with these systems in a manner that resembles everyday speech. However, these methods have not yet been tested on building semantic ontologies. This paper introduces a novel workflow and tool for enabling users to ask questions about a specific building's data, using natural language and receive answers automatically generated by GPT-4o. Our approach integrates semantic ontologies with advanced LLM capabilities to automate three critical steps: (1) generating SPARQL queries to retrieve time series references from ontological models, (2) extracting the corresponding time series data from the Building Automation System, and (3) performing computations and visualizations tailored to the user's query. The proposed method simplifies access to BAS data, allowing both domain experts and non-specialists to conduct sophisticated analyses without needing extensive technical knowledge of semantic web technologies. By demonstrating this pipeline, we facilitate more accessible and scalable data-driven decision-making in building operations and management.

Mulayim, Ozan Baris↗

Extraction and Analysis of Time Series Data from Building Automation Systems Using Large Language Models

Semantic schemas like Haystack 4, Brick and ASHRAE standard 223 enable the structured, standardized, and machine-readable representation of building data, facilitating interoperability, data integration, and advanced analytics. However, extracting information from these models requires specialized expertise in SPARQL and other programming languages, skills that are not commonly found among building professionals. Recent advancements in Large Language Models (LLMs), such as ChatGPT, enable the construction of queries using natural language, making it easier for individuals to interact with these systems in a manner that resembles everyday speech. However, these methods have not yet been tested on building semantic ontologies. This paper introduces a novel workflow and tool for enabling users to ask questions about a specific building's data, using natural language and receive answers automatically generated by GPT-4o. Our approach integrates semantic ontologies with advanced LLM capabilities to automate three critical steps: (1) generating SPARQL queries to retrieve time series references from ontological models, (2) extracting the corresponding time series data from the Building Automation System, and (3) performing computations and visualizations tailored to the user's query. The proposed method simplifies access to BAS data, allowing both domain experts and non-specialists to conduct sophisticated analyses without needing extensive technical knowledge of semantic web technologies. By demonstrating this pipeline, we facilitate more accessible and scalable data-driven decision-making in building operations and management.

Mulayim, Ozan Baris↗

AI-enabled Lorentz microscopy for quantitative imaging of nanoscale magnetic spin textures

The manipulation and control of nanoscale magnetic spin textures are of rising interest as they are potential foundational units in next-generation computing paradigms. Achieving this requires a quantitative understanding of the spin texture behavior under external stimuli using in situ experiments. Lorentz transmission electron microscopy (LTEM) enables real-space imaging of spin textures at the nanoscale, but quantitative characterization of in situ data is extremely challenging. Here, we present an AI-enabled phase-retrieval method based on integrating a generative deep image prior with an image formation forward model for LTEM. Our approach uses a single out-of-focus image for phase retrieval and achieves significantly higher accuracy and robustness to noise compared to existing methods. Furthermore, our method is capable of isolating sample heterogeneities from magnetic contrast, as shown by application to simulated and experimental data. This approach allows quantitative phase reconstruction of in situ data and can also enable near real-time quantitative magnetic imaging.

36 MATERIALS SCIENCE↗

RAG for FLAG: AI Assistance for a Physics Code

Artificial intelligence (AI) has quickly become an important tool in scientific research, where significant efforts are underway to develop tools that will expedite the research process. One area of particular impact is scientific software, which can be particularly complex, and therefore time consuming to learn and use effectively. AI assistants are increasingly helping to streamline the process by performing tasks such as interactively answering user questions or suggesting solutions. Los Alamos National Laboratory (LANL) develops several advanced scientific codes, such as FLAG, which can be used to run multiphysics simulations. With this study, our goal was to develop an AI assistant for FLAG that could help make the process of understanding the software and running physics simulations more efficient. To develop an AI assistant for FLAG, we used a method called retrieval-augmented generation (RAG), which is a technique that uses information from relevant data sources to enhance the accuracy of large language models (LLMs). We used the FLAG user manual and other FLAG documentation as the knowledge base for the RAG system. When a user provides a query, RAG retrieves relevant sections from the knowledge base in response, then uses those excerpts to generate grounded and contextually rich answers. We found that our AI assistant was able to provide context aware answers and source references to user queries. To evaluate performance, we developed a set of 40 benchmark questions and compared the accuracy of the responses to those of two standard LLMs without retrieval. Our AI assistant significantly outperformed the standard LLMs at answering FLAG-related questions, with an 82.5% accuracy rate, compared to 47.5% for both of the standard LLMs. This has the potential to make the process of learning and using FLAG much easier, especially for new users. Ultimately, it supports LANL’s broader mission by empowering scientists and engineers to focus more on discovery and analysis rather than on navigating complex software systems.

97 MATHEMATICS AND COMPUTING↗