Search NASA⌕ Search

SEARCH · Search NASA

Results for “data retrieval”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Towards a RAG-based summarization for the Electron Ion Collider

Abstract The complexity and sheer volume of information — encompassing documents, papers, data, and other resources — from large-scale experiments demand significant time and effort to navigate, making the task of accessing and utilizing these varied forms of information daunting, particularly for new collaborators and early-career scientists.To tackle this issue, a Retrieval Augmented Generation (RAG)-based Summarization AI for EIC (RAGS4EIC) is under development. This AI-Agent not only condenses information but also effectively references relevant responses, offering substantial advantages for collaborators. Our project involves a two-step approach: first, querying a comprehensive vector database containing all pertinent experiment information; second, utilizing a Large Language Model (LLM) to generate concise summaries enriched with citations based on user queries and retrieved data. We describe the evaluation methods that use RAG assessments (RAGAs) scoring mechanisms to assess the effectiveness of responses. Furthermore, we describe the concept of prompt template based instruction-tuning which provides flexibility and accuracy in summarization. Importantly, the implementation relies on LangChain [1], which serves as the foundation of our entire workflow. This integration ensures efficiency and scalability, facilitating smooth deployment and accessibility for various user groups within the Electron Ion Collider (EIC) community. This innovative AI-driven framework not only simplifies the understanding of vast datasets but also encourages collaborative participation, thereby empowering researchers. As a demonstration, a web application has been developed to explain each stage of the RAG Agent development in detail. The application can be accessed athttps://rags4eic-ai4eic.streamlit.app.[A tagged version of the source code can be found inhttps://github.com/ai4eic/EIC-RAG-Project/releases/tag/AI4EIC2023_PROCEEDING.]

Instruments & Instrumentation↗

Enhanced Microbial Detection Capabilities by a Rapid Portable Instrument

We present data describing a progression of continuing technology development - from expanding the detection capabilities of the current PTS unit to re-outfitting the instrument with a protein microarray increasing the number of detectable compounds. To illustrate the adaptability of the cartridge format, on-orbit operations data from the ISS demonstrate the detection of the fungal cell wall compound beta-glucan using applicable LOCAD-PTS cartridges. LOCAD-PTS is a handheld device consisting of a spectrophotometer, an onboard pumping mechanism, and data storage capabilities. A suite of interchangeable cartridges lined with four distinct capillaries allow a hydrated sample to mix with necessary reagents in the channels before being pumped to the optical well for spectrophotometric analysis. The reagents housed in one type of cartridge trigger a reaction based on the Limulus Amebocyte Lysate (LAL) assay, which results in the release of paranitroaniline dye. The dye is measured using a 395 nm filter. The LAL assay detects the Gram-negative bacterial cell wall molecule, endotoxin or lipopolysaccharide (LPS). The more dye released, the greater the concentration of endotoxin in the sample. Sampling, quantitative analysis, and data retrieval require less than 20 minutes. This is significantly faster than standard culture-based methods, which require at least a 24 hour incubation period.Using modified cartridges, we demonstrate the detection of Gram negative bacteria with protein microarray technology. Additionally, we provide data from multiple field tests where both standard and advanced PTS technologies were used. These tests investigate the transfer of target microbial molecules from one surface to another. Collectively, these data demonstrate that the new cartridges expand the number of compounds detected by LOCAD-PTS, while maintaining the rapid, in situ analysis characteristic of the instrument. The unit provides relevant data for verifying sterile sample collection protocols, which are critical for conducting accurate scientific experiments during future missions to the Moon and Mars.

Morris, Heather↗

Easy, Scalable Subsetting of GEDI Point Clouds

The GEDI Subsetter, a Python tool developed for NASA’s Multi-mission Algorithm and Analysis Platform (MAAP), optimizes the accessibility and visualization of GEDI point clouds by enabling users to efficiently subset data in a convenient, scalable manner. Complex science data often requires users to learn new software skills and handle many large files. Handling and cleaning large data sets is tedious and error-prone. These challenges significantly impede analysis. One of the goals of NASA's MAAP is to provide a platform that lowers the barrier to conducting research and analysis at scale. When a group of MAAP users wanted to conduct above-ground biomass estimation using GEDI data, we found that their existing workflow for leveraging GEDI data suffered from the barriers mentioned above. Furthermore, their workflow did not scale easily beyond a small number of granules. We found that existing tools related to GEDI data retrieval and subsetting were too limiting, so the GEDI Subsetter was written to support MAAP users’ needs. Being able to run many subsetting jobs simultaneously in the MAAP, and parallelizing the code itself, has led to significant speed improvements in obtaining relevant data, reducing subsetting time from hours to minutes. MAAP users can now more quickly and easily obtain only the data relevant to their research, by choosing which GEDI collection they want to work with (L1A, L2A, L2B, or L4A), and how they want to subset it, by specifying an area of interest, a temporal range, and relevant attributes. This has significantly reduced the feedback loop for users, allowing them to much more quickly subset GEDI data and begin their analysis. Although the GEDI Subsetter originally targeted users of the MAAP, it is generalized such that it can also be used outside of the MAAP and includes a command-line interface for convenience. Furthermore, with minor modifications, it should be possible to use it with non-GEDI data as the general pattern should be applicable to other sparse/track-based sensors.

Charles Daniels↗

Updates to the ATLAS Data Carousel Project

The High Luminosity upgrade to the LHC (HL-LHC) is expected to deliver scientific data at the multi-exabyte scale. In order to address this unprecedented data storage challenge, the ATLAS experiment launched the Data Carousel project in 2018. Data Carousel is a tape-driven workflow whereby bulk production campaigns with input data resident on tape are executed by staging and promptly processing a sliding window to disk buffer such that only a small fraction of inputs are pinned on disk at any one time. Data Carousel is now in production for ATLAS in Run3. In this paper, we provide updates on recent Data Carousel R&D projects, including data-on-demand and tape smart writing. Data-on-demand removes from disk data that has not been accessed for a predefined period, when users request them, they will be either staged from tape or recreated by following the original production steps. Tape smart writing employs intelligent algorithms for file placement on tape in order to retrieve data back more efficiently, which is our long term strategy to achieve optimal tape usage in Data Carousel.

42 ENGINEERING↗

Interoperability of Heliophysics Virtual Observatories

If you'd like to find interrelated heliophysics (also known as space and solar physics) data for a research project that spans, for example, magnetic field data and charged particle data from multiple satellites located near a given place and at approximately the same time, how easy is this to do? There are probably hundreds of data sets scattered in archives around the world that might be relevant. Is there an optimal way to search these archives and find what you want? There are a number of virtual observatories (VOs) now in existence that maintain knowledge of the data available in subdisciplines of heliophysics. The data may be widely scattered among various data centers, but the VOs have knowledge of what is available and how to get to it. The problem is that research projects might require data from a number of subdisciplines. Is there a way to search multiple VOs at once and obtain what is needed quickly? To do this requires a common way of describing the data such that a search using a common term will find all data that relate to the common term. This common language is contained within a data model developed for all of heliophysics and known as the SPASE (Space Physics Archive Search and Extract) Data Model. NASA has funded the main part of the development of SPASE but other groups have put resources into it as well. How well is this working? We will review the use of SPASE and how well the goal of locating and retrieving data within the heliophysics community is being achieved. Can the VOs truly be made interoperable despite being developed by so many diverse groups?

Thieman, J.↗

Tracking-Data-Conversion Tool

Object Oriented Data Technology (OODT) is a software framework for creating a Web-based system for exchange of scientific data that are stored in diverse formats on computers at different sites under the management of scientific peers. OODT software consists of a set of cooperating, distributed peer components that provide distributed peer-topeer (P2P) services that enable one peer to search and retrieve data managed by another peer. In effect, computers running OODT software at different locations become parts of an integrated data-management system.

Flora-Adams, Dana↗

Software Framework for Peer Data-Management Services

Object Oriented Data Technology (OODT) is a software framework for creating a Web-based system for exchange of scientific data that are stored in diverse formats on computers at different sites under the management of scientific peers. OODT software consists of a set of cooperating, distributed peer components that provide distributed peer-to-peer (P2P) services that enable one peer to search and retrieve data managed by another peer. In effect, computers running OODT software at different locations become parts of an integrated data-management system.

Hughes, John↗

The emulsion chamber technology experiment

Photographic emulsion has the unique property of recording tracks of ionizing particles with a spatial precision of 1 micron, while also being capable of deployment over detector areas of square meters or 10's of square meters. Detectors are passive, their cost to fly in Space is a fraction of that of instruments of similar collecting. A major problem in their continued use has been the labor intensiveness of data retrieval by traditional microscope methods. Two factors changing the acceptability of emulsion technology in space are the astronomical costs of flying large electronic instruments such as ionization calorimeters in Space, and the power and low cost of computers, a small revolution in the laboratory microscope data-taking. Our group at UAH made measurements of the high energy composition and spectra of cosmic rays. The Marshall group has also specialized in space radiation dosimetry. Ionization calorimeters, using alternating layers of lead and photographic emulsion, to measure particle energies up to 10(exp 15) eV were developed. Ten balloon flights were performed with them. No such calorimeters have ever flown in orbit. In the ECT program, a small emulsion chamber was developed and will be flown on the Shuttle mission OAST-2 to resolve the principal technological questions concerning space exposures. These include assessments of: (1) pre-flight and orbital exposure to background radiation, including both self-shielding and secondary particle generation; the practical limit to exposure time in space can then be determined; (2) dynamics of stack to optimize design for launch and weightlessness; and (3) thermal and vacuum constraints on emulsion performance. All these effects are cumulative and affect our ability to perform scientific measurements but cannot be adequately predicted by available methods.

Gregory, John C.↗

Tracking the Hunga Tonga-Hunga Ha’apai Eruption Stratospheric Aerosol and Trace Gas Plumes Using Machine Learning

On January 15, 2022, the Hunga Tonga-Hunga Ha’apai (hereafter, Hunga Tonga) submarine volcano had an explosive eruption that thrusted ash, gases, and water vapor through the troposphere into the stratosphere and mesosphere. Previous studies manually tracked the aerosol and trace gas plumes over time across different positions in the southern hemisphere. Using data retrieved from low earth orbiting satellite instruments (e.g., OMPS, OMI, and CALIPSO), this research demonstrates how open-source machine learning (ML) models, like Meta’s Segment Anything Model (SAM), with prompt engineering can perform automatic plume tracking following the Hunga Tonga eruption. This extensible methodology, and modular data processing and modeling pipeline using NASA Earthdata and Openscapes, establishes a framework for systematically and rapidly studying extreme events, including volcanic eruptions and large-scale wildfires. By combining advanced machine learning techniques, such as SAM’s zero-shot learning, with large volumes of remote sensing data, this work demonstrates how AI and open science can accelerate research and generate actionable results. The tools and technologies presented here can help translate earth science to action from NASA’s current and future Earth observing satellite missions (e.g., the Atmosphere Observing System (AOS)), and assist researchers and stakeholders in understanding, mapping, and responding to natural disasters and extreme events in a changing world.

David M. Giles↗

The Sensitivity of SeaWiFS Ocean Color Retrievals to Aerosol Amount and Type

As atmospheric reflectance dominates top-of-the-atmosphere radiance over ocean, atmospheric correction is a critical component of ocean color retrievals. This paper explores the operational Sea-viewing Wide Field-of-View Sensor (SeaWiFS) algorithm atmospheric correction with approximately 13 000 coincident surface-based aerosol measurements. Aerosol optical depth at 440 nm (AOD(sub 440)) is overestimated for AOD below approximately 0.1-0.15 and is increasingly underestimated at higher AOD; also, single-scattering albedo (SSA) appears overestimated when the actual value less than approximately 0.96.AOD(sub 440) and its spectral slope tend to be overestimated preferentially for coarse-mode particles. Sensitivity analysis shows that changes in these factors lead to systematic differences in derived ocean water-leaving reflectance (Rrs) at 440 nm. The standard SeaWiFS algorithm compensates for AOD anomalies in the presence of nonabsorbing, medium-size-dominated aerosols. However, at low AOD and with absorbing aerosols, in situ observations and previous case studies demonstrate that retrieved Rrs is sensitive to spectral AOD and possibly also SSA anomalies. Stratifying the dataset by aerosol-type proxies shows the dependence of the AOD anomaly and resulting Rrs patterns on aerosol type, though the correlation with the SSA anomaly is too subtle to be quantified with these data. Retrieved chlorophyll-a concentrations (Chl) are affected in a complex way by Rrs differences, and these effects occur preferentially at high and low Chl values. Absorbing aerosol effects are likely to be most important over biologically productive waters near coasts and along major aerosol transport pathways. These results suggest that future ocean color spacecraft missions aiming to cover the range of naturally occurring and anthropogenic aerosols, especially at wavelengths shorter than 440 nm, will require better aerosol amount and type constraints.

single scattering albedo↗

Site B - NREL ASSIST (SN11) Thermodynamic Retrievals TROPoe / Derived Data

This dataset contains daily files with thermodynamic profiles retrieved with the optimal estimation physical retrieval TROPoe v0.12 (Turner and Löhnert 2014; Turner and Blumberg 2019; Turner and Löhnert 2021). The profiles are retrieved every 10 minutes from instantaneous observations from the NREL ASSIST-II (SN 11) infrared spectrometer. Observations are noise-filtered but not averaged in time to minimize errors due to non-uniform clouds. Additional input data in TROPoe are cloud base height (CBH), which is a combined data product that uses data from ceilometers at sites A1 and H and scanning lidars from ARM sites C1 and E37. The CBH is weighted inversely proportionally to the distance to the respective site to take into account the spatial variability of clouds (see https://github.com/StefanoWind/ASSIST_analysis/blob/main/awaken_processing/combine_cbh.py). The full pipeline for running the retrieval is available at https://github.com/StefanoWind/TROPoe_processor. Met data was not ingested. In addition to these temporally resolved input data, TROPoe requires an a priori dataset (prior) that provides mean climatological estimates of thermodynamic profiles and specifies how temperature and humidity covary with height as an input (for details see, e.g., Djalalova et al. 2022). The prior is a key component of the retrieval and provides a constraint on the ill-posed inversion problem. A monthly prior was computed from operational radiosonde launches at ARM SGP, OK.

17 WIND ENERGY↗

Site G - NREL ASSIST (SN10) Thermodynamic Retrievals TROPoe / Derived Data

This dataset contains daily files with thermodynamic profiles retrieved with the optimal estimation physical retrieval TROPoe v0.12 (Turner and Löhnert 2014; Turner and Blumberg 2019; Turner and Löhnert 2021). The profiles are retrieved every 10 minutes from instantaneous observations from the NREL ASSIST-II (SN 10) infrared spectrometer. Observations are noise-filtered but not averaged in time to minimize errors due to non-uniform clouds. Additional input data in TROPoe are cloud base height (CBH), which is a combined data product that uses data from ceilometers at sites A1 and H and scanning lidars from ARM sites C1 and E37. The CBH is weighted inversely proportionally to the distance to the respective site to take into account the spatial variability of clouds (see https://github.com/StefanoWind/ASSIST_analysis/blob/main/awaken_processing/combine_cbh.py). The full pipeline for running the retrieval is available at https://github.com/StefanoWind/TROPoe_processor. Met data was not ingested. In addition to these temporally resolved input data, TROPoe requires an a priori dataset (prior) that provides mean climatological estimates of thermodynamic profiles and specifies how temperature and humidity covary with height as an input (for details see, e.g., Djalalova et al. 2022). The prior is a key component of the retrieval and provides a constraint on the ill-posed inversion problem. A monthly prior was computed from operational radiosonde launches at ARM SGP, OK.

17 WIND ENERGY↗

Site C1a - NREL ASSIST (SN12) Thermodynamic Retrievals TROPoe / Derived Data

This dataset contains daily files with thermodynamic profiles retrieved with the optimal estimation physical retrieval TROPoe v0.12 (Turner and Löhnert 2014; Turner and Blumberg 2019; Turner and Löhnert 2021). The profiles are retrieved every 10 minutes from instantaneous observations from the NREL ASSIST-II (SN 12) infrared spectrometer. Observations are noise-filtered but not averaged in time to minimize errors due to non-uniform clouds. Additional input data in TROPoe are cloud base height (CBH), which is a combined data product that uses data from ceilometers at sites A1 and H and scanning lidars from ARM sites C1 and E37. The CBH is weighted inversely proportionally to the distance to the respective site to take into account the spatial variability of clouds (see https://github.com/StefanoWind/ASSIST_analysis/blob/main/awaken_processing/combine_cbh.py). The full pipeline for running the retrieval is available at https://github.com/StefanoWind/TROPoe_processor. Met data was not ingested. In addition to these temporally resolved input data, TROPoe requires an a priori dataset (prior) that provides mean climatological estimates of thermodynamic profiles and specifies how temperature and humidity covary with height as an input (for details see, e.g., Djalalova et al. 2022). The prior is a key component of the retrieval and provides a constraint on the ill-posed inversion problem. A monthly prior was computed from operational radiosonde launches at ARM SGP, OK.

17 WIND ENERGY↗

A New Machine Learning Based Analysis for Improving Satellite Retrieved Atmospheric Composition Data: OMI SO2 as an Example

Despite recent progress, satellite retrievals of anthropogenic SO2 still suffer from relatively low signal-tonoise ratios. In this study, we demonstrate a new machine learning data analysis method to improve the quality of satellite SO2 products. In the absence of large ground-truth datasets for SO2, we start from SO2 slant column densities (SCDs) retrieved from the Ozone Monitoring Instrument (OMI) using a data-driven, physically based algorithm and calculate the ratio between the SCD and the root mean square (rms) of the fitting residuals for each pixel. To build the training data, we select presumably clean pixels with small SCD / rms ratios (SRRs) and set their target SCDs to zero. For polluted pixels with relatively large SRRs, we set the target to the original retrieved SCDs. We then train neural networks (NNs) to reproduce the target SCDs using predictors including SRRs for individual pixels, solar zenith, viewing zenith and phase angles, scene reflectivity, and O3 column amounts, as well as the monthly mean SRRs. For data analysis, we employ two NNs: (1) one trained daily to produce analyzed SO2 SCDs for polluted pixels each day and (2) the other trained once every month to produce analyzed SCDs for less polluted pixels for the entire month. Test results for 2005 show that our method can significantly reduce noise and artifacts over background regions. Over polluted areas, the monthly mean NN-analyzed and original SCDs generally agree to within ±15 %, indicating that our method can retain SO2 signals in the original retrievals except for large volcanic eruptions. This is further confirmed by running both the NN-analyzed and original SCDs through a topdown emission algorithm to estimate the annual SO2 emissions for ∼ 500 anthropogenic sources, with the two datasets yielding similar results. We also explore two alternative approaches to the NN-based analysis method. In one, we employ a simple linear interpolation model to analyze the original SCD retrievals. In the other, we develop a PCA–NN algorithm that uses OMI measured radiances, transformed and dimension-reduced with a principal component analysis (PCA) technique, as inputs to NNs for SO2 SCD retrievals. While the linear model and the PCA–NN algorithm can reduce retrieval noise, they both underestimate SO2 over polluted areas. Overall, the results presented here demonstrate that our new data analysis method can significantly improve the quality of existing OMI SO2 retrievals. The method can potentially be adapted for other sensors and/or species and enhance the value of satellite data in air quality research and applications.

Can Li↗

An iterative bidirectional gradient boosting approach for CVR baseline estimation

Here this paper presents a novel Iterative Bidirectional Gradient Boosting Model (IBi-GBM) for estimating the baseline of Conservation Voltage Reduction (CVR) programs. In contrast to many existing methods, we treat CVR baseline estimation as a missing data retrieval problem. The approach involves dividing the load and its corresponding temperature profiles into three periods: pre-CVR, CVR, and post-CVR. To restore the missing load profile during the CVR period, the method employs a three-step process. First, a forward-pass GBM is executed using data from the pre-CVR period as inputs. Subsequently, a backward-pass GBM is applied using data from the post-CVR period. The two restored load profiles are reconciled, considering pre-calculated weights derived from forecasting accuracy, and only the leftmost and rightmost points are retained. The newly restored points are then included as inputs for the subsequent iteration. This iterative procedure continues until the original load data in the CVR period is fully restored. We develop IBi-GBM using actual smart meter and Supervisory Control and Data Acquisition (SCADA) data. Our results demonstrate that IBi-GBM exhibits robust performance across various data resolutions and in different seasons and outperforms existing methods by achieving a 1-2% reduction in normalized Root Mean Square Error (nRMSE).

42 ENGINEERING↗

Controlled Vocabularies Boost International Participation and Normalization of Searches

The Global Change Master Directory's (GCMD) science staff set out to document Earth science data and provide a mechanism for it's discovery in fulfillment of a commitment to NASA's Earth Science progam and to the Committee on Earth Observation Satellites' (CEOS) International Directory Network (IDN.) At the time, whether to offer a controlled vocabulary search or a free-text search was resolved with a decision to support both. The feedback from the user community indicated that being asked to independently determine the appropriate 'English" words through a free-text search would be very difficult. The preference was to be 'prompted' for relevant keywords through the use of a hierarchy of well-designed science keywords. The controlled keywords serve to 'normalize' the search through knowledgeable input by metadata providers. Earth science keyword taxonomies were developed, rules for additions, deletions, and modifications were created. Secondary sets of controlled vocabularies for related descriptors such as projects, data centers, instruments, platforms, related data set link types, and locations, along with free-text searches assist users in further refining their search results. Through this robust 'search and refine' capability in the GCMD users are directed to the data and services they seek. The next step in guiding users more directly to the resources they desire is to build a 'reasoning' capability for search through the use of ontologies. Incorporating twelve sets of Earth science keyword taxonomies has boosted the GCMD S ability to help users define and more directly retrieve data of choice.

Olsen, Lola M.↗

Seasat-A oceanographic data system and users

The Seasat-A system is reviewed with emphasis on data retrieval, processing and dissemination plans. Attention is paid to the sensors of the Seasat satellite including the compressed pulse radar altimeter, the coherent synthetic aperture imaging radar, the microwave wind scatterometer, and the scanning visible/infrared radiometer. Particular emphasis is placed on a particular set of experiments: the Navy's Fleet Numerical Weather Control will receive Seasat data in real time and process the data into products that will be used in weather and sea condition forecasts.

Mccandless, S. W., Jr.↗

The Network Information Management System (NIMS) in the Deep Space Network

In an effort to better manage enormous amounts of administrative, engineering, and management data that is distributed worldwide, a study was conducted which identified the need for a network support system. The Network Information Management System (NIMS) will provide the Deep Space Network with the tools to provide an easily accessible source of valid information to support management activities and provide a more cost-effective method of acquiring, maintaining, and retrieval data.

Wales, K. J.↗