Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data Science”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 667 records · Page 37

User Scientific Data Systems: Experience Report

This paper presents an abbreviated history of NASA science data management system development over the past ten years by selecting two case studies, each representative of a distinct era of science data management systems.

Scientific Data↗

The 1990 annual statistics and highlights report

The National Space Science Data Center (NSSDC) has archived over 6 terabytes of space and Earth science data accumulated over nearly 25 years. It now expects these holdings to nearly double every two years. The science user community needs rapid access to this archival data and information about data. The NSSDC has been set on course to provide just that. Five years ago the NSSDC came on line, becoming easily reachable for thousands of scientists around the world through electronic networks it managed and other international electronic networks to which it connected. Since that time, the data center has developed and implemented over 15 interactive systems, operational nearly 24 hours per day, and is reachable through DECnet, TCP/IP, X25, and BITnet communication protocols. The NSSDC is a clearinghouse for the science user to find data needed through the Master Directory system whether it is at the NSSDC or deposited in over 50 other archives and data management facilities around the world. Over 13,000 users accessed the NSSDC electronic systems, during the past year. Thousands of requests for data have been satisfied, resulting in the NSSDC's sending out a volume of data last year that nearly exceeded a quarter of its holdings. This document reports on some of the highlights and distribution statistics for most of the basic NSSDC operational services for fiscal year 1990. It is intended to be the first of a series of annual reports on how well NSSDC is doing in supporting the space and Earth science user communities.

Green, James L.↗

Introduction to Air Traffic Management

The presentation introduces students and faculty to air traffic management with focus on air traffic data for data-science. Starting with the common attributes of transportation systems — highway transportation, air transportation and data transportation, the initial set of slides discuss the purpose of data-science in air traffic management, reasons why air traffic management is challenging, and the multidisciplinary nature of air traffic management research. The history of flight from 1903 — Wright Flyer — to 1987 — formation of the National Air Traffic Controllers Association — is briefly discussed. The national airspace system is described in terms of airports in the U. S., air traffic control facilities (flight service stations, terminal, enroute and system command center), airspace geometry (sectors, airways and navaids), governing regulations and directives, airspace classification (Class A through G), special use airspace, visual flight rules and instrument flight rules. The contents of a flight-plan are described. Weather briefing is discussed. The surveillance equipment used for surface, terminal area and enroute are described, and the aircraft states obtained using the surveillance data are listed. Airline operations control functions — schedule development, flight planning, resource scheduling and flight following — are noted. Next, the roles and responsibilities of air traffic controllers and traffic flow managers are discussed. Separation standards and conflict resolution techniques are outlined. Finally, traffic flow management techniques are reviewed with an illustrative example.

Air Traffic Management↗

Introduction to Air Traffic Management

The presentation introduces students and faculty to air traffic management with focus on air traffic data for data-science. Starting with the common attributes of transportation systems — highway transportation, air transportation and data transportation, the initial set of slides discuss the purpose of data-science in air traffic management, reasons why air traffic management is challenging, and the multidisciplinary nature of air traffic management research. The history of flight from 1903 — Wright Flyer — to 1987 — formation of the National Air Traffic Controllers Association — is briefly discussed. The national airspace system is described in terms of airports in the U. S., air traffic control facilities (flight service stations, terminal, enroute and system command center), airspace geometry (sectors, airways and navaids), governing regulations and directives, airspace classification (Class A through G), special use airspace, visual flight rules and instrument flight rules. The contents of a flight-plan are described. Weather briefing is discussed. The surveillance equipment used for surface, terminal area and enroute are described, and the aircraft states obtained using the surveillance data are listed. Airline operations control functions — schedule development, flight planning, resource scheduling and flight following — are noted. Next, the roles and responsibilities of air traffic controllers and traffic flow managers are discussed. Separation standards and conflict resolution techniques are outlined. Finally, traffic flow management techniques are reviewed with an illustrative example.

Air Traffic Management↗

Search Enhancements using Natural Language Processing Techniques

NASA Goddard Earth Sciences Data and Information Services Center (GESDISC) is one of the 12 NASA Science Mission Directorate Data Centers. The main goal of GESDISC is to provide earth science data, information, and services to the earth science data community. Consequently, data discovery is at the center of our mission and our search engine is the primary tool for our users to interact, find, and access our data. Existing search approaches are largely focused on hard-matching of keywords in the search query with dataset metadata. Here we propose to expand the search by introducing a complementary natural language processing (NLP) search. At the heart of our proposed NLP search, we trained a joint embedding using scientific text corpus and a curated set of dataset metadata. The embedding learns the association between words in our dataset metadata and those of the scientific text corpus. This enables us to go beyond simple hard-matching of a query and data set metadata and have a notion of “similarity” between the search query and the datasets. We further integrated our NLP search into the Elastic Search (ES) framework leveraging similarity search capabilities offered through the “dense_vector” field type. Our preliminary evaluations show that our proposed NLP search has the potential to be utilized to complement the existing search engine and serve as a base for a dataset recommendation system.

Armin Mehrabian↗

Processing Spacecraft Data Without Confusion

Producing multiple versions of the same data product for the same time frame with the same remotely sensed inputs can be a recipe for disaster. Yet, amidst the commotion of satellite launch and early operations (LEO), such data processing is needed. After LEO, the situation gets worse. Processing newly arriving data ("forward processing") is augmented with reprocessing and algorithm development, comparison, evaluation, and testing -- often happening all at the same time. The problem can be analyzed in three main parts -- maintaining multiple versions of algorithms and data so that end-product users are not overwhelmed,allocating computer resources efficiently, and simplifying production operations so that va st amounts of data can be processed with minimal staff and fewer errors. OMIDAPS provides a framework for execution of algorithms that transform lower level data acquired by OMI on NASA's Aura satellite into higher level science data products. In contrast to traditional science data processing systems, we address all parts of the problem with an innovative approach allowing multiple data processing to run within a single physical system. The data products, imports, exports, and execution planning are all segregated into distinct "ArchiveSets." This paper describes reasons for multiple concurrent productions on a typical satellite data processing project using OMI as an example. It describes the virtual data processing system concept and its advantages over separate physical processing strings. It explores the specific implementation of the virtual systems within OMIDAPS and discusses some of the implications of our approach and describes how virtual processing is used to accomplish the overall mission of OMI data processing.

Tilmes, Curt↗

Soil Moisture Active Passive Mission L4_SM Data Product Assessment (Version 2 Validated Release)

During the post-launch SMAP calibration and validation (Cal/Val) phase there are two objectives for each science data product team: 1) calibrate, verify, and improve the performance of the science algorithm, and 2) validate the accuracy of the science data product as specified in the science requirements and according to the Cal/Val schedule. This report provides an assessment of the SMAP Level 4 Surface and Root Zone Soil Moisture Passive (L4_SM) product specifically for the product's public Version 2 validated release scheduled for 29 April 2016. The assessment of the Version 2 L4_SM data product includes comparisons of SMAP L4_SM soil moisture estimates with in situ soil moisture observations from core validation sites and sparse networks. The assessment further includes a global evaluation of the internal diagnostics from the ensemble-based data assimilation system that is used to generate the L4_SM product. This evaluation focuses on the statistics of the observation-minus-forecast (O-F) residuals and the analysis increments. Together, the core validation site comparisons and the statistics of the assimilation diagnostics are considered primary validation methodologies for the L4_SM product. Comparisons against in situ measurements from regional-scale sparse networks are considered a secondary validation methodology because such in situ measurements are subject to up-scaling errors from the point-scale to the grid cell scale of the data product. Based on the limited set of core validation sites, the wide geographic range of the sparse network sites, and the global assessment of the assimilation diagnostics, the assessment presented here meets the criteria established by the Committee on Earth Observing Satellites for Stage 2 validation and supports the validated release of the data. An analysis of the time average surface and root zone soil moisture shows that the global pattern of arid and humid regions are captured by the L4_SM estimates. Results from the core validation site comparisons indicate that "Version 2" of the L4_SM data product meets the self-imposed L4_SM accuracy requirement, which is formulated in terms of the ubRMSE: the RMSE (Root Mean Square Error) after removal of the long-term mean difference. The overall ubRMSE of the 3-hourly L4_SM surface soil moisture at the 9 km scale is 0.035 cubic meters per cubic meter requirement. The corresponding ubRMSE for L4_SM root zone soil moisture is 0.024 cubic meters per cubic meter requirement. Both of these metrics are comfortably below the 0.04 cubic meters per cubic meter requirement. The L4_SM estimates are an improvement over estimates from a model-only SMAP Nature Run version 4 (NRv4), which demonstrates the beneficial impact of the SMAP brightness temperature data. L4_SM surface soil moisture estimates are consistently more skillful than NRv4 estimates, although not by a statistically significant margin. The lack of statistical significance is not surprising given the limited data record available to date. Root zone soil moisture estimates from L4_SM and NRv4 have similar skill. Results from comparisons of the L4_SM product to in situ measurements from nearly 400 sparse network sites corroborate the core validation site results. The instantaneous soil moisture and soil temperature analysis increments are within a reasonable range and result in spatially smooth soil moisture analyses. The O-F residuals exhibit only small biases on the order of 1-3 degrees Kelvin between the (re-scaled) SMAP brightness temperature observations and the L4_SM model forecast, which indicates that the assimilation system is largely unbiased. The spatially averaged time series standard deviation of the O-F residuals is 5.9 degrees Kelvin, which reduces to 4.0 degrees Kelvin for the observation-minus-analysis (O-A) residuals, reflecting the impact of the SMAP observations on the L4_SM system. Averaged globally, the time series standard deviation of the normalized O-F residuals is close to unity, which would suggest that the magnitude of the modeled errors approximately reflects that of the actual errors. The assessment report also notes several limitations of the "Version 2" L4_SM data product and science algorithm calibration that will be addressed in future releases. Regionally, the time series standard deviation of the normalized O-F residuals deviates considerably from unity, which indicates that the L4_SM assimilation algorithm either over- or under-estimates the actual errors that are present in the system. Planned improvements include revised land model parameters, revised error parameters for the land model and the assimilated SMAP observations, and revised surface meteorological forcing data for the operational period and underlying climatological data. Moreover, a refined analysis of the impact of SMAP observations will be facilitated by the construction of additional variants of the model-only reference data. Nevertheless, the “Version 2” validated release of the L4_SM product is sufficiently mature and of adequate quality for distribution to and use by the larger science and application communities.

SMAP L4_SM↗

AI Curation Methods for NASA Scientific Data

The NASA Open Science Data Repository (OSDR) serves as a central hub for sharing and accessing NASA's vast collection of scientific data, supporting researchers across diverse fields. To enhance the efficiency, accuracy, and accessibility of this data, we are leveraging advanced artificial intelligence (AI) techniques as part of the AI for Curation project. By integrating large language models (LLMs) into our data curation workflow, we aim to streamline the entire process—from data submission to user interaction. This initiative focuses on improving key areas, including data ingestion, curation, and user engagement with curated datasets, impacting multiple domains and a wide user base. First, we are developing tools that can automatically parse data in various formats, using LLMs to convert unstructured data into structured, standardized formats. This reduces the manual effort required for curation, allowing curators to focus on more critical scientific analyses. Additionally, AI and machine learning (ML) models are being implemented to automate data validation and verification, ensuring the highest standards of data quality and reliability. Finally, we are creating a conversational AI agent to interact with the curated scientific studies in OSDR, helping users easily navigate the repository and access relevant data. By enhancing data discoverability and accessibility, these advancements will foster new research opportunities and promote the principles of open science.

Walter Alvarado↗

Enabling Space Biological Knowledge Discovery Through Image and Video Data Sharing

Increased biomedical risks associated with deep space crewed missions (cis-Lunar, Mars transit/surface) require development of health countermeasures, novel ecosystem support, risk modeling, and fundamental space biological knowledge discovery. Molecular-omics, physiological-phenotypic-behavioral, and environmental-radiation telemetry data from space biological and health studies are needed for reuse by scientists to address these tasks. The data as well as space-relevant biospecimens are being made more findable, accessible, interoperable, and reusable through NASA’s Open Science Data Repository (OSDR). This new OSDR umbrella grouping includes NASA GeneLab, the NASA Ames Life Sciences Data Archive (ALSDA), and the NASA Biological Institutional Scientific Collection. The OSDR system design appropriately handles metadata and processed-tabular results from ALSDA studies collected from space experiments. But raw and processed ALSDA bioimage and video datasets require an expansion of OSDR’s data architecture to handle ingestion, curation, and egress. The academic-industry bioimaging field saw a scientific renaissance in the past several years through leveraging open-source software, international collaborations, machine learning, and other open science/programming approaches. As crewed missions and more biological experiments are on the deep space horizon, OSDR is embracing data stewardship through listening to feedback from subject matter experts and designing an expanded architecture which is appropriate for NASA’s goals to enable analysis and reuse of bioimaging and video data for the public science community.Discovery Through Image and Video Data Sharing

space biology↗

PACE Technical Report Series, Volume 6: Data Product Requirements and Error Budgets Consensus Document

This chapter summarizes ocean color science data product requirements for the Plankton, Aerosol, Cloud,ocean Ecosystem (PACE) mission's Ocean Color Instrument (OCI) and observatory. NASA HQ delivered Level-1 science data product requirements to the PACE Project, which encompass data products to be produced and their associated uncertainties. These products and uncertainties ultimately determine the spectral nature of OCI and the performance requirements assigned to OCI and the observatory. This chapter ultimately serves to provide context for the remainder of this volume, which describes tools developed that allocate these uncertainties into their components, including allowable OCI systematic and random uncertainties, observatory geo location uncertainties, and geophysical model uncertainties.

Cetinic, Ivona↗

Developing Fluorescence-Based Sensors to Support Rare Earth Element Separation

Rare earth elements (REEs) are essential to most renewable energy technologies. Unfortunately, as we transition to sustainable energy production, the demand for REEs is rapidly growing well beyond current rates of production. As a result, novel means of efficient, scalable, and easily adaptable methods for processing primary and recycle feedstocks are needed. Development and integration of sensors for highly selective in-line monitoring can support more efficient design and testing of such novel separation processes, as well as more cost-effective deployment of those separation flowsheets. Work here will explore the application of fluorescence spectroscopy, a highly sensitive and selective technique, to quantify multiple lanthanides in complex mixtures including known interferents or quenching agents. Results include identification of the optimal excitation wavelength and the limit of detection of various rare earth elements as well as the performance of data-science-based quantification approaches in streams where “unknowns” are present. Overall, the data science tools in conjunction with optical sensor data were able to quantify analytes in the presence of other lanthanides which can be anticipated in the actual industrial stream. Here we include characterization of lanthanides in a microfluidic device similar to those used in new process development. This study demonstrates the capability of utilizing fluorescence spectroscopy to quantify analytes in a complicated solution matrix, suggesting this is a successful approach for in-line monitoring to optimize the separation efficiency in an industrial stream.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

NASA's EOSDIS, Trust and Certification

NASA's Earth Observing System Data and Information System (EOSDIS) has been in operation since August 1994, managing most of NASA's Earth science data from satellites, airborne sensors, filed campaigns and other activities. Having been designated by the Federal Government as a project responsible for production, archiving and distribution of these data through its Distributed Active Archive Centers (DAACs), the Earth Science Data and Information System Project (ESDIS) is responsible for EOSDIS, and is legally bound by the Office of Management and Budgets circular A-130, the Federal Records Act. It must follow the regulations of the National Institute of Standards and Technologies (NIST) and National Archive and Records Administration (NARA). It must also follow the NASA Procedural Requirement 7120.5 (NASA Space Flight Program and Project Management). All these ensure that the data centers managed by ESDIS are trustworthy from the point of view of efficient and effective operations as well as preservation of valuable data from NASA's missions. Additional factors contributing to this trust are an extensive set of internal and external reviews throughout the history of EOSDIS starting in the early 1990s. Many of these reviews have involved external groups of scientific and technological experts. Also, independent annual surveys of user satisfaction that measure and publish the American Customer Satisfaction Index (ACSI), where EOSDIS has scored consistently high marks since 2004, provide an additional measure of trustworthiness. In addition, through an effort initiated in 2012 at the request of NASA HQ, the ESDIS Project and 10 of 12 DAACs have been certified by the International Council for Science (ICSU) World Data System (WDS) and are members of the ICSUWDS. This presentation addresses questions such as pros and cons of the certification process, key outcomes and next steps regarding certification. Recently, the ICSUWDS and Data Seal of Approval (DSA) organizations merged their Core Trustworthy Data Repositories Requirements and require that members be recertified every three years. Given the rigor with which NASA manages the ESDIS Project and the DAACs, the recertification through WDSDSA, while involving some additional work, is a relatively simple process.

Certification↗

Earth Observing Data System Data and Information System (EOSDIS) Overview

The National Aeronautics and Space Administration (NASA) acquires and distributes an abundance of Earth science data on a daily basis to a diverse user community worldwide. The NASA Big Earth Data Initiative (BEDI) is an effort to make the acquired science data more discoverable, accessible, and usable. This presentation will provide a brief introduction to the Earth Observing System Data and Information System (EOSDIS) project and the nature of advances that have been made by BEDI to other Federal Users.

Earth Science↗

Planetary Data Workshop, Part 1

The community of planetary scientists addresses two general problems regarding planetary science data: (1) important data sets are being permanently lost; and (2) utilization is constrainted by difficulties in locating and accessing science data and supporting information necessary for its use. A means to correct the problems, provide science and functional requirements for a systematic and phased approach, and suggest technologies and standards appropriate to the solution were explored.

Source record↗

Limiting Data Friction by Reducing Data Download Using Spatiotemporally Aligned Data Organization Through STARE

Current data processing practice limits the volume and variety of relevant geoscience data that can practically be applied to important problems. File archives in centralized data centers are the principal means by which Earth Science data are accessed. This approach, however, requires laborious search, retrieval, and eventual customization/adaptation for the data to be used. Such fractionation makes it even more difficult to share outcomes, i.e. research artifacts and data products, hampering reusability and repeatability, since end users generally have their own research agenda and preferences as well as scarce resources. Thus, while finding and downloading data files from central data centers are already costly for end users working in their own field, using data products from other disciplines rapidly becomes prohibitive. This curtails scientific productivity, limits avenues of study, and endangers quality and reproducibility. The Spatio-Temporal Adaptive Resolution Encoding (STARE) is a unifying scheme that facilitates the indexing, access, and fusion of diverse Earth Science data. STARE implements an innovative encoding of geo-spatiotemporal information, originally developed for aligning datasets with diverse spatiotemporal characteristics in an array database. The spatial component of STARE recursively quadfurcates a root polyhedron, producing a hierarchical scheme for addressing geographic locations and regions. The temporal component of STARE uses conventional date-time units as an indexing hierarchy. The additional encoding of spatial and temporal resolution information in STARE enables comparisons and conditional selections across diverse datasets. Moreover, spatiotemporal set-operations, e.g. union and intersection, are mapped to efficient integer operations with STARE. Applied to existing data models (point, grid, spacecraft swath) and corresponding granules, STARE indexes provide a streamlined description usable as geo-spatiotemporal metadata. When coupled with large scale, distributed hardware and software, STARE-based data access reduces pre-analysis data preparation costs by offering a convenient means to align different datasets spatiotemporally without specialized effort in parallel computing or distributed data management.

Kuo, Kwo-Sen↗

Data Summarization and Inference at Scale

This is the final report for the DOE ASCR grant SC-0022260, Data Summarization and Inference at Scale, PI: Alex Pothen, Purdue University. The goal of the project was to solve data-intensive and compute-intensive problems in the physical sciences, engineering, information science, data science, etc. by designing and implementing new algorithms that could work with a subset of the data. The four subgoals were: (a) The solution of problems where the data is too large to be stored in the memory of a computer. In this streaming model of computation, the data arrives as a stream of elements to the computer, each element is processed as it arrives, and a decision is made to discard the data or to store it; only a small subset of the data proportional to the size of the output solution is stored, and when all the data has been streamed, a solution to the problem is computed from the stored subset. (b) The use of machine learning methods to compute solutions to data-intensive problems. The use of GPUs is critical to obtain high performance on machine learning tasks, but their memory sizes are smaller relative to that of CPUs. For large-scale problems, the data is sampled many times, and small samples are used with repetition, for robustness, to compute solutions to inference tasks. This sampling reduces the memory required to solve the problem, but attention is needed to avoid slow convergence to the solutions, and reduced accuracy of inference. We propose submodular optimization, Large Language Models, and physics-informed neural networks to enable GPU computations here. (c) Modeling and visualization of high-dimensional data using interpretable features. Clinical proteomic data sets from immunology for the detection of cancer and other diseases are temporal and high-dimensional, and algorithms for visualizing these data sets using clinically interpretable features are lacking. We propose methods that compute distances based on the optimal transportation problem and graph edit distances to address this problem. We also propose the use of optimal transport-based distances, spatial statistics, and network structure to classify image data sets, We apply these algorithms to electron micrographs of the peripheral nervous system in the digestive tract. (d) The design of data-intensive algorithms on emerging architectures, specifically, noisy, intermediate-scale quantum (NISQ) devices. Quantum computers offer the possibility of exploring large solution spaces due to the principle of superposition, but current quantum computers are limited by few qubits, short coherence times due to noise, poor interconections among the qubits, etc. We propose the use of the divide and conquer paradigm to solve large-scale problems, wherein collections of small subproblems are solved on the quantum devices, and the solutions to the subproblems are integrated into a solution for the original problem on a classical computer.

97 MATHEMATICS AND COMPUTING↗