Search NASASearch

SEARCH · Search NASA

Results for “data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Do Citizen Science Intense Observation Periods Increase Data Usability? A Deep Dive of the NASA GLOBE Clouds Data Set With Satellite Comparisons

The Global Learning and Observations to Benefit the Environment (GLOBE) citizen science program has recently conducted a series of month-long intensive observation periods (IOPs), asking the public to submit daily reports on cloud and sky conditions from all regions of Earth. This provides a wealth of crowdsourced observations from the ground, which complements other conventional scientific cloud data. In addition, the GLOBE reports are matched in space and time with geostationary and low Earth orbit satellites, which allows for a straightforward comparison of cloud properties, and minimizes the biases associated with mismatched sampling between participants and satellites. The matched GLOBE dataset is used to calculate the mean observed cloud cover by atmospheric level both worldwide and by region. The overall magnitudes of cloud cover between the GLOBE participants and the matched satellites agree within 10%, which is notable given the distinctly different natures of the data sources. The mean vertical cloud profiles show GLOBE reporting more low-level clouds and fewer high-level clouds than satellites. The low cloud disagreement is likely related to satellites missing low clouds when high clouds block their view. Conversely, the high cloud disagreement is related primarily to cloud opacity, as satellites may miss some optically thin clouds. Monte Carlo testing shows the results to be robust, and the tripled amount of IOP data reduces uncertainty by half. These findings also highlight ways in which citizen science IOP data may be used to support scientific research while accounting for their unique properties. Plain Language Summary: Citizen science is becoming an increasingly prominent aspect of scientific research, and so it important to study how citizen science data can be used effectively. For example, The GLOBE Program has recently conducted a series of special data-collecting events, or “challenges”, which gathered large numbers of reports on cloud and sky conditions. Because NASA GLOBE Clouds matches the participant reports with cloud observations from satellites, we can use these data to get a combined view of clouds from above and below. When looking at the average cloud cover for different atmospheric levels across Earth, we find that the GLOBE participants and the satellites agree quite closely. This is a surprising and fascinating find, given how different in nature volunteer ground reports are to satellite measurements. However, there are some small but notable disagreements between GLOBE participants and satellites about the distribution of cloud cover at different levels. In addition, by testing the data for uncertainty, we show that the results from the GLOBE data are reliable, and that more public participation improves the reliability. So, by carefully designing the analysis methodology, and by testing for the uncertainty of the data, citizen science can make a meaningful contribution to scientific research.

J. Brant Dodson

FAIRLinked: Data FAIRification Tools for Materials Data Science

FAIRLinked is a software package created to support the FAIRification of materials science data, ensuring proper alignment with FAIR principles: Findable, Accessible, Interoperable, and Reusable. It is built to be compatible with MDS-Onto, an ontology designed to capture the semantics of various types of materials data, enabling integration and sharing across different research workflows. The package is subdivided into three subpackages: InterfaceMDS, RDFTableConversion, and QBWorkflow. The first subpackage, InterfaceMDS allows users to search for terms using either string search or various filters, explore different domains and subdomains, and add terms to MDS-Onto. RDFTableConversion is used for serialization and deserialization of data from CSV into JSONLDs and vice versa in a way that captures the semantics of the data using MDS-Onto. Lastly, QBWorkflow is a serialization and deserialization workflow that incorporates RDF Data Cube vocabulary, useful for working with multidimensional datasets. By offering these packages, FAIRLinked lowers the barrier of creating FAIR, machine-actionable data for researchers in the materials science community.

FAIR

Re-Organizing Earth Observation Data Storage to Support Temporal Analysis of Big Data

The Earth Observing System Data and Information System archives many datasets that are critical to understanding long-term variations in Earth science properties. Thus, some of these are large, multi-decadal datasets. Yet the challenge in long time series analysis comes less from the sheer volume than the data organization, which is typically one (or a small number of) time steps per file. The overhead of opening and inventorying complex, API-driven data formats such as Hierarchical Data Format introduces a small latency at each time step, which nonetheless adds up for datasets with O(10^6) single-timestep files. Several approaches to reorganizing the data can mitigate this overhead by an order of magnitude: pre-aggregating data along the time axis (time-chunking); storing the data in a highly distributed file system; or storing data in distributed columnar databases. Storing a second copy of the data incurs extra costs, so some selection criteria must be employed, which would be driven by expected or actual usage by the end user community, balanced against the extra cost.

data storage

The NASA Open Science Data Repository: Biomedical Data, Analysis Tools, and Informatic Collaborations

Increased biomedical risks and challenges associated with deep space missions require knowledge discovery, health countermeasures, and biomedical support capabilities. Maximally open-access and reusable data is needed by developers, scientists, and engineers to develop these systems. The NASA Open Science Data Repository (OSDR) is a maximally open access and FAIR database (ie., findable, accessible, interoperable, and reusable), and meets various scientific, technical, and operational needs. It offers users and submitters the ability to upload, download, search, share, analyze, cite, and visualize data across ‘omics, physiological, phenotypic, payload, hardware, behavioral, bioimaging, video, and environmental monitoring telemetry datasets. OSDR is an expanded database, based upon the successes of NASA GeneLab. OSDR has >460 studies with datasets covering model organisms to non-NASA human astronauts. There are ~12 datasets from the Inspiration 4 (I4) mission, spanning metagenomics, comprehensive metabolic panels, clonal hematopoiesis, spatial transcriptomics, proteomics, and cytokine panels. In the interest of data privacy, two I4 datasets with raw files relating to the epitranscriptome, and a new request feature is live in OSDR (with a backend review process established) developed from industry norms. OSDR is collecting and curating biomedical human data from a new sub-orbital research flight and is open to more space life science/biomedical submissions from the international and commercial sectors. OSDR also recently began a collaboration with the European Space Agency (ESA) to collect and curate >200 terabytes of human and model organism data. The OSDR submission portal is designed to ingest and curate ~25 ‘omics and ~50 physiological-phenotypic-imaging assay data types. Tools available for OSDR users include: 1) an Environmental Data Application to compare radiation, CO2, relative humidity, temperature, and other telemetry across missions and subjects, 2) the RadLab database, a collaboration between NASA, ESA, the German and Italian Space Agencies, and the Bulgarian Academy of Sciences, and 3) a Multi-study visualization tool which enables users to look across and combine ‘omics datasets. There are ~600 volunteer OSDR Analysis Working Group (AWG) members providing feedback on scientific data/metadata standards and collaborating to mine-reuse OSDR in research. OSDR/GeneLab has enabled ~60 publications reusing data as of October 2023.

space biology

Data Efficiency Assessment of Generative Adversarial Networks for Critical Heat Flux Synthetic Data Generation

This study investigates the application of generative artificial intelligence techniques, particularly conditional generative adversarial networks (cGAN), in real-world engineering contexts, with a specific focus on synthetic data generation for critical heat flux (CHF). Utilizing a dataset comprising more than 20,000 real experimental CHF measurements, we conduct a series of experiments to examine cGAN’s behavior. These experiments encompass varying sizes of the training dataset, training cGAN on data from diverse experimental sources to generate new data on unseen experimental setups, and assessing the impact of excluding various input features on cGAN’s data generation accuracy. Our findings underscore the pronounced data dependency of cGAN for reliable performance, with decreased efficacy observed with smaller training dataset sizes. Notably, cGAN exhibits varying performance when trained on data from different experiments, with superior predictive capabilities observed for certain experiment sources compared to others. For instance, when cGAN was trained on data from Smolin et al.’s experiments or Zenkevich et al., it exhibited relatively good performance in generating the data from Becker et al., Kirillov et al., and Alekseev et al. experiments. In contrast, when trained with Alekseev et al.’s data and tasked with generating other experimental setups, cGAN showed notably poor performance. In both scenarios, cGAN’s performance was inferior compared to training on samples from all experiments concurrently. A feature importance analysis highlights the significant influence of parameters such as mass flux and heated length on accurate CHF generation, while other parameters like diameter and pressure have less impact. Inlet temperature is identified as a moderating factor by cGAN.

22 - GENERAL STUDIES OF NUCLEAR REACTORS

Trust Not Verify? The Critical Need for Data Curation Standards in Materials Informatics

The importance of data curation has been recognized in multiple areas of research; however, the discussion of this important issue is only beginning to emerge in materials science. In this Perspective, we highlight the benefits of using the standardized data curation protocols in materials science and discuss current gaps in accurate and reproducible data reporting using case studies drawn from high-impact materials science papers and well-known databases such as the Crystallography Open Database (COD) and the Cambridge Structural Database (CSD). We argue that both experimental and computational materials scientists need to embrace a culture of rigorous data curation as part of modern research data management. We propose a sample data curation pipeline for materials chemistry and illustrate its use by creating two new materials chemistry databases. Here, we hope that this perspective will serve to catalyze further discussion and promote the continuous development of rigorous data curation practices within the materials science research community. We posit that adherence to best practices of data curation will promote and enhance the reliability, reproducibility, and integrity of materials research and enable the development of reliable AI and machine learning models that critically depend on the use of quality data.

Chemical structure

Fine-Root Ecology Database (FRED): A Global Collection of Root Trait Data with Coincident Site, Vegetation, Edaphic, and Climatic Data, Version 4.

To address the need for a centralized root trait database, we compiled the Fine-Root Ecology Database (FRED) from published and unpublished data sources. We have continued to add to the FRED database since the release of FRED 1.0 in 2017, followed by 2.0 in 2018, and 3.0 in 2021. This new release of FRED 4.0 now has 213,941 observations of 238 root traits, for a combined total of roughly 3.4 million data fields for root traits and ancillary data together. FRED 4.0 has 39.8% more root trait observations than FRED 3.0 and a 34.4% increase in unique data sources. This release of FRED 4.0 also includes significant increases in geographic regions that have long been underrepresented in global datasets, notably in the tropical low latitudes. Ancillary data on associated site, vegetation, edaphic, and climatic conditions from across the globe have also increased concurrently with root trait observations. FRED is focused on fine roots (traditionally defined as roots less than 2 mm in diameter), as coarse roots are studied using different methodology, often at very different scales, and have different traits and trait interpretations. Despite this fine-root focus, FRED accepts data collected from roots of all sizes and contains observations of many root classes including coarse roots. Data collection will continue for the foreseeable future. The FRED4_Entire_Database_2026.csv file is the flat csv data file for FRED 4.0, and the FRED4_dd.csv file is the data dictionary of all columns available in FRED, including column IDs, column names, definitions, and unit (where applicable).

54 ENVIRONMENTAL SCIENCES

Study of data collection platform concepts: Data collection system user requirements

The overall purpose of the survey was to provide real world data on user requirements. The intent was to assess data collection system user requirements by questioning actual potential users rather than speculating on requirements. The end results of the survey are baseline requirements models for both a data collection platform and a data collection system. These models were derived from the survey results. The real value of these models lies in the fact that they are based on actual user requirements as delineated in the survey questionnaires. Some users desire data collection platforms of small size and light weight. These sizes and weights are beyond the present state of the art. Also, the survey provided a wealth of information on the nature and constituency of the data collection user community as well as information on user applications for data collection systems. Finally, the data sheds light on the generalized platform concept. That is, the diversity of user requirements shown in the data indicates the difficulty that can be anticipated in attempting to implement such a concept.

Source record

The use of LANDSAT-4 MSS digital data in temporal data sets and the evaluation of scene-to-scene registration accuracy

The MSS sensor on LANDSAT 4 is, in certain performance aspects, different from those on LANDSATS 1 through 3. These differences created some concern in the NASA research community as to whether individual data sets can be registered accurately enough to produce acceptable data sets for multitemporal data analysis. The use of LANDSAT 4 MSS digital data in temporal data sets is examined and a method is presented for estimating temporal registration accuracy based on the use of an X-Y digitizer and grey tone electrostatic plots. Results indicate that the RMS temporal registration errors are not significantly different from the temporal data sets generated using LANDSAT 4 and LANDSAT 2 data (33.35 meters) and the temporal data set constructed from two LANDSAT 2 data sets (33.61 meters). A derivation of the model used to evaluate the temporal registration is included.

Anderson, J. E.

Shuttle Imaging Radar-A (SIR-A) data as a complement to Landsat Multispectral Scanner (MSS) data

Principal components analysis and supervised classifications were performed on two dates of Landsat multispectral scanner (MSS) data registered to one date of Shuttle Imaging Radar-A (SIR-A) data in a wheat-growing area of New South Wales, Australia. The purpose was to evaluate SIR-A data as a complement to Landsat MSS data in an agricultural environment. The SIR-A data was filtered using a 7 x 7 pixel moving window median filter. Principal components analysis indicated the SIR-A data were discriminating between trees and agricultural fields. Supervised classifications using wheat, pasture, trees, and idle classes resulted in increased accuracies for wheat and pasture and slightly decreased accuracies for trees and idle for the Landsat MSS/SIR-A registered data sets over the Landsat MSS alone. Overall classification accuracies were unchanged for one date and substantially increased for the other when the SIR-A data were added to the Landsat MSS data.

Henninger, D. L.

The use of Landsat-4 MSS digital data in temporal data sets and the evaluation of scene-to-scene registration accuracy

The MSS sensor on Landsat 4 is, in certain performance aspects, diferent from those of Landsats 1 through 3. These differences created some concern in the NASA research community as to whether individual data sets can be registered accurately enough to produce acceptable data sets for multitemporal data analysis. The use of Landsat 4 MSS digital data in temporal data sets is examined and a method is presented for estimating temporal registration accuracy based on the use of an X-Y digitizer and grey tone electrostatic plots. Results indicate that the RMS temporal registration errors are not significantly different from the temporal data sets generated using Landsat 4 and Landsat 2 data (33.35 meters) and the temporal data set constructed from two Landsat 2 data sets (33.61 meters). A derivation of the model used to evaluate the temporal registration is included.

Anderson, J. E.

Earth observing system. Data and information system. Volume 2A: Report of the EOS Data Panel

The purpose of this report is to provide NASA with a rationale and recommendations for planning, implementing, and operating an Earth Observing System data and information system that can evolve to meet the Earth Observing System's needs in the 1990s. The Earth Observing System (Eos), defined by the Eos Science and Mission Requirements Working Group, consists of a suite of instruments in low Earth orbit acquiring measurements of the Earth's atmosphere, surface, and interior; an information system to support scientific research; and a vigorous program of scientific research, stressing study of global-scale processes that shape and influence the Earth as a system. The Eos data and information system is conceived as a complete research information system that would transcend the traditional mission data system, and include additional capabilties such as maintaining long-term, time-series data bases and providing access by Eos researchers to relevant non-Eos data. The Working Group recommends that the Eos data and information system be initiated now, with existing data, and that the system evolve into one that can meet the intensive research and data needs that will exist when Eos spacecraft are returning data in the 1990s.

Source record

Satellite data management for effective data access

The management of data generated from satellite missions has not always led to effective access of that data by the scientific community. NASA has tried to alleviate this problem for ocean scientists, by initiating a program, the NASA Ocean Data System (NODS). The menu-based user interface that NODS employs allows a user to make request and receive answers within a short time of accessing the system. A catalog system, which holds information about oceanographic data sets may be queried to determine the suitability of a particular data set. Once a candidate data set is found, the user is directed to the person or place which actually holds the data. NODS also has an archive system that holds data from ocean-observing satellites. The archive may be queried to obtain a manageable data subset that can be delivered in a useful form.

Hogan, Patrick D.

Global data bases on distribution, characteristics and methane emission of natural wetlands: Documentation of archived data tape

Global digital data bases on the distribution and environmental characteristics of natural wetlands, compiled by Matthews and Fung (1987), were archived for public use. These data bases were developed to evaluate the role of wetlands in the annual emission of methane from terrestrial sources. Five global 1 deg latitude by 1 deg longitude arrays are included on the archived tape. The arrays are: (1) wetland data source, (2) wetland type, (3) fractional inundation, (4) vegetation type, and (5) soil type. The first three data bases on wetland locations were published by Matthews and Fung (1987). The last two arrays contain ancillary information about these wetland locations: vegetation type is from the data of Matthews (1983) and soil type from the data of Zobler (1986). Users should consult original publications for complete discussion of the data bases. This short paper is designed only to document the tape, and briefly explain the data sets and their initial application to estimating the annual emission of methane from natural wetlands. Included is information about array characteristics such as dimensions, read formats, record lengths, blocksizes and value ranges, and descriptions and translation tables for the individual data bases.

Matthews, Elaine

A geometric comparison of video camera-captured raster data to vector-parented raster data generated by the X-Y digitizing table

The relative accuracy of a georeferenced raster data set captured by the Megavision 1024XM system using the Videk Megaplus CCD cameras is compared to a georeferenced raster data set generated from vector lines manually digitized through the ELAS software package on a Summagraphics X-Y digitizer table. The study also investigates the amount of time necessary to fully complete the rasterization of the two data sets, evaluating individual areas such as time necessary to generate raw data, time necessary to edit raw data, time necessary to georeference raw data, and accuracy of georeferencing against a norm. Preliminary results exhibit a high level of agreement between areas of the vector-parented data and areas of the captured file data where sufficient control points were chosen. Maps of 1:20,000 scale were digitized into raster files of 5 meter resolution per pixel and overall error in RMS was estimated at less than eight meters. Such approaches offer time and labor-saving advantages as well as increasing the efficiency of project scheduling and enabling the digitization of new types of data.

Swalm, C.

The role of data management in discipline-independent data visualization

The common data format (CDF) is described in terms of its support applications for the database management of visualization systems. The CDF is a self-describing data abstraction technique for the storage and manipulation of multidimensional data that are based on block structures. The discipline-independent approach is designed to manage, manipulate, archive, display, and analyze data, and can be applied to heterogeneous equipment communicating different data structures over networks. An improved CDF version incorporates a hyperplane access allowing random aggregate access to subdimensional blocks within a multidimensional variable. The visualization pipeline is also discussed, which controls the flow of data and permits the visualization of different classes of data representation techniques. The system is found to accommodate a large variety of scientific data structures and large disk-based data sets.

Treinish, Lloyd A.

NASDA's earth observation satellite data archive policy for the earth observation data and information system (EOIS)

NASDA's new Advanced Earth Observing Satellite (ADEOS) is scheduled for launch in August, 1996. ADEOS carries 8 sensors to observe earth environmental phenomena and sends their data to NASDA, NASA, and other foreign ground stations around the world. The downlink data bit rate for ADEOS is 126 MB/s and the total volume of data is about 100 GB per day. To archive and manage such a large quantity of data with high reliability and easy accessibility it was necessary to develop a new mass storage system with a catalogue information database using advanced database management technology. The data will be archived and maintained in the Master Data Storage Subsystem (MDSS) which is one subsystem in NASDA's new Earth Observation data and Information System (EOIS). The MDSS is based on a SONY ID1 digital tape robotics system. This paper provides an overview of the EOIS system, with a focus on the Master Data Storage Subsystem and the NASDA Earth Observation Center (EOC) archive policy for earth observation satellite data.

Sobue, Shin-ichi