Search NASASearch

SEARCH · Search NASA

Results for “data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Enhancing Discoverability and Management of Atmospheric Data at Scale: Solutions from the ARM Data Center

The Atmospheric Radiation Measurement (ARM) is a multi-laboratory and multi-institutional U.S. Department of Energy (DOE) Office of Science National User Facility. The ARM Data Center (ADC), located at Oak Ridge National Laboratory, collects, archives, and shares vast atmospheric data crucial for climate research. The ADC manages over 7 PB of data from 460 instruments worldwide, processing it into more than 11,000 diverse data products using the Network Common Data Form (NetCDF) for machine-independent accessibility. The primary challenge addressed in this paper is the efficient management and distribution of vast and diverse datasets essential for the climate research community, enhancing accessibility through advanced tools like Data Discovery. The ADC has developed advanced infrastructure and software architecture to handle the continuous influx of heterogeneous data to enhance data discoverability, resulting in increased scientific collaboration. In 2023, users from over 34 countries downloaded and utilized ARM data, resulting in 1,455 publications. The ADC’s efforts have significantly improved the discoverability and usability of atmospheric data, fostering extensive scientific research and collaboration. This paper details the solutions implemented by the ADC team for efficient data discovery and distribution, and it demonstrates ARM’s capability of staging processed data for scientific analysis.

Shah, Chirag [ORNL] (ORCID:0000000203145737)

Performance Analysis of Data Processing in Distributed File Systems with Near Data Processing

In the era of big data, the escalating volume and velocity of data generation pose significant challenges in data processing. Traditional systems like Spark and Hadoop manage the increasing amount and velocity of data by improving data placement and processing speeds. However, they face inherent limitations due to the essential data movement required for processing. In this paper, we explore the Skyhook framework, a novel extension of the Ceph distributed system, which significantly reduces the need for data movement. We present an extensive case study using the Skyhook framework, applying it with the TPC-H and K-means clustering algorithms. More specifically, we leverage the TPC-H benchmark to distinguish between CPU-intensive and I/O-intensive tasks. We explore the integration of K-means clustering into SQL, coupled with a near-data processing system to offload the computational burden of the K-means clustering algorithm to storage nodes. We conduct a comprehensive performance evaluation of distributed data processing applications across three processing approaches: traditional layout (baseline), optimized layout, and near-data processing. Additionally, we introduce the use of the FIO tool to simulate real-world system workloads, enabling the measurement of performance metrics such as average latency and CPU utilization. Our research is a significant advance in understanding how to optimize data processing systems to meet the demands of the modern data landscape.

Hou, Shiyue

Super-Resolution for Renewable Energy Resource Data with Wind from Reanalysis Data and Application to Ukraine

With a potentially increasing share of the electricity grid relying on wind to provide generating capacity and energy, there is an expanding global need for historically accurate, spatiotemporally continuous, high-resolution wind data. Conventional downscaling methods for generating these data based on numerical weather prediction have a high computational burden and require extensive tuning for historical accuracy. In this work, we present a novel deep learning-based spatiotemporal downscaling method using generative adversarial networks (GANs) for generating historically accurate high-resolution wind resource data from the European Centre for Medium-Range Weather Forecasting Reanalysis version 5 data (ERA5). In contrast to previous approaches, which used coarsened high-resolution data as low-resolution training data, we use true low-resolution simulation outputs. We show that by training a GAN model with ERA5 as the low-resolution input and Wind Integration National Dataset Toolkit (WTK) data as the high-resolution target, we achieved results comparable in historical accuracy and spatiotemporal variability to conventional dynamical downscaling. This GAN-based downscaling method additionally reduces computational costs over dynamical downscaling by two orders of magnitude. We applied this approach to downscale 30 km, hourly ERA5 data to 2 km, 5 min wind data for January 2000 through December 2023 at multiple hub heights over Ukraine, Moldova, and part of Romania. With WTK coverage limited to North America from 2007–2013, this is a significant spatiotemporal generalization. The geographic extent centered on Ukraine was motivated by stakeholders and energy-planning needs to rebuild the Ukrainian power grid in a decentralized manner. This 24-year data record is the first member of the super-resolution for renewable energy resource data with wind from the reanalysis data dataset (Sup3rWind).

17 WIND ENERGY

Pilot climate data system: A state-of-the-art capability in scientific data management

The Pilot Climate Data System (PCDS) was developed by the Information Management Branch of NASA's Goddard Space Flight Center to manage a large collection of climate-related data of interest to the research community. The PCDS now provides uniform data catalogs, inventories, access methods, graphical displays and statistical calculations for selected NASA and non-NASA data sets. Data manipulation capabilities were developed to permit researchers to easily combine or compare data. The current capabilities of the PCDS include many tools for the statistical survey of climate data. A climate researcher can examine any data set of interest via flexible utilities to create a variety of two- and three-dimensional displays, including vector plots, scatter diagrams, histograms, contour plots, surface diagrams and pseudo-color images. The graphics and statistics subsystems employ an intermediate data storage format which is data-set independent. Outside of the graphics system there exist other utilities to select, filter, list, compress, and calculate time-averages and variances for any data of interest. The PCDS now fully supports approximately twenty different data sets and is being used on a trial basis by several different in-house research grounds.

Smith, P. H.

Guidelines for submitting data to the National Space Science Data Center

The mission of the National Space Science Data Center (NSSDC) is to disseminate space science data for further analysis beyond that provided by the principal investigators (PIs) or team leaders (TLs) and their coworkers. Consequently, the NSSDC is responsible for the acquisition, organization, storage, retrieval, announcement, and distribution of scientific data obtained mainly from satellites and spacecraft. Any scientist may acquired data from the NSSDC and use them in further studies, either alone or in conjunction with data from ground-based or spacecraft experiments. With the responsibility for archiving data is the concomitant responsibility for distributing the documentation necessary to make those data usable. Since the group most knowledgeable about a particular experiment and its data is the PI or TL and his coworkers, and since the NSSDC cannot possibly supply the qualified personnel needed to write this documentation comprehensively, it is the responsibility of the PI or TL to provide the essential documentation. The NSSDC will support this effort by defining what is needed, by reviewing what is provided, and by reproducing and distributing the resulting documentation with the data. For a high-use data set, the NSSDC may publish the documentation as a Data Users Note; for a low-use data set, the NSSDC may distribute a Xerox, microfilm, or microfiche copy of the documentation.

Source record

Residual acceleration data on IML-1: Development of a data reduction and dissemination plan

The main thrust of our work in the third year of contract NAG8-759 was the development and analysis of various data processing techniques that may be applicable to residual acceleration data. Our goal is the development of a data processing guide that low gravity principal investigators can use to assess their need for accelerometer data and then formulate an acceleration data analysis strategy. The work focused on the flight of the first International Microgravity Laboratory (IML-1) mission. We are also developing a data base management system to handle large quantities of residual acceleration data. This type of system should be an integral tool in the detailed analysis of accelerometer data. The system will manage a large graphics data base in the support of supervised and unsupervised pattern recognition. The goal of the pattern recognition phase is to identify specific classes of accelerations so that these classes can be easily recognized in any data base. The data base management system is being tested on the Spacelab 3 (SL3) residual acceleration data.

Rogers, Melissa J. B.

Determination of an Optimal Commercial Data Bus Architecture for a Flight Data System

NASA/Marshall Space Flight Center (MSFC) is continually looking for methods to reduce cost and schedule while keeping the quality of work high. MSFC is NASA's lead center for space transportation and microgravity research. When supporting NASA's programs several decisions concerning the avionics system must be made. Usually many trade studies must be conducted to determine the best ways to meet the customer's requirements. When deciding the flight data system, one of the first trade studies normally conducted is the determination of the data bus architecture. The schedule, cost, reliability, and environments are some of the factors that are reviewed in the determination of the data bus architecture. Based on the studies, the data bus architecture could result in a proprietary data bus or a commercial data bus. The cost factor usually removes the proprietary data bus from consideration. The commercial data bus's range from Versa Module Eurocard (VME) to Compact PCI to STD 32 to PC 104. If cost, schedule and size are prime factors, VME is usually not considered. If the prime factors are cost, schedule, and size then Compact PCI, STD 32 and PC104 are the choices for the data bus architecture. MSFC's center director has funded a study from his discretionary fund to determine an optimal low cost commercial data bus architecture. The goal of the study is to functionally and environmentally test Compact PCI, STD 32 and PC 104 data bus architectures. This paper will summarize the results of the data bus architecture study.

Crawford, Kevin

Extending the LWS Data Environment: Distributed Data Processing and Analysis

The final stages of this work saw changes to the original framework, as well as the completion and integration of several data processing services. Initially, it was thought that a peer-to-peer architecture was necessary to make this work possible. The peer-to-peer architecture provided many benefits including the dynamic discovery of new services that would be continually added. A prototype example was built and while it showed promise, a major disadvantage was seen in that it was not easily integrated into the existing data environment. While the peer-to-peer system worked well for finding and accessing distributed data processing services, it was found that its use was limited by the difficulty in calling it from existing tools and services. After collaborations with members of the data community, it was determined that our data processing system was of high value and that a new interface should be pursued in order for the community to take full advantage of it. As such; the framework was modified from a peer-to-peer architecture to a more traditional web service approach. Following this change multiple data processing services were added. These services include such things as coordinate transformations and sub setting of data. Observatory (VHO), assisted with integrating the new architecture into the VHO. This allows anyone using the VHO to search for data, to then pass that data through our processing services prior to downloading it. As a second attempt at demonstrating the new system, a collaboration was established with the Collaborative Sun Earth Connector (CoSEC) group at Lockheed Martin. This group is working on a graphical user interface to the Virtual Observatories and data processing software. The intent is to provide a high-level easy-to-use graphical interface that will allow access to the existing Virtual Observatories and data processing services from one convenient application. Working with the CoSEC group we provided access to our data processing tools from within their software. This now allows the CoSEC community to take advantage of our services and also demonstrates another means of accessing our system.

Narock, Thomas

Transitioning NPOESS Data to Weather Offices: The SPoRT Paradigm with EOS Data

Real-time satellite information provides one of many data sources used by NWS weather forecast offices (WFOs) to diagnose current weather conditions and to assist in short-term forecast preparation. While GOES satellite data provides relatively coarse spatial resolution coverage of the continental U.S. on a 10-15 minute repeat cycle, polar orbiting imagery has the potential to provide snapshots of weather conditions at high-resolution in many spectral channels. Additionally, polar orbiting sounding data can provide additional information on the thermodynamic structure of the atmosphere in data sparse regions of at asynoptic observation times. The NASA Short-term Prediction Research and Transition (SPoRT) project has demonstrated the utility of polar orbiting MODIS and AIRS data on the Terra and Aqua satellites to improve weather diagnostics and short-term forecasting on the regional and local scales. SPoRT scientists work directly forecasters at selected WFOS in the Southern Region (SR) to help them ingest these unique data streams into their AWIPS system, understand how to use the data (through on-site and distance learn techniques), and demonstrate the utility of these products to address significant forecast problems. This process also prepares forecasters for the use of similar observational capabilities from NPOESS operational sensors. NPOESS environmental data records (EDRs) from the Visible 1 Infrared Imager I Radiometer Suite (VIIRS), the Cross-track Infrared Sounder (CrlS) and Advanced Technology Microwave Sounder (ATMS) instruments and additional value-added products produced by NESDIS will be available in near real-time and made available to WFOs to extend their use of NASA EOS data into the NPOESS era. These new data streams will be integrated into the NWs's new AWIPS II decision support tools. The AWIPS I1 system to be unveiled in WFOs in 2009 will be a JAVA-based decision support system which preserves the functionality of the existing systems and offers unique development opportunities for new data sources and applications in the Service Orientated Architecture ISOA) environment. This paper will highlight some of the SPoRT activities leading to the integration of VllRS and CrIS/ATMS data into the display capabilities of these new systems to support short-term forecasting problems at WFOs.

Jedlovec, Gary

NASA's Global Change Master Directory: Discover and Access Earth Science Data Sets, Related Data Services, and Climate Diagnostics

NASA's Global Change Master Directory provides the scientific community with the ability to discover, access, and use Earth science data, data-related services, and climate diagnostics worldwide. The GCMD offers descriptions of Earth science data sets using the Directory Interchange Format (DIF) metadata standard; Earth science related data services are described using the Service Entry Resource Format (SERF); and climate visualizations are described using the Climate Diagnostic (CD) standard. The DIF, SERF and CD standards each capture data attributes used to determine whether a data set, service, or climate visualization is relevant to a user's needs. Metadata fields include: title, summary, science keywords, service keywords, data center, data set citation, personnel, instrument, platform, quality, related URL, temporal and spatial coverage, data resolution and distribution information. In addition, nine valuable sets of controlled vocabularies have been developed to assist users in normalizing the search for data descriptions. An update to the GCMD's search functionality is planned to further capitalize on the controlled vocabularies during database queries. By implementing a dynamic keyword "tree", users will have the ability to search for data sets by combining keywords in new ways. This will allow users to conduct more relevant and efficient database searches to support the free exchange and re-use of Earth science data. http://gcmd.nasa.gov/

Aleman, Alicia

Examining Dense Data Usage near the Regions with Severe Storms in All-Sky Microwave Radiance Data Assimilation and Impacts on GEOS Hurricane Analyses

Many numerical weather prediction (NWP) centers assimilate radiances affected by clouds and precipitation from microwave sensors, with the expectation that these data can provide critical constraints on meteorological parameters in dynamically sensitive regions to make significant impacts on forecast accuracy for precipitation. The Global Modeling and Assimilation Office (GMAO) at NASA Goddard Space Flight Center assimilates all-sky microwave radiance data from various microwave sensors such as all-sky GPM Microwave Imager (GMI) radiance in the Goddard Earth Observing System (GEOS) atmospheric data assimilation system (ADAS), which includes the GEOS atmospheric model, the Gridpoint Statistical Interpolation (GSI) atmospheric analysis system, and the Goddard Aerosol Assimilation System (GAAS). So far, most of NWP centers apply same large data thinning distances, that are used in clear-sky radiance data to avoid correlated observation errors, to all-sky microwave radiance data. For example, NASA GMAO is applying 145 km thinning distances for most of satellite radiance data including microwave radiance data in which all-sky approach is implemented. Even with these coarse observation data usage in all-sky assimilation approach, noticeable positive impacts from all-sky microwave data on hurricane track forecasts were identified in GEOS-5 system. The motivation of this study is based on the dynamic thinning distance method developed in our all-sky framework to use of denser data in cloudy and precipitating regions due to relatively small spatial correlations of observation errors. To investigate the benefits of all-sky microwave radiance on hurricane forecasts, several hurricane cases selected between 2016-2017 are examined. The dynamic thinning distance method is utilized in our all-sky approach to understand the sources and mechanisms to explain the benefits of all-sky microwave radiance data from various microwave radiance sensors like Advanced Microwave Sounder Unit (AMSU-A), Microwave Humidity Sounder (MHS), and GMI on GEOS-5 analyses and forecasts of various hurricanes.

Kim, Min-Jeong

NASA Earth Science Data Systems: Open Data, Services and Software

Open, Public, Electronic and Necessary (OPEN) Government Data Act, which requires all non-sensitive government data to be made available in open and machine-readable formats by default is part of the overall Foundations for Evidence-Based Policymaking (FEBP) Act passed in late 2018. This town hall will bring together data officers and policy makers from NOAA, EPA, NASA and others to discuss the impact of the act on data management strategies going forward. In 1994 NASA's Earth Science Division committed to an open data policy for all civilian Earth satellite data. NASA's Earth Observing System Data and Information System (EOSDIS) became the first large scale data system to facilitate public access to global Earth system data and information. This presentation reviews key elements of EOSDIS data policy and data management activities that support the OPEN Government Data Act.

Moses, John F.

Extravehicular Activity Mission System Software (EMSS) - Enabling Human Planetary Exploration Data Within The Broader Planetary Data Ecosystem

The planetary science community is once again on the verge of generating, capturing and analyzing human planetary exploration data, this time via the Artemis program. Artemis missions will involve robotic missions in addition to human extravehicular activity (EVA) where crew will be generating scientific data [1]. Present-day robotic mission data expectations for data archiving involves ingesting data into the Planetary Data System (PDS), but how might PDS be leveraged/adapted/ready (or not) for human spaceflight mission data, particularly EVA data that includes non-scientific data that provides important context to the scientific data gathered on the lunar surface? This question has broader implications than what this abstract can answer, but we wanted to pose the question to 1) get conversations started and 2) highlight how operations software data handling could play a role in overall data curation.

M J Miller

Automated Data Accountability for Missions in Mars Rover Data

As the Mars Curiosity Rover transmits data to the JPL Ground Data System (GDS), it frequently observes data loss and corruption, requiring re-transmits from the rover and Ground Data System Analysts (GDSA) to monitor the downlink process. As new missions are launched, the GDSA team redistributes analysts to these new missions, causing shortages in previous missions. The GDSA team can significantly benefit from the automation and optimization of the downlink process of telemetry data. In fact, there is a need for a better understanding of why the data is corrupted, so that the GDSA team can best determine the root cause of the issues in the GDS. This paper presents machine learning and deep learning based approaches to automate and optimize the detection of data loss. We first created a pipeline to automatically accumulate data from the telemetry databases (MAROS, Telemetry Data Storage, and GDS Elastic Search Database) in the downlink process. With our newly created datasets, we perform feature selection to supplement the GDSA understanding of the downlink process and provide supplemental analysis on the importance of different features. We implement various machine learning and deep learning based models, including support vector machines, ensemble methods, and deep neural networks and evaluate their accuracies in identifying whether a downlink process is complete or incomplete. We utilize fast hyperparameter optimization methods that allow our models to quickly be re-trained, allowing them to quickly be tuned and optimized on daily incoming data in real time. This hyperparameter optimization also allows our methods to be quickly integrated into other JPL missions. Our results show that our best-performing machine learning and deep learning based models outperform the existing GDSA detection software by 6 accuracy points and can aid analysts by providing insights into the data accountability problem. Since these various machine learning and deep learning approaches vary significantly in interpretability, we provide a discussion on the tradeoffs between their performance and trustworthiness in helping detect issues in data transmission.

Divsalar, Dariush

NASA’s Human Data Repositories: An In Depth Look at the New Data Request Process

As NASA transitions its focus to travel back to the moon and on to new destinations, the need to ensure the capture, analysis, and application of research and medical data is of greater urgency than at any other previous time. In this era of limited resources and challenging schedules, the Human Research Program (HRP), based at NASA’s Johnson Space Center (JSC), recognizes the need to extract the greatest possible amount of information from the data already captured. To this end, the HRP Chief Scientist Office (CSO), HRP Program Planning and Control (PP&C) Office, and the Space Medicine Operations Division have been working together to make reuse of both research data and medical monitoring data more accessible to the user community through the Life Science Data Archive (LSDA) and the Lifetime Surveillance of Astronaut Health (LSAH) Repositories. The task of both LSDA and LSAH repositories is to acquire, preserve, and distribute retrospective research (LSDA) and medical (LSAH) data and information both within the NASA community and to the science community at large, for knowledge discovery, retrospective analysis, and planning of future research studies. An additional goal is to encourage collaboration with non-NASA institutions also faced with enhancing human performance in extreme environments. In September 2022, the LSDA website and its contents transitioned to a new NASA Life Sciences Portal (https://nlsp.nasa.gov/explore/lsdahome). This site continues to feature publicly releasable information such as non-attributable datasets, experiment descriptions (from Project Mercury to ISS, as well as from multiple flight analog missions), descriptions of medical monitoring data, and LSAH newsletters (1992 - 2022). The website also provides an updated portal to request additional research and medical data not accessible from the public website. This presentation will provide an in-depth look at the new system as it relates to finding and requesting retrospective data. We will also detail processes from making a request to delivering data for different types of data requests (i.e., attributable, or non-attributable). This includes descriptions of various approval boards, what information and actions the requestor is responsible for, and key milestones in making data available for reuse.

D. M. Thomas

Towards A Flexible Data Fusion Tool Incorporating Model, Satellite, Regulatory Monitor and Low-Cost Sensor Data for Air Quality Estimation and Forecasting

Air quality managers, researchers, and concerned community scientists around the world have a variety of sources for air quality information, ranging from traditional regulatory monitoring networks and atmospheric chemistry models to remote sensing data products and low-cost sensor networks. However, the ability to incorporate data from these disparate sources and synthesize a comprehensive overview of the local air quality situation remains a considerable barrier for many end-users. This presentation will outline a tool, currently in development, which will address this need using a flexible data fusion approach. The tool will make use of air quality forecast model outputs (primarily from the NASA GEOS-CF composition forecast modeling system), satellite remote sensing data (from instruments including MODIS, VIIRS, TROPOMI, plus TEMPO for the US when available), and in-situ data from official regulatory and/or low-cost networks where these are available. The ability to incorporate data from low-cost sensor networks will be a key feature of the tool; it will make use of other available data sources to calibrate the low-cost sensor data on a regional scale, then use these calibrated low-cost sensor data for localized updating to resolve finer-scale air quality patterns. Development of this tool is taking place with the help of national and international partners and end-user groups, coordinated through the US EPA and the United Nations Environment Programme (UNEP). The tool is being developed on the Google Earth Engine cloud computing platform to facilitate integration of diverse data sources and free access by a broad community of end-users. Stewardship of the tool will be passed to US EPA and UNEP to support future activities with end-users in the US and around the world, and the tool itself will remain freely accessible. We hope that this tool will lower the barrier to entry for various user groups worldwide, including community scientists, who struggle to integrate disparate data sources to gain insight into their local air quality situations. This presentation will cover the early stages of the development of the tool, including the underlying methods and some pilot case studies in integrating low-cost sensor data.

global models

RadLab and the Environmental Data Application Dashboard: Graphical and Programming Interfaces for Interrogation of Space Telemetry Data

Sensors on the International Space Station (ISS) and multiple spacecraft elsewhere in Earth orbit and in deep space continuously monitor and collect environmental data, transmitting this information back to Earth. These data include ionizing radiation and, on the ISS, CO2, relative humidity levels, and temperature, and are of great importance to space biology research. Ionizing radiation in particular has been established in ground-based experiments as being correlated with increased risk of carcinogenesis and cardiovascular and neurological effects. Looking ahead to future long duration crewed missions beyond low Earth orbit, the ability to study how factors including CO2 levels, light cycle, temperature modulate the response to ionizing radiation and microgravity is essential. To date, access to these data has been fragmented across space agencies, spacecraft, and databases. To address this issue, NASA’s Open Science Data Repository (osdr.nasa.gov) has developed two Web applications: the Environmental Data Application (EDA) and a radiation-specific RadLab. Each consists of an API (application programming interface) and an associated GUI (graphical user interface) that provide single points of access to the data. To date, OSDR has focused on the sensors from payloads and radiation detectors located on the ISS. The Web applications process telemetry information and associated data, such as spacecraft location and orientation, from multiple international databases. The applications’ request syntax enables users to interrogate these data by craft, sensor type, time range, radiation type (galactic cosmic rays, solar particle events, the contribution of the South Atlantic Anomaly), facilitating arbitrary comparisons of original source data at varying time resolutions. The applications provide programmatic access for use in computational pipelines and GUIs for data visualization and exploration, making these data FAIR (Findable, Accessible, Interoperable, and Reusable), complementing the biological data contained in OSDR, and providing the space science community with a valuable resource for scientific analyses.

radiation

Big-data Efficient and Automated Science Transfer (BEAST): An Open-Source Software Architecture for Arc Jet Data Management, Modeling, and Automation

Big-data Efficient and Automated Science Transfer (BEAST) was conceived to address the existing ground testing data management of the NASA Ames arc jet facilities (e.g., manually entered Excel files and USB drive data transfers). These data management practices were seen as a choke point for future thermal protection system (TPS) development as they limit statistical tracking, resolution of diagnostics, coordination between video/time series, data throughput, and data processing speed/efficiency. Consequently, BEAST was developed to provide a new data infrastructure with streamlined data collection, processing, transfer, and analysis. This new framework also seeks to implement the FAIR principles of data stewardship: Findable, Accessible, Interoperable, and Reusable. The BEAST framework is based on a combination of the Python Django web framework and the Python data stack to provide a monolithic, open-source platform for data management, automation, and machine learning. This architecture was chosen for maintainability and scalability for a small, in-house development team. This paper will describe the application framework, deployment, and discuss the benefits and future plans for the system.

Data management