Search NASA⌕ Search

SEARCH · Search NASA

Results for “Amazon S3”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Accessing Data Stored in Amazon S3 Using the Hyrax OPeNDAP Server

For three years we have been investigating data storage and retrieval in the Amazon Cloud, striving to optimize both storage structures and (cache-enhanced retrieval processes. These optimizations are hidden from data users, who simply employ the Data Access Protocol (DAP), a popular, Web-based Application Programmer Interface (API) developed by OPeNDAP. We present our collected findings and their realization in the Hyrax data server. Specific techniques include caching data on spinning disk; accessing sharded data in place; optimizing the organizational structures for data within the flat key-value space of Simple Storage Service (S3); and employing an API extension that permits simultaneous operations on many datasets in a single request.

Data access↗

SatCORPS Global Cloud Composite (GCC): the Design and Delivery of A High Quality, High Resolution, Global Cloud Product Available in Near-Real Time

The NASA Satellite ClOud and Radiation Property retrieval System (SatCORPS) supports the development of an analysis ready and cloud-optimized data transformation pipeline and geospatial service enablement of a global cloud composite (GCC) product derived from global geostationary satellite imagery. This geospatial service will be available at high temporal and spatial resolution via the SatCORPS web mapping application for visualization and analysis as well as direct ingestion to common geospatial software and custom programming. The resulting global cloud composite products from the processing pipeline can then be geospatially-service enabled as ArcGIS Image Services and Open Geospatial Consortium (OGC) Web Mapping/Coverage Services for visualization and analysis via a web mapping application and common geospatial software. Near real time global observations are created through the composition of five geostationary satellites that provides modelling and forecasting communities with the capability to provide high quality and timely information to start the projection process. The Global Cloud Composite product combines information from geostationary satellites, GOES-16, GOES-17, Himawari-8, Meteosat-11 and Meteosat-9 to create a single global composite netcdf file and images using the different products within the netcdf file. The SatCORPS team, though our Global Cloud Composite (GCC) product and web-based visualization tools including Geographical Information System (GIS) services provide near real time global cloud product information to both automated processes and traditional web users that is timely and high quality derived from geostationary satellites. The Global Cloud Composite product takes advantage of the scalable processing resources provided by the AWS batch service to provide new composites every thirty minutes. Because information from each of the low earth orbiting satellites is available on schedules tuned to the specific satellite, the processing algorithm temporally composites the final dataset as each satellite’s information becomes available. The SatCORPS team has leveraged our experience using Amazon Web Services (AWS) to build a low latency high availability tool that allows end users both human and automated to acquire high quality and high-resolution Geostationary Earth Orbiting (GEO) information at zero cost to the end user. This presentation will describe how we architected and implemented the service as well as lessons learned based on our experiences both developing and operating the system. The lessons learned include how we integrated multiple services including Amazon Batch, Amazon S3 and Amazon Lambda service to create a low cost but high-performance processing system that is capable of identifying and processing the most appropriate satellite overpass information into global cloud composites. We will also describe our web-based tools including our Geographic Information System that can be used for visualization and analysis. The products from the processing can be geospatially-service enabled as ArcGIS Image Services and Open Geospatial Consortium (OGC) Web Mapping/Coverage Services for visualization and analysis via a web mapping application and common geospatial software. The SatCORPS Global Composite Cloud product provides sophisticated global composited cloud research products with very low latency that we see that as filling a rapidly growing need in the research and modelling community with no up-front nor ongoing costs associated with downloading or using the information.

AWS AMCE SMCE GCC SATCORPS GLOBAL CLOUD COMPOSITE ↗

Federated Access from DOE Labs to Distributed Storage in the EIC Era of Computing

The Electron Ion Collider (EIC) collaboration and future experiment is a unique scientific ecosystem within Nuclear Physics as the experiment starts right off as a crosscollaboration between Brookhaven National Lab (BNL) & Jefferson Lab (JLab). As a result, this muti-lab computing model tries at best to provide services accessible from anywhere by anyone who is part of the collaboration. While the computing model for the EIC is not finalized, it is anticipated that the computational and storage resources will be made accessible to a wide range of collaborators across the world. The use of federated ID seems to be a critical element to the strategy of providing such services, allowing seamless access to each lab site computing resources. However, providing Federated access to a Federated storage is not a trivial matter and has its share of technical challenges. In this contribution, we focus on the steps we took towards the deployment of a distributed object storage system that integrates with Amazon S3 and Federated ID. We will first cover for and explain the first stage storage solutions provided to the EIC during the detector design phase. Our initial test deployment consisted of Lustre storage using MinIO, hence providing an S3 interface. High Availability load balancers were added later to provide the initial scalability it lacked. Performance of that system will be shown. While this embryonic solution worked well, it had many limitations. Looking ahead, the Ceph object storage is considered a top-of-the-line solution in the storage community - since the Ceph Object Gateway is compatible with the Amazon S3 API out of the box, our next phase will use a native S3 storage. Our Ceph deployment will consist of erasure coded storage nodes to maximize storage potential along with multiple Ceph Object Gateways for redundant access. We will compare performance of our next stage implementations. Finally, we will present how to leverage OpenID Connect with the Ceph Object Gateway’s to enable Federated ID access. We hope this contribution will serve the community needs as we move forward with cross-lab collaborations and the need for Federated ID access to distributed compute facilities.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

OPeNDAP Clients, Aggregation and S3

In this talk, we will discuss our work for testing OPeNDAP client access of data stored in the Amazon S3 cloud storage using a set of common analysis tools including Panoply, Jupyter Notebooks with Python xarray, NCO command line tool package, ArcGIS, and GDAL. We will also discuss our ongoing work on improving performance in Hyrax aggregation functionality.

Amazon S3↗

HDF5 Roadmap 2019-2020

In this talk we will give an overview of the new features of the upcoming HDF5 release 1.12.0, and outline the HDF5 roadmap for the next year. We will demonstrate new open source file drivers to access HDF5 files via Amazon Simple Storage Service (Amazon S3) and on Hadoop Distributed File system (HDFS). We will use this presentation to get feedback on the HDF5 roadmap from the ESDIS users and application developers.

Cloud↗

Comparison of Radiosonde Datasets: SondeHub and Integrated Global Radiosonde Archive

SondeHub aggregates radiosonde telemetry data uploaded from community-run radiosonde receiver stations. This radiosonde telemetry dataset is open-source, available to anyone through Amazon S3. There are also other public radiosonde datasets such as National Centers for Environmental Information (NCEI)’s Integrated Global Radiosonde Archive (IGRA). While there are many similarities between the two datasets, there are many differences as well due to the nature of the two datasets: one is community-run, while the other is managed by a government agency. This report presents the result of analyzing and comparing the two datasets.

54 ENVIRONMENTAL SCIENCES↗

NASA POWER: Providing Analysis-Ready, Cloud-Optimized Data for AI /ML Training and Applications in Earth Science

As global demand for sustainable development grows, the integration of Earth Observation (EO) data into decision making frameworks has become a primary objective for the scientific community. The NASA Prediction of Worldwide Energy Resources (POWER) project serves as a bridge between NASA EO data and the specialized needs of the renewable energy, sustainable infrastructure and agroclimatology communities. In this poster presentation we will present an overview of POWER data products and services along with its use in diverse research to decision-making workflows. By providing over 40 years of high-resolution historical, hourly and daily solar and meteorological data, POWER transforms satellite observations and global model reanalysis into actionable, Analysis-Ready Dataset (ARD). Currently, the project delivers over 250 industry-friendly parameters to the users from different NASA datasets like CERES SYN1Deg, MERRA-2, and IMERG alongside downscaled CMIP6 climate model data, fulfilling over 16 million requests from 50,000 unique users monthly. To ensure data quality and traceability, these parameters are rigorously validated against the ground-based observations from the Baseline Surface Radiation Network (BSRN) and the Global Surface Summary of the Day (GSOD) – these results will be discussed in the presentation. A newly introduced web-based PaRameter Uncertainty ViEwer (PRUVE) tool will be presented that provides an online validation platform to the users that benchmarks satellite-based and assimilation data products against these surface measurements. To reduce technical barriers to data adoption, POWER data is accessible through RESTful APIs, ESRI ArcGIS Image Services, a web-based Data Access Viewer tool, allowing users to visualize, validate and apply the dataset. For efficient data delivery POWER data is cloud-optimized into Zarr datastore accessible through NASA managed Amazon S3 ensures high-performance allowing users to integrate EO directly into operational pipelines. These customized services will be presented. Use cases from application will be presented from the energy sector - such as for design of generation systems, performance monitoring of solar power plants, in infrastructure sector- optimizing building energy efficiency and thermal comfort, in agriculture – such as driving crop simulation and yield forecasting models to enable climate resilient farming. Furthermore, the shift toward machine learning (ML) in EO research that has positioned POWER as a key provider for training datasets which will be discussed. Use-cases will be presented to showcase how NASA data is enabling the development of predictive tools for climate variability and resource management. The poster will present POWER’s future plans including technology development to enhance data traceability and reproducibility and improving I/O performance to support the rapid integration of new EO products, ensuring that POWER remains a robust scalable backend for the evolving landscape of AI-driven Earth Science. Additionally, POWER is developing an AI Agent and an MCP-Server to enable industry AI-Agentic workflows.

Neha Khadka↗

Towards Efficient Scientific Data Management Using Cloud Storage

A software prototype allows users to backup and restore data to/from both public and private cloud storage such as Amazon's S3 and NASA's Nebula. Unlike other off-the-shelf tools, this software ensures user data security in the cloud (through encryption), and minimizes users operating costs by using space- and bandwidth-efficient compression and incremental backup. Parallel data processing utilities have also been developed by using massively scalable cloud computing in conjunction with cloud storage. One of the innovations in this software is using modified open source components to work with a private cloud like NASA Nebula. Another innovation is porting the complex backup to- cloud software to embedded Linux, running on the home networking devices, in order to benefit more users.

He, Qiming↗

Hosting Hyrax in the Cloud While Preserving the User Experience

Cloud object stores enable data providers to reduce costs and provide uniform data servers across organizational divisions. This talk will address how new capabilities provided by cloud computing can be used with existing software so that users get the benefits provided by cloud technology (access to data at lower cost, e.g.) without disrupting their existing workflow(s). We will address ways to reduce response latency often associated with Cloud object stores and ways to simulate the hierarchy of a POSIX file system using a key-value object store. This work is based an on-going NASA EED2 project to enhance the Hyrax data server for operation in the Amazon Web Services Cloud.

Hyrax server↗

Changes in Characteristics of Future Climate Across the U.S.: Time Series Analysis of Climate Model Data by NASA POWER

NASA’s Prediction of Worldwide Energy Resource (POWER) project facilitates the use of NASA Earth Science data holdings within the energy, agricultural, and building heating/cooling design industries. POWER packages solar and meteorological data at various temporal levels from several NASA projects in a user friendly GIS-enabled web services system (https://power.larc.nasa.gov). Data users can access these data either through an intuitive data viewer, image services fully integrable with GIS analysis, connections in the cloud through an Amazon Web Services S3 Bucket, or fully customizable access through an API. Data provided by POWER has been used to remotely monitor solar array fields and integrated in a sizing tool for off-grid solar and storage systems. POWER data has also been coupled with key building decision tools to support design and retrofitting of building energy systems for energy efficiency and reduction of greenhouse gases. POWER is now developing capabilities to provide time series of the projected future evolution of surface quantities important to future energy production and use, such as heating/cooling degree days, temperature, wind speed, and downwelling solar flux. We present here a range of possible future changes in these quantities at locations throughout the continental United States. We show how both average and extreme values of the quantities will evolve from present-day to future climate conditions. We plan to provide these projections for users in the energy and sustainable energy communities.

Bradley M. Hegyi↗

NASA POWER: Providing Present and Future Climate Services Based on NASA Data for the Energy, Agricultural, and Sustainable Buildings Communities

NASA’s Prediction of Worldwide Energy Resource (POWER) project facilitates the use of NASA Earth Science data holdings within the renewable energy, agricultural, and building heating/cooling design industries. POWER packages solar and meteorological data at various temporal levels from several NASA projects in a user friendly GIS-enabled web services system (https://power.larc.nasa.gov). Data users can access these data either through an intuitive data viewer, image services fully integrable with GIS analysis, connections in the cloud through an Amazon Web Services S3 Bucket, or fully customizable access through an API. Data provided by POWER has been successfully used by decision makers to support actions that address climate change. For example, POWER data has been used to remotely monitor solar array fields and integrated in a sizing tool for off-grid solar and storage systems. POWER data has also been coupled with key building decision tools to support design and retrofitting of building energy systems for energy efficiency and reduction of greenhouse gases. POWER is now developing climate services to provide time series of the projected future evolution of key quantities that interest our users, such as heating/cooling degree days, temperature, wind speed, and downwelling solar flux. We demonstrate the potential of the new climate services by presenting here a range of possible future changes in these quantities at different NASA centers across the continental United States. These data services are based on downscaled climate model data from the NASA Earth Exchange Global Daily Downscaled Projections (NEX-GDDP) data set. We highlight the important insights that new climate services can provide. Our climate services will help our user communities quantify the impacts of climate change to support their key decisions in planning for the future, both inside and outside the Federal Government, especially for decisions in renewable energy and in building heating and cooling.

Bradley Hegyi↗

Leveraging the Cloud for Robust and Efficient Lunar Image Processing

The Lunar Mapping and Modeling Project (LMMP) is tasked to aggregate lunar data, from the Apollo era to the latest instruments on the LRO spacecraft, into a central repository accessible by scientists and the general public. A critical function of this task is to provide users with the best solution for browsing the vast amounts of imagery available. The image files LMMP manages range from a few gigabytes to hundreds of gigabytes in size with new data arriving every day. Despite this ever-increasing amount of data, LMMP must make the data readily available in a timely manner for users to view and analyze. This is accomplished by tiling large images into smaller images using Hadoop, a distributed computing software platform implementation of the MapReduce framework, running on a small cluster of machines locally. Additionally, the software is implemented to use Amazon's Elastic Compute Cloud (EC2) facility. We also developed a hybrid solution to serve images to users by leveraging cloud storage using Amazon's Simple Storage Service (S3) for public data while keeping private information on our own data servers. By using Cloud Computing, we improve upon our local solution by reducing the need to manage our own hardware and computing infrastructure, thereby reducing costs. Further, by using a hybrid of local and cloud storage, we are able to provide data to our users more efficiently and securely. 12 This paper examines the use of a distributed approach with Hadoop to tile images, an approach that provides significant improvements in image processing time, from hours to minutes. This paper describes the constraints imposed on the solution and the resulting techniques developed for the hybrid solution of a customized Hadoop infrastructure over local and cloud resources in managing this ever-growing data set. It examines the performance trade-offs of using the more plentiful resources of the cloud, such as those provided by S3, against the bandwidth limitations such use encounters with remote resources. As part of this discussion this paper will outline some of the technologies employed, the reasons for their selection, the resulting performance metrics and the direction the project is headed based upon the demonstrated capabilities thus far.

Cloud Computing↗

High Performance Access to Archival Data Stored in HDF4 and HDF5 on Cloud Object Stores Without Reformatting the Files

Cloud computing offers numerous advantages for users of extensive Earth science data collections. These benefits encompass direct online access to data files and granules from any location, scalable access supporting parallel computing workflows, and flexible computing tools enabling innovative experimentation with processing techniques. However, older archival file formats designed for distinct computing systems hinder efficient access to decade-long time-series data when compared to data stored in modern cloud-optimized formats like Web Object Stores (WOS), exemplified by Amazon Web Services’ Simple Storage Service (S3). We describe DMR++ (Dataset Metadata Response plus plus), a technology facilitating efficient access to HDF5 (Hierarchical Data Format, version 5) and HDF4 files stored on WOS systems without requiring data reformatting. DMR++ achieves performance comparable to technologies like Zarr while preserving the original file structure, a substantial benefit considering the vast quantity of archival files held by organizations such as NASA. Moreover, DMR++ typically outperforms cloud-optimized versions of HDF5. Essentially an XML (Extensible Markup Language) document usually stored alongside the described data, DMR++ can also be generated on-the-fly but is generally created during data staging to the WOS. Archival files that use HDF4/5 often store large arrays of numerical data. The data in these files is often compressed, typically reducing their size by a factor of four or more. To achieve efficient access to portions of those arrays, they are 'chunked' into smaller sub-arrays, each individually compressed. The chunk size is a compromise, where spinning disks can efficiently access data in smaller chunks while S3 favors larger chunks. A simple optimization of aggregating smaller chunks that are stored adjacently, transferring them in a single access and then individually decompressing them will improve performance. NASA data pose an additional challenge: special Application Programmer Interface (API) libraries are often needed to compute some variables. These libraries are incompatible with WOS environments. Our solution involves storing computed values in the DMR++ document or a companion file, making them accessible like other variables and eliminating the need for specialized APIs. We outline specific optimizations for both satellite grid and swath data stored in HDF4-EOS2 (Earth Observing System).

James Gallagher↗

Forming Aggregations using Virtual Sharding: Lessons Learned from Simple Scalable Storage (S3)

Data aggregation is the ability to combine separate datasets to form a single new logical dataset provides users with a powerful abstraction. The advantage of an aggregate dataset is that the users are freed from having to understand, and incorporate into their workflow, knowledge about the (ad hoc) organization of the constituent datasets. However, aggregating large numbers of files can be computationally complex with data server systems performing many repetitive operations. As part of the authors work on subsetting data stored on Amazon Web Service (AWS) Simple Storage Service (S3), we developed technology to read portions of otherwise monolithic data files. This enables the formation of virtual shards for user in subsetting data stored in HDF5 (hierarchical data format, version 5) files. This same tool can be used to form aggregations that combine data stored in many HDF5 files when those files are stored on S3. The nature of the virtual sharding and the algorithm that exploits it for subsetting is such that it can also be used for aggregation with the need for many of the repetitive operations required by the per file aggregation techniques. We will present timing information that demonstrates the flexibility of this approach. However, the lessons learned is that while this is a useful result in and of itself, these very same techniques can be applied in other contexts where data are stored in services and on media other than S3. For example, this same technique can be applied to data stored on spinning disk. Pushing the envelope for S3 forced a reexamination of our data access techniques which lead to unexpected positive benefits.

Gallagher, James↗

Trade Study: Storing NASA HDF5/netCDF-4 Data in the Amazon Cloud and Retrieving Data Via Hyrax Server Data Server

This study explored three candidate architectures with different types of objects and access paths for serving NASA Earth Science HDF5 data via Hyrax running on Amazon Web Services (AWS). We studied the cost and performance for each architecture using several representative Use-Cases. The objectives of the study were: Conduct a trade study to identify one or more high performance integrated solutions for storing and retrieving NASA HDF5 and netCDF4 data in a cloud (web object store) environment. The target environment is Amazon Web Services (AWS) Simple Storage Service (S3). Conduct needed level of software development to properly evaluate solutions in the trade study and to obtain required benchmarking metrics for input into government decision of potential follow-on prototyping. Develop a cloud cost model for the preferred data storage solution (or solutions) that accounts for different granulation and aggregation schemes as well as cost and performance trades.We will describe the three architectures and the use cases along with performance results and recommendations for further work.

AWS cost↗

Task 28: Web Accessible APIs in the Cloud Trade Study

This study explored three candidate architectures for serving NASA Earth Science Hierarchical Data Format Version 5 (HDF5) data via Hyrax running on Amazon Web Services (AWS). We studied the cost and performance for each architecture using several representative Use-Cases. The objectives of the project are: Conduct a trade study to identify one or more high performance integrated solutions for storing and retrieving NASA HDF5 and Network Common Data Format Version 4 (netCDF4) data in a cloud (web object store) environment. The target environment is Amazon Web Services (AWS) Simple Storage Service (S3).Conduct needed level of software development to properly evaluate solutions in the trade study and to obtain required benchmarking metrics for input into government decision of potential follow-on prototyping. Develop a cloud cost model for the preferred data storage solution (or solutions) that accounts for different granulation and aggregation schemes as well as cost and performance trades.

cost model↗

Use of Schema on Read in Earth Science Data Archives

Traditionally, NASA Earth Science data archives have file-based storage using proprietary data file formats, such as HDF and HDF-EOS, which are optimized to support fast and efficient storage of spaceborne and model data as they are generated. The use of file-based storage essentially imposes an indexing strategy based on data dimensions. In most cases, NASA Earth Science data uses time as the primary index, leading to poor performance in accessing data in spatial dimensions. For example, producing a time series for a single spatial grid cell involves accessing a large number of data files. With exponential growth in data volume due to the ever-increasing spatial and temporal resolution of the data, using file-based archives poses significant performance and cost barriers to data discovery and access. Storing and disseminating data in proprietary data formats imposes an additional access barrier for users outside the mainstream research community. At the NASA Goddard Earth Sciences Data Information Services Center (GES DISC), we have evaluated applying the schema-on-read principle to data access and distribution. We used Apache Parquet to store geospatial data, and have exposed data through Amazon Web Services (AWS) Athena, AWS Simple Storage Service (S3), and Apache Spark. Using the schema-on-read approach allows customization of indexing spatially or temporally to suit the data access pattern. The storage of data in open formats such as Apache Parquet has widespread support in popular programming languages. A wide range of solutions for handling big data lowers the access barrier for all users. This presentation will discuss formats used for data storage, frameworks with This presentation will discuss formats used for data storage, frameworks with support for schema-on-read used for data access, and common use cases covering data usage patterns seen in a geospatial data archive.

cloud applications↗

Extending Rucio with modern cloud storage support

Rucio is a software framework designed to facilitate scientific collaborations in efficiently organising, managing, and accessing extensive volumes of data through customizable policies. The framework enables data distribution across globally distributed locations and heterogeneous data centres, integrating various storage and network technologies into a unified federated entity. Rucio offers advanced features like distributed data recovery and adaptive replication, and it exhibits high scalability, modularity, and extensibility. Originally developed to meet the requirements of the high-energy physics experiment ATLAS, Rucio has been continuously expanded to support LHC experiments and diverse scientific communities. Recent R&D projects within these communities have evaluated the integration of both private and commercially-provided cloud storage systems, leading to the development of additional functionalities for seamless integration within Rucio. Furthermore, the underlying systems, FTS and GFAL/Davix, have been extended to cater to specific use cases. This contribution focuses on the technical aspects of this work, particularly the challenges encountered in building a generic interface for self-hosted cloud storage, such as MinIO or CEPH S3 Gateway, and established providers like Google Cloud Storage and Amazon Simple Storage Service. Additionally, the integration of decentralised clouds like SEAL is explored. Key aspects, including authentication and authorisation, direct and remote access, throughput and cost estimation, are highlighted, along with shared experiences in daily operations.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗