Search NASASearch

SEARCH · Search NASA

Results for “cloud storage”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Towards Efficient Scientific Data Management Using Cloud Storage

A software prototype allows users to backup and restore data to/from both public and private cloud storage such as Amazon's S3 and NASA's Nebula. Unlike other off-the-shelf tools, this software ensures user data security in the cloud (through encryption), and minimizes users operating costs by using space- and bandwidth-efficient compression and incremental backup. Parallel data processing utilities have also been developed by using massively scalable cloud computing in conjunction with cloud storage. One of the innovations in this software is using modified open source components to work with a private cloud like NASA Nebula. Another innovation is porting the complex backup to- cloud software to embedded Linux, running on the home networking devices, in order to benefit more users.

He, Qiming

Using Cloud-Based Storage Technologies for Earth Science Data

Cloud based infrastructure may offer several key benefits of scalability, built in redundancy and reduced total cost of ownership as compared with a traditional data center approach. However, most of the tools and software systems developed for NASA data repositories were not developed with a cloud based infrastructure in mind and do not fully take advantage of commonly available cloud-based technologies. Object storage services are provided through all the leading public (Amazon Web Service, Microsoft Azure, Google Cloud, etc.) and private (Open Stack) clouds, and may provide a more cost-effective means of storing large data collections online. We describe a system that utilizes object storage rather than traditional file system based storage to vend earth science data. The system described is not only cost effective, but shows superior performance for running many different analytics tasks in the cloud. To enable compatibility with existing tools and applications, we outline client libraries that are API compatible with existing libraries for HDF5 and NetCDF4. Performance of the system is demonstrated using clouds services running on Amazon Web Services.

Data

Leveraging the Cloud for Robust and Efficient Lunar Image Processing

The Lunar Mapping and Modeling Project (LMMP) is tasked to aggregate lunar data, from the Apollo era to the latest instruments on the LRO spacecraft, into a central repository accessible by scientists and the general public. A critical function of this task is to provide users with the best solution for browsing the vast amounts of imagery available. The image files LMMP manages range from a few gigabytes to hundreds of gigabytes in size with new data arriving every day. Despite this ever-increasing amount of data, LMMP must make the data readily available in a timely manner for users to view and analyze. This is accomplished by tiling large images into smaller images using Hadoop, a distributed computing software platform implementation of the MapReduce framework, running on a small cluster of machines locally. Additionally, the software is implemented to use Amazon's Elastic Compute Cloud (EC2) facility. We also developed a hybrid solution to serve images to users by leveraging cloud storage using Amazon's Simple Storage Service (S3) for public data while keeping private information on our own data servers. By using Cloud Computing, we improve upon our local solution by reducing the need to manage our own hardware and computing infrastructure, thereby reducing costs. Further, by using a hybrid of local and cloud storage, we are able to provide data to our users more efficiently and securely. 12 This paper examines the use of a distributed approach with Hadoop to tile images, an approach that provides significant improvements in image processing time, from hours to minutes. This paper describes the constraints imposed on the solution and the resulting techniques developed for the hybrid solution of a customized Hadoop infrastructure over local and cloud resources in managing this ever-growing data set. It examines the performance trade-offs of using the more plentiful resources of the cloud, such as those provided by S3, against the bandwidth limitations such use encounters with remote resources. As part of this discussion this paper will outline some of the technologies employed, the reasons for their selection, the resulting performance metrics and the direction the project is headed based upon the demonstrated capabilities thus far.

Cloud Computing

Acquisition of and Access to Research Omics Data

Omics data are essential for understanding the myriad and complex effects of space environments on humans. To assure maximum benefit from these kinds of data, the NASA Human Research Program Data Management Plan stipulates that human omics data should be archived within and accessed through the NASA Life Sciences Portal (NLSP). The NLSP has the capability to acquire and provision access to omics (and other kinds of) research results for individual and ad-hoc groups of subjects at the direction of institutional review boards, or other authorizing bodies or individuals, per institutional, program and investigation-specific policies and procedures. However, because some single-subject omics data, like CT scans and other kinds of large, complex biomedical data, could be used to identify heretofore unknown risks to the subject’s health, or, in certain cases, be used to identify a subject, NASA Policy Directive 7170.1 describes various policies regarding the management of and access to “research genetic testing” data, which includes many kinds of omics data. For example, NPD 7170.1 prohibits access to human research genetic data by NASA personnel who make employment decisions for the subjects from whom the data were obtained. To meet the objective of acquiring research omics data for NLSP in compliance with the policies in NPD 7170.1 and other applicable NASA policies, we designed NOMADS (the NLSP Omics Multimodal Acquisition of Data System), a new component that supports the transfer of large research data files, including research genetic testing data, using one of several different transfer mechanisms. The choice of mechanism is made by the submitter of the data, with guiding information from the system, and is likely to often be determined in large part by the nature and source location of the data. For example, for small files where the source data files are not already stored in a cloud storage system, users are likely to prefer to transfer their data to the NLSP via a web browser. Conversely, for large sets of files already organized and stored in a cloud storage system, users may opt for NOMAD’s cloud-to-cloud transfer method. All omics datasets targeted for the NASA Life Sciences Data Archive must pass a variety of quality checks to ensure data integrity and adherence to the standards defined by the LSDA Data Submission Guidelines (DSG) (see https://nlsp.nasa.gov/explore/lsdahome/datasubmit). These include requirements that data are consistent with open standards established by the omics community. Non-compliant data will not be accepted however archivists are available to advise submitters on how to revise data submissions and re-submit until compliance is achieved. Following compliance with the LSDA DSG, omics data next undergo a variety of additional quality checks to ensure the data meet omics community standards. Domain specific Omics data quality control tools and techniques are continually evolving and linked to the advancements in omics assays utilized and thus, the tools and techniques utilized by the LSDA for data quality control and validation will need to be sustained accordingly. All human omics data will be access controlled according to the policies described above, and requiring IRB approval for any additional access grants once the data are acquired (including access for analysis using the NLSP workspace tools).

Omics

Acquisition of and Access to Research Omics Data

Omics data are essential for understanding the myriad and complex effects of space environments on humans. To assure maximum benefit from these kinds of data, the NASA Human Research Program Data Management Plan stipulates that human omics data should be archived within and accessed through the NASA Life Sciences Portal (NLSP). The NLSP has the capability to acquire and provision access to omics (and other kinds of) research results for individual and ad-hoc groups of subjects at the direction of institutional review boards, or other authorizing bodies or individuals, per institutional, program and investigation-specific policies and procedures. However, because some single-subject omics data, like CT scans and other kinds of large, complex biomedical data, could be used to identify heretofore unknown risks to the subject’s health, or, in certain cases, be used to identify a subject, NASA Policy Directive 7170.1 describes various policies regarding the management of and access to “research genetic testing” data, which includes many kinds of omics data. For example, NPD 7170.1 prohibits access to human research genetic data by NASA personnel who make employment decisions for the subjects from whom the data were obtained. To meet the objective of acquiring research omics data for NLSP in compliance with the policies in NPD 7170.1 and other applicable NASA policies, we designed NOMADS (the NLSP Omics Multimodal Acquisition of Data System), a new component that supports the transfer of large research data files, including research genetic testing data, using one of several different transfer mechanisms. The choice of mechanism is made by the submitter of the data, with guiding information from the system, and is likely to often be determined in large part by the nature and source location of the data. For example, for small files where the source data files are not already stored in a cloud storage system, users are likely to prefer to transfer their data to the NLSP via a web browser. Conversely, for large sets of files already organized and stored in a cloud storage system, users may opt for NOMAD’s cloud-to-cloud transfer method. All omics datasets targeted for the NASA Life Sciences Data Archive must pass a variety of quality checks to ensure data integrity and adherence to the standards defined by the LSDA Data Submission Guidelines (DSG) (see https://nlsp.nasa.gov/explore/lsdahome/datasubmit). These include requirements that data are consistent with open standards established by the omics community. Non-compliant data will not be accepted however archivists are available to advise submitters on how to revise data submissions and re-submit until compliance is achieved. Following compliance with the LSDA DSG, omics data next undergo a variety of additional quality checks to ensure the data meet omics community standards. Domain specific Omics data quality control tools and techniques are continually evolving and linked to the advancements in omics assays utilized and thus, the tools and techniques utilized by the LSDA for data quality control and validation will need to be sustained accordingly. All human omics data will be access controlled according to the policies described above, and requiring IRB approval for any additional access grants once the data are acquired (including access for analysis using the NLSP workspace tools).

Omics

Data Management in the Continuum: Cross-facility Object-based Data Transfers

Scientific workflows are evolving from relying on a monolithic storage subsystem at a single High-Performance Computing (HPC) facility to using geographically distributed file systems, repositories, and cloud storage. As a result, storing, accessing, transferring, and managing scientific data have become highly complex and prone to performance inefficiencies. This paper delves into these challenges by exploring an optimized end-to-end interface designed to seamlessly connect various local and remote storage systems, enabling efficient data movement of objects across HPC–Cloud and HPC–HPC environments. We showcase this capability through an object-focused data management runtime system, discuss the effects of relaxed consistency semantics in distributed object scenarios, and illustrate its application in an earthquake simulation workflow. Besides reducing the amount of data by selectively transferring regions of interest, our facility-local results achieved a speedup of 45 × over an optimized HDF5 usage and 15 × over the HDF5 with caching by using the new interface in PDC-XF.

Bez, Jean Luca

SatCORPS Hybridized Cloud Product Data Storage: The Design of a Hybrid Data Repository That Leverages the Strengths of the Cloud and the Data Center

There is a strong demand for the near real time NASA Langley Satellite ClOud and Radiation Property retrieval System (SatCORPS) products. As important as real-time information, archived copies of the products form the basis for targeted research focusing on specific events or conditions. To make these SatCORPS products available for downloading, the SatCORPS group has developed a number of tools and technologies to create a hybrid data storage system that leverages the strengths of both cloud and on-premises resources. In this work, we describe the technologies the group uses to marshal disparate data repositories and materialize them into a single searchable overview and give a broad description of the organization of the dataset. As with any implementation, the strengths, weaknesses and constraints surrounding the components establish priorities and provide insight where trade-offs are necessary. We further describe the design and architecture underpinning our hybrid data repository and delivery system.

AWS

Datum: A Scientific Metadata Catalog

The data catalog market is currently flooded with a myriad of different products, but none serve the scientific community well. There are cloud-native tools like Databricks, Snowflake,to on-premise solutions like Collibra and Datahub. The common failing of all these tools however, is their inability to serve the scientific data community directly. Most catalogs are targeted towards financial, health, or user data - not sensor or scientific domain data. They also prioritize integrations that often don’t exist or are just starting to be used in the scientific realm - all while ignoring common scientific tools and file types. Datum is a catalog which targets the scientific data directly, including the tools and networks in which those tools are used. We work with the producers and consumers of the data where they are, targeting cloud and on-premise with a focus on classified networks. Datum is an Erlang/Elixir application. Technical Features Note: The features listed below are still under development and may change, slightly, upon final delivery of the product. File Formats - Datum has the ability to read additional metadata and provides processing pipelines for the following file formats: Plain Text, PDF, LaTeX, HTML, Open Document Format (.odt), XML, CSV/TSV (and other standard delimiters), OpenDocument Database and Spreadsheets, Geo-Referenced TIFF, Common Data Format, HDF/HDF5, LabView TDMS, Excel, DeltaTables, Parquet, Apache Iceberg, Apache Hudi and many others. Metadata Collection - Scanners for the local and networked file systems and cloud storage providers. Network integration with common databases such as MSSQL and MySQL. User Plugin System - Users are able to provide either file processing, metadata extraction, or sampling plugins in the programming language of their choice. Authentication/Authorization -: OIDC integration, SCIM provisioning and EntraID integration out of the box. Full user and group management system with a “least privilege” operating mode. Governance - Customizable data governance platform; dictate and enforce required metadata, enforce data embargos, and enforce user agreements and NDAs before data access. Ability to create health checks on data, rejecting abandoned or poorly curated data and automatically removing it from the search index. Ability for users to submit corrections. Search - Semantic search is a first class citizen. No licenses to expensive, external software required. Integrated use of vectors and vector-based search allows for AI agent integration at all levels of operation. Metadata Model - Display and control data’s lineage and connections to other data and data directories. Data is modeled after a filesystem - an organization instantly recognizable and navigable by most any user. CLI and SDK - Ships with a Command Line Interface (CLI) tool and with a fully-featured Python SDK. This allows for rapid and programmatic use of Datum by every level of user. Minimal Infrastructure - Datum ships as a single executable file and can be run on any operating system and most CPU architectures. Datum has no reliance on external databases, search indexing tools, or other outside services - and it runs equally well on edge computing devices, cloud services, or in a clustered HPC environment.

darrington, john

NASA Tech Briefs, April 2013

Topics covered include: Fully Integrated, Miniature, High-Frequency Flow Probe Utilizing MEMS Leadless SOI Technology; Nanoscale Surface Plasmonics Sensor With Nanofluidic Control; Advanced Dispersed Fringe Sensing Algorithm for Coarse Phasing Segmented Mirror Telescopes; Neural Network Back-Propagation Algorithm for Sensing Hypergols; Bulk Moisture and Salinity Sensor; Change-Based Satellite Monitoring Using Broad Coverage and Targetable Sensing; Circularly Polarized Microwave Antenna Element with Very Low Off-Axis Cross-Polarization; Ultra-Low Heat-Leak, High-Temperature Superconducting Current Leads for Space Applications; Flash Cracking Reactor for Waste Plastic Processing; An Automated Safe-to-Mate (ASTM) Tester; Wireless Chalcogenide Nanoionic-Based Radio-Frequency Switch; Compute Element and Interface Box for the Hazard Detection System; DOT Transmit Module; Composite Aerogel Multifoil Protective Shielding; Li-Ion Electrolytes with Improved Safety and Tolerance to High-Voltage Systems; Polymer-Reinforced, Non-Brittle, Lightweight Cryogenic Insulation; Controlled, Site-Specific Functionalization of Carbon Nanotubes with Diazonium Salts; Regenerable Sorbent for CO2 Removal; Sprayable Aerogel Bead Compositions With High Shear Flow Resistance and High Thermal Insulation Value; Lexan Linear Shaped Charge Holder with Magnets and Backing Plate; Robotic Ankle for Omnidirectional Rock Anchors; Wind, Wave, and Tidal Energy Without Power Conditioning; An Active Heater Control Concept to Meet IXO Type Mirror Module Thermal-Structural Distortion Requirement; Waterless Clothes-Cleaning Machine; Integrated Electrical Wire Insulation Repair System; LVGEMS Time-of-Flight Mass Spectrometry on Satellites; Surface Inspection Tool for Optical Detection of Surface Defects; Per-Pixel, Dual-Counter Scheme for Optical Communications; Certification-Based Process Analysis; Surface Navigation Using Optimized Waypoints and Particle Swarm Optimization; Smart-Divert Powered Descent Guidance to Avoid the Backshell Landing Dispersion Ellipse; Estimating Foreign-Object-Debris Density from Photogrammetry Data; Adaptive Sampling of Spatiotemporal Phenomena with Optimization Criteria; Building a 2.5D Digital Elevation Model From 2D Imagery; Eyes on the Earth 3D; Target Trailing With Safe Navigation for Maritime Autonomous Surface Vehicles; Adams-Based Rover Terramechanics and Mobility Simulator - ARTEMIS; ISTP CDF Skeleton Editor; Uplink Summary Generator (ULSGEN) Version 1.0; Robotics On-Board Trainer (ROBoT); Software Engineering Tools for Scientific Models; Automatic Data Filter Customization Using a Genetic Algorithm; Tracker Toolkit; Towards Efficient Scientific Data Management Using Cloud Storage; On a Formal Tool for Reasoning About Flight Software Cost Analysis; A Nanostructured Composites Thermal Switch Controls Internal and External Short Circuit in Lithium Ion Batteries; Spacecraft Crew Cabin Condensation Control; and Functional Near-Infrared Spectroscopy Signals Measure Neuronal Activity in the Cortex.

Source record

OPeNDAP Clients, Aggregation and S3

In this talk, we will discuss our work for testing OPeNDAP client access of data stored in the Amazon S3 cloud storage using a set of common analysis tools including Panoply, Jupyter Notebooks with Python xarray, NCO command line tool package, ArcGIS, and GDAL. We will also discuss our ongoing work on improving performance in Hyrax aggregation functionality.

Amazon S3

Climates of Warm Earth-Like Planets III: Fractional Habitability from a Water Cycle Perspective

The habitable fraction of a planet's surface is important for the detectability of surface biosignatures. The extent and distribution of habitable areas are influenced by external parameters that control the planet's climate, atmospheric circulation, and hydrological cycle. We explore these issues using the ROCKE-3D general circulation model, focusing on terrestrial water fluxes and thus the potential for the existence of complex life on land. Habitability is examined as a function of insolation and planet rotation for an Earth-like world with zero obliquity and eccentricity orbiting the Sun. We assess fractional habitability using an aridity index that measures the net supply of water to the land. Earth-like planets become "superhabitable" (a larger habitable surface area than Earth) as insolation and day-length increase because their climates become more equable, reminiscent of past warm periods on Earth when complex life was abundant and widespread. The most slowly rotating, most highly irradiated planets, though, occupy a hydrological regime unlike any on Earth, with extremely warm, humid conditions at high latitudes but little rain and subsurface water storage. Clouds increasingly obscure the surface as insolation increases, but visibility improves for modest increases in rotation period. Thus, moderately slowly rotating rocky planets with insolation near or somewhat greater than modern Earth's appear to be promising targets for surface characterization by a future direct imaging mission.

Del Genio, Anthony D.

Climates of Warm Earth-like Planets. III. Fractional Habitability from a Water Cycle Perspective

The habitable fraction of a planet's surface is important for the detectability of surface biosignatures. The extent and distribution of habitable areas are influenced by external parameters that control the planet's climate, atmospheric circulation, and hydrological cycle. We explore these issues using the ROCKE-3D general circulation model, focusing on terrestrial water fluxes and thus the potential for the existence of complex life on land. Habitability is examined as a function of insolation and planet rotation for an Earth-like world with zero obliquity and eccentricity orbiting the Sun. We assess fractional habitability using an aridity index that measures the net supply of water to the land. Earth-like planets become "superhabitable" (a larger habitable surface area than Earth) as insolation and day-length increase because their climates become more equable, reminiscent of past warm periods on Earth when complex life was abundant and widespread. The most slowly rotating, most highly irradiated planets, though, occupy a hydrological regime unlike any on Earth, with extremely warm, humid conditions at high latitudes but little rain and subsurface water storage. Clouds increasingly obscure the surface as insolation increases, but visibility improves for modest increases in rotation period. Thus, moderately slowly rotating rocky planets with insolation near or somewhat greater than modern Earth's appear to be promising targets for surface characterization by a future direct imaging mission.

Exoplanet atmospheric variability

Leveraging Cloud Computing to Improve Storage Durability, Availability, and Cost for MER Maestro

The Maestro for MER (Mars Exploration Rover) software is the premiere operation and activity planning software for the Mars rovers, and it is required to deliver all of the processed image products to scientists on demand. These data span multiple storage arrays sized at 2 TB, and a backup scheme ensures data is not lost. In a catastrophe, these data would currently recover at 20 GB/hour, taking several days for a restoration. A seamless solution provides access to highly durable, highly available, scalable, and cost-effective storage capabilities. This approach also employs a novel technique that enables storage of the majority of data on the cloud and some data locally. This feature is used to store the most recent data locally in order to guarantee utmost reliability in case of an outage or disconnect from the Internet. This also obviates any changes to the software that generates the most recent data set as it still has the same interface to the file system as it did before updates

Chang, George W.

Utilizing HDF4 File Content Maps for the Cloud

We demonstrate a prototype study that HDF4 file content map can be used for efficiently organizing data in cloud object storage system to facilitate cloud computing. This approach can be extended to any binary data formats and to any existing big data analytics solution powered by cloud computing because HDF4 file content map project started as long term preservation of NASA data that doesn't require HDF4 APIs to access data.

Elastic Search

Notes on a storage manager for the Clouds kernel

The Clouds project is research directed towards producing a reliable distributed computing system. The initial goal is to produce a kernel which provides a reliable environment with which a distributed operating system can be built. The Clouds kernal consists of a set of replicated subkernels, each of which runs on a machine in the Clouds system. Each subkernel is responsible for the management of resources on its machine; the subkernal components communicate to provide the cooperation necessary to meld the various machines into one kernel. The implementation of a kernel-level storage manager that supports reliability is documented. The storage manager is a part of each subkernel and maintains the secondary storage residing at each machine in the distributed system. In addition to providing the usual data transfer services, the storage manager ensures that data being stored survives machine and system crashes, and that the secondary storage of a failed machine is recovered (made consistent) automatically when the machine is restarted. Since the storage manager is part of the Clouds kernel, efficiency of operation is also a concern.

Pitts, David V.

Electron cloud predictions for the Hadron Storage Ring of the Electron-Ion Collider and planned mitigations

This paper reports on a collection of electron cloud studies to determine the electron cloud threshold for different sections along the beampipe of the Hadron Storage Ring (HSR) for the Electron-Ion Collider (EIC), presents the results of a study of the interaction of the beam with the electron clouds, and discusses the limitation of potential solutions like scrubbing and Landau damping.

43 PARTICLE ACCELERATORS

Using Big Data Technologies with Earth Science Data in HDF5: HDF5 Scalable Solutions

HDF5 (Hierarchical Data Format 5) is open-source, high-performance software that consists of an abstract data model, library, and fileformat used for storing and managing extremely large and/or complex data collections. NASA Earth Observing System (EOS) Data and Information Systems use HDF5 as an archival format to store remote sensing data from EOS satellites. HDF5 is also used to store other types of Geoscience and Strophysical data, e.g., seismic data and data from Low-Frequency Array (LOFAR) radio telescopes. Data stored in HDF5 has reached tens of petabytes and is growing at an accelerated rate.With the growing amout of HDF5 Earth Science data to analyze and process, scientists need to adopt big data technologies including new storage paradigms such as cloud and object storage. To run models and perform data analysis they also need to utilizied efficient and diverse ways to access data, from high-performance computing's (HPC) Message Passing Interface (MPI) I/O and deep memory hierarchies (DMH) to non-HPC frameworks such as Apache Hadoop, Spark, and Drill. The HDF Group continually works to enable usage of big data technologies in HDF software.

Knox, Larry

The Third International Cloud Condensation Nuclei Workshop

Twenty-five instruments were tested, including size characterization devices and two Aitken counters. The test aerosols were supplied to the instruments by an on-line generation system, thereby eliminating the need for storage bags. Cloud condensation chambers and haze chambers are highlighted.

Kocmond, W. C.