Search NASASearch

SEARCH · Search NASA

Results for “Data migration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Storage system architectures and their characteristics

Not all users storage requirements call for 20 MBS data transfer rates, multi-tier file or data migration schemes, or even automated retrieval of data. The number of available storage solutions reflects the broad range of user requirements. It is foolish to think that any one solution can address the complete range of requirements. For users with simple off-line storage requirements, the cost and complexity of high end solutions would provide no advantage over a more simple solution. The correct answer is to match the requirements of a particular storage need to the various attributes of the available solutions. The goal of this paper is to introduce basic concepts of archiving and storage management in combination with the most common architectures and to provide some insight into how these concepts and architectures address various storage problems. The intent is to provide potential consumers of storage technology with a framework within which to begin the hunt for a solution which meets their particular needs. This paper is not intended to be an exhaustive study or to address all possible solutions or new technologies, but is intended to be a more practical treatment of todays storage system alternatives. Since most commercial storage systems today are built on Open Systems concepts, the majority of these solutions are hosted on the UNIX operating system. For this reason, some of the architectural issues discussed focus around specific UNIX architectural concepts. However, most of the architectures are operating system independent and the conclusions are applicable to such architectures on any operating system.

Sarandrea, Bryan M.

dCache: The Storage System of Choice for Data-Intensive Applications

The ever-increasing volumes of data produced by modern scientific facilities like EuXFEL and LHC put significant stress on data management infrastructure operated by laboratories and research centers. The challenges to be addressed span the entire data life cycle, from ingest and efficient data analysis to long-term preservation, typically involving large tape libraries. dCache, a storage system developed in collaboration between the Deutsches Elektronen-Synchrotron (DESY), Fermi National Accelerator Laboratory, and Nordic e-Infrastructure Collaboration (NeIC), is designed to manage a large number of disk servers and to facilitate transparent data migration to and from archival storage. Its multifaceted approach offers a unified method to support a variety of scientific use cases with the same storage infrastructure, including high-throughput data ingest, data sharing over wide area networks, efficient access from HPC clusters, and long-term data preservation on tertiary storage. Initially developed for high energy physics (HEP) experiments, dCache is now used by various scientific communities, including astrophysics, biomedical research, and life sciences, each having specific requirements. This paper presents architecture, deployment strategies, performance and scalability enhancements, and recent advancements in dCache addressing the needs of scientific communities. Finally, we touch on the development and release process, ensuring the software’s high quality.

DCache

Optimal Sensor Mobility Design for Target Tracking with Distributed Sensing, Communication and Computing Infrastructure

The paper presents an airborne target tracking approach with data fusion from distributed stationary and mobile sensor network transmitted through a wireless/wired communication network and using a distributed computing infrastructure. To obtain high quality measurements, sensor flight platforms' positions and orientations are optimized with respect to the available target's position and velocity estimates, and a formation control strategy is applied to derive desired trajectories for the flight platforms. As sensors move, the communication network topology switching is determined using the worst-case approach to handle uncertainties in the target's and sensors' positions, which alters communication delays for sensors' data transmission to computing centers. An optimal data migration algorithm is periodically applied to determine these communication delays for each sensor and the optimal location with minimum end-to-end latency for the target tracking algorithm execution, which is based on adaptive information fusion from multi-modal mobile and stationary sensors inside an Extended Kalman Filter framework. The approach is validated in a desktop simulation environment using synthetic sensor data generated for a simulated target's flight.

Distributed sensing

Using Docker Containers to Extend Reproducibility Architecture for the NASA Earth Exchange (NEX)

NASA Earth Exchange (NEX) is a data, supercomputing and knowledge collaboratory that houses NASA satellite, climate and ancillary data where a focused community can come together to address large-scale challenges in Earth sciences. As NEX has been growing into a petabyte-size platform for analysis, experiments and data production, it has been increasingly important to enable users to easily retrace their steps, identify what datasets were produced by which process chains, and give them ability to readily reproduce their results. This can be a tedious and difficult task even for a small project, but is almost impossible on large processing pipelines. We have developed an initial reproducibility and knowledge capture solution for the NEX, however, if users want to move the code to another system, whether it is their home institution cluster, laptop or the cloud, they have to find, build and install all the required dependencies that would run their code. This can be a very tedious and tricky process and is a big impediment to moving code to data and reproducibility outside the original system. The NEX team has tried to assist users who wanted to move their code into OpenNEX on Amazon cloud by creating custom virtual machines with all the software and dependencies installed, but this, while solving some of the issues, creates a new bottleneck that requires the NEX team to be involved with any new request, updates to virtual machines and general maintenance support. In this presentation, we will describe a solution that integrates NEX and Docker to bridge the gap in code-to-data migration. The core of the solution is saemi-automatic conversion of science codes, tools and services that are already tracked and described in the NEX provenance system, to Docker - an open-source Linux container software. Docker is available on most computer platforms, easy to install and capable of seamlessly creating and/or executing any application packaged in the appropriate format. We believe this is an important step towards seamless process deployment in heterogeneous environments that will enhance community access to NASA data and tools in a scalable way, promote software reuse, and improve reproducibility of scientific results.

earth exchange

olcf/hsi_xfer

A wrapper for the `hsi` tool that provides a simple user interface to allow the most efficient and least stressful (to the system) way to migrate data off of HPSS filesystems

Wynne, JamesRiley [Oak Ridge National Laboratory (

State of the art survey of network operating systems development

The results of the State-of-the-Art Survey of Network Operating Systems (NOS) performed for Goddard Space Flight Center are presented. NOS functional characteristics are presented in terms of user communication data migration, job migration, network control, and common functional categories. Products (current or future) as well as research and prototyping efforts are summarized. The NOS products which are revelant to the space station and its activities are evaluated.

Source record

An analysis of file migration in a UNIX supercomputing environment

The super computer center at the National Center for Atmospheric Research (NCAR) migrates large numbers of files to and from its mass storage system (MSS) because there is insufficient space to store them on the Cray supercomputer's local disks. This paper presents an analysis of file migration data collected over two years. The analysis shows that requests to the MSS are periodic, with one day and one week periods. Read requests to the MSS account for the majority of the periodicity; as write requests are relatively constant over the course of a week. Additionally, reads show a far greater fluctuation than writes over a day and week since reads are driven by human users while writes are machine-driven.

Miller, Ethan L.

Volume serving and media management in a networked, distributed client/server environment

The E-Systems Modular Automated Storage System (EMASS) is a family of hierarchical mass storage systems providing complete storage/'file space' management. The EMASS volume server provides the flexibility to work with different clients (file servers), different platforms, and different archives with a 'mix and match' capability. The EMASS design considers all file management programs as clients of the volume server system. System storage capacities are tailored to customer needs ranging from small data centers to large central libraries serving multiple users simultaneously. All EMASS hardware is commercial off the shelf (COTS), selected to provide the performance and reliability needed in current and future mass storage solutions. All interfaces use standard commercial protocols and networks suitable to service multiple hosts. EMASS is designed to efficiently store and retrieve in excess of 10,000 terabytes of data. Current clients include CRAY's YMP Model E based Data Migration Facility (DMF), IBM's RS/6000 based Unitree, and CONVEX based EMASS File Server software. The VolSer software provides the capability to accept client or graphical user interface (GUI) commands from the operator's console and translate them to the commands needed to control any configured archive. The VolSer system offers advanced features to enhance media handling and particularly media mounting such as: automated media migration, preferred media placement, drive load leveling, registered MediaClass groupings, and drive pooling.

Herring, Ralph H.

Use of HSM with Relational Databases

Hierarchical storage management (HSM) systems have evolved to become a critical component of large information storage operations. They are built on the concept of using a hierarchy of storage technologies to provide a balance in performance and cost. In general, they migrate data from expensive high performance storage to inexpensive low performance storage based on frequency of use. The predominant usage characteristic is that frequency of use is reduced with age and in most cases quite rapidly. The result is that HSM provides an economical means for managing and storing massive volumes of data. Inherent in HSM systems is system managed storage, where the system performs most of the work with minimum operations personnel involvement. This automation is generally extended to include: backup and recovery, data duplexing to provide high availability, and catastrophic recovery through use of off-site storage.

Breeden, Randall

A New Approach to Parallel Dynamic Partitioning for Adaptive Unstructured Meshes

Classical mesh partitioning algorithms were designed for rather static situations, and their straightforward application in a dynamical framework may lead to unsatisfactory results, e.g., excessive data migration among processors. Furthermore, special attention should be paid to their amenability to parallelization. In this paper, a novel parallel method for the dynamic partitioning of adaptive unstructured meshes is described. It is based on a linear representation of the mesh using self-avoiding walks.

Heber, Gerd

Beyond a Terabyte File System

The Numerical Aerodynamics Simulation Facility's (NAS) CRAY C916/1024 accesses a "virtual" on-line file system, which is expanding beyond a terabyte of information. This paper will present some options to fine tuning Data Migration Facility (DMF) to stretch the online disk capacity and explore the transitions to newer devices (STK 4490, ER90, RAID).

Powers, Alan K.

Parallel Processing of Adaptive Meshes with Load Balancing

Many scientific applications involve grids that lack a uniform underlying structure. These applications are often also dynamic in nature in that the grid structure significantly changes between successive phases of execution. In parallel computing environments, mesh adaptation of unstructured grids through selective refinement/coarsening has proven to be an effective approach. However, achieving load balance while minimizing interprocessor communication and redistribution costs is a difficult problem. Traditional dynamic load balancers are mostly inadequate because they lack a global view of system loads across processors. In this paper, we propose a novel and general-purpose load balancer that utilizes symmetric broadcast networks (SBN) as the underlying communication topology, and compare its performance with a successful global load balancing environment, called PLUM, specifically created to handle adaptive unstructured applications. Our experimental results on an IBM SP2 demonstrate that the SBN-based load balancer achieves lower redistribution costs than that under PLUM by overlapping processing and data migration.

Das, Sajal K.

Experiences From NASA/Langley's DMSS Project

There is a trend in institutions with high performance computing and data management requirements to explore mass storage systems with peripherals directly attached to a high speed network. The Distributed Mass Storage System (DMSS) Project at the NASA Langley Research Center (LaRC) has placed such a system into production use. This paper will present the experiences, both good and bad, we have had with this system since putting it into production usage. The system is comprised of: 1) National Storage Laboratory (NSL)/UniTree 2.1, 2) IBM 9570 HIPPI attached disk arrays (both RAID 3 and RAID 5), 3) IBM RS6000 server, 4) HIPPI/IPI3 third party transfers between the disk array systems and the supercomputer clients, a CRAY Y-MP and a CRAY 2, 5) a "warm spare" file server, 6) transition software to convert from CRAY's Data Migration Facility (DMF) based system to DMSS, 7) an NSC PS32 HIPPI switch, and 8) a STK 4490 robotic library accessed from the IBM RS6000 block mux interface. This paper will cover: the performance of the DMSS in the following areas: file transfer rates, migration and recall, and file manipulation (listing, deleting, etc.); the appropriateness of a workstation class of file server for NSL/UniTree with LaRC's present storage requirements in mind the role of the third party transfers between the supercomputers and the DMSS disk array systems in DMSS; a detailed comparison (both in performance and functionality) between the DMF and DMSS systems LaRC's enhancements to the NSL/UniTree system administration environment the mechanism for DMSS to provide file server redundancy the statistics on the availability of DMSS the design and experiences with the locally developed transparent transition software which allowed us to make over 1.5 million DMF files available to NSL/UniTree with minimal system outage

Source record

Applications of MERRA-2 data for avian migration, biomass burning, and dusty atmospheric rivers

Three different applications of MERRA-2 data are presented. 1) Using radar data, we introduced a new concept for spatial patterns of bird migration across the contiguous U.S. This approach allowed us to use MERRA-2 data and learn that remote forcing in the tropical Pacific—through a chain of processes including atmospheric Rossby wave trains— controls the climatic conditions, associated with bird migration in North America. 2) We showed that emissions from biomass burning in the Congo Basin are partly controlled by the low-level winds, which are in turn associated with the intensity of the subtropical high in the Indian Ocean. Using back-trajectory analysis, we found that these emissions combined with their transport mechanism explain the interannual variability of black carbon in West Africa. 3) Our analysis showed that atmospheric rivers in the Middle East contribute to both heavy flood and dust transport within their corridor. We also found that warm advection and rain-on-snow effect of dusty atmospheric rivers further enhance the chance of flood through rapid snowmelt processes.

Amin Dezfuli

Applications of MERRA-2 Data for Avian Migration, Biomass Burning, and Dusty Atmospheric Rivers

Three different applications of MERRA-2 data are presented. 1) Using radar data, we introduced a new concept for spatial patterns of bird migration across the contiguous U.S. This approach allowed us to use MERRA-2 data and learn that remote forcing in the tropical Pacific—through a chain of processes including atmospheric Rossby wave trains— controls the climatic conditions, associated with bird migration in North America. 2) We showed that emissions from biomass burning in the Congo Basin are partly controlled by the low-level winds, which are in turn associated with the intensity of the subtropical high in the Indian Ocean. Using back-trajectory analysis, we found that these emissions combined with their transport mechanism explain the interannual variability of black carbon in West Africa. 3) Our analysis showed that atmospheric rivers in the Middle East contribute to both heavy flood and dust transport within their corridor. We also found that warm advection and rain-on-snow effect of dusty atmospheric rivers further enhance the chance of flood through rapid snowmelt processes.

Amin Dezfuli

Hydrological Data at the NASA GES DISC: Current Capabilities and New Opportunities

The NASA Goddard Earth Sciences Data and Information Services Center (GES DISC) is one of twelve NASA Earth science data centers that document, process, archive and distribute data from Earth observation missions and projects. GES DISC maintains an archive of several hydrology datasets, including the Land Data Assimilation Systems (LDAS) and the Gravity Recovery and Climate Experiment (GRACE) Data Assimilation for Drought Monitoring (GRACE-DA-DM) data products. These datasets include model output of heat fluxes, rain, snow, soil temperature, soil moisture, and runoff; and observational forcing data, including surface pressure, temperature, precipitation, downward shortwave and longwave radiation, humidity, and wind. The temporal resolution of the hydrology data at GES DISC ranges from hourly to monthly, and spatial resolutions range from 0.1° to 1.0°. The GES DISC provides services which enable users to aggregate, temporally and spatially subset, regrid, and visualize archived data including the GES DISC Subsetter, Hydrology Data Rods, and the Geospatial Interactive Online Visualization and Analysis Infrastructure (GIOVANNI). The Hydrology Data Rods service optimally reorganizes large hydrological data sets as extended time series, providing more efficient access for the hydrological community. The time series data (aka “data rods”) were integrated into hydrology community tools, such as the Data Rods Explorer on HydroShare. Furthermore, the GES DISC is in the process of migrating its data and services to the cloud. Hydrological data available at the GES DISC are now available in the Amazon Web Services (AWS) cloud (us-west-2 region) providing users Direct S3 data access and the capability for cloud computing operations. In this presentation, the hydrology data products and services currently available at the GES DISC will be summarized. Also discussed are the migration to the cloud, user support through this transition, and the status of migrating the data rods service to the cloud.

Ashley Heath

A Note on Interfacing Object Warehouses and Mass Storage Systems for Data Mining Applications

Data mining is the automatic discovery of patterns, associations, and anomalies in data sets. Data mining requires numerically and statistically intensive queries. Our assumption is that data mining requires a specialized data management infrastructure to support the aforementioned intensive queries, but because of the sizes of data involved, this infrastructure is layered over a hierarchical storage system. In this paper, we discuss the architecture of a system which is layered for modularity, but exploits specialized lightweight services to maintain efficiency. Rather than use a full functioned database for example, we use light weight object services specialized for data mining. We propose using information repositories between layers so that components on either side of the layer can access information in the repositories to assist in making decisions about data layout, the caching and migration of data, the scheduling of queries, and related matters.

Grossman, Robert L.

Technology Assessment of High Capacity Data Storage Systems: Can We Avoid a Data Survivability Crisis?

In a recent address at the California Science Center in Los Angeles, Vice President Al Gore articulated a Digital Earth Vision. That vision spoke to developing a multi-resolution, three-dimensional visual representation of the planet into which we can roam and zoom into vast quantities of embedded geo-referenced data. The vision was not limited to moving through space, but also allowing travel over a time-line, which can be set for days, years, centuries, or even geological epochs. A working group of Federal Agencies, developing a coordinated program to implement the Vice President's vision, developed the definition of the Digital Earth as a visual representation of our planet that enables a person to explore and interact with the vast amounts of natural and cultural geo-referenced information gathered about the Earth. One of the challenges identified by the agencies was whether the technology existed that would be available to permanently store and deliver all the digital data that enterprises might want to save for decades and centuries. Satellite digital data is growing by Moore's Law as is the growth of computer generated data. Similarly, the density of digital storage media in our information-intensive society is also increasing by a factor of four every three years. The technological bottleneck is that the bandwidth for transferring data is only growing at a factor of four every nine years. This implies that the migration of data to viable long-term storage is growing more slowly. The implication is that older data stored on increasingly obsolete media are at considerable risk if they cannot be continuously migrated to media with longer life times. Another problem occurs when the software and hardware systems for which the media were designed are no longer serviced by their manufacturers. Many instances exist where support for these systems are phased out after mergers or even in going out of business. In addition, survivability of older media can suffer from physical breakdown of components (e.g. tapes simply lose their magnetic properties after a long time in storage). As a result, a potential data survivability crisis is emerging. The scale of the crisis is comparable to that facing the Social Security System. Sometime in one or two decades, the exponential growth of data will become so great that many enterprises will not be able to migrate through their data to more permanent media during the lifetime of the media on which it resides. This will result in significant losses of data and their resultant impacts. To avoid this crisis, we need to plan and devote greater financial and intellectual resources are needed for the development and refinement of new storage media and migration technologies in order to preserve all data any organization determines worth saving permanently. This talk will explore technological solutions and suggested recommendations to address this technological data crisis.

Halem, Milton