Search NASA⌕ Search

SEARCH · Search NASA

Results for “Archive”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Data archive for NO(y) from observations and construction and testing of airborne instrument for simultaneous measurement of NO, NO2, NO(y), and O3

The compilation and archiving of NO(x) and NO(y) measurements began in mid-March 1994. Since the submission of the first report, data summaries have been obtained for the TROPOZ 2, STRATOZ 3, OCTA and TOR/Schauinsland campaigns, and the full data sets will become a part of this archive in the near future. Climatologies of NO(x) and NO(y) have been developed from these and previously archived data sets, including the available GTE campaigns (ABLE-2A, B, -3A, B, CITE-2, -3, TRACE-A, PEM WEST-A) and AASE 1 and 2. The data have been grouped by season and altitude (boundary layer and 3 km ranges in the free troposphere). Maps showing median values of midday NO, NO(x) and NO(y) have been produced for each season for the boundary layer and 3 km ranges of the free troposphere. The statistics of the data (median, mean, and standard deviation, central 67% and 90%) have also been determined, and are shown in representative figures included in this report.

Carroll, Mary Anne↗

Analysis of the access patterns at GSFC distributed active archive center

The Goddard Space Flight Center (GSFC) Distributed Active Archive Center (DAAC) has been operational for more than two years. Its mission is to support existing and pre Earth Observing System (EOS) Earth science datasets, facilitate the scientific research, and test Earth Observing System Data and Information System (EOSDIS) concepts. Over 550,000 files and documents have been archived, and more than six Terabytes have been distributed to the scientific community. Information about user request and file access patterns, and their impact on system loading, is needed to optimize current operations and to plan for future archives. To facilitate the management of daily activities, the GSFC DAAC has developed a data base system to track correspondence, requests, ingestion and distribution. In addition, several log files which record transactions on Unitree are maintained and periodically examined. This study identifies some of the users' requests and file access patterns at the GSFC DAAC during 1995. The analysis is limited to the subset of orders for which the data files are under the control of the Hierarchical Storage Management (HSM) Unitree. The results show that most of the data volume ordered was for two data products. The volume was also mostly made up of level 3 and 4 data and most of the volume was distributed on 8 mm and 4 mm tapes. In addition, most of the volume ordered was for deliveries in North America although there was a significant world-wide use. There was a wide range of request sizes in terms of volume and number of files ordered. On an average 78.6 files were ordered per request. Using the data managed by Unitree, several caching algorithms have been evaluated for both hit rate and the overhead ('cost') associated with the movement of data from near-line devices to disks. The algorithm called LRU/2 bin was found to be the best for this workload, but the STbin algorithm also worked well.

Johnson, Theodore↗

NASDA's earth observation satellite data archive policy for the earth observation data and information system (EOIS)

NASDA's new Advanced Earth Observing Satellite (ADEOS) is scheduled for launch in August, 1996. ADEOS carries 8 sensors to observe earth environmental phenomena and sends their data to NASDA, NASA, and other foreign ground stations around the world. The downlink data bit rate for ADEOS is 126 MB/s and the total volume of data is about 100 GB per day. To archive and manage such a large quantity of data with high reliability and easy accessibility it was necessary to develop a new mass storage system with a catalogue information database using advanced database management technology. The data will be archived and maintained in the Master Data Storage Subsystem (MDSS) which is one subsystem in NASDA's new Earth Observation data and Information System (EOIS). The MDSS is based on a SONY ID1 digital tape robotics system. This paper provides an overview of the EOIS system, with a focus on the Master Data Storage Subsystem and the NASDA Earth Observation Center (EOC) archive policy for earth observation satellite data.

Sobue, Shin-ichi↗

Towards the Interoperability of Web, Database, and Mass Storage Technologies for Petabyte Archives

At the San Diego Supercomputer Center, a massive data analysis system (MDAS) is being developed to support data-intensive applications that manipulate terabyte sized data sets. The objective is to support scientific application access to data whether it is located at a Web site, stored as an object in a database, and/or storage in an archival storage system. We are developing a suite of demonstration programs which illustrate how Web, database (DBMS), and archival storage (mass storage) technologies can be integrated. An application presentation interface is being designed that integrates data access to all of these sources. We have developed a data movement interface between the Illustra object-relational database and the NSL UniTree archival storage system running in a production mode at the San Diego Supercomputer Center. With this interface, an Illustra client can transparently access data on UniTree under the control of the Illustr DBMS server. The current implementation is based on the creation of a new DBMS storage manager class, and a set of library functions that allow the manipulation and migration of data stored as Illustra 'large objects'. We have extended this interface to allow a Web client application to control data movement between its local disk, the Web server, the DBMS Illustra server, and the UniTree mass storage environment. This paper describes some of the current approaches successfully integrating these technologies. This framework is measured against a representative sample of environmental data extracted from the San Diego Ba Environmental Data Repository. Practical lessons are drawn and critical research areas are highlighted.

Moore, Reagan↗

The Challenges Facing Science Data Archiving on Current Mass Storage Systems

This paper discusses the desired characteristics of a tape-based petabyte science data archive and retrieval system required to store and distribute several terabytes (TB) of data per day over an extended period of time, probably more than 115 years, in support of programs such as the Earth Observing System Data and Information System (EOSDIS). These characteristics take into consideration not only cost effective and affordable storage capacity, but also rapid access to selected files, and reading rates that are needed to satisfy thousands of retrieval transactions per day. It seems that where rapid random access to files is not crucial, the tape medium, magnetic or optical, continues to offer cost effective data storage and retrieval solutions, and is likely to do so for many years to come. However, in environments like EOS these tape based archive solutions provide less than full user satisfaction. Therefore, the objective of this paper is to describe the performance and operational enhancements that need to be made to the current tape based archival systems in order to achieve greater acceptance by the EOS and similar user communities.

Peavey, Bernard↗

A Complete Public Archive for the Einstein Imaging Proportional Counter

Consistent with our proposal to the Astrophysics Data Program in 1992, we have completed the design, construction, documentation, and distribution of a flexible and complete archive of the data collected by the Einstein Imaging Proportional Counter. Along with software and data delivered to the High Energy Astrophysics Science Archive Research Center at Goddard Space Flight Center, we have compiled and, where appropriate, published catalogs of point sources, soft sources, hard sources, extended sources, and transient flares detected in the database along with extensive analyses of the instrument's backgrounds and other anomalies. We include in this document a brief summary of the archive's functionality, a description of the scientific catalogs and other results, a bibliography of publications supported in whole or in part under this contract, and a list of personnel whose pre- and post-doctoral education consisted in part in participation in this project.

Helfand, David J.↗

International Ultraviolet Explorer Final Archive

CSC processed IUE images through the Final Archive Data Processing System. Raw images were obtained from both NDADS and the IUEGTC optical disk platters for processing on the Alpha cluster, and from the IUEGTC optical disk platters for DECstation processing. Input parameters were obtained from the IUE database. Backup tapes of data to send to VILSPA were routinely made on the Alpha cluster. IPC handled more than 263 requests for priority NEWSIPS processing during the contract. Staff members also answered various questions and requests for information and sent copies of IUE documents to requesters. CSC implemented new processing capabilities into the NEWSIPS processing systems as they became available. In addition, steps were taken to improve efficiency and throughput whenever possible. The node TORTE was reconfigured as the I/O server for Alpha processing in May. The number of Alpha nodes used for the NEWSIPS processing queue was increased to a maximum of six in measured fashion in order to understand the dependence of throughput on the number of nodes and to be able to recognize when a point of diminishing returns was reached. With Project approval, generation of the VD FITS files was dropped in July. This action not only saved processing time but, even more significantly, also reduced the archive storage media requirements, and the time required to perform the archiving, drastically. The throughput of images verified through CDIVS and processed through NEWSIPS for the contract period is summarized below. The number of images of a given dispersion type and camera that were processed in any given month reflects several factors, including the availability of the required NEWSIPS software system, the availability of the corresponding required calibrations (e.g., the LWR high-dispersion ripple correction and absolute calibration), and the occurrence of reprocessing efforts such as that conducted to incorporate the updated SWP sensitivity-degradation correction in May.

Source record↗

Evaluation of DVD-R for Archival Applications

For more than a decade, CD-ROM and CD-R have provided an unprecedented level of reliability, low cost and cross-platform compatibility to support federal data archiving and distribution efforts. However, it should be remembered that years of effort were required to achieve the standardization that has supported the growth of the CD industry. Incompatibilities in the interpretation of the ISO-9660 standard on different operating systems had to be dealt with, and the imprecise specifications in the Orange Book Part n and Part Hi led to incompatibilities between CD-R media and CD-R recorders. Some of these issues were presented by the authors at Optical Data Storage '95. The major current problem with the use of CD technology is the growing volume of digital data that needs to be stored. CD-ROM collections of hundreds of volumes and CD-R collections of several thousand volumes are becoming almost too cumbersome to be useful. The emergence of Digital Video Disks Recorder (DVD-R) technology promises to reduce the number of discs required for archive applications by a factor of seven while providing improved reliability. It is important to identify problem areas for DVD-R media and provide guidelines to manufacturers, file system developers and users in order to provide reliable data storage and interchange. The Data Distribution Laboratory (DDL) at NASA's Jet Propulsion Laboratory began its evaluation of DVD-R technology in early 1998. The initial plan was to obtain a DVD-Recorder for preliminary testing, deploy reader hardware to user sites for compatibility testing, evaluate the quality and longevity of DVD-R media and develop proof-of-concept archive collections to test the reliability and usability of DVD-R media and jukebox hardware.

Martin, Michael D.↗

Did I Say Terabyte? I Meant Petabyte: Data Archiving in the Era of SDO

Two years ago (Gurman 1999, Bull, AAS, 31, 955), we discussed the treatment of archives of the order of 10 Tbyte per year from solar physics missions in the period 2004 - 2006 (e.g. Solar-B and STEREO). By early 2007, we expect that the Sun-Earth Connections community will have to deal with data sets from the Solar Dynamics Observatory (SDO) of order I Tbyte per day. As in the previous work, we examine several alternatives for dealing with data flow and service on a fire-hose scale, and show that off-the shelf, network-attached storage can provide an inexpensive and scaleable solution. We discuss some of the differences between an SDO data archive, as well as the range of requirements for data integrity, disaster recovery, &c. in various scenarios for archive concentration or distribution.

Gurman, Joseph B.↗

New Energetic Radio Pulsars: An Archival X-Ray Survey

This ADP grant was to analyze archival X-ray data obtained in the direction of radio pulsars that were recently discovered as part of the Parkes Multibeam Pulsar Search, which was done using the 64-m Parkes radio telescope in Australia. The survey discovered nearly 700 pulsars, of which roughly three dozen were possible candidates for the detection of X-ray emission. Our team looked at 30 of the most interesting candidates. In most cases, there was insufficient data in the archive to conclude anything. However in several cases, there were interesting archival observations. In three cases, a detailed analysis proved scientifically interesting, and two publications have resulted.

Source record↗

Preservation and Enhancement of the Spacewatch Data Archives

In March of 1998, the asteroid 1997 XF11 was announced to be potentially hazardous after being tracked over 90 days. A potential two year wait for confirming observations was shortened to under 24 hours because of the existence of archived photographic prediscovery images. Spacewatch was a pioneer in using CCD scanning and possesses a valuable digital archive of its scans. Unfortunately these data are aging on magnetic tape and will soon be lost. Since 1990, the Spacewatch project gathered some 1.5 Terabytes of scan data covering roughly 75,000 degrees of sky to a limiting magnitude of V = 21.5. The data have not yet been mined for all of their asteroids for scientific studies and orbit determination. Spacewatch's real-time motion detection program MODP was constrained by the computers of the era to use simplified image processing algorithms at a reduced efficiency. Jedicke and Herron estimated MODP's efficiency at finding asteroids to be approximately 60 percent to V=18 and improving somewhat thereafter. This lead to a substantial bias correction in their analyses. Larsen has developed a MODP replacement capable in excess of 90 percent efficiency in the same range and able to push a magnitude fainter in completeness. We propose a program of post-processing and re-archiving Spacewatch data. Our scans would be transferred from tape to CD-ROMs and converted to FITS images -- establishing a consistent data format and media for both past and future Spacewatch observations. Larsen's MODP replacement would mine these data for previously undetected motions, which would be made available to the Minor Planet Center and our ongoing asteroid population studies. A searchable observation record would be made generally available for prediscovery work. We estimate the net asteroid yield of this proposal is equivalent to three full years of Spacewatch operations.

Larsen, Jeffrey A.↗

Large Scale Data Mining to Improve Usability of Data: An Intelligent Archive Testbed

Research in certain scientific disciplines - including Earth science, particle physics, and astrophysics - continually faces the challenge that the volume of data needed to perform valid scientific research can at times overwhelm even a sizable research community. The desire to improve utilization of this data gave rise to the Intelligent Archives project, which seeks to make data archives active participants in a knowledge building system capable of discovering events or patterns that represent new information or knowledge. Data mining can automatically discover patterns and events, but it is generally viewed as unsuited for large-scale use in disciplines like Earth science that routinely involve very high data volumes. Dozens of research projects have shown promising uses of data mining in Earth science, but all of these are based on experiments with data subsets of a few gigabytes or less, rather than the terabytes or petabytes typically encountered in operational systems. To bridge this gap, the Intelligent Archives project is establishing a testbed with the goal of demonstrating the use of data mining techniques in an operationally-relevant environment. This paper discusses the goals of the testbed and the design choices surrounding critical issues that arose during testbed implementation.

Ramapriyan, Hampapuram↗

NASA Remote Sensing Data in Earth Sciences: Processing, Archiving, Distribution, Applications at the GES DISC

The NASA Goddard Earth Sciences Data and Information Services Center (GES DISC) is one of the major Distributed Active Archive Centers (DAACs) archiving and distributing remote sensing data from the NASA's Earth Observing System. In addition to providing just data, the GES DISC/DAAC has developed various value-adding processing services. A particularly useful service is data processing a t the DISC (i.e., close to the input data) with the users' algorithms. This can take a number of different forms: as a configuration-managed algorithm within the main processing stream; as a stand-alone program next to the on-line data storage; as build-it-yourself code within the Near-Archive Data Mining (NADM) system; or as an on-the-fly analysis with simple algorithms embedded into the web-based tools (to avoid downloading unnecessary all the data). The existing data management infrastructure at the GES DISC supports a wide spectrum of options: from data subsetting data spatially and/or by parameter to sophisticated on-line analysis tools, producing economies of scale and rapid time-to-deploy. Shifting processing and data management burden from users to the GES DISC, allows scientists to concentrate on science, while the GES DISC handles the data management and data processing at a lower cost. Several examples of successful partnerships with scientists in the area of data processing and mining are presented.

Leptoukh, Gregory G.↗

A Testbed Demonstration of an Intelligent Archive in a Knowledge Building System

The last decade's influx of raw data and derived geophysical parameters from several Earth observing satellites to NASA data centers has created a data-rich environment for Earth science research and applications. While advances in hardware and information management have made it possible to archive petabytes of data and distribute terabytes of data daily to a broad community of users, further progress is necessary in the transformation of data into information, and information into knowledge that can be used in particular applications in order to realize the full potential of these valuable datasets. In examining what is needed to enable this progress in the data provider environment that exists today and is expected to evolve in the next several years, we arrived at the concept of an Intelligent Archive in context of a Knowledge Building System (IA/KBS). Our prior work and associated papers investigated usage scenarios, required capabilities, system architecture, data volume issues, and supporting technologies. We identified six key capabilities of an IA/KBS: Virtual Product Generation, Significant Event Detection, Automated Data Quality Assessment, Large-Scale Data Mining, Dynamic Feedback Loop, and Data Discovery and Efficient Requesting. Among these capabilities, large-scale data mining is perceived by many in the community to be an area of technical risk. One of the main reasons for this is that standard data mining research and algorithms operate on datasets that are several orders of magnitude smaller than the actual sizes of datasets maintained by realistic earth science data archives. Therefore, we defined a test-bed activity to implement a large-scale data mining algorithm in a pseudo-operational scale environment and to examine any issues involved. The application chosen for applying the data mining algorithm is wildfire prediction over the continental U.S. This paper reports a number of observations based on our experience with this test-bed. While proof-of-concept for data mining scalability and utility has been a major goal for the research reported here, it was not the only one. The other five capabilities of an WKBS named above have been considered as well, and an assessment of the implications of our experience for these other areas will also be presented. The lessons learned through the testbed effort and presented in this paper will benefit technologists, scientists, and system operators as they consider introducing IA/KBS capabilities into production systems.

Ramapriyan, Hampapuram↗

Cassini/Huygens Program Archive Plan for Science Data

The purpose of this document is to describe the Cassini/Huygens science data archive system which includes policy, roles and responsibilities, description of science and supplementary data products or data sets, metadata, documentation, software, and archive schedule and methods for archive transfer to the NASA Planetary Data System (PDS).

NASA Planetary Science Data System↗

Observations on Cost Modeling and Performance Measurement of Long Term Archives

This paper describes a prototype suite of Excel-based tools that could be used for estimating lifecycle costs for newly planned or modified long-term archival facilities. These tools may also prove valuable for monitoring the long-term performance of such facilities once operational. Cost estimation is by analogy, using statistical curve-fitting techniques across a database of comparable data activities. The database currently includes 29 operational data centers ranging from small (2 FTEs) to large (66 FTEs), and is readily expandable to include additional activities specifically involving data preservation and added value. Each comparable data center is described in terms of its staffing, throughput workload, archival and distribution requirements, levels of user service, overall complexity, degree of automation, and other data, comprising 94 distinct descriptors in all. The descriptors were developed by normalizing heterogeneous data from the various centers and mapping them into an Excel framework consistent with the OAIS reference model. A user-friendly tool is provided for generating input to and updating the comparables database. This tool can also be used to benchmark the performance (in terms of cost versus throughput) of an operational data center, and to update the staffing, cost and workload data on a periodic basis. The comparables database could thus provide a history of staffing and throughput over time, as a means of performance monitoring and providing feedback for continuous improvement. Ancillary tools are also provided for performing "what-if' cost exercises for planning purposes, and for graphical display of data and results. We provide a high-level description of the tools; present our experiences and observations on gathering the information and maintaining the database; and discuss how this tool set might be applied to long term archives.

Fontaine, Kathy↗

Characterizing Space Environments with Long-Term Space Plasma Archive Resources

A significant scientific benefit of establishing and maintaining long-term space plasma data archives is the ready access the archives afford to resources required for characterizing spacecraft design environments. Space systems must be capable of operating in the mean environments driven by climatology as well as the extremes that occur during individual space weather events. Long- term time series are necessary to obtain quantitative information on environment variability and extremes that characterize the mean and worst case environments that may be encountered during a mission. In addition, analysis of large data sets are important to scientific studies of flux limiting processes that provide a basis for establishing upper limits to environment specifications used in radiation or charging analyses. We present applications using data from existing archives and highlight their contributions to space environment models developed at Marshall Space Flight Center including the Chandra Radiation Model, ionospheric plasma variability models, and plasma models of the L2 space environment.

Minow, Joseph I.↗

Restoration and PDS Archive of Apollo Lunar Rock Sample Data

In 2008, scientists at the Johnson Space Center (JSC) Lunar Sample Laboratory and Image Science & Analysis Laboratory (under the auspices of the Astromaterials Research and Exploration Science Directorate or ARES) began work on a 4-year project to digitize the original film negatives of Apollo Lunar Rock Sample photographs. These rock samples together with lunar regolith and core samples were collected as part of the lander missions for Apollos 11, 12, 14, 15, 16 and 17. The original film negatives are stored at JSC under cryogenic conditions. This effort is data restoration in the truest sense. The images represent the only record available to scientists which allows them to view the rock samples when making a sample request. As the negatives are being scanned, they are also being formatted and documented for permanent archive in the NASA Planetary Data System (PDS) archive. The ARES group is working collaboratively with the Imaging Node of the PDS on the archiving.

Garcia, P. A.↗