Managing the on-board data storage, acknowledgement and retransmission system for Spitzer
UNKNOWN
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
UNKNOWN
A combination of a local network, a mass storage system, and an autonomous set processor serving as a data/storage management machine is described. Its characteristics include: content-accessible data bases usable from all connected devices; efficient storage/access of large data bases; simple and direct programming with data manipulation and storage management handled by the set processor; simple data base design and entry from source representation to set processor representation with no predefinition necessary; capability available for user sort/order specification; significant reduction in tape/disk pack storage and mounts; flexible environment that allows upgrading hardware/software configuration without causing major interruptions in service; minimal traffic on data communications network; and improved central memory usage on large processors.
The NASA EOS Data and Information System (EOSDIS) Core System (ECS) will contain one of the largest data management systems ever built - the ECS Science and Data Processing System (SDPS). SDPS is designed to support long term Global Change Research by acquiring, producing, and storing earth science data, and by providing efficient means for accessing and manipulating that data. The first two releases of SDPS, Release A and Release B, will be operational in 1997 and 1998, respectively. Release B will be deployed at eight Distributed Active Archiving Centers (DAAC's). Individual DAAC's will archive different collections of earth science data, and will vary in archive capacity. The storage and management of these data collections is the responsibility of the SDPS Data Server subsystem. It is anticipated that by the year 2001, the Data Server subsystem at the Goddard DAAC must support a near-line data storage capacity of one petabyte. The development of SDPS is a system integration effort in which COTS products will be used in favor of custom components in very possible way. Some software and hardware capabilities required to meet ECS data volume and storage management requirements beyond 1999 are not yet supported by available COTS products. The ECS project will not undertake major custom development efforts to provide these capabilities. Instead, SDPS and its Data Server subsystem are designed to support initial implementations with current products, and provide an evolutionary framework that facilitates the introduction of advanced COTS products as they become available. This paper provides a high-level description of the Data Server subsystem design from a COTS integration standpoint, and discussed some of the major issues driving the design. The paper focuses on features of the design that will make the system scalable and adaptable to changing technologies.
A brief example of the use of formal methods techniques in the specification of a software system is presented. The report is part of a larger effort targeted at defining a formal methods pilot project for NASA. One possible application domain that may be used to demonstrate the effective use of formal methods techniques within the NASA environment is presented. It is not intended to provide a tutorial on either formal methods techniques or the application being addressed. It should, however, provide an indication that the application being considered is suitable for a formal methods by showing how such a task may be started. The particular system being addressed is the Structured File Services (SFS), which is a part of the Data Storage and Retrieval Subsystem (DSAR), which in turn is part of the Data Management System (DMS) onboard Spacestation Freedom. This is a software system that is currently under development for NASA. An informal mathematical development is presented. Section 3 contains the same development using Penelope (23), an Ada specification and verification system. The complete text of the English version Software Requirements Specification (SRS) is reproduced in Appendix A.
The CMS[1] experiment manages a large-scale data infrastructure, currently handling over 200 PB of disk and 500 PB of tape storage and transferring more than 1 PB of data per day on average between various WLCG[2] sites. Utilizing Rucio[3] for high-level data management, FTS[4] for data transfers, and a variety of storage and network technologies at the sites, CMS confronts inevitable challenges due to the system’s growing scale and evolving nature. Key challenges include managing transfer and storage failures, optimizing data distribution across different storages based on production and analysis needs, implementing necessary technology upgrades and migrations, and efficiently handling user requests. The data management team has established comprehensive monitoring to supervise this system and has successfully addressed many of these challenges. The team’s efforts aim to ensure data availability and protection, minimize failures and manual interventions, maximize transfer throughput and resource utilization, and provide reliable user support. This paper details the operational experience of CMS with its data management system in recent years, focusing on the encountered challenges, the effective strategies employed to overcome them and the ongoing challenges as we prepare for future demands.
The engineering-specified requirements for integrated information processing by means of the Integrated Programs for Aerospace-Vehicle Design (IPAD) system are presented. A data model is described and is based on the design process of a typical aerospace vehicle. General data management requirements are specified for data storage, retrieval, generation, communication, and maintenance. Information management requirements are specified for a two-component data model. In the general portion, data sets are managed as entities, and in the specific portion, data elements and the relationships between elements are managed by the system, allowing user access to individual elements for the purpose of query. Computer program management requirements are specified for support of a computer program library, control of computer programs, and installation of computer programs into IPAD.
The need to marshal the extensive data that figure in design problems is underlined. An approach to managing engineering data for use in a computerized integrated design system is described. The approach is embodied in an experimental integrated software system that is used to demonstrate and evaluate such data management functions as storage, retrieval, query, manipulation, and modification. The development and organization of the experimental integrated design system and the application of the system to selected test problems are discussed, together with insights into data management issues gained from the study.
Large-scale scientific collaborations like ATLAS, Belle II, CMS, DUNE, and others involve hundreds of research institutes and thousands of researchers spread across the globe. These experiments generate petabytes of data, with volumes soon expected to reach exabytes. Consequently, there is a growing need for computation, including structured data processing from raw data to consumer-ready derived data, extensive Monte Carlo simulation campaigns, and a wide range of end-user analysis. To manage these computational and storage demands, centralized workflow and data management systems are implemented. However, decisions regarding data placement and payload allocation are often made disjointly and via heuristic means. A significant obstacle in adopting more effective heuristic or AI-driven solutions is the absence of a quick and reliable introspective dynamic model to evaluate and refine alternative approaches. In this study, we aim to develop such an interactive system using real-world data. By examining job execution records from the PanDA workflow management system, we have pinpointed key performance indicators such as queuing time, error rate, and the extent of remote data access. The dataset includes five months of activity. Additionally, we are creating a generative AI model to simulate time series of payloads, which incorporate visible features like category, event count, and submitting group, as well as hidden features like the total computational load—derived from existing PanDA records and computing site capabilities. These hidden features, which are not visible to job allocators, whether heuristic or AI-driven, influence factors such as queuing times and data movement.
The goals are to conduct a research and development program aimed at determining the most effective way to do SETI within the constraints of current technology and estimated budgets. The general search strategy adopted is that which is recommended by the SETI Science Working Group. The strategy for an all sky survey for SETI was further developed over the last year. Scan patterns, scan rates, and signal detection algorithms were developed. Spectral power measurement instrumentation was tested at the Venus Station of the Goldstone Deep Space Communication Complex. A specially designed radio frequency interference (RFI) measurement system was built and installed at the Venus Station. A data base management system for storage and retrieval of the RFI data was partially implemented on a VAX 750 computer at the Jet Propulsion Laboratory.
A major challenge facing data processing centers today is data management. This includes the storage of large volumes of data and access to it. Current media storage for large data volumes is typically off line and frequently off site in warehouses. Access to data archived in this fashion can be subject to long delays, errors in media selection and retrieval, and even loss of data through misplacement or damage to the media. Similarly, designers responsible for architecting systems capable of continuous high-speed recording of large volumes of digital data are faced with the challenge of identifying technologies and configurations that meet their requirements. Past approaches have tended to evaluate the combination of the fastest tape recorders with the highest capacity tape media and then to compromise technology selection as a consequence of cost. This paper discusses an architecture that addresses both of these challenges and proposes a cost effective solution based on robots, high speed helical scan tape drives, and large-capacity media.
Giovanni is an exploration tool at the NASA Goddard Earth Sciences Data Information Services Center (GES DISC), providing 22 analysis and visualization services for over 1600 Earth Science data variables. Owing to its popularity, Giovanni has experienced a consistent growth in overall demand, with periodic usage spikes attributed to trainings by education organizations, extensive data analysis in response to natural disasters, preparations for science meetings, etc. Furthermore, the new generation of spaceborne sensors and high resolution models have resulted in an exponential growth in data volume with data distributed across the traditional boundaries of data centers. Seamless exploration of data (without users having to worry about data center boundaries) has been a key recommendation of the GES DISC User Working Group. These factors have required new strategies for delivering acceptable performance. The cloud-based Giovanni, built on Amazon Web Services (AWS), evaluates (1) AWS native solutions to provide a scalable, serverless architecture; (2) open standards for data storage in the Cloud; (3) a cost model for operations; and (4) end-user performance. Our preliminary findings indicate that the use of serverless architecture has a potential to significantly reduce development and operational cost of Giovanni. The combination of using AWS managed services, storage of data in open standards, and schema-on-read data access strategy simplifies data access and analytics, in addition to making data more accessible to the end users of Giovanni through popular programming languages.
Spatiotemporal data has evolved in scale due to augmented use in cross-domain applications. Simultaneously, there is substantial growth in the availability of Geographic Information Systems (GIS) data provided by the United States Geological Survey (USGS) along with other federal, state, county, or local agencies through open-data portals and public access APIs. However, data availability does not equate with accessibility. Large-scale analyses and applications require robust, performant data management with co-location of data storage and computing. The insufficiency of data management infrastructure compels researchers to adopt ad hoc project- specific GIS data storage solutions (e.g., copying data to High-Performance computer file systems). As an ad hoc storage strategy does not scale, it hampers cross-domain analyses causing difficulty in data reuse and utilizing existing code bases. Furthermore, GIS data is complex and requires expertise to analyze and manipulate due to its intricate data structures and data-specific projection transformations. Despite the challenges, we recognize that derived GIS data products, e.g., satellite or LIDAR-based images, can be used in downstream applications such as AI by domain, but non-GIS experts. To address the data needs and overcome the challenges, we are working towards a GIS Data Platform focused on efficient data storage, data discovery and access, and an API to enable common workflows. We propose a knowledge-graph (KG) approach for data discovery, whereby datasets are semantically linked to higher- level constructs such as projects and research areas. The semantic data links enable researchers to explore datasets in a top-down approach by specifying relevant and meaningful terms (assists in finding hidden data). An advantage is that the nodes and edges in a knowledge graph create built-in semantic documentation. Deeper spatiotemporal connections between data sources can be encoded via Graph Neural Networks (GNN) (Zhang et al., 2021). The KG approach can be extended to integrate the data itself in a Virtual KG (VKG). Our work will derive inspiration from large-scale VKG efforts that have been undertaken or are currently underway as part of the OpenStreetMap project (Ding et al., 2021). For DOE Data Days, we share the proposed geospatial data platform hybrid (cloud/on-prem) architecture, our work-to-date on storing, retrieving, and transforming LiDAR and raster data relevant to two important NREL use-cases, including the Renewable Energy Potential (reV) Model, and present our proposal for a KG based data discovery engine.
The future of humanity’s presence beyond Earth depends on the successful commercialization of space. For commercialization to succeed, companies need cost-efficient architectures to support their business models and minimize risks for human capital, design, development, and operations. An ongoing challenge to any space enterprise is the reality that terrestrial network technologies are insufficient to provide reliable communications between assets in space. Whether you need to ensure your valuable data is safely transmitted to the ground or reliably delivered between platforms in orbit, ensuring data integrity over intermittent communication links is a necessity. Current solutions to space communications rely heavily on manual recording, storing, and retrieval of data from spacecraft. The current standard in space communication protocols, Consultative Committee for Space Data Systems (CCSDS) Space Packet standard, is reliant on inflexible network architectures based around mission-critical infrastructure to ensure data delivery. However, by automating the recording, storing, retrieval, and verification of data with Delay Tolerant Networks (DTN), the operator is freed from the dependence on manual data management and expensive mission critical infrastructure. NASA has been developing delay tolerant systems since the late 1990’s. Multiple DTN implementations have been established during that time, each suited to different use cases. Most notably, the DTN deployment for the International Space Station (ISS) includes demonstration of two DTN technologies: Interplanetary Overlay Network (ION) and Delay Tolerant Network Marshall Enterprise (DTNME). Beyond ISS, there are even more NASA DTN deployments being considered. Now that DTN implementations are maturing, it is appropriate to reflect upon these decades of work, review the integration and performance of the existing ISS deployment, and explore the future possibilities for DTN deployment industry-wide. The ISS DTN deployment is a complex architecture consisting of different DTN implementations for the onboard and ground network environments. The ION DTN implementation is being used in the on-board network. The Huntsville Operations Support Center (HOSC) DTN implementation, DTNME, is used by the ground network supporting ISS and will soon be a second onboard gateway too. The two implementations work cooperatively to provide high fidelity data services to flight operations users and payload developers across the globe. Though the two implementations yield a quality service, limitations are evident. Data rate, data storage, and device management are constrained by the services themselves and the complex nature of the deployment. Evolution of operations concepts will improve system capabilities and stability, but significant improvement will require additional development to the implementations themselves and to the overall deployment architecture. Taking advantage of the ongoing development and operation of the ISS DTN service will be central to the success of the future evolutions of NASA DTN deployments while demonstrating the benefits of DTN’s low-cost reliable data communication protocols for the growing commercial space industry. A broad effort on DTN integration and support is necessary to promote expansion beyond existing applications. NASA is developing several useful DTN implementations across a number of different systems: ION, DTNME, High-Rate DTN (HDTN), Bundle Protocol Library (BPLib), and others. To prevent fragmentation, DTN implementation teams need to communicate, collaborate, and integrate with one another to build a solid operational foundation for new DTN deployments. The establishment of a group that can assist new DTN users with understanding the purpose of each DTN implementation, provide best practices, and serve as a general knowledge base is paramount. Potential use of DTN on Gateway and other future NASA missions further drives the need for streamlined communication between DTN implementation teams. A well-integrated and highly engaged NASA DTN working group should help provide system architects the best DTN solutions for future commercial space efforts. This paper will first review the history of DTN implementations, explore the shortcoming of current space networking solutions given available limits in technology, and therefore establish the need for Delay Tolerant Networking in space communications. Secondly, the authors will explore NASA’s array of DTN implementations and highlight their usefulness to space applications. Thirdly, this paper will establish general DTN implementation distinguishing factors. Fourthly, the authors will discuss attempts to create a generic DTN comparison matrix, and the authors will review potential future topics in DTN innovation and collaboration, highlighting several key future efforts. Finally, this paper will describe how the institution of a NASA DTN Working Group will benefit DTN adoption across the governmental and commercial space sector. The goal of this paper is to encourage enthusiasm for DTN, share strategies for improving DTN on both current and future applications, promote the collaboration of DTN implementation groups within the international space operations community, and open the conversations about DTN, priorities, complexities, and innovation to the wider spaceflight industry.
The future of humanity’s presence beyond Earth depends on the successful commercialization of space. For commercialization to succeed, companies need cost-efficient architectures to support their business models and minimize risks for human capital, design, development, and operations. An ongoing challenge to any space enterprise is the reality that terrestrial network technologies are insufficient to provide reliable communications between assets in space. Whether you need to ensure your valuable data is safely transmitted to the ground or reliably delivered between platforms in orbit, ensuring data integrity over intermittent communication links is a necessity. Current solutions to space communications rely heavily on manual recording, storing, and retrieval of data from spacecraft. The current standard in space communication protocols, Consultative Committee for Space Data Systems (CCSDS) Space Packet standard, is reliant on inflexible network architectures based around mission-critical infrastructure to ensure data delivery. However, by automating the recording, storing, retrieval, and verification of data with Delay Tolerant Networks (DTN), the operator is freed from the dependence on manual data management and expensive mission critical infrastructure. NASA has been developing delay tolerant systems since the late 1990’s. Multiple DTN implementations have been established during that time, each suited to different use cases. Most notably, the DTN deployment for the International Space Station (ISS) includes demonstration of two DTN technologies: Interplanetary Overlay Network (ION) and Delay Tolerant Network Marshall Enterprise (DTNME). Beyond ISS, there are even more NASA DTN deployments being considered. Now that DTN implementations are maturing, it is appropriate to reflect upon these decades of work, review the integration and performance of the existing ISS deployment, and explore the future possibilities for DTN deployment industry-wide. The ISS DTN deployment is a complex architecture consisting of different DTN implementations for the onboard and ground network environments. The ION DTN implementation is being used in the on-board network. The Huntsville Operations Support Center (HOSC) DTN implementation, DTNME, is used by the ground network supporting ISS and will soon be a second onboard gateway too. The two implementations work cooperatively to provide high fidelity data services to flight operations users and payload developers across the globe. Though the two implementations yield a quality service, limitations are evident. Data rate, data storage, and device management are constrained by the services themselves and the complex nature of the deployment. Evolution of operations concepts will improve system capabilities and stability, but significant improvement will require additional development to the implementations themselves and to the overall deployment architecture. Taking advantage of the ongoing development and operation of the ISS DTN service will be central to the success of the future evolutions of NASA DTN deployments while demonstrating the benefits of DTN’s low-cost reliable data communication protocols for the growing commercial space industry. A broad effort on DTN integration and support is necessary to promote expansion beyond existing applications. NASA is developing several useful DTN implementations across a number of different systems: ION, DTNME, High-Rate DTN (HDTN), Bundle Protocol Library (BPLib), and others. To prevent fragmentation, DTN implementation teams need to communicate, collaborate, and integrate with one another to build a solid operational foundation for new DTN deployments. The establishment of a group that can assist new DTN users with understanding the purpose of each DTN implementation, provide best practices, and serve as a general knowledge base is paramount. Potential use of DTN on Gateway and other future NASA missions further drives the need for streamlined communication between DTN implementation teams. A well-integrated and highly engaged NASA DTN working group should help provide system architects the best DTN solutions for future commercial space efforts. This paper will first review the history of DTN implementations, explore the shortcoming of current space networking solutions given available limits in technology, and therefore establish the need for Delay Tolerant Networking in space communications. Secondly, the authors will explore NASA’s array of DTN implementations and highlight their usefulness to space applications. Thirdly, this paper will establish general DTN implementation distinguishing factors. Fourthly, the authors will discuss attempts to create a generic DTN comparison matrix, and the authors will review potential future topics in DTN innovation and collaboration, highlighting several key future efforts. Finally, this paper will describe how the institution of a NASA DTN Working Group will benefit DTN adoption across the governmental and commercial space sector. The goal of this paper is to encourage enthusiasm for DTN, share strategies for improving DTN on both current and future applications, promote the collaboration of DTN implementation groups within the international space operations community, and open the conversations about DTN, priorities, complexities, and innovation to the wider spaceflight industry.
Statement of Problem Instruments for high speed and high capacity in-situ data identification, classification and storage capabilities are needed by NASA for the information management and analysis of extremely large volume of data sets in future space exploration, space habitation and utilization, in addition to the various missions to planet-earth programs. Parameters such as communication delays, limited resources, and inaccessibility of human manipulation require more intelligent, compact, low power, and light weight information management and data storage techniques. New and innovative algorithms and architecture using photonics will enable us to meet these challenges. The technology has applications for other government and public agencies.
The study of solar neutrinos presents significant opportunities in astrophysics, nuclear physics, and particle physics. However, the low-energy nature of these neutrinos introduces considerable challenges to isolate them from background events, requiring detectors with low-energy threshold, high spatial and energy resolutions, and low data rate. We present the study of solar neutrinos with a kiloton-scale liquid argon detector located underground, instrumented with a pixel readout using the Q-Pix technology. We explore the potential of using volume fiducialization, directional topological information, light signal coincidence, and pulse-shape discrimination to enhance solar neutrino sensitivity. We find that discriminating neutrino signals below 5 MeV is very difficult. However, we show that these methods are useful for the detection of solar neutrinos when external backgrounds are sufficiently understood and when the detector is built using low-background techniques. When building a workable background model for this study, we identify 𝛾 background from the cavern walls and from capture of 𝛼 particles in radon decay chains as both critical to solar neutrino sensitivity and significantly underconstrained by existing measurements. Finally, we highlight that the main advantage of the use of Q-Pix for solar neutrino studies lies in its ability to enable the continuous readout of all low-energy events with minimal data rates and manageable storage for further off-line analyses.
High-performance computing (HPC) applications have traditionally relied on parallel file systems and file transfer services to manage data movement and storage. Alternative approaches have been proposed that use direct communications between application components, trading persistence and fault tolerance for speed. Event-driven architectures, as popularized in enterprise contexts, present a compelling middle ground, avoiding the performance cost and API constraints of parallel file systems while retaining persistence and offering impedance matching between application components. However, adapting streaming frameworks to HPC workloads requires addressing challenges unique to HPC systems. This paper investigates the potential for a streaming framework designed for HPC infrastructures and use cases. We introduce Mofka, a persistent event-streaming framework designed specifically for HPC environments. Mofka combines the capabilities of a traditional streaming service with optimizations tailored to the HPC context, such as support for massively multicore nodes, efficient scaling for large producer-consumer workflows, RDMA-enabled high-performance network communications, specialized network fabrics with multiple links per node, and efficient handling of large scientific data payloads. Built using the Mochi suite of HPC data service components, Mofka provides a lightweight, modular, and high-performance solution for persistent streaming in HPC systems. We present the architecture of Mofka and evaluate its performance against Kafka and Redpanda using benchmarks on diverse platforms, including Argonne's Polaris and Oak Ridge's Frontier supercomputers, showing up to 8× improvement in throughput in some scenarios. We then demonstrate its utility in several real-world applications: a tomographic reconstruction pipeline, a workflow for the discovery of metal-organic frameworks for carbon capture, and the instrumentation of Dask workflows for provenance tracking and performance analysis.
The first steps taken to ensure the controlled evolution of existing facilities toward greater interoperability and sharing of resources among NASA-supported earth science and applications data systems (ESADS) are described. Recommendations made by the various panels during the 1987 ESADS Workshop are presented. The panels were concerned with directories and catalogs, data archives, data manipulation software, computational facilities, data storage media, database management, and networking. Consideration was also given to the tracking and tuning of overall development and management coordination issues.