Search NASA⌕ Search

SEARCH · Search NASA

Results for “data storage data management”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

An Autonomous MCP Bridge to Rucio: Enhancing Data Management Accessibility for High Energy Physics

The Rucio Data Management System [1] is an important tool used by High Energy Physics experiments, including those at Fermi National Accelerator Laboratory, to store and manage exabyte-scale scientific datasets. Despite its central role in coordinating data across globally distributed storage sites, Rucio's command line interface (CLI) presents a steep learning curve, and makes it difficult for scientists to navigate through. To solve this issue, a containerized Model Context Protocol (MCP) [2] server was built that connects Large Language Models directly to Rucio, allowing AI agents to handle data tasks by using simple, natural language rather than memorized terminal commands. The core engineering focus of this project was moving the server away from slow terminal commands that require text parsing and replacing them with a native Python Client API toolset and a planned REST API framework. Moving to the Python API handles data operations directly in memory, which helps clear up formatting errors, provides the AI with clean, structured JSON data and speeds up tool execution. To prove that the system actually works, a benchmarking pipeline was also built with various questions to test the AI across four different model configurations. The questions included finding data scopes, tracking down specific datasets, and checking replication rules. Through benchmarking, early runs showed that with raw terminal text, the model would get confused and stuck, whereas switching to the Python API to feed the AI clean, structured data yielded massive improvement. By creating an intelligent and autonomous bridge to a storage network, this project shows how AI can be implemented in scientific data management, which ultimately helps scientists at Fermilab spend less time sorting through data and more time focusing on their experiments and analysis.

Akella, Kashyap [William Rainey Harper Coll.]↗

Service Management Database for DSN Equipment

This data- and event-driven persistent storage system leverages the use of commercial software provided by Oracle for portability, ease of maintenance, scalability, and ease of integration with embedded, client-server, and multi-tiered applications. In this role, the Service Management Database (SMDB) is a key component of the overall end-to-end process involved in the scheduling, preparation, and configuration of the Deep Space Network (DSN) equipment needed to perform the various telecommunication services the DSN provides to its customers worldwide. SMDB makes efficient use of triggers, stored procedures, queuing functions, e-mail capabilities, data management, and Java integration features provided by the Oracle relational database management system. SMDB uses a third normal form schema design that allows for simple data maintenance procedures and thin layers of integration with client applications. The software provides an integrated event logging system with ability to publish events to a JMS messaging system for synchronous and asynchronous delivery to subscribed applications. It provides a structured classification of events and application-level messages stored in database tables that are accessible by monitoring applications for real-time monitoring or for troubleshooting and analysis over historical archives.

Zendejas, Silvino↗

SWARM: Reimagining scientific workflow management systems in a distributed world

Modern scientific workflows process massive amounts of data from diverse instruments and sensors, leveraging geographically distributed, heterogeneous compute and storage resources—from leadership-class systems to edge devices—connected by high-performance networks. The diversity of resources introduces challenges in harnessing their full potential, with resilience issues arising across applications, system software, networks, storage, and hardware. Today, workflow management systems (WMS) coordinate the execution of computation and data management tasks across target resources. However, WMS’s centralized nature makes them vulnerable to faults and scalability issues that may result in failures of entire computational campaigns. In conclusion, this paper introduces a novel agentic framework for workflow management, fully distributing and decentralizing the WMS functions and modeling them as swarm intelligence agents infused with advanced artificial intelligence solutions and traditional distributed computing algorithms that can make coordinated decisions in the presence of failures of the underlying cyberinfrastructure.

Swarm intelligence↗

The NASA Ames Research Center Institutional Scientific Collection: History, Best Practices and Scientific Opportunities

The NASA Ames Life Sciences Institutional Scientific Collection (ISC), which is composed of the Ames Life Sciences Data Archive (ALSDA) and the Biospecimen Storage Facility (BSF), is managed by the Space Biosciences Division and has been operational since 1993. The ALSDA is responsible for archiving information and animal biospecimens collected from life science spaceflight experiments and matching ground control experiments. Both fixed and frozen spaceflight and ground tissues are stored in the BSF within the ISC. The ALSDA also manages a Biospecimen Sharing Program, performs curation and long-term storage operations, and makes biospecimens available to the scientific community for research purposes via the Life Science Data Archive public website (https:lsda.jsc.nasa.gov). As part of our best practices, a viability testing plan has been developed for the ISC, which will assess the quality of archived samples. We expect that results from the viability testing will catalyze sample use, enable broader science community interest, and improve operational efficiency of the ISC. The current viability test plan focuses on generating disposition recommendations and is based on using ribonucleic acid (RNA) integrity number (RIN) scores as a criteria for measurement of biospecimen viablity for downstream functional analysis. The plan includes (1) sorting and identification of candidate samples, (2) conducting a statiscally-based power analysis to generate representaive cohorts from the population of stored biospecimens, (3) completion of RIN analysis on select samples, and (4) development of disposition recommendations based on the RIN scores. Results of this work will also support NASA open science initiatives and guides development of the NASA Scientific Collections Directive (a policy on best practices for curation of biological collections). Our RIN-based methodology for characterizing the quality of tissues stored in the ISC since the 1980s also creates unique scientific opportunities for temporal assessment across historical missions. Support from the NASA Space Biology Program and the NASA Human Research Program is gratefully acknowledged.

ALSDA↗

The Space Station Data Management System - Avionics that integrate

The Space Station Data Management System (DMS) comprises the networked computers, mass storage, workstations, and instrumentation interfaces required to support onboard systems and payload operations. This paper gives an overview of the current DMS architecture and discusses its role as onboard integrator in four of its major functional areas: (1) data communication; (2) data processing; (3) data administration, storage and retrieval; and (4) data presentation at the human-computer interface.

Whitelaw, Virginia↗

Data management support for selected climate data sets using the climate data access system

The functional capabilities of the Goddard Space Flight Center (GSFC) Climate Data Access System (CDAS), an interactive data storage and retrieval system, and the archival data sets which this system manages are discussed. The CDAS manages several climate-related data sets, such as the First Global Atmospheric Research Program (GARP) Global Experiment (FGGE) Level 2-b and Level 3-a data tapes. CDAS data management support consists of three basic functions: (1) an inventory capability which allows users to search or update a disk-resident inventory describing the contents of each tape in a data set, (2) a capability to depict graphically the spatial coverage of a tape in a data set, and (3) a data set selection capability which allows users to extract portions of a data set using criteria such as time, location, and data source/parameter and output the data to tape, user terminal, or system printer. This report includes figures that illustrate menu displays and output listings for each CDAS function.

Reph, M. G.↗

Expert system development for probabilistic load simulation

A knowledge based system LDEXPT using the intelligent data base paradigm was developed for the Composite Load Spectra (CLS) project to simulate the probabilistic loads of a space propulsion system. The knowledge base approach provides a systematic framework of organizing the load information and facilitates the coupling of the numerical processing and symbolic (information) processing. It provides an incremental development environment for building generic probabilistic load models and book keeping the associated load information. A large volume of load data is stored in the data base and can be retrieved and updated by a built-in data base management system. The data base system standardizes the data storage and retrieval procedures. It helps maintain data integrity and avoid data redundancy. The intelligent data base paradigm provides ways to build expert system rules for shallow and deep reasoning and thus provides expert knowledge to help users to obtain the required probabilistic load spectra.

Ho, H.↗

EOSDIS: Archive and Distribution Systems in the Year 2000

Earth Science Enterprise (ESE) is a long-term NASA research mission to study the processes leading to global climate change. The Earth Observing System (EOS) is a NASA campaign of satellite observatories that are a major component of ESE. The EOS Data and Information System (EOSDIS) is another component of ESE that will provide the Earth science community with easy, affordable, and reliable access to Earth science data. EOSDIS is a distributed system, with major facilities at seven Distributed Active Archive Centers (DAACs) located throughout the United States. The EOSDIS software architecture is being designed to receive, process, and archive several terabytes of science data on a daily basis. Thousands of science users and perhaps several hundred thousands of non-science users are expected to access the system. The first major set of data to be archived in the EOSDIS is from Landsat-7. Another EOS satellite, Terra, was launched on December 18, 1999. With the Terra launch, the EOSDIS will be required to support approximately one terabyte of data into and out of the archives per day. Since EOS is a multi-mission program, including the launch of more satellites and many other missions, the role of the archive systems becomes larger and more critical. In 1995, at the fourth convening of NASA Mass Storage Systems and Technologies Conference, the development plans for the EOSDIS information system and archive were described. Five years later, many changes have occurred in the effort to field an operational system. It is interesting to reflect on some of the changes driving the archive technology and system development for EOSDIS. This paper principally describes the Data Server subsystem including how the other subsystems access the archive, the nature of the data repository, and the mass-storage I/O management. The paper reviews the system architecture (both hardware and software) of the basic components of the archive. It discusses the operations concept, code development, and testing phase of the system. Finally, it describes the future plans for the archive.

Behnke, Jeanne↗

A conceptual design for an integrated data base management system for remote sensing data

The requirements of potential users were considered in the design of an integrated data base management system, developed to be independent of any specific computer or operating system, and to be used to support investigations in weather and climate. Ultimately, the system would expand to include data from the agriculture, hydrology, and related Earth resources disciplines. An overview of the system and its capabilities is presented. Aspects discussed cover the proposed interactive command language; the application program command language; storage and tabular data maintained by the regional data base management system; the handling of data files and the use of system standard formats; various control structures required to support the internal architecture of the system; and the actual system architecture with the various modules needed to implement the system. The concepts on which the relational data model is based; data integrity, consistency, and quality; and provisions for supporting concurrent access to data within the system are covered in the appendices.

Maresca, P. A.↗

On-board data management study for EOPAP

The requirements, implementation techniques, and mission analysis associated with on-board data management for EOPAP were studied. SEASAT-A was used as a baseline, and the storage requirements, data rates, and information extraction requirements were investigated for each of the following proposed SEASAT sensors: a short pulse 13.9 GHz radar, a long pulse 13.9 GHz radar, a synthetic aperture radar, a multispectral passive microwave radiometer facility, and an infrared/visible very high resolution radiometer (VHRR). Rate distortion theory was applied to determine theoretical minimum data rates and compared with the rates required by practical techniques. It was concluded that practical techniques can be used which approach the theoretically optimum based upon an empirically determined source random process model. The results of the preceding investigations were used to recommend an on-board data management system for (1) data compression through information extraction, optimal noiseless coding, source coding with distortion, data buffering, and data selection under command or as a function of data activity, (2) for command handling, (3) for spacecraft operation and control, and (4) for experiment operation and monitoring.

Davisson, L. D.↗

A Novel and Scalable Method for Microencapsulating Salt Hydrate Phase Change Materials in Core–Shell Fibers

Phase change materials (PCMs) are in high demand for applications such as thermal energy storage in buildings, electronics cooling, and thermal management of electric vehicle batteries and data centers. Among these materials, salt hydrate PCMs are particularly attractive due to their high thermal energy storage capacity and low cost. However, they suffer from two major issues: leakage in the melted phase and phase segregation during phase transitions. Microencapsulation is the primary process capable of addressing both of these challenges. However, there is no reliable or scalable method available for microencapsulating salt hydrate PCMs. As a result, the full potential of salt hydrates for building and data center applications has yet to be realized. In this work, we present an innovative method for the microencapsulation of salt hydrate PCMs using a co‐axial pushing technique. This process creates core–shell fibers, with the salt hydrate as the core and a polymer as the shell. Our approach demonstrates strong potential for scalable microencapsulation of salt hydrate PCMs. In conclusion, achieving scalability could enable their widespread use in applications such as data center cooling, battery thermal management, and building climate control.

Sharma, Jaswinder [Oak Ridge National Laboratory (↗

Best Practices for Nuclear Experiment Data Preservation at Idaho National Laboratory: A Guide for Researchers and Reactor Operators

Preserving experimental data is essential for supporting advancements in nuclear science and ensuring the longevity of Idaho National Laboratory's contributions to reactor technology and safety. This report provides a comprehensive guide to best practices for experimental data management and preservation, focusing on standardized data formats, redundancy in storage, metadata documentation, and alignment with international standards. By following these recommendations, experimentalists and reactor operators can enhance the accessibility, reproducibility, and utility of critical datasets for regulatory review, validation computational methods, and future research.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Data Grid Management Systems

The "Grid" is an emerging infrastructure for coordinating access across autonomous organizations to distributed, heterogeneous computation and data resources. Data grids are being built around the world as the next generation data handling systems for sharing, publishing, and preserving data residing on storage systems located in multiple administrative domains. A data grid provides logical namespaces for users, digital entities and storage resources to create persistent identifiers for controlling access, enabling discovery, and managing wide area latencies. This paper introduces data grids and describes data grid use cases. The relevance of data grids to digital libraries and persistent archives is demonstrated, and research issues in data grids and grid dataflow management systems are discussed.

Moore, Reagan W.↗

Smart CO2 Transport-Route Planning Tool: Providing Data and Insights for Accelerating Carbon Transport & Storage Deployment

Overview presentation given at the 2024 FECM / NETL Carbon Management Research Project Review Meeting on NETL's Bipartisan Infrastructure Law-funded Smart CO2 Transport-Route Planning Tool and associated geodatabase. This machine learning informed, data-driven public resource was designed to inform regulators, industry, and researchers plan and develop safe and efficient transport routes across the country.

Romeo, Lucy↗

Towards the Interoperability of Web, Database, and Mass Storage Technologies for Petabyte Archives

At the San Diego Supercomputer Center, a massive data analysis system (MDAS) is being developed to support data-intensive applications that manipulate terabyte sized data sets. The objective is to support scientific application access to data whether it is located at a Web site, stored as an object in a database, and/or storage in an archival storage system. We are developing a suite of demonstration programs which illustrate how Web, database (DBMS), and archival storage (mass storage) technologies can be integrated. An application presentation interface is being designed that integrates data access to all of these sources. We have developed a data movement interface between the Illustra object-relational database and the NSL UniTree archival storage system running in a production mode at the San Diego Supercomputer Center. With this interface, an Illustra client can transparently access data on UniTree under the control of the Illustr DBMS server. The current implementation is based on the creation of a new DBMS storage manager class, and a set of library functions that allow the manipulation and migration of data stored as Illustra 'large objects'. We have extended this interface to allow a Web client application to control data movement between its local disk, the Web server, the DBMS Illustra server, and the UniTree mass storage environment. This paper describes some of the current approaches successfully integrating these technologies. This framework is measured against a representative sample of environmental data extracted from the San Diego Ba Environmental Data Repository. Practical lessons are drawn and critical research areas are highlighted.

Moore, Reagan↗

Leveraging the Cloud for Robust and Efficient Lunar Image Processing

The Lunar Mapping and Modeling Project (LMMP) is tasked to aggregate lunar data, from the Apollo era to the latest instruments on the LRO spacecraft, into a central repository accessible by scientists and the general public. A critical function of this task is to provide users with the best solution for browsing the vast amounts of imagery available. The image files LMMP manages range from a few gigabytes to hundreds of gigabytes in size with new data arriving every day. Despite this ever-increasing amount of data, LMMP must make the data readily available in a timely manner for users to view and analyze. This is accomplished by tiling large images into smaller images using Hadoop, a distributed computing software platform implementation of the MapReduce framework, running on a small cluster of machines locally. Additionally, the software is implemented to use Amazon's Elastic Compute Cloud (EC2) facility. We also developed a hybrid solution to serve images to users by leveraging cloud storage using Amazon's Simple Storage Service (S3) for public data while keeping private information on our own data servers. By using Cloud Computing, we improve upon our local solution by reducing the need to manage our own hardware and computing infrastructure, thereby reducing costs. Further, by using a hybrid of local and cloud storage, we are able to provide data to our users more efficiently and securely. 12 This paper examines the use of a distributed approach with Hadoop to tile images, an approach that provides significant improvements in image processing time, from hours to minutes. This paper describes the constraints imposed on the solution and the resulting techniques developed for the hybrid solution of a customized Hadoop infrastructure over local and cloud resources in managing this ever-growing data set. It examines the performance trade-offs of using the more plentiful resources of the cloud, such as those provided by S3, against the bandwidth limitations such use encounters with remote resources. As part of this discussion this paper will outline some of the technologies employed, the reasons for their selection, the resulting performance metrics and the direction the project is headed based upon the demonstrated capabilities thus far.

Cloud Computing↗

Data management in NOAA

NOAA has 11 terabytes of digital data stored on 240,000 computer tapes. There are an additional 100 terabytes (TB) of geostationary satellite data stored in digital form on specially configured SONY U-Matic video tapes at the University of Wisconsin. There are over 90,000,000 non-digital form records in manuscript, film, printed, and chart form which are not easily accessible. The three NOAA Data Centers service 6,000 requests per year and publish 5,000 bulletins which are distributed to 40,000 subscribers. Seventeen CD-ROM's have been produced. Thirty thousand computer tapes containing polar satellite data are being copied to 12 inch WORM optical disks for research applications. The present annual data accumulation rate of 10 TB will grow to 30 TB in 1994 and to 100 TB by the year 2000. The present storage and distribution technologies with their attendant support systems will be overwhelmed by these increases if not improved. Increased user sophistication coupled with more precise measurement technologies will demand better quality control mechanisms, especially for those data maintained in an indefinite archive. There is optimism that the future will offer improved media technologies to accommodate the volumes of data. With the advanced technologies, storage and performance monitoring tools will be pivotal to the successful long-term management of data and information.

Callicott, William M.↗