Search NASASearch

SEARCH · Search NASA

Results for “data storage data management”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Preservation Environments

The long-term preservation of digital entities requires mechanisms to manage the authenticity of massive data collections that are written to archival storage systems. Preservation environments impose authenticity constraints and manage the evolution of the storage system technology by building infrastructure independent solutions. This seeming paradox, the need for large archives, while avoiding dependence upon vendor specific solutions, is resolved through use of data grid technology. Data grids provide the storage repository abstractions that make it possible to migrate collections between vendor specific products, while ensuring the authenticity of the archived data. Data grids provide the software infrastructure that interfaces vendor-specific storage archives to preservation environments.

Moore, Reagan W.

dCache: The Storage System of Choice for Data-Intensive Applications

The ever-increasing volumes of data produced by modern scientific facilities like EuXFEL and LHC put significant stress on data management infrastructure operated by laboratories and research centers. The challenges to be addressed span the entire data life cycle, from ingest and efficient data analysis to long-term preservation, typically involving large tape libraries. dCache, a storage system developed in collaboration between the Deutsches Elektronen-Synchrotron (DESY), Fermi National Accelerator Laboratory, and Nordic e-Infrastructure Collaboration (NeIC), is designed to manage a large number of disk servers and to facilitate transparent data migration to and from archival storage. Its multifaceted approach offers a unified method to support a variety of scientific use cases with the same storage infrastructure, including high-throughput data ingest, data sharing over wide area networks, efficient access from HPC clusters, and long-term data preservation on tertiary storage. Initially developed for high energy physics (HEP) experiments, dCache is now used by various scientific communities, including astrophysics, biomedical research, and life sciences, each having specific requirements. This paper presents architecture, deployment strategies, performance and scalability enhancements, and recent advancements in dCache addressing the needs of scientific communities. Finally, we touch on the development and release process, ensuring the software’s high quality.

DCache

Object storage model for CMS data

In CMS, data access and management is organized around the data-tier model: a static definition of what subset of event information is available in a particular dataset, realized as a collection of files. In previous work, we have proposed a novel data management model that obviates the need for data tiers by exploding files into individual event data product objects. In this work, we estimate the potential savings in data volume based on user analysis patterns.

Smith, Nick

FJET Database Project: Extract, Transform, and Load

The Data Mining & Knowledge Management team at Kennedy Space Center is providing data management services to the Frangible Joint Empirical Test (FJET) project at Langley Research Center (LARC). FJET is a project under the NASA Engineering and Safety Center (NESC). The purpose of FJET is to conduct an assessment of mild detonating fuse (MDF) frangible joints (FJs) for human spacecraft separation tasks in support of the NASA Commercial Crew Program. The Data Mining & Knowledge Management team has been tasked with creating and managing a database for the efficient storage and retrieval of FJET test data. This paper details the Extract, Transform, and Load (ETL) process as it is related to gathering FJET test data into a Microsoft SQL relational database, and making that data available to the data users. Lessons learned, procedures implemented, and programming code samples are discussed to help detail the learning experienced as the Data Mining & Knowledge Management team adapted to changing requirements and new technology while maintaining flexibility of design in various aspects of the data management project.

excel vba

Optimizing tertiary storage organization and access for spatio-temporal datasets

We address in this paper data management techniques for efficiently retrieving requested subsets of large datasets stored on mass storage devices. This problem represents a major bottleneck that can negate the benefits of fast networks, because the time to access a subset from a large dataset stored on a mass storage system is much greater that the time to transmit that subset over a network. This paper focuses on very large spatial and temporal datasets generated by simulation programs in the area of climate modeling, but the techniques developed can be applied to other applications that deal with large multidimensional datasets. The main requirement we have addressed is the efficient access of subsets of information contained within much larger datasets, for the purpose of analysis and interactive visualization. We have developed data partitioning techniques that partition datasets into 'clusters' based on analysis of data access patterns and storage device characteristics. The goal is to minimize the number of clusters read from mass storage systems when subsets are requested. We emphasize in this paper proposed enhancements to current storage server protocols to permit control over physical placement of data on storage devices. We also discuss in some detail the aspects of the interface between the application programs and the mass storage system, as well as a workbench to help scientists to design the best reorganization of a dataset for anticipated access patterns.

Chen, Ling Tony

Simulating Secure Data Exchange and Storage for Urban Air Mobility Environments

Urban Air Mobility (UAM) defines an environment for managing operations of vertical takeoff and landing (VTOL) and short takeoff and landing (STOL) vehicles in an urban environment. Within a UAM environment, UAM operators manage fleets of vehicles, relying on Providers of Services for UAM (PSUs) for managing flights in a region of airspace. Flight plan deconfliction is primarily performed by the Discovery and Synchronization Service (DSS), and the Federal Aviation Administration (FAA) maintains control over the UAM space via the FAA-Industry Exchange Protocol (FIDXP). UAM is a federated environment with many different entities owning and operating vehicles, PSUs, and other services. These entities often need to interoperate or access data generated by other organizations. This paper demonstrates the feasibility of using blockchain to facilitate a secure data exchange and storage for this flight information in a UAM environment. In particular, this paper is focused on flight plans and telemetry data. A blockchain network was developed with a set of smart contracts for managing relevant flight data. Hyperledger Fabric was chosen as it is performent, scalable, and allows organizations to reuse existing public key infrastructure (PKI) for identity management. A set of simulated UAM services were also developed. These services propose flight plans and negotiate with other UAM services for airspace access. All interactions between UAM services, as well as vehicle telemetry data, is recorded onto the blockchain. Vehicle telemetry data is generated by a vehicle flight simulation service. This paper successfully demonstrates the feasibility of using blockchain as a secure data exchange and storage mechanism in a UAM environment.

UAM

The Kepler End-to-End Data Pipeline: From Photons to Far Away Worlds

The Kepler mission is described in overview and the Kepler technique for discovering exoplanets is discussed. The design and implementation of the Kepler spacecraft, tracing the data path from photons entering the telescope aperture through raw observation data transmitted to the ground operations team is described. The technical challenges of operating a large aperture photometer with an unprecedented 95 million pixel detector are addressed as well as the onboard technique for processing and reducing the large volume of data produced by the Kepler photometer. The technique and challenge of day-to-day mission operations that result in a very high percentage of time on target is discussed. This includes the day to day process for monitoring and managing the health of the spacecraft, the annual process for maintaining sun on the solar arrays while still keeping the telescope pointed at the fixed science target, the process for safely but rapidly returning to science operations after a spacecraft initiated safing event and the long term anomaly resolution process.The ground data processing pipeline, from the point that science data is received on the ground to the presentation of preliminary planetary candidates and supporting data to the science team for further evaluation is discussed. Ground management, control, exchange and storage of Kepler's large and growing data set is discussed as well as the process and techniques for removing noise sources and applying calibrations to intermediate data products.

data archiving

Mic-hackathon 2024: hackathon on machine learning for electron and scanning probe microscopy

Microscopy is one of the primary sources of information on materials structure and functionality at the nanometer and atomic scales. The data generated through microscopy is often contained in well-structured datasets, enriched with extensive metadata and sample histories, although not always with the same level of detail or storage format. The broad incorporation of data management plans by major funding agencies ensures the preservation and accessibility of this data. However, deriving insights from these rich datasets remains challenging due to the lack of established code ecosystems, standardized benchmarks, and integration strategies. Correspondingly, the efficiency of data usage is very low, and time expenditures at the analysis stage are enormous. In addition to post-acquisition data analysis, the emergence of application programming interfaces by major microscope manufacturers now creates opportunities for real-time ML-based data analytics to enable automated decision making, and particularly ML-agent controlled real-time microscope operation. Despite these opportunities, there is a significant gap in integrating the ML community with the broader microscopy community, limiting the value that these methods bring to physics and materials discovery and materials optimization. Hackathons address these challenges by fostering collaboration between ML experts and microscopy professionals, encouraging the development of innovative solutions that leverage ML for microscopy and preparing the workforce of the future both for microscopy-intensive domains areas, instrument manufacturers, and ML scientists interested in real world applications for fundamental research, materials optimization, and manufacturing. The hackathon generated benchmark datasets and digital twins of microscopes that further contribute to the development of the field and establish data analysis ecosystems. All the codes can be found at GitHub(https://github.com/KalininGroup/Mic-hackathon-2024-codes-publication/tree/1.0.0.1) and Zenodo (https://zenodo.org/records/15579940).

97 MATHEMATICS AND COMPUTING

Pushing the limits of NAND technology scaling with ferroelectrics

Artificial intelligence (AI) continues to drive transformative advancements across various industries. The data-intensive nature of AI training (and inferencing) has resulted in the generation of unprecedented volumes of data with machine-generated content surpassing human-generated data by more than 100-fold in 2025. Efficiently managing this data influx necessitates advanced digital storage technologies. However, traditional NAND flash memory, which is critical for supporting data flows in AI systems—alongside high-bandwidth memory, for AI training—faces fundamental scaling limitations as it approaches the 1000-layer milestone, encompassing more than 40 trillion transistors. This article delves into the potential of hafnia-based ferroelectric materials as a breakthrough solution to these challenges. Recent advancements indicate that the intrinsic limitations of ferroelectric field-effect transistors (FEFETs) can be mitigated through material and device-level engineering. These advancements enable FEFETs to meet the stringent density, reliability, and scalability requirements of future three-dimensional NAND technology. The role of ferroelectrics in addressing NAND scaling challenges and expanding storage capabilities presents a promising avenue for meeting the storage demands of the AI-driven era.

3D NAND

Spaceborne optical disk controller development

The current status and potential applications of an optical-disk buffer (ODB) memory system being developed by an interagency consortium including NASA and the USAF are reviewed. The design goals for the ODB include usable capacity 1 Tb, maximum data rate 1.6 Gb/s, read error rate less than 10 to the -12th, time to initial access less than 100 ms, and unlimited read/write cycles. Present efforts focus on a brassboard ODB which employs 12 14-inch magnetooptic disks and 24 nine-diode read/write heads. A typical space application of an optical disk mass memory system (ODMMS) is discussed: as communications buffer, temporary storage, and/or multiuser I/O buffer for data management on the Space Station Earth Observing System. Environmental, operational, system-architecture, and functional-separation factors; critical design issues; and standardization questions for spaceborne ODMMSs are examined in detail.

Shull, Thomas A.

Network accessible multi-terabyte archive

The viewgraphs of a discussion on Network Accessible Multi-terabyte Archive presented at the National Space Science Data Center (NSSDC) Mass Storage Workshop is included. Topics covered in the presentation include the rotary storage system (RSS) including RSS data access, data management, administration, and hardware and software architecture.

Rybczynski, Fred

A Spacebased Ocean Surface Exchange Data Analysis System

Emerging technologies have provided unprecedented opportunities to transform information into knowledge and disseminate them in a much faster, cheaper, and userfriendly mode. We have set up a system to produce and disseminate high level (gridded) ocean surface wind data from the NASA Scatterometer and European Remote Sensing missions. The data system is being expanded to produce real-time gridded ocean surface winds from an improved sensor SeaWinds on the Quikscat Mission. The wind field will be combined with hydrologic parameters from the Tropical Rain Measuring Mission to monitor evolving weather systems and natural hazard in real time. It will form the basis for spacebased Ocean Surface Exchange Data Analysis System (SOSEDAS) which will include the production of ocean surface momentum, heat, and water fluxes needed for interdisciplinary studies of ocean-atmosphere interaction. Various commercial or non-commercial software tools have been compared and selected in terms of their ability in database management, remote data accessing, graphical interface, data quality, storage needs and transfer speed, etc. Issues regarding system security and user authentication, distributed data archiving and accessing, strategy to compress large-volume geophysical and satellite data/image. and increasing transferring speed are being addressed. A simple and easy way to access information and derive knowledge from spacebased data of multiple missions is being provided. The evolving 'knowledge system' will provide relevant infrastructure to address Earth System Science, make inroads in educating an informed populace, and illuminate decision and policy making.

Tang, Wenqing

Collectives for Multiple Resource Job Scheduling Across Heterogeneous Servers

Efficient management of large-scale, distributed data storage and processing systems is a major challenge for many computational applications. Many of these systems are characterized by multi-resource tasks processed across a heterogeneous network. Conventional approaches, such as load balancing, work well for centralized, single resource problems, but breakdown in the more general case. In addition, most approaches are often based on heuristics which do not directly attempt to optimize the world utility. In this paper, we propose an agent based control system using the theory of collectives. We configure the servers of our network with agents who make local job scheduling decisions. These decisions are based on local goals which are constructed to be aligned with the objective of optimizing the overall efficiency of the system. We demonstrate that multi-agent systems in which all the agents attempt to optimize the same global utility function (team game) only marginally outperform conventional load balancing. On the other hand, agents configured using collectives outperform both team games and load balancing (by up to four times for the latter), despite their distributed nature and their limited access to information.

Tumer, K.

Optical mass memory system (AMM-13). AMM-13 system segment specification

The performance, design, development, and test requirements for an optical mass data storage and retrieval system prototype (AMM-13) are established. This system interfaces to other system segments of the NASA End-to-End Data System via the Data Base Management System segment and is designed to have a storage capacity of 10 to the 13th power bits (10 to the 12th power bits on line). The major functions of the system include control, input and output, recording of ingested data, fiche processing/replication and storage and retrieval.

Bailey, G. A.

GSE, data management system programmers/User' manual

The GSE data management system is a computerized program which provides for a central storage source for key data associated with the mechanical ground support equipment (MGSE). Eight major sort modes can be requested by the user. Attributes that are printed automatically with each sort include the GSE end item number, description, class code, functional code, fluid media, use location, design responsibility, weight, cost, quantity, dimensions, and applicable documents. Multiple subsorts are available for the class code, functional code, fluid media, use location, design responsibility, and applicable document categories. These sorts and how to use them are described. The program and GSE data bank may be easily updated and expanded.

Schlagheck, R. A.

Mass storage at NSA

The need to manage large amounts of data on robotically controlled devices has been critical to the mission of this Agency for many years. In many respects this Agency has helped pioneer, with their industry counterparts, the development of a number of products long before these systems became commercially available. Numerous attempts have been made to field both robotically controlled tape and optical disk technology and systems to satisfy our tertiary storage needs. Custom developed products were architected, designed, and developed without vendor partners over the past two decades to field workable systems to handle our ever increasing storage requirements. Many of the attendees of this symposium are familiar with some of the older products, such as: the Braegen Automated Tape Libraries (ATL's), the IBM 3850, the Ampex TeraStore, just to name a few. In addition, we embarked on an in-house development of a shared disk input/output support processor to manage our every increasing tape storage needs. For all intents and purposes, this system was a file server by current definitions which used CDC Cyber computers as the control processors. It served us well and was just recently removed from production usage.

Shields, Michael F.

Petabyte Class Storage at Jefferson Lab (CEBAF)

By 1997, the Thomas Jefferson National Accelerator Facility will collect over one Terabyte of raw information per day of Accelerator operation from three concurrently operating Experimental Halls. When post-processing is included, roughly 250 TB of raw and formatted experimental data will be generated each year. By the year 2000, a total of one Petabyte will be stored on-line. Critical to the experimental program at Jefferson Lab (JLab) is the networking and computational capability to collect, store, retrieve, and reconstruct data on this scale. The design criteria include support of a raw data stream of 10-12 MB/second from Experimental Hall B, which will operate the CEBAF (Continuous Electron Beam Accelerator Facility) Large Acceptance Spectrometer (CLAS). Keeping up with this data stream implies design strategies that provide storage guarantees during accelerator operation, minimize the number of times data is buffered allow seamless access to specific data sets for the researcher, synchronize data retrievals with the scheduling of postprocessing calculations on the data reconstruction CPU farms, as well as support the site capability to perform data reconstruction and reduction at the same overall rate at which new data is being collected. The current implementation employs state-of-the-art StorageTek Redwood tape drives and robotics library integrated with the Open Storage Manager (OSM) Hierarchical Storage Management software (Computer Associates, International), the use of Fibre Channel RAID disks dual-ported between Sun Microsystems SMP servers, and a network-based interface to a 10,000 SPECint92 data processing CPU farm. Issues of efficiency, scalability, and manageability will become critical to meet the year 2000 requirements for a Petabyte of near-line storage interfaced to over 30,000 SPECint92 of data processing power.

Chambers, Rita

Immutable Secure Data Exchange and Storage for Urban Air Mobility Environments

The Urban Air Mobility (UAM) environment is derived from the Unmanned Traffic Management (UTM) concept of operations. Within the environment, UAM operators work independently to manage aerial vehicles in the urban environment. Providers of Services (PSU), UAM operators, and Supplemental Data Service Providers provide services to support flight operations within the UAM environment. The intent of this work is to leverage a permissioned blockchain approach, to, simulate secure data exchange and storage for UAM environments. Blockchain technologies can be used for identity management of vehicles, people, and systems.

Blockchain