Search NASA⌕ Search

SEARCH · Search NASA

Results for “Distributed Data Management”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23

Cumulus: NASA Archives in the Cloud

NASA's Earth Observing System Data and Information System (EOSDIS) houses nearly 30PBs of critical Earth Science data and with upcoming missions is expected to balloon to between 200PBs-300PBs over the next seven years. The magnitude of data collected makes it infeasible to download data and process it locally, forcing us to re-think how we store and work with earth science data. NASA has looked to the cloud to address this, building its open source Cumulus software to manage the ingest of diverse data in a wide variety of formats into the cloud and provide services to manage and access the data. In this talk, we will describe how Cumulus provides common features needed to manage a cloud archive in the realm of ingest, data stewardship, and cost-controlled distribution of science data to users and services.

EOSDIS↗

NASA Archives and Data Stewardship in the Cloud with Cumulus

NASA's Earth Observing System Data and Information System (EOSDIS) houses nearly 30PB (petabytes) of critical Earth Science data and with upcoming missions is expected to balloon to between 200PBs-300PBs over the next seven years. The magnitude of data collected makes it infeasible to download data and process it locally, forcing us to re-think how we store and work with earth science data. NASA has looked to the cloud to address this, building its open source Cumulus software to manage the ingest of diverse data in a wide variety of formats into the cloud and provide services to manage and access the data. In this talk, we will describe how Cumulus provides common features needed to manage a cloud archive in the realm of ingest, data stewardship, and cost-controlled distribution of science data to users and services.

NASA↗

Monitoring Disasters by Use of Instrumented Robotic Aircraft

Efforts are under way to develop data-acquisition, data-processing, and data-communication systems for monitoring disasters over large geographic areas by use of uninhabited aerial systems (UAS) robotic aircraft that are typically piloted by remote control. As integral parts of advanced, comprehensive disaster- management programs, these systems would provide (1) real-time data that would be used to coordinate responses to current disasters and (2) recorded data that would be used to model disasters for the purpose of mitigating the effects of future disasters and planning responses to them. The basic idea is to equip UAS with sensors (e.g., conventional video cameras and/or multispectral imaging instruments) and to fly them over disaster areas, where they could transmit data by radio to command centers. Transmission could occur along direct line-of-sight paths and/or along over-the-horizon paths by relay via spacecraft in orbit around the Earth. The initial focus is on demonstrating systems for monitoring wildfires; other disasters to which these developments are expected to be applicable include floods, hurricanes, tornadoes, earthquakes, volcanic eruptions, leaks of toxic chemicals, and military attacks. The figure depicts a typical system for monitoring a wildfire. In this case, instruments aboard a UAS would generate calibrated thermal-infrared digital image data of terrain affected by a wildfire. The data would be sent by radio via satellite to a data-archive server and image-processing computers. In the image-processing computers, the data would be rapidly geo-rectified for processing by one or more of a large variety of geographic-information- system (GIS) and/or image-analysis software packages. After processing by this software, the data would be both stored in the archive and distributed through standard Internet connections to a disaster-mitigation center, an investigator, and/or command center at the scene of the fire. Ground assets (in this case, firefighters and/or firefighting equipment) would also be monitored in real time by use of Global Positioning System (GPS) units and radio communication links between the assets and the UAS. In this scenario, the UAS would serve as a data-relay station in the sky, sending packets of information concerning the locations of assets to the image-processing computer, wherein this information would be incorporated into the geo-rectified images and maps. Hence, the images and maps would enable command-center personnel to monitor locations of assets in real time and in relation to locations affected by the disaster. Optionally, in case of a disaster that disrupted communications, the UAS could be used as an airborne communication relay station to partly restore communications to the affected area. A prototype of a system of this type was demonstrated in a project denoted the First Response Experiment (Project FiRE). In this project, a controlled outdoor fire was observed by use of a thermal multispectral scanning imager on a UAS that delivered image data to a ground station via a satellite uplink/ downlink telemetry system. At the ground station, the image data were geo-rectified in nearly real time for distribution via the Internet to firefighting managers. Project FiRE was deemed a success in demonstrating several advances essential to the eventual success of the continuing development effort.

Wegener, Steven S.↗

Demonstrating Acquisition of Real-Time Thermal Data Over Fires Utilizing UAVs

A disaster mitigation demonstration, designed to integrate remote-piloted aerial platforms, a thermal infrared imaging payload, over-the-horizon (OTH) data telemetry and advanced image geo-rectification technologies was initiated in 2001. Project FiRE incorporates the use of a remotely piloted Uninhabited Aerial Vehicle (UAV), thermal imagery, and over-the-horizon satellite data telemetry to provide geo-corrected data over a controlled burn, to a fire management community in near real-time. The experiment demonstrated the use of a thermal multi-spectral scanner, integrated on a large payload capacity UAV, distributing data over-the-horizon via satellite communication telemetry equipment, and precision geo-rectification of the resultant data on the ground for data distribution to the Internet. The use of the UAV allowed remote-piloted flight (thereby reducing the potential for loss of human life during hazardous missions), and the ability to "finger and stare" over the fire for extended periods of time (beyond the capabilities of human-pilot endurance). Improved bit-rate capacity telemetry capabilities increased the amount, structure, and information content of the image data relayed to the ground. The integration of precision navigation instrumentation allowed improved accuracies in geo-rectification of the resultant imagery, easing data ingestion and overlay in a GIS framework. We focus on these technological advances and demonstrate how these emerging technologies can be readily integrated to support disaster mitigation and monitoring strategies regionally and nationally.

Ambrosia, Vincent G.↗

Performance Analysis of Data Processing in Distributed File Systems with Near Data Processing

In the era of big data, the escalating volume and velocity of data generation pose significant challenges in data processing. Traditional systems like Spark and Hadoop manage the increasing amount and velocity of data by improving data placement and processing speeds. However, they face inherent limitations due to the essential data movement required for processing. In this paper, we explore the Skyhook framework, a novel extension of the Ceph distributed system, which significantly reduces the need for data movement. We present an extensive case study using the Skyhook framework, applying it with the TPC-H and K-means clustering algorithms. More specifically, we leverage the TPC-H benchmark to distinguish between CPU-intensive and I/O-intensive tasks. We explore the integration of K-means clustering into SQL, coupled with a near-data processing system to offload the computational burden of the K-means clustering algorithm to storage nodes. We conduct a comprehensive performance evaluation of distributed data processing applications across three processing approaches: traditional layout (baseline), optimized layout, and near-data processing. Additionally, we introduce the use of the FIO tool to simulate real-world system workloads, enabling the measurement of performance metrics such as average latency and CPU utilization. Our research is a significant advance in understanding how to optimize data processing systems to meet the demands of the modern data landscape.

Hou, Shiyue↗

Bridging Control and Deployment: A Cross-Layer Analysis of Scalable Building Cluster Control

Building cluster control has emerged as a promising approach for enabling flexible and coordinated operation of distributed building systems, yet its transition from pilot demonstrations to routine grid-interactive operation remains limited. This paper argues that this gap cannot be explained by control algorithms alone. Instead, it arises from interacting barriers in communication infrastructure, data and semantic interoperability, uncertainty management, stakeholder participation, market design, and policy support. Accordingly, the paper reviews both technical and non-technical barriers to building cluster control. Technical challenges include heterogeneous devices and protocols, communication latency and reliability, distributed decision-making, and uncertainty propagation across aggregated loads. Non-technical barriers include user participation, stakeholder coordination, incentive allocation, and data governance. Existing solution approaches are synthesized, including semantic interoperability frameworks, edge and hierarchical communication architectures, distributed and transactive control strategies, uncertainty-aware optimization, policy mechanisms, and market reforms. Based on this analysis, two research directions are identified: testing infrastructures that can evaluate control performance under realistic multi-building conditions, and abstraction methods that allow building clusters to interact with other energy sectors through standardized flexibility representations. Overall, the paper provides a structured review of how building cluster control can move from isolated demonstrations toward reproducible, market-compatible, and grid-relevant implementation.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Making Better Use of Satellite Data: The Satellite Needs Working Group

The U.S. Group on Earth Observations (USGEO) initiated in 2016 the Satellite Needs Working Group (SNWG) to identify and communicate the Earth observation needs of U.S. federal agencies. The SNWG identifies such needs through a biennial survey followed by interviews and follow-up discussions by the satellite Earth data providers of the U.S. Government: the National Aeronautics and Space Administration (NASA), the National Oceanic and Atmospheric Administration (NOAA), and the U.S. Geological Survey (USGS). Solutions and services are identified that leverage current or upcoming satellite missions to meet the identified needs; implementation of services that are estimated to significantly increase the level of satisfaction of multiple U.S. agencies are funded by NASA. The SNWG process has resulted in the implementation of numerous services that have impacted operations of not only U.S. agencies but of academic and international institutions as well. Notable examples include the Harmonized Landsat Sentinel-2 project that leverages European Space Agency (ESA) and NASA satellite assets to generate a global, analysis-ready, surface reflectance product with a temporal resolution of two days; the Airborne Data Management Group that curates and provides access to relevant resources, information and data from existing and past NASA field campaigns; and the Dynamic Surface Water Extent product which consists of harmonized but independent water extents derived from both from optical and radar data. These and other products are hosted at the NASA Distributed Active Archive Centers (DAACs) for free and open access. The SNWG Management Office at NASA’s Interagency Implementation and Advanced Concepts Team (IMPACT) manages the implementation of selected solutions and, importantly, NASA’s response to the needs of federal agencies through a Stakeholder Engagement Program. The impact of implemented services and solutions is hampered without efforts to build capacity around the use of the services. The Stakeholder Engagement Program ensures relevant training and outreach to SNWG agencies in collaboration with each solution implementation team.

remote sensing↗

Terminal area energy management regime investigations utilizing an 0.030-scale model (47-0) of the space shuttle vehicle orbiter configuration 140A/B/C/R in the Ames Research Center 11 x 11 foot transonic wind tunnel (OH/48)

Data obtained in a wind tunnel test were examined to: (1) obtain pressure distributions, forces and moments over the vehicle 5 Orbiter in the terminal area energy management (TAEM) and approach phases of flight; (2) obtain elevon and rudder hinge moments in the TAEM and approach phases of flight; (3) obtain body flap and elevon loads for verification of loads balancing with integrated pressure distributions; and (4) obtain pressure distributions near the short OMS pods in the high subsonic, transonic and low supersonic Mach number regimes. Testing was conducted over a Mach number range from 0.6 to 1.4 with Reynolds number variations from 7.57 x 1 million to 2.74 x 1 million per foot. Model angle of attack was varied from -4 to 16 degrees and angles of sideslip ranged from -8 to 8 degrees.

Hawthorne, P. J.↗

Experiences From NASA/Langley's DMSS Project

There is a trend in institutions with high performance computing and data management requirements to explore mass storage systems with peripherals directly attached to a high speed network. The Distributed Mass Storage System (DMSS) Project at the NASA Langley Research Center (LaRC) has placed such a system into production use. This paper will present the experiences, both good and bad, we have had with this system since putting it into production usage. The system is comprised of: 1) National Storage Laboratory (NSL)/UniTree 2.1, 2) IBM 9570 HIPPI attached disk arrays (both RAID 3 and RAID 5), 3) IBM RS6000 server, 4) HIPPI/IPI3 third party transfers between the disk array systems and the supercomputer clients, a CRAY Y-MP and a CRAY 2, 5) a "warm spare" file server, 6) transition software to convert from CRAY's Data Migration Facility (DMF) based system to DMSS, 7) an NSC PS32 HIPPI switch, and 8) a STK 4490 robotic library accessed from the IBM RS6000 block mux interface. This paper will cover: the performance of the DMSS in the following areas: file transfer rates, migration and recall, and file manipulation (listing, deleting, etc.); the appropriateness of a workstation class of file server for NSL/UniTree with LaRC's present storage requirements in mind the role of the third party transfers between the supercomputers and the DMSS disk array systems in DMSS; a detailed comparison (both in performance and functionality) between the DMF and DMSS systems LaRC's enhancements to the NSL/UniTree system administration environment the mechanism for DMSS to provide file server redundancy the statistics on the availability of DMSS the design and experiences with the locally developed transparent transition software which allowed us to make over 1.5 million DMF files available to NSL/UniTree with minimal system outage

Source record↗

Bridging the Gap on Data and Analysis for Distribution System Planning: Information That Utilities Can Provide Regulators, State Energy Offices and Other Stakeholders

Electric utilities conduct planning annually to ensure their distribution system meets technical standards, policies, and regulations; addresses forecasted grid conditions; satisfies customer needs; and advances utility priorities. The plan identifies grid deficiencies, analyzes potential solutions, and prioritizes capital investments and other expenditures. About 20 U.S. states and jurisdictions require regulated utilities to file some type of distribution system plan with the public utility commission for review. Requirements for sharing distribution system data and analyses vary widely, from few specific requirements to a detailed list of information that must be provided. While utilities conduct extensive analysis to develop distribution system plans, in most jurisdictions regulators and stakeholders do not know what data are available and how the utility uses the data in planning and investing. This report aims to bridge the gap by increasing understanding of the types of data and analyses utilities employ to develop distribution system plans and how the information affects their decision-making. The report describes information that states and stakeholders can ask for related to 11 data categories: -Forecasting loads and distributed energy resources (DERs) -Scenario analysis -Worst-performing circuits -Asset management strategy -Hosting capacity analysis -Value of DERs -Grid needs assessment -Cost-effectiveness framework for investments -Distribution system investment strategy and implementation -Geotargeted programs -Non-wires alternatives procurements.

24 POWER TRANSMISSION AND DISTRIBUTION↗

A relational data-knowledge base system and its potential in developing a distributed data-knowledge system

A new approach used in constructing a rational data knowledge base system is described. The relational database is well suited for distribution due to its property of allowing data fragmentation and fragmentation transparency. An example is formulated of a simple relational data knowledge base which may be generalized for use in developing a relational distributed data knowledge base system. The efficiency and ease of application of such a data knowledge base management system is briefly discussed. Also discussed are the potentials of the developed model for sharing the data knowledge base as well as the possible areas of difficulty in implementing the relational data knowledge base management system.

Rahimian, Eric N.↗

An automated calibration laboratory for flight research instrumentation: Requirements and a proposed design approach

NASA's Dryden Flight Research Facility (Ames-Dryden), operates a diverse fleet of research aircraft which are heavily instrumented to provide both real time data for in-flight monitoring and recorded data for postflight analysis. Ames-Dryden's existing automated calibration (AUTOCAL) laboratory is a computerized facility which tests aircraft sensors to certify accuracy for anticipated harsh flight environments. Recently, a major AUTOCAL lab upgrade was initiated; the goal of this modernization is to enhance productivity and improve configuration management for both software and test data. The new system will have multiple testing stations employing distributed processing linked by a local area network to a centralized database. The baseline requirements for the new AUTOCAL lab and the design approach being taken for its mechanization are described.

Oneill-Rood, Nora↗

A hierarchical distributed control model for coordinating intelligent systems

A hierarchical distributed control (HDC) model for coordinating cooperative problem-solving among intelligent systems is described. The model was implemented using SOCIAL, an innovative object-oriented tool for integrating heterogeneous, distributed software systems. SOCIAL embeds applications in 'wrapper' objects called Agents, which supply predefined capabilities for distributed communication, control, data specification, and translation. The HDC model is realized in SOCIAL as a 'Manager'Agent that coordinates interactions among application Agents. The HDC Manager: indexes the capabilities of application Agents; routes request messages to suitable server Agents; and stores results in a commonly accessible 'Bulletin-Board'. This centralized control model is illustrated in a fault diagnosis application for launch operations support of the Space Shuttle fleet at NASA, Kennedy Space Center.

Adler, Richard M.↗

An automated calibration laboratory - Requirements and design approach

NASA's Dryden Flight Research Facility (Ames-Dryden), operates a diverse fleet of research aircraft which are heavily instrumented to provide both real time data for in-flight monitoring and recorded data for postflight analysis. Ames-Dryden's existing automated calibration (AUTOCAL) laboratory is a computerized facility which tests aircraft sensors to certify accuracy for anticipated harsh flight environments. Recently, a major AUTOCAL lab upgrade was initiated; the goal of this modernization is to enhance productivity and improve configuration management for both software and test data. The new system will have multiple testing stations employing distributed processing linked by a local area network to a centralized database. The baseline requirements for the new AUTOCAL lab and the design approach being taken for its mechanization are described.

O'Neil-Rood, Nora↗

Simple, Scalable, Script-Based Science Processor (S4P)

The development and deployment of data processing systems to process Earth Observing System (EOS) data has proven to be costly and prone to technical and schedule risk. Integration of science algorithms into a robust operational system has been difficult. The core processing system, based on commercial tools, has demonstrated limitations at the rates needed to produce the several terabytes per day for EOS, primarily due to job management overhead. This has motivated an evolution in the EOS Data Information System toward a more distributed one incorporating Science Investigator-led Processing Systems (SIPS). As part of this evolution, the Goddard Earth Sciences Distributed Active Archive Center (GES DAAC) has developed a simplified processing system to accommodate the increased load expected with the advent of reprocessing and launch of a second satellite. This system, the Simple, Scalable, Script-based Science Processor (S42) may also serve as a resource for future SIPS. The current EOSDIS Core System was designed to be general, resulting in a large, complex mix of commercial and custom software. In contrast, many simpler systems, such as the EROS Data Center AVHRR IKM system, rely on a simple directory structure to drive processing, with directories representing different stages of production. The system passes input data to a directory, and the output data is placed in a "downstream" directory. The GES DAAC's Simple Scalable Script-based Science Processing System is based on the latter concept, but with modifications to allow varied science algorithms and improve portability. It uses a factory assembly-line paradigm: when work orders arrive at a station, an executable is run, and output work orders are sent to downstream stations. The stations are implemented as UNIX directories, while work orders are simple ASCII files. The core S4P infrastructure consists of a Perl program called stationmaster, which detects newly arrived work orders and forks a job to run the appropriate executable (registered in a configuration file for that station). Although S4P is written in Perl, the executables associated with a station can be any program that can be run from the command line, i.e., non-interactively. An S4P instance is typically monitored using a simple Graphical User Interface. However, the reliance of S4P on UNIX files and directories also allows visibility into the state of stations and jobs using standard operating system commands, permitting remote monitor/control over low-bandwidth connections. S4P is being used as the foundation for several small- to medium-size systems for data mining, on-demand subsetting, processing of direct broadcast Moderate Resolution Imaging Spectroradiometer (MODIS) data, and Quick-Response MODIS processing. It has also been used to implement a large-scale system to process MODIS Level 1 and Level 2 Standard Products, which will ultimately process close to 2 TB/day.

Lynnes, Christopher↗

A Unified Level of Service Model for NASA Earth Science Data Stewardship

During the past year, the Interagency Implementation and Concepts Team (IMPACT) reviewed existing service models in use at various NASA data centers in an effort to produce a unified, cohesive, and comprehensive Level of Service Model for all of NASA Distributed Active Archive Centers (DAACs). NASA DAACs are responsible for ensuring NASA Earth Science data are accurately and securely ingested, distributed, supported, and preserved. The term “Service" as used here refers to the spectrum of data management activities and outputs provided by DAACs in support of the data cared fo by each data cente. The unified Level-of- Service (LoS) model described in this presentation utilizes both the NASA-defined data product category and the data processing level to easily identify an appropriate level-of-service to be applied to a data product throughout the full data life cycle. This LoS model is to be used by DAACs when appraising incoming data in order to determine the appropriate and required services to provide. The LoS model utilizes a 3-level system in which services build upon previous levels and thereby require greater commitment and effort both on the part of the DAAC personnel and the data producer at the highest level. The LoS model description also contains examples of ways to communicate with data producers and data users what services can be expected, thereby bringing more consistent user experiences across the enterprise. In this presentation, we will outline the features of the LoS model and describe how it relates to the FAIR data practices and the NOAA Maturity Matrix model.

Smith, Deborah↗

Discrete Event Simulation-Based Timeline Validation Using R2U2

The Gateway Vehicle Systems Manager (VSM), the top-level software control system in a distributed, hierarchical Autonomous System Management Architecture is, like most modern spacecraft software control systems, heavily data-driven. For example, schedules (timelines) will be developed on the ground and, due to the high degree of autonomy, contain complex procedures involving conditional branching, variable timing, and resource contention resolution. In order to verify that an uploaded timeline will function correctly, it is necessary to explore the feasible set of possible executions. While it is possible to test a timeline using a mission simulation, the complexity of the system and duration of a timeline limits the number of trials and therefore the test coverage. To address this problem, the VSM team is using a discrete event system model that can rapidly generate from a timeline sets of event sequences using Monte Carlo techniques. To achieve rapid and trustworthy checking of the event sequences, we use an offline version of the runtime model checking tool R2U2. This presentation describes the approach the VSM team is using to implement the discrete event simulation and evaluate event sequences using R2U2. The presentation will discuss: 1. Description of the timelines by VSM in the context of VSM operations 2. Expansion of a timeline into a sequence of atomic events 3. Adjustment, in the Monte Carlo environment, of an event sequence to account for uncertainty, external events, and failures 4. Definition of R2U2 input and mission-time linear temporal logic files 5. Generation and use of R2U2 verdict sequences 6. Lessons learned and future work

Verification↗

Towards G2G: Systems of Technology Database Systems

We present an approach and methodology for developing Government-to-Government (G2G) Systems of Technology Database Systems. G2G will deliver technologies for distributed and remote integration of technology data for internal use in analysis and planning as well as for external communications. G2G enables NASA managers, engineers, operational teams and information systems to "compose" technology roadmaps and plans by selecting, combining, extending, specializing and modifying components of technology database systems. G2G will interoperate information and knowledge that is distributed across organizational entities involved that is ideal for NASA future Exploration Enterprise. Key contributions of the G2G system will include the creation of an integrated approach to sustain effective management of technology investments that supports the ability of various technology database systems to be independently managed. The integration technology will comply with emerging open standards. Applications can thus be customized for local needs while enabling an integrated management of technology approach that serves the global needs of NASA. The G2G capabilities will use NASA s breakthrough in database "composition" and integration technology, will use and advance emerging open standards, and will use commercial information technologies to enable effective System of Technology Database systems.

Maluf, David A.↗