Search NASA⌕ Search

SEARCH · Search NASA

Results for “Distributed Data Management”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21

Framework for Processing Citizens Science Data for Applications to NASA Earth Science Missions

Citizen science (or crowdsourcing) has drawn much high-level recent and ongoing interest and support. It is poised to be applied, beyond the by-now fairly familiar use of, e.g., Twitter for natural hazards monitoring, to science research, such as augmenting the validation of NASA earth science mission data. This interest and support is seen in the 2014 National Plan for Civil Earth Observations, the 2015 White House forum on citizen science and crowdsourcing, the ongoing Senate Bill 2013 (Crowdsourcing and Citizen Science Act of 2015), the recent (August 2016) Open Geospatial Consortium (OGC) call for public participation in its newly-established Citizen Science Domain Working Group, and NASA's initiation of a new Citizen Science for Earth Systems Program (along with its first citizen science-focused solicitation for proposals). Over the past several years, we have been exploring the feasibility of extracting from the Twitter data stream useful information for application to NASA precipitation research, with both "passive" and "active" participation by the twitterers. The Twitter database, which recently passed its tenth anniversary, is potentially a rich source of real-time and historical global information for science applications. The time-varying set of "precipitation" tweets can be thought of as an organic network of rain gauges, potentially providing a widespread view of precipitation occurrence. The validation of satellite precipitation estimates is challenging, because many regions lack data or access to data, especially outside of the U.S. and in remote and developing areas. Mining the Twitter stream could augment these validation programs and, potentially, help tune existing algorithms. Our ongoing work, though exploratory, has resulted in key components for processing and managing tweets, including the capabilities to filter the Twitter stream in real time, to extract location information, to filter for exact phrases, and to plot tweet distributions. The key step is to process the "precipitation" tweets to be compatible with satellite-retrieved precipitation data. These key components for processing and managing "precipitation" tweets (and additional ones to be developed) are not limited to precipitation, nor are they limited to the Twitter social medium. Indeed, to maximize the value of our work for NASA earth science programs, these components should be generalized and be part of an overall framework for processing citizen science data for science research. In this paper, we outline such a framework.

earth science satellite data↗

A study of a space communication system for the control and monitoring of the electric distribution system. Volume 2: Supporting data and analyses

It is technically feasible to design a satellite communication system to serve the United States electric utility industry's needs relative to load management, real-time operations management, remote meter reading and to determine the costs of various elements of the system. The functions associated with distribution automation and control and communication system requirements are defined. Factors related to formulating viable communication concepts, the relationship of various design factors to utility operating practices, and the results of the cost analysis are discussed The system concept and several ways in which the concept could be integrated into the utility industry are described.

Vaisnys, A.↗

A Data Science and Machine Learning Platform Supporting Large Particle Accelerator Control and Diagnostics Applications Final Report: SBIR Initial Phase II DE-SC0022583

The Machine Learning Data Platform (MLDP) is a product providing full-stack support for data science, Machine Learning, and Artificial Intelligence (ML/AI) applications at particle accelerator and large experimental physics facilities. It supports ML/AI applications from front-end, high-speed acquisition of heterogeneous, time-series data, through data archiving and management, to back-end analysis. The MLDP embodies a “data-science ready” platform for data analysis and ML/AI applications in diagnosis, modelling, control, and optimization of these facilities. It provides data scientists and applications a consistent, datacentric interface to archive data standardizing implementation and deployment of ML/AI algorithms to different operations configurations within the same facility, or between facilities. Being an open-source, public-domain project, the MLDP is intended for broadest possible impact by increasing accessibility and minimizing the required expertise for installation and operation. The MLDP can also be deployed at user facilities for experimental data collection, archiving, and analysis. It is capable of acquisition and archiving of heterogeneous data from experimental equipment (e.g., images, arrays, structures, etc.) along with system hardware configurations (e.g., scalars, tables), control system process variables, and any metadata required for provenance. Thus, the MLDP can manage experimental data through its entire lifecycle, from acquisition and archiving, through analysis and investigation, to release and final publication.

43 PARTICLE ACCELERATORS↗

Advanced data management system architectures testbed

The objective of the Architecture and Tools Testbed is to provide a working, experimental focus to the evolving automation applications for the Space Station Freedom data management system. Emphasis is on defining and refining real-world applications including the following: the validation of user needs; understanding system requirements and capabilities; and extending capabilities. The approach is to provide an open, distributed system of high performance workstations representing both the standard data processors and networks and advanced RISC-based processors and multiprocessor systems. The system provides a base from which to develop and evaluate new performance and risk management concepts and for sharing the results. Participants are given a common view of requirements and capability via: remote login to the testbed; standard, natural user interfaces to simulations and emulations; special attention to user manuals for all software tools; and E-mail communication. The testbed elements which instantiate the approach are briefly described including the workstations, the software simulation and monitoring tools, and performance and fault tolerance experiments.

Grant, Terry↗

Collecting and Processing Earth Science Data Metrics at NASA ESDIS

Since the launch of Terra satellite in 1999, the number of Earth Science remote sensing data products created and distributed by NASA's Earth Observing System (EOS) Data and Information System (EOSDIS) has increased from a few hundred to nearly ten thousand. NASA's Earth Science Data and Information System (ESDIS) Metrics System (EMS) collects metrics on data ingest, archive, and distribution by its Distributed Active Archive Centers (DAACs) and the Science Investigator-led Systems (SIPS), known as Data Providers. These metrics are critical in helping NASA management as well as data producers in resource planning and gaining a wide range of knowledge of data users and data usage.EMS receives flat files, or log files of data archive, ingest, and distribution either in their raw format, such as Apache web logs, or text files of log records formatted by the Data Providers. Tens of millions of records are processed each day to extract metrics on data products, user information, distribution protocols and services, and so on. The metrics are then made available to designated parties.This presentation provides an overview of the EMS processing workflow and improvement efforts made in recent years to handle ever-increasing number of data records and new metrics requirements, discusses several key steps including mapping log records to data products and identifying user communities along with geo-distribution, and demonstrates typical metrics capabilities produced by the EMS system. Challenges and potential approaches to improve the system are also discussed.

Pan, Jianfu↗

Autonomous Information Unit for Fine-Grain Data Access Control and Information Protection in a Net-Centric System

As communication and networking technologies advance, networks will become highly complex and heterogeneous, interconnecting different network domains. There is a need to provide user authentication and data protection in order to further facilitate critical mission operations, especially in the tactical and mission-critical net-centric networking environment. The Autonomous Information Unit (AIU) technology was designed to provide the fine-grain data access and user control in a net-centric system-testing environment to meet these objectives. The AIU is a fundamental capability designed to enable fine-grain data access and user control in the cross-domain networking environments, where an AIU is composed of the mission data, metadata, and policy. An AIU provides a mechanism to establish trust among deployed AIUs based on recombining shared secrets, authentication and verify users with a username, X.509 certificate, enclave information, and classification level. AIU achieves data protection through (1) splitting data into multiple information pieces using the Shamir's secret sharing algorithm, (2) encrypting each individual information piece using military-grade AES-256 encryption, and (3) randomizing the position of the encrypted data based on the unbiased and memory efficient in-place Fisher-Yates shuffle method. Therefore, it becomes virtually impossible for attackers to compromise data since attackers need to obtain all distributed information as well as the encryption key and the random seeds to properly arrange the data. In addition, since policy can be associated with data in the AIU, different user access and data control strategies can be included. The AIU technology can greatly enhance information assurance and security management in the bandwidth-limited and ad hoc net-centric environments. In addition, AIU technology can be applicable to general complex network domains and applications where distributed user authentication and data protection are necessary. AIU achieves fine-grain data access and user control, reducing the security risk significantly, simplifying the complexity of various security operations, and providing the high information assurance across different network domains.

Chow, Edward T.↗

Climatespark: an In-Memory Distributed Computing Framework for Big Climate Data Analytics

The unprecedented growth of climate data creates new opportunities for climate studies, and yet big climate data pose a grand challenge to climatologists to efficiently manage and analyze big data. The complexity of climate data content and analytical algorithms increases the difficulty of implementing algorithms on high performance computing systems. This paper proposes an in-memory, distributed computing framework, ClimateSpark, to facilitate complex big data analytics and time-consuming computational tasks. Chunking data structure improves parallel I/O efficiency, while a spatiotemporal index is built for the chunks to avoid unnecessary data reading and preprocessing. An integrated, multi-dimensional, array-based data model (ClimateRDD) and ETL operations are developed to address big climate data variety by integrating the processing components of the climate data lifecycle. ClimateSpark utilizes Spark SQL and Apache Zeppelin to develop a web portal to facilitate the interaction among climatologists, climate data, analytic operations and computing resources (e.g., using SQL query and Scala/Python notebook). Experimental results show that ClimateSpark conducts different spatiotemporal data queries/analytics with high efficiency and data locality. ClimateSpark is easily adaptable to other big multiple- dimensional, array-based datasets in various geoscience domains.

Hu, Fei↗

Early-EOS data and information system

NASA's Earth Observing System (EOS), an integral part of the U.S. Global Change Research Program, will provide simultaneous observations from a suite of instruments in low-earth orbit. The EOS Data and Information System (EOSDIS) will handle the data from those instruments, as well as provide access to observations and related information from other earth science missions. The Early-EOSDIS Program will provide initial improved support for global change research by building upon present capabilities and data, and will establish a working prototype EOSDIS for selected archiving, distribution, and information management functions by mid-1994.

Ludwig, George H.↗

Integrating International Space Station payload operations

The payload operations support for the International Space Station (ISS) payload is reported on, describing payload activity planning, payload operations control, payload data management and overall operations integration. The operations concept employed is based on the distribution of the payload operations responsibility between the researchers and ISS partners. The long duration nature of the ISS mission dictates the geographical distribution of the payload operations activities between the different national centers. The coordination and integration of these operations will be assured by NASA's Payload Operations Integration Center (POIC). The prime objective of the POIC is the achievement of unified operations through communication and collaboration.

Noneman, Steven R.↗

Participation in the Cluster Magnetometer Consortium for the Cluster Mission

Prof. M. G. Kivelson (UCLA) and Dr. R. C. Elphic (LANL) are Co-investigators on the Cluster Magnetometer Consortium (CMC) that provided the fluxgate magnetometers and associated mission support for the Cluster Mission. The CMC designated UCLA as the site with primary responsibility for the inter-calibration of data from the four spacecraft and the production of fully corrected data critical to achieving the mission objectives. UCLA was also charged with distributing magnetometer data to the U.S. Co-investigators. UCLA also supported the Technical Management Team, which was responsible for the detailed design of the instrument and its interface. In this final progress report we detail the progress made by the UCLA team in achieving the mission objectives.

Kivelson, Margaret↗

A Framework to Demonstrate a DNP3 Interface With a CIM-Based Data Integration Platform: Preprint

The contemporary electrical grid is characterized by its complexity and abundance of data. A control-rich environment supported by information and communication technologies within an Advanced Distribution Management System (ADMS) presents a viable and cost-effective option for utility companies aiming to implement advanced real-time analytical schemes for monitoring and remotely controlling distribution feeders. Modular platform-based approaches to distribution operations require a structured framework for acquiring field device measurements, performing analytics, converting the setpoint to the correct protocol, and sending it on the appropriate communications network to the field devices. We present the development and deployment of an application service to integrate an open-source standardsbased platform with an ADMS test bed with field devices using the Distributed Network Protocol (DNP3) for data exchange. The step-by-step procedure for establishing the DNP3-Master service on an open-source distribution platform is outlined, comprehensively explaining the Master setup process. Moreover, sample use case results highlight the capabilities of the DNP3- Master service setup. Results demonstrate the scalability and configurability of the DNP3-Master service, making it adaptable for integration with other relevant applications, thus providing potential opportunities for real-world field trials and real-time assessments.

ADMS↗

The Application of Remotely Sensed Data and Models to Benefit Conservation and Restoration Along the Northern Gulf of Mexico Coast

New data, tools, and capabilities for decision making are significant needs in the northern Gulf of Mexico and other coastal areas. The goal of this project is to support NASA s Earth Science Mission Directorate and its Applied Science Program and the Gulf of Mexico Alliance by producing and providing NASA data and products that will benefit decision making by coastal resource managers and other end users in the Gulf region. Data and research products are being developed to assist coastal resource managers adapt and plan for changing conditions by evaluating how climate changes and urban expansion will impact land cover/land use (LCLU), hydrodynamics, water properties, and shallow water habitats; to identify priority areas for conservation and restoration; and to distribute datasets to end-users and facilitating user interaction with models. The proposed host sites for data products are NOAA s National Coastal Data Development Center Regional Ecosystem Data Management, and Mississippi-Alabama Habitat Database. Tools will be available on the Gulf of Mexico Regional Collaborative website with links to data portals to enable end users to employ models and datasets to develop and evaluate LCLU and climate scenarios of particular interest. These data will benefit the Mobile Bay National Estuary Program in ongoing efforts to protect and restore the Fish River watershed and around Weeks Bay National Estuarine Research Reserve. The usefulness of data products and tools will be demonstrated at an end-user workshop.

Quattrochi, Dale↗

Distribution to the Astronomy Community of the Compressed Digitized Sky Survey

The Space Telescope Science Institute has compressed an all-sky collection of ground-based images and has printed the data on a two volume, 102 CD-ROM disc set. The first part of the survey (containing images of the southern sky) was published in May 1994. The second volume (containing images of the northern sky) was published in January 1995. Software which manages the image retrieval is included with each volume. The Astronomical Society of the Pacific (ASP) is handling the distribution of the lOx compressed data and has sold 310 sets as of October 1996. ASP is also handling the distribution of the recently published 100x version of the northern sky survey which is publicly available at a low cost. The target markets for the 100x compressed data set are the amateur astronomy community, educational institutions, and the general public. During the next year, we plan to publish the first version of a photometric calibration database which will allow users of the compressed sky survey to determine the brightness of stars in the images.

Postman, Marc↗

An Autonomous MCP Bridge to Rucio: Enhancing Data Management Accessibility for High Energy Physics

The Rucio Data Management System [1] is an important tool used by High Energy Physics experiments, including those at Fermi National Accelerator Laboratory, to store and manage exabyte-scale scientific datasets. Despite its central role in coordinating data across globally distributed storage sites, Rucio's command line interface (CLI) presents a steep learning curve, and makes it difficult for scientists to navigate through. To solve this issue, a containerized Model Context Protocol (MCP) [2] server was built that connects Large Language Models directly to Rucio, allowing AI agents to handle data tasks by using simple, natural language rather than memorized terminal commands. The core engineering focus of this project was moving the server away from slow terminal commands that require text parsing and replacing them with a native Python Client API toolset and a planned REST API framework. Moving to the Python API handles data operations directly in memory, which helps clear up formatting errors, provides the AI with clean, structured JSON data and speeds up tool execution. To prove that the system actually works, a benchmarking pipeline was also built with various questions to test the AI across four different model configurations. The questions included finding data scopes, tracking down specific datasets, and checking replication rules. Through benchmarking, early runs showed that with raw terminal text, the model would get confused and stuck, whereas switching to the Python API to feed the AI clean, structured data yielded massive improvement. By creating an intelligent and autonomous bridge to a storage network, this project shows how AI can be implemented in scientific data management, which ultimately helps scientists at Fermilab spend less time sorting through data and more time focusing on their experiments and analysis.

Akella, Kashyap [William Rainey Harper Coll.]↗

Distribution of Cost Growth in Robotic Space Science Missions

Cost growth characterization is a critical factor for effective cost risk analysis and project planning. This study analyzed low level budget changes in Jet Propulsion Laboratory-managed space science missions, which occurred during the development of the project. The data was then curve fit, according to cost distribution categories, to provide a reference set of distribution parameters with sufficient granularity to effectively model cost growth in robotic space science missions.

cost↗

Integrated Distribution Planning

The contemporary distribution planning landscape is comprised of an increasing number of factors that require integration into the engineering of the modern electric grid. Expectations for electric utilities to accommodate heightened awareness of stakeholders' interest in things like decarbonization, resilience and equity are growing. As these interests are formed into objectives, many jurisdictions will experience increasing levels of load modifying technologies like DER, building and industrial electrification and electric vehicles which prove not only to challenge the capabilities of the grid; but the processes by which planning for it is traditionally done. Other related factors that strain the conventional distribution planning mold are the swelling amount and sources of data associated with these technologies and the need it creates for improved capabilities in the processes and tools that manage it. As the complexity of the distribution system expands, so will the distribution system's effects on the transmission and generation systems that it is a part of. Forecasting distribution system load and DER are examples of areas where this complexity will manifest, and harmonizing distribution forecasting with transmission and generation forecasting requires higher amounts of intentionality as these typically separate processes become a solitary one. Of course, core activities do not cease as a utility begins to integrate these other factors, and in this webinar we explore specifics of how distribution planning can be expected to evolve as progress towards Integrated Distribution System Planning is made.

24 POWER TRANSMISSION AND DISTRIBUTION↗

NASA Occupational Health Program FY98 Self-Assessment

The NASA Functional Management Review process requires that each NASA Center conduct self-assessments of each functional area. Self-Assessments were completed in June 1998 and results were presented during this conference session. During FY 97 NASA Occupational Health Assessment Team activities, a decision was made to refine the NASA Self-Assessment Process. NASA Centers were involved in the ISO registration process at that time and wanted to use the management systems approach to evaluate their occupational health programs. This approach appeared to be more consistent with NASA's management philosophy and would likely confer status needed by Senior Agency Management for the program. During FY 98 the Agency Occupational Health Program Office developed a revised self-assessment methodology based on the Occupational Health and Safety Management System developed by the American Industrial Hygiene Association. This process was distributed to NASA Centers in March 1998 and completed in June 1998. The Center Self Assessment data will provide an essential baseline on the status of OHP management processes at NASA Centers. That baseline will be presented to Enterprise Associate Administrators and DASHO on September 22, 1998 and used as a basis for discussion during FY 99 visits to NASA Centers. The process surfaced several key management system elements warranting further support from the Lead Center. Input and feedback from NASA Centers will be essential to defining and refining future self assessment efforts.

Brisbin, Steven G.↗

Integrated System Health Management: Foundational Concepts, Approach, and Implementation.

Implementation of integrated system health management (ISHM) capability is fundamentally linked to the management of data, information, and knowledge (DIaK) with the purposeful objective of determining the health of a system. It is akin to having a team of experts who are all individually and collectively observing and analyzing a complex system, and communicating effectively with each other in order to arrive to an accurate and reliable assessment of its health. We present concepts, procedures, and a specific approach as a foundation for implementing a credible ISHM capability. The capability stresses integration of DIaK from all elements of a system. The intent is also to make possible implementation of on-board ISHM capability, in contrast to a remote capability. The information presented is the result of many years of research, development, and maturation of technologies, and of prototype implementations in operational systems (rocket engine test facilities). The paper will address the following topics: 1. ISHM Model of a system 2. Detection of anomaly indicators. 3. Determination and confirmation of anomalies. 4. Diagnostic of causes and determination of effects. 5. Consistency checking cycle. 6. Management of health information 7. User Interfaces 8. Example implementation ISHM has been defined from many perspectives. We define it as a capability that might be achieved by various approaches. We describe a specific approach that has been matured throughout many years of development, and pilot implementations. ISHM is a capability that is achieved by integrating data, information, and knowledge (DIaK) that might be distributed throughout the system elements (which inherently implies capability to manage DIaK associated with distributed sub-systems). DIaK must be available to any element of a system at the right time and in accordance with a meaningful context. ISHM Functional Capability Level (FCL) is measured by how well a system performs the following functions: (1) detect anomalies, (2) diagnose causes, (3) predict future anomalies/failures, and (4) provide the user with an integrated awareness about the condition of every element in the system and guide user decisions.

Figueroa, Fernando↗