Search NASA⌕ Search

SEARCH · Search NASA

Results for “federated computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Three Dimensional Computer Graphics Federates for the 2012 Smackdown Simulation

The Simulation Interoperability Standards Organization (SISO) Smackdown is a two-year old annual event held at the 2012 Spring Simulation Interoperability Workshop (SIW). A primary objective of the Smackdown event is to provide college students with hands-on experience in developing distributed simulations using High Level Architecture (HLA). Participating for the second time, the University of Alabama in Huntsville (UAHuntsville) deployed four federates, two federates simulated a communications server and a lunar communications satellite with a radio. The other two federates generated 3D computer graphics displays for the communication satellite constellation and for the surface based lunar resupply mission. Using the Light-Weight Java Graphics Library, the satellite display federate presented a lunar-texture mapped sphere of the moon and four Telemetry Data Relay Satellites (TDRS), which received object attributes from the lunar communications satellite federate to drive their motion. The surface mission display federate was an enhanced version of the federate developed by ForwardSim, Inc. for the 2011 Smackdown simulation. Enhancements included a dead-reckoning algorithm and a visual indication of which communication satellite was in line of sight of Hadley Rille. This paper concentrates on these two federates by describing the functions, algorithms, HLA object attributes received from other federates, development experiences and recommendations for future, participating Smackdown teams.

Fordyce, Crystal↗

Designing resilient IoT and Edge Computing with federated tinyML

The rapid growth of the Internet of Things (IoT) and Edge Computing (EC) has brought significant conveniences to modern society but has also greatly expanded the cyber attack surfaces, particularly as these technologies are being increasingly integrated into critical systems such as power grids, healthcare, and smart homes. Here, to improve IoT/EC’s cybersecurity posture, we leveraged Artificial Intelligence (AI) and Machine Learning (ML) by employing tinyML to monitor voluminous IoT data for cyber threats while addressing devices’ resource constraints, and utilizing Federated Learning (FL) to share local detection knowledge across the system while preserving privacy. Building on our three-layer architecture combining tinyML and FL to enhance autonomous cyber attack detection, this paper demonstrated that the architecture improves detection accuracy, reduces resource consumption, and enables lightweight, secure IoT device monitoring. These results were validated using the public N-BaIoT dataset as well as real IoT network traffic data collected under multiple attack scenarios from our testbeds. Additionally, we introduced an enhanced FL methodology with a novel preprocessing stage, including federated feature selection and global preprocessor construction, to address IoT/EC data heterogeneity. We developed a physical IoT testbed for attack simulations and data collection, implemented a tinyML-powered detector for realistic model validation, and also built a virtual testbed for scalable evaluations of FL models across diverse network environments.

Cognitive cyber↗

CAFE AU LAIT: Compute-Aware Federated Augmented Low-Rank AI Training

Federated finetuning is crucial for unlocking the knowledge embedded in pretrained Large Language Models (LLMs) when data are geographically distributed across clients. Unlike finetuning with data from a single institution, federated finetuning allows collaboration across multiple institutions, enabling the utilization of diverse and decentralized datasets while preserving data privacy. Given the high computing costs of LLM training and the emphasis on energy efficiency in Federated Learning (FL), Low-Rank Adaptation (LoRA) has emerged as a widely adopted algorithm due to its significantly reduced number of trainable parameters. However, this assumes that all data silos have the necessary computing resources to compute local updates of LLMs. Nevertheless, in practice, the computing resources across clients are highly heterogeneous: while some may have access to hundreds of GPUs, others might have limited or no GPU access. Recently, federated finetuning using synthetic data has been proposed, allowing clients to participate in a collaborative training run without training LLMs locally. However, our experimental results reveal a performance gap between models trained using synthetic data and those trained using local updates. Motivated by the observed heterogeneity in computing resources and the performance gap, we propose a novel two-stage algorithm that leverages the storage and computing capabilities of a strong server. In the first stage, under the coordination of the strong server, clients with limited computing resources collaborate to generate synthetic data, which is transferred to and stored on the strong server. In the second stage, the strong server uses this synthetic data on behalf of the resource-constrained clients to perform federated LoRA finetuning alongside clients with sufficient computing resources. This approach ensures that all clients can participate in the finetuning process. Experimental results demonstrate that incorporating local updates from even a small fraction of clients improves performance compared to using synthetic data for all clients. Furthermore, we incorporate the Gaussian mechanism in both stages to guarantee client-level differential privacy.

Wang, Jiayi [ORNL]↗

Virtual Infrastructure Twins: Software Testing Platforms for Computing-Instrument Ecosystems

Science ecosystems are being built by federating computing systems and instruments located at geographically distributed sites over wide-area networks. These computing-instrument ecosystems are expected to support complex workflows that incorporate remote, automated AI-driven science experiments. Their realization, however, requires various designs to be explored and software components to be developed, in order to support the orchestration of distributed computations and experiments. It is often too expensive, infeasible, or disruptive for the entire ecosystem to be available during the typically long software development and testing periods. We propose a Virtual Infrastructure Twin (VIT) of the ecosystem that emulates its network and computing components, and incorporates its instrument software simulators. It provides a software environment nearly identical to the ecosystem to support early development and testing, and design space exploration. We present a brief overview of previous digital infrastructure twins that culminated in the VIT concept, including (i) the virtual science network environment for developing software-defined networking solutions, and (ii) the virtual federated science instrument environment for testing the federation software stack and remote instrument control software. We briefly describe VITs for Nion microscope steering and access to GPU systems.

Rao, Nageswara↗

Virtual Infrastructure Twins: Software Testing Platforms for Computing-Instrument Ecosystems

Science ecosystems are being built by federating computing systems and instruments located at geographically distributed sites over wide-area networks. These computing-instrument ecosystems are expected to support complex workflows that incorporate remote, automated AI-driven science experiments. Their realization, however, requires various designs to be explored and software components to be developed, in order to support the orchestration of distributed computations and experiments. It is often too expensive, infeasible, or disruptive for the entire ecosystem to be available during the typically long software development and testing periods. We propose a Virtual Infrastructure Twin (VIT) of the ecosystem that emulates its network and computing components, and incorporates its instrument software simulators. It provides a software environment nearly identical to the ecosystem to support early development and testing, and design space exploration. We present a brief overview of previous digital infrastructure twins that culminated in the VIT concept, including (i) the virtual science network environment for developing software-defined networking solutions, and (ii) the virtual federated science instrument environment for testing the federation software stack and remote instrument control software. We briefly describe VITs for Nion microscope steering and access to GPU systems.

Rao, Nageswara↗

Game-Theoretic Strategies for Cyber-Physical Infrastructures Under Component Disruptions

In this work, networked infrastructures of recursively defined systems composed of discrete cyber and physical components are considered. The components of basic systems at the finest levels can be disrupted by cyber or physical means, and can be reinforced to survive at certain costs. A problem of ensuring the infrastructure performance is formulated as a game between a provider and an attacker, who probabilistically choose components to reinforce and attack, respectively. The disruptions of this infrastructure are characterized using the aggregate failure correlation function that specifies the conditional failure probability of the infrastructure given the failure of an individual system at that level. The survival probabilities of basic systems satisfy simple product-form, first-order differential equations expressed in terms of the multiplier functions. The utility functions of the provider and attacker are composed of the reward and cost terms, both expressed in terms of the component reinforcement and attack probabilities. The Nash equilibrium of this game is characterized, along with the sensitivity functions of the survival probabilities of basic systems that highlight their dependence individually on the cost-benefit terms, the correlation functions, and the multiplier functions. These results are illustrated using simplified models of a distributed cloud servers infrastructure, a 5G data network infrastructure, a high performance computing federation, and a smart energy grid infrastructure.

42 ENGINEERING↗

Impact of remote sensing upon the planning, management, and development of water resources

A survey of the principal water resource users was conducted to determine the impact of new remote data streams on hydrologic computer models. The analysis of the responses and direct contact demonstrated that: (1) the majority of water resource effort of the type suitable to remote sensing inputs is conducted by major federal water resources agencies or through federally stimulated research, (2) the federal government develops most of the hydrologic models used in this effort; and (3) federal computer power is extensive. The computers, computer power, and hydrologic models in current use were determined.

Loats, H. L.↗

RuralAI in Tomato Farming: Integrated Sensor System, Distributed Computing, and Hierarchical Federated Learning for Crop Health Monitoring

Precision horticulture is evolving due to scalable sensor deployment and machine learning (ML) integration. These advancements boost the operational efficiency of individual farms, balancing the benefits of analytics with autonomy requirements. However, given concerns that affect wide geographic regions (e.g., climate change), there is a need to apply models that span farms. Federated learning (FL) has emerged as a potential solution. FL enables decentralized ML across different farms without sharing private data. Traditional FL assumes simple two-tier network topologies and, thus, falls short of operating on more complex networks found in real-world agricultural scenarios. Networks vary across crops and farms and encompass various sensor data modes, extending across jurisdictions. New hierarchical FL (HFL) approaches are needed for more efficient and context-sensitive model sharing, accommodating regulations across multiple jurisdictions. Here, we present the RuralAI architecture deployment for tomato crop monitoring, featuring sensor field units for soil, crop, and weather data collection. HFL with personalization is used to offer localized and adaptive insights. Model management, aggregation, and transfers are facilitated via a flexible approach, enabling seamless communication between local devices, edge nodes, and the cloud.

60 APPLIED LIFE SCIENCES↗

Secure Federated Learning Across Heterogeneous Cloud and High-Performance Computing Resources: A Case Study on Federated Fine-Tuning of LLaMA 2

Federated learning enables multiple data owners to collaboratively train robust machine learning models without transferring large or sensitive local datasets by only sharing the parameters of the locally trained models. Here, in this article, we elaborate on the design of our Advanced Privacy-Preserving Federated Learning (APPFL) framework, which streamlines end-to-end secure and reliable federated learning experiments across cloud computing facilities and high-performance computing resources by leveraging Globus Compute, a distributed function as a service platform, and Amazon Web Services. We further demonstrate the use case of APPFL in fine-tuning an LLaMA 2 7B model using several cloud resources and supercomputers.

97 MATHEMATICS AND COMPUTING↗

Cognitive IoT and Edge Computing for Intrusion Detection with Federated TinyML

Internet of Things (IoT) and Edge Computing (EC) are rapidly becoming an integral part of the modern society. By 2030, there is estimated to be over 40 billion active and connected IoT devices [1]. This rapid progress also comes with a significant implication on cybersecurity. Back-end infrastructure and systems have a much broader attack than they did previously due to vulnerable IoT/EC devices being connected to wireless networks. This expanding attack surface is a growing concern because IoT/EC are increasingly being used in critical systems such as power grids, health care, and smart homes. To effectively address a problem of this scale, cognitive cyber methods—which can autonomously detect and react to cyber attacks as they develop—are needed. To address this, we bring Artificial Intelligence (AI) and Machine Learning (ML) to IoT/EC devices, using tinyML to monitor voluminous IoT data against cyber threats, and using Federated Learning (FL) to share local detection knowledge across the system while preserving privacy. We propose a novel three-layer architecture: (1) an IoT layer for tinyML-based inference, (2) an edge layer for ML model training, and (3) a cloud layer for FL operations. Using the publicly available 11-class N-BaIoT dataset [2], we demonstrate that this architecture mitigates resource constraints at the IoT layer while improving detection accuracy over standard two-layer designs. An outlier-resistant scaler, feature reduction, and quantization enable the tinyML model to maintain detection accuracy with a reduced model size. Additionally, federated learning that only utilizes the intersection (across heterogenous devices) of the reduced feature set achieves superior detection accuracy compared to locally trained models.

Li, Mingyan [ORNL] (ORCID:0009000569532640)↗

Emulation Framework for Distributed Large-Scale Systems Integration

Recent trends in systems engineering include integration of very large-scale systems, which entails significant challenges when they are geographically dispersed. In these scenarios, intelligent integration of distributed large-scale systems requires significant coordination among hardware elements as well as all software components. The approach of integrated systems (both computing platform and experimental equipment) for end-to-end orchestration is called federation. Virtual frameworks can aid in the testing, assessment, and implementation of a functional system of interconnected resources. We present an emulation framework that replicates the software environments of multi-site federations of computing systems and instruments. Our emulation framework allows systems engineers to reduce developmentcost and avoid disruptions to production infrastructure. Our framework was effectively used to develop and test software modules for various tasks including container orchestration and instrument access. For performance assessment, however, the emulated framework is severely limited in providing accurate network and IO measurements at 10 Gbps and higher data rates. The data transfer performance profiles estimated using these emulated measurements are usually inaccurate for high bandwidth and high latency connections, since emulation does not accurately reflect the critical network transport dynamics.We utilize measurements from a physical testbed with hardware network emulators to obtain data transfer profiles that closely match the expected profiles for the emulated federations. We show the effectiveness of our approach by an illustrative example of integrated (federated) multi-site ultra large-scale systems that are connected via high speed wide area networks.

Imam, Neena↗

Automated, reliable, and efficient continental-scale replication of 7.3 petabytes of computational simulation data: A case study

We report on our experiences replicating 7.3 petabytes (PB) of Earth System Grid Federation (ESGF) computational simulation data from Lawrence Livermore National Laboratory (LLNL) in California to Argonne National Laboratory (ANL) in Illinois and Oak Ridge National Laboratory (ORNL) in Tennessee—a task motivated by a need for increased reliability, capacity, and performance. This task presented significant challenges: the need to move 29 million files twice under time pressure from aging storage hardware; a source file system bottleneck limiting throughput to 1.5 GB/s; frequent site maintenance windows; and the need for complete reliability at scale. We addressed these challenges using a simple replication tool that invoked Globus to transfer large bundles of files while tracking progress in a database, dynamically rerouting transfers to work around maintenance periods and file system limitations. Under the covers, Globus organized transfers to make efficient use of the high-speed Energy Sciences network (ESnet) and the data transfer nodes deployed at participating sites, and also addressed security, integrity checking, and recovery from a variety of transient failures. This success demonstrates the considerable benefits that can accrue from the adoption of performant data replication infrastructure. The replication tool is available at https://github.com/esgf2-us/data-replication-tools.

Globus↗

NASA Tech Briefs, July 1995

Topics include: mechanical components, electronic components and circuits, electronic systems, physical sciences, materials, computer programs, mechanics, machinery, manufacturing/fabrication, mathematics and information sciences, book and reports, and a special section of Federal laboratory computing Tech Briefs.

Source record↗

Impact of remote sensing upon the planning, management, and development of water resources

Principal water resources users were surveyed to determine the impact of remote data streams on hydrologic computer models. Analysis of responses demonstrated that: most water resources effort suitable to remote sensing inputs is conducted through federal agencies or through federally stimulated research; and, most hydrologic models suitable to remote sensing data are federally developed. Computer usage by major water resources users was analyzed to determine the trends of usage and costs for the principal hydrologic users/models. The laws and empirical relationships governing the growth of the data processing loads were described and applied to project the future data loads. Data loads for ERTS CCT image processing were computed and projected through the 1985 era.

Castruccio, P. A.↗

Performance Evaluation Tools for Next Generation Scalable Computing Platforms

The Federal High Performance and Communications (HPCC) Program continue to focus on R&D in a wide range of high performance computing and communications technologies. Using its accomplishments in the past four years as building blocks towards a Global Information Infrastructure (GII), an Implementation Plan that identifies six Strategic Focus Areas for R&D has been proposed. This white paper argues that a new generation of system software and programming tools must be developed to support these focus areas, so that the R&D we invest today can lead to technology pay-off a decade from now. The Global Computing Infrastructure (GCI) in the Year 2000 and Beyond would consists of thousands of powerful computing nodes connected via high-speed networks across the globe. Users will be able to obtain computing in formation services the GCI with the ease of using a plugging a toaster into the electrical outlet on the wall anywhere in the country. Developing and managing the GO requires performance prediction and monitoring capabilities that do not exist. Various accomplishments in this field today must be integrated and expanded to support this vision.

Yan, Jerry C.↗

Using high-performance networks to enable computational aerosciences applications

One component of the U.S. Federal High Performance Computing and Communications Program (HPCCP) is the establishment of a gigabit network to provide a communications infrastructure for researchers across the nation. This gigabit network will provide new services and capabilities, in addition to increased bandwidth, to enable future applications. An understanding of these applications is necessary to guide the development of the gigabit network and other high-performance networks of the future. In this paper we focus on computational aerosciences applications run remotely using the Numerical Aerodynamic Simulation (NAS) facility located at NASA Ames Research Center. We characterize these applications in terms of network-related parameters and relate user experiences that reveal limitations imposed by the current wide-area networking infrastructure. Then we investigate how the development of a nationwide gigabit network would enable users of the NAS facility to work in new, more productive ways.

Johnson, Marjory J.↗

Multi-Kernel Adaptive Support Vector Machine for Scalable Predictive Maintenance

Application of data-driven solutions across an industry is challenging, since the data are often stored locally, and increasing privacy and security concerns restrict access to the data. In addition, it is highly unlikely that all potential data patterns are captured in a single data source. Because it is highly unlikely that all potential data patterns are captured in a single data source, machine learning (ML) models developed from a single source cannot be robust enough. An alternative is to train the ML model at each source and develop a distributed knowledge discovery and aggregation approach to build global knowledge. In this paper, we develop and demonstrate a distributed ML model, federated transfer learning (FTL), using a multi-kernel-based adaptive support vector machine (MK-A-SVM). For federated learning (FL), the multi-kernel (MK) approach enables feature-specific model aggregation under data heterogeneity; whereas for transfer learning (TL) the adaptive model enables utilization of an aggregated model from a different task. The proposed approach is validated using nuclear power plant (NPP) vertical motor-driven pump data to predict the health condition of vertical motor-driven pumps as an anomaly detection. The efficiency of the proposed approach is also quantified and compared with neural network.

42 ENGINEERING↗

Scientific computing plan for the ECCE detector at the Electron Ion Collider

The Electron Ion Collider (EIC) is the next generation of precision QCD facility to be built at Brookhaven National Laboratory in conjunction with Thomas Jefferson National Laboratory. There are a significant number of software and computing challenges that need to be overcome at the EIC. During the EIC detector proposal development period, the ECCE consortium began identifying and addressing these challenges in the process of producing a complete detector proposal based upon detailed detector and physics simulations. Here, in this document, the software and computing efforts to produce this proposal are discussed; furthermore, the computing and software model and resources required for the future of ECCE are described.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗