Search NASA⌕ Search

SEARCH · Search NASA

Results for “Distributed Computing Resources”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Optimization of distributed compute resources utilization in the CMS Global Pool

The CMS Submission Infrastructure is the primary system for managing computing resources for CMS workflows, including data processing, simulation, and analysis. It integrates geographically distributed resources from Grid, HPC, and cloud providers into federated pools managed by HTCondor and Glidein- WMS, for a total of around 500k CPU cores. This system dynamically manages workloads based on priorities defined by the collaboration. Additionally, CMS scheduling strategies must be flexible to handle multiple concurrent workloads while considering changing processing demands and resource availability from various providers.Efficient utilization of vast amounts of distributed compute resources is a key element for the success of the scientific programs of the LHC experiments. Optimizing the system is essential to maximize resource efficiency and fully utilize the distributed computing power. The CMS Submission Infrastructure team thus systematically investigates sources of inefficiency in workload scheduling to reduce their impact. In addition, a strategy of pilot overloading has been introduced to compensate for other inefficiency sources, thereby optimizing resource utilization and enhancing computational throughput.

Mascheroni, Marco [UC, San Diego (main)]↗

Accelerating science: The usage of commercial clouds in ATLAS Distributed Computing

The ATLAS experiment at CERN is one of the largest scientific machines built to date and will have ever growing computing needs as the Large Hadron Collider collects an increasingly larger volume of data over the next 20 years. ATLAS is conducting R&D projects on Amazon Web Services and Google Cloud as complementary resources for distributed computing, focusing on some of the key features of commercial clouds: lightweight operation, elasticity and availability of multiple chip architectures. The proof of concept phases have concluded with the cloud-native, vendoragnostic integration with the experiment’s data and workload management frameworks. Google Cloud has been used to evaluate elastic batch computing, ramping up ephemeral clusters of up to O(100k) cores to process tasks requiring quick turnaround. Amazon Web Services has been exploited for the successful physics validation of the Athena simulation software on ARM processors. We have also set up an interactive facility for physics analysis allowing endusers to spin up private, on-demand clusters for parallel computing with up to 4 000 cores, or run GPU enabled notebooks and jobs for machine learning applications. The success of the proof of concept phases has led to the extension of the Google Cloud project, where ATLAS will study the total cost of ownership of a production cloud site during 15 months with 10k cores on average, fully integrated with distributed grid computing resources and continue the R&D projects.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Exploiting Kubernetes to Simplify the Deployment and Management of the Multi-purpose CMS Pilot Job Factory

GlideinWMS, a widely utilized workload management system in high-energy physics (HEP) research, serves as the backbone for efficient job provisioning across distributed computing resources. It is utilized by various experiments and organizations, including CMS, OSG, Dune, and FIFE, to create HTCondor pools as large as 600k cores. In particular, a shared factory service historically deployed at UCSD has been configured to interface with more than 500 routes to compute clusters. As part of our team’s initiative to modernize infrastructure and enhance scalability, we undertook the migration of the GlideinWMS factory service into the Kubernetes environment. Leveraging the flexibility and orchestration capabilities of Kubernetes, we successfully deployed the factory service within the OSG Tiger Kubernetes cluster. The major benefits Kubernetes gives us is it streamlines the management and monitoring of the factory infrastructure, and improves fault tolerance through its resilient deployment strategies. Through this case study, we aim to share insights, challenges, and best practices encountered during the migration process. Our experience underscores the benefits of embracing containerization and Kubernetes orchestration for HEP computing infrastructure, paving the way for scalability and resilience in distributed computing environments.

Dost, Jeffrey Michael [UC, San Diego (main)]↗

torc (Torc Workflow Management System) [SWR-24-127]

This software package orchestrates execution of a workflow of jobs on distributed computing resources. It is optimized for use on HPCs with Slurm, but also can be used in the cloud and on local computers. Please refer to the documentation at https://nrel.github.io/torc

Thom, Daniel [National Renewable Energy Laboratory↗

Panorama 360 (Final Report)

This is the final technical report for the DOE-funded Panorama 360 project. Panorama 360 provided a resource for the collection, analysis, and sharing of performance data about end-to-end scientific workflows executing on DOE facilities. The work focused on workflows that include experimental data generation at DOE facilities. The main activities of Panorama 360 include the development of: 1. A distributed repository that stores different types of workflow execution data (e.g., point and time series performance traces at fine- and coarse-grained levels); 2. A set of open-source data capture, curation, and publishing tools fully integrated with a state-of-the-art workflow management system that automates data ingestion to the repository and enables users to discover, query, and process data from the repository; 3. A set of analysis algorithms and machine learning based tools to perform analysis and characterization of the gathered data, which can be used to detect anomalous performance or system faults; and 4. Best practices and recommendations for workflow evaluation, analysis, execution, and architectures.

97 MATHEMATICS AND COMPUTING↗

A framework to enhance disaster debris estimation with AI and aerial photogrammetry

This study addresses the critical need to enhance disaster preparedness and response, focusing on hurricane impact assessment and debris estimation. Accurate assessments in this context are critical for post-event search-and-rescue (SAR) operations and resource distribution. Recent computing advancements are revealing the potential of unmanned aerial vehicles (UAVs) and artificial intelligence (AI) technologies in collecting data and assisting with post-hurricane reconnaissance. However, the use of AI and UAV photogrammetry for accurate disaster impact analysis remains underexplored. To this end, this study proposes a damage and debris analysis framework harnessing reality capture through aerial imagery and photogrammetry. Within this framework, a region-based neural network is leveraged to detect debris locations in aerial imagery with favorable performance. In a testbed within the Beaumont-Port Arthur region, in Southeast Texas, this study performs 3D reality captures of the built environment. Since the accuracy of the 3D reality capture is of importance in research areas associated with time-sensitive disaster response, we further investigate the optimal 2D aerial imagery overlap ratio required to generate a sufficiently accurate 3D model for disaster impact analysis and debris volume estimation. Results indicate that, in the case of aerial imagery for infrastructure systems, a minimum of 60 % overlap is recommended for damage assessment and debris analysis. In contrast, for flat green areas, a minimum of 50 % overlap is adequate. Overall, for disaster response applications, our study reveals that an overlap ratio between 60 % and 70 % is optimal for achieving a balance between time efficiency and data quality in aerial data collection. Furthermore, these quantitative recommendations are crucial for enabling efficient disaster response efforts. Additionally, our study outcomes will improve disaster impact analysis and facilitating timely and effective response strategies.

Artificial intelligence↗

Utilizing Distributed Heterogeneous Computing with PanDA in ATLAS

In recent years, advanced and complex analysis workflows have gained increasing importance in the ATLAS experiment at CERN, one of the large scientific experiments at LHC. Support for such workflows has allowed users to exploit remote computing resources and service providers distributed worldwide, overcoming limitations on local resources and services. The spectrum of computing options keeps increasing across the Worldwide LHC Computing Grid (WLCG), volunteer computing, high-performance computing, commercial clouds, and emerging service levels like Platform-as-a-Service (PaaS), Container-as-a-Service (CaaS) and Function-as-a-Service (FaaS), each one providing new advantages and constraints. Users can significantly benefit from these providers, but at the same time, it is cumbersome to deal with multiple providers, even in a single analysis workflow with fine-grained requirements coming from their applications’ nature and characteristics. In this paper, we will first highlight issues in geographically-distributed heterogeneous computing, such as the insulation of users from the complexities of dealing with remote providers, smart workload routing, complex resource provisioning, seamless execution of advanced workflows, workflow description, pseudointeractive analysis, and integration of PaaS, CaaS, and FaaS providers. We will also outline solutions developed in ATLAS with the Production and Distributed Analysis (PanDA) system and future challenges for LHC Run4.

97 MATHEMATICS AND COMPUTING↗

A Review of Edge Computing Technology and Its Applications in Power Systems

Recent advancements in network-connected devices have led to a rapid increase in the deployment of smart devices and enhanced grid connectivity, resulting in a surge in data generation and expanded deployment to the edge of systems. Classic cloud computing infrastructures are increasingly challenged by the demands for large bandwidth, low latency, fast response speed, and strong security. Therefore, edge computing has emerged as a critical technology to address these challenges, gaining widespread adoption across various sectors. This paper introduces the advent and capabilities of edge computing, reviews its state-of-the-art architectural advancements, and explores its communication techniques. A comprehensive analysis of edge computing technologies is also presented. Furthermore, this paper highlights the transformative role of edge computing in various areas, particularly emphasizing its role in power systems. It summarizes edge computing applications in power systems that are oriented from the architectures, such as power system monitoring, smart meter management, data collection and analysis, resource management, etc. Additionally, the paper discusses the future opportunities of edge computing in enhancing power system applications.

24 POWER TRANSMISSION AND DISTRIBUTION↗

High-Q cavity interface for color centers in thin film diamond

Quantum information technology offers the potential to realize unprecedented computational resources via secure channels distributing entanglement between quantum computers. Diamond, as a host to optically-accessible spin qubits, is a leading platform to realize quantum memory nodes needed to extend such quantum links. Photonic crystal (PhC) cavities enhance light-matter interaction and are essential for an efficient interface between spins and photons that are used to store and communicate quantum information respectively. Here, we demonstrate one- and two-dimensional PhC cavities fabricated in thin-film diamonds, featuring quality factors (Q) of 1.8 × 10 5 and 1.6 × 10 5 , respectively, the highest Qs for visible PhC cavities realized in any material. Importantly, our fabrication process is simple and high-yield, based on conventional planar fabrication techniques, in contrast to the previous with complex undercut processes. We also demonstrate fiber-coupled 1D PhC cavities with high photon extraction efficiency, and optical coupling between a single SiV center and such a cavity at 4 K achieving a Purcell factor of 18. The demonstrated photonic platform may fundamentally improve the performance and scalability of quantum nodes and expedite the development of related technologies.

97 MATHEMATICS AND COMPUTING↗

Distributed Machine Learning Workflow with PanDA and iDDS in LHC ATLAS

Machine Learning (ML) has become one of the important tools for High Energy Physics analysis. As the size of the dataset increases at the Large Hadron Collider (LHC), and at the same time the search spaces become bigger and bigger in order to exploit the physics potentials, more and more computing resources are required for processing these ML tasks. In addition, complex advanced ML workflows are developed in which one task may depend on the results of previous tasks. How to make use of vast distributed CPUs/GPUs in WLCG for these big complex ML tasks has become a popular research area. In this paper, we present our efforts enabling the execution of distributed ML workflows on the Production and Distributed Analysis (PanDA) system and intelligent Data Delivery Service (iDDS). First, we describe how PanDA and iDDS deal with large-scale ML workflows, including the implementation to process workloads on diverse and geographically distributed computing resources. Next, we report real-world use cases, such as HyperParameter Optimization, Monte Carlo Toy confidence limits calculation, and Active Learning. Finally, we conclude with future plans.

97 MATHEMATICS AND COMPUTING↗

Diaspora: Resilience-enabling services for science from HPC to edge

Scientific applications of interest to DOE must increasingly engage distributed resources (e.g., instruments, remote computers, data stores, edge devices) and deliver more stringent levels of service (e.g., uninterrupted processing of experiment data streams). In such systems, state is distributed and components can fail in many ways, often silently, making application resilience a major concern. Addressing the resilience needs of such applications requires methods for gaining knowledge of resources and applications and for translating that knowledge into action. We are working on addressing these needs in the context of multi-messenger astronomy, where detecting and responding to unusual transient events in multiple cosmic messengers (gravitational wave, electromagnetic, high- energy particles) from different instruments leads to a federated learning problem.

47 OTHER INSTRUMENTATION↗

Cybersecurity Assessment for a Behind-the-Meter Solar PV System: A Use Case for the DER-CF

The world's energy production is shifting toward lower-cost, cleaner, more efficient, and sustainable sources. The increasing numbers of distributed energy resources (DERs) are allowing for the rapid transformation of electric grids toward achieving the goal of energy decarbonization. Along with cleaner and more efficient energy, however, we must also aim for a secure energy future. Solar photovoltaic (PV) systems are an important part of this transition. This paper discusses a cybersecurity risk assessment for behind-the-meter DERs using a solar PV system as a use case of the Distributed Energy Resource Cybersecurity Framework (DER-CF) developed by the National Renewable Energy Laboratory. This poster presents a conference paper on the risk assessment processes and summarizes the DER-CF's use case recommendations to strengthen the cybersecurity posture of the electric grid.

cybersecurity↗

SWARM: Reimagining scientific workflow management systems in a distributed world

Modern scientific workflows process massive amounts of data from diverse instruments and sensors, leveraging geographically distributed, heterogeneous compute and storage resources—from leadership-class systems to edge devices—connected by high-performance networks. The diversity of resources introduces challenges in harnessing their full potential, with resilience issues arising across applications, system software, networks, storage, and hardware. Today, workflow management systems (WMS) coordinate the execution of computation and data management tasks across target resources. However, WMS’s centralized nature makes them vulnerable to faults and scalability issues that may result in failures of entire computational campaigns. In conclusion, this paper introduces a novel agentic framework for workflow management, fully distributing and decentralizing the WMS functions and modeling them as swarm intelligence agents infused with advanced artificial intelligence solutions and traditional distributed computing algorithms that can make coordinated decisions in the presence of failures of the underlying cyberinfrastructure.

Swarm intelligence↗

How Distributed Energy Resources Can Support Resilience in Utility Distribution Networks

The goal of this webinar is to engage with electric utilities in the Midwest, particularly small public utilities, to understand the industry's needs for science tools to plan for winter resilience in the future, designing tools that will benefit electric power resilience in all communities. Michigan Tech leads this project with partners from multiple academic, government, and industry groups and asked NLR to present on DERs and laboratory tools and resources.

24 POWER TRANSMISSION AND DISTRIBUTION↗

REopt: Energy Decision Support [Slides]

The REopt presentation, developed for the Energy Technology Innovation Partnership Project, provides an overview of the REopt tool. It covers tool's capabilities, how to use the tool, resources, and applications.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Constant-Overhead Fault-Tolerant Bell-Pair Distillation Using High-Rate Codes

We present a fault-tolerant Bell-pair distillation scheme achieving constant overhead through high-rate quantum low-density parity-check (qLDPC) codes. Our approach maintains a constant distillation rate equal to the code rate while requiring no additional overhead beyond the physical qubits of the code. Full circuit-level analysis demonstrates fault-tolerance for input Bell-pair infidelities below a threshold ∼10%, readily achievable with near-term capabilities. Unlike previous proposals, our scheme keeps the output Bell pairs encoded in qLDPC codes at each node, eliminating unencoding overhead and enabling direct use in distributed quantum applications through recent advances in qLDPC computation. These results establish qLDPC-based distillation as a practical route toward resource-efficient quantum networks and distributed quantum computing.

quantum communication, protocols & technology↗

PSU ESI Review

A guide to developing an Energy Service Interface (ESI) was created as part of the Grid Modernization Laboratory Consortium 2.5.2 ESI project. The approach applies device-agnostic and service-oriented ESI principles and leverages documents such as the Interoperability Maturity Model and Common Grid Service Definitions to provide a methodology to review, develop, and update standards and profiles to engage distributed energy resources (DER) to provide grid services. This document evaluates the ESI developed by Portland State University’s Power Engineering Group under the Electric Grid of Things project funded by the U.S. Department of Energy. The evaluation explores the compliance of this specific implementation with the GMLC ESI principles to provide an example of an ESI profile and gap analysis.

24 POWER TRANSMISSION AND DISTRIBUTION↗

A Privacy-Aware Federated Learning Framework for Distributed Energy Resource Analytics in Constrained Environments

To be resilient against extreme weather events, the rural communities in Puerto Rico are leveraging distributed energy resources (DER). However, computing frameworks sup-porting the grid in critical decision-making are still largely centralized. Sensitive consumer data are transmitted over the Internet or cellular networks to a secondary or tertiary node. It guarantees better situational awareness at the cost of a wider attack surface, jeopardizing user privacy, as more DER come online. Cloud, Edge, and Fog computing all require data aggregation at some level. This paper introduces a privacy-aware federated learning framework that leverages the Fog model by pushing analytics all the way to the DER and load assets. These local models train on individual asset data and transmit only learned parameters (such as weights) over secure communications to a global decision-maker. By abstracting personally identifiable consumer data without impacting decision optimality, this framework better aligns with distributed power generation paradigm.

Sundararajan, Aditya↗