Search NASA⌕ Search

SEARCH · Search NASA

Results for “flexible computing workloads”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Flexible AI Models for Grid Resilience

The rapid growth in size and complexity of artificial intelligence (AI) and machine learning (ML) models has led to increased energy demands, posing a threat to the reliability of the existing power grid. This project addresses the challenge of highly intermittent and energy-intensive inference workloads by (1) developing fidelity-adaptive neural networks capable of dynamic response to grid conditions and (2) integrating these networks with power flow simulations to assess their impact on power grid reliability. We will explore both top-down and bottom-up approaches to create hierarchies of submodels that provide a controlled trade-off between power draw and prediction accuracy. The top-down method utilizes NN pruning to reduce a flagship model into progressively smaller, energy-efficient variants. The bottom-up approach employs geometrically principled weight setting strategies to construct depth-efficient models from the ground up. A real-time hardware-in-the-loop (HIL) platform will be developed to simulate a scaled AC power grid, integrating live AI workload power draw and enabling dynamic model switching in response to grid feedback. This work will provide a novel framework for evaluating the impact of flexible AI/ML workloads on grid performance and establish new methodologies for energy-aware computing in data centers. The outcomes will demonstrate that adaptive AI/ML can play a critical role in improving grid stability while advancing NREL's leadership in energy-efficient computing research.

24 POWER TRANSMISSION AND DISTRIBUTION↗

XaaS: Acceleration as a Service to Enable Productive High-Performance Cloud Computing

High-performance computing (HPC) and the cloud have evolved independently, specializing their innovations into performance or productivity. Acceleration as a Service (XaaS) is a recipe to empower both fields with a shared execution platform that provides transparent access to computing resources, regardless of the underlying cloud or HPC service provider. Bridging HPC and cloud advancements, XaaS presents a unified architecture built on performance-portable containers. Here, our converged model concentrates on low-overhead, high-performance communication and computing, targeting resource-intensive workloads from climate simulations to machine learning. XaaS lifts the restricted allocation model of Function as a Service (FaaS), allowing users to benefit from the flexibility and efficient resource utilization of serverless computing while supporting long-running and performance-sensitive workloads from HPC.

97 MATHEMATICS AND COMPUTING↗

Flexible silicon photonic architecture for accelerating distributed deep learning

The increasing size and complexity of deep learning (DL) models have led to the wide adoption of distributed training methods in datacenters (DCs) and high-performance computing (HPC) systems. However, communication among distributed computing units (CUs) has emerged as a major bottleneck in the training process. In this study, we propose Flex-SiPAC, a flexible silicon photonic accelerated compute cluster designed to accelerate multi-tenant distributed DL training workloads. Flex-SiPAC takes a co-design approach that combines a silicon photonic hardware platform with a tailored collective algorithm, optimized to leverage the unique physical properties of the architecture. The hardware platform integrates a novel wavelength-reconfigurable transceiver design and a micro-resonator-based wavelength-reconfigurable switch, enabling the system to achieve flexible bandwidth steering in the wavelength domain. The collective algorithm is designed to support reconfigurable topologies, enabling efficient all-reduce communications that are commonly used in DL training. The feasibility of the Flex-SiPAC architecture is demonstrated through two testbed experiments. First, an optical testbed experiment demonstrates the flexible routing of wavelengths by shuffling an array of input wavelengths using a custom-designed spatial-wavelength selective switch. Second, a four-GPU testbed running two DL workloads shows a 23% improvement in job completion time compared to a similarly sized leaf-spine topology. We further evaluate Flex-SiPAC using large-scale simulations, which show that Flex-SiPAC is able to reduce the communication time by 26% to 29% compared to state-of-the-art compute clusters under representative collective operations.

Wu, Zhenguo (ORCID:0000000322847985)↗

Anthropometric considerations for a 4-axis side-arm flight controller

A data base on multiaxis side-arm flight controls was generated. The rapid advances in fly-by-light technology, automatic stability systems, and onboard computers have combined to create flexible flight control systems which could reduce the workload imposed on the operator by complex new equipment. This side-arm flight controller combines four controls into one unit and should simplify the pilot's task. However, the use of a multiaxis side-arm flight controller without complete cockpit integration may tend to increase the pilot's workload.

Debellis, W. B.↗

Cockpit Adaptive Automation and Pilot Performance

The introduction of high-level automated systems in the aircraft cockpit has provided several benefits, e.g., new capabilities, enhanced operational efficiency, and reduced crew workload. At the same time, conventional 'static' automation has sometimes degraded human operator monitoring performance, increased workload, and reduced situation awareness. Adaptive automation represents an alternative to static automation. In this approach, task allocation between human operators and computer systems is flexible and context-dependent rather than static. Adaptive automation, or adaptive task allocation, is thought to provide for regulation of operator workload and performance, while preserving the benefits of static automation. In previous research we have reported beneficial effects of adaptive automation on the performance of both pilots and non-pilots of flight-related tasks. For adaptive systems to be viable, however, such benefits need to be examined jointly in the context of a single set of tasks. The studies carried out under this project evaluated a systematic method for combining different forms of adaptive automation. A model for effective combination of different forms of adaptive automation, based on matching adaptation to operator workload was proposed and tested. The model was evaluated in studies using IFR-rated pilots flying a general-aviation simulator. Performance, subjective, and physiological (heart rate variability, eye scan-paths) measures of workload were recorded. The studies compared workload-based adaptation to to non-adaptive control conditions and found evidence for systematic benefits of adaptive automation. The research provides an empirical basis for evaluating the effectiveness of adaptive automation in the cockpit. The results contribute to the development of design principles and guidelines for the implementation of adaptive automation in the cockpit, particularly in general aviation, and in other human-machine systems. Project goals were met or exceeded. The results of the research extended knowledge of automation-related performance decrements in pilots and demonstrated the positive effects of adaptive task allocation. In addition, several practical implications for cockpit automation design were drawn from the research conducted. A total of 12 articles deriving from the project were published.

Parasuraman, Raja↗

Integrating AI Data Centers with the Power Grid

The rapid expansion of artificial intelligence (AI) has triggered an unprecedented surge in electricity demand, with US data center energy use projected to double or triple 2023 levels by 2028. This exponential growth places strain on grid infrastructure, which can hinder timely construction of desired computing capacity. To bridge this supply-demand gap, utilities and AI developers are increasingly turning to demand flexibility, a strategy that incentivizes shifting or reducing power use during peak periods of grid stress. Data centers are uniquely equipped for flexible operations due to their digital workloads, built-in redundancy, and onsite energy assets. This article outlines four primary mechanisms to enable data center flexibility: computational load flexibility (shifting tasks temporally or geographically), flexible use of core facility infrastructure adjustments, energy storage utilization, and onsite electricity generation. To encourage adoption, utilities are deploying new tariff designs, including voluntary interruptible service riders, mandated flexibility requirements, and streamlined interconnection processes for flexible loads. For the highly capitalized and rapidly growing AI industry, the primary motivators for embracing these strategies are expediting facility interconnection, satisfying emerging regulatory mandates, and mitigating community resistance. While demand flexibility cannot substitute the long-term need for new bulk power generation, it serves as an essential, immediate solution for enabling near-term deployment. By transforming data centers from grid stressors into stabilizing assets, flexible operations can ensure reliable grid integration, ease market pressures, and support a resilient power system.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Flexible User-Defined Domain Decomposition in Kilometer-Scale E3SM Land Model Simulation

The Energy Exascale Earth System Model (E3SM) Land Model (ELM) has been extended to kilometer-scale (km-ELM) resolutions, enabling high-fidelity simulations of terrestrial processes at 1 km x 1 km grid spacing. In ELM, domain decomposition partitions the computational domain across processors, ensuring efficient parallel execution. Currently, round-robin decomposition is applied, providing a straightforward way to distribute computational workload. As ELM continues evolving at the kilometer-scale (km-scale), particularly with integrating lateral flow modeling, decomposition strategies must also account for the increased workload and data movement. This paper introduces a flexible user-defined domain decomposition framework, allowing users to customize domain partitioning based on application requirements. The impact of different decomposition strategies is evaluated across various applications concerning computation, communication, and I/O. Results demonstrate that while 1D partitioning yields superior I/O performance, k-nearest neighbors (KNN) clustering effectively reduces inter-process communication overhead. This study lays the groundwork for scalable partitioning in large-scale land surface simulations, enhancing next-generation Earth system modeling.

Wang, Dali [ORNL] (ORCID:0000000168065108)↗

Future Generation High Performance Computing Center (FG-HPCC): RFI Technical Considerations

Lawrence Livermore National Security, LLC (LLNS) is interested in receiving information about technologies that could be available in the 2029-2030 timeframe that may serve to enable the vision for a Future Generation High Performance Computing (HPC) Center (FG-HPCC) described in this document. The future HPC Center vision has been conceived to meet the future mission needs of the Advanced Simulation and Computing (ASC) Program within the National Nuclear Security Administration (NNSA). LLNS envisions a center composed not of many independent clusters, but of heterogeneous elements accessible to users as a single system. The capabilities will be integrated to create a scalable, flexible, yet tightly coupled computing center capable of integrated HPC, AI, and cloud-like workloads.

97 MATHEMATICS AND COMPUTING↗

Optimization of distributed compute resources utilization in the CMS Global Pool

The CMS Submission Infrastructure is the primary system for managing computing resources for CMS workflows, including data processing, simulation, and analysis. It integrates geographically distributed resources from Grid, HPC, and cloud providers into federated pools managed by HTCondor and Glidein- WMS, for a total of around 500k CPU cores. This system dynamically manages workloads based on priorities defined by the collaboration. Additionally, CMS scheduling strategies must be flexible to handle multiple concurrent workloads while considering changing processing demands and resource availability from various providers.Efficient utilization of vast amounts of distributed compute resources is a key element for the success of the scientific programs of the LHC experiments. Optimizing the system is essential to maximize resource efficiency and fully utilize the distributed computing power. The CMS Submission Infrastructure team thus systematically investigates sources of inefficiency in workload scheduling to reduce their impact. In addition, a strategy of pilot overloading has been introduced to compensate for other inefficiency sources, thereby optimizing resource utilization and enhancing computational throughput.

Mascheroni, Marco [UC, San Diego (main)]↗

Automated Software for Manned Spacecraft - Bridging the Gap from Sci Fi to Reality

With a voice command or a few taps on the console, the spacecraft pivots on a dime at high velocity and gently docks to an orbiting space platform. This is the image most people have of the complex software computations and integrated hardware performance necessary for a spacecraft to successfully perform an automated launch, rendezvous, and docking. Today’s reality is that while computer operations are advancing rapidly, science fiction over-simplifies and over-sells current capabilities. This paper discusses the integration of spacecraft computer automation into the operation of one of the United States’ new Commercial Crew vehicles - the Boeing CST-100 Starliner. Lessons learned by the Boeing Mission Operations team, a private-public partnership with NASA, from conceptual design through real-time operation of the first test flight will be discussed. Focus will center on how operations has learned to use the automated software to their advantage while also knowing how to adjust the automation in response to spacecraft or mission anomalies. One goal of advanced spacecraft automation is the ability to reduce both the crew workload and the ground control footprint while at the same time increasing spacecraft and mission flexibility. Historically, crewed spacecraft required a large number of operators on the ground to use a plethora of tools to compute nominal and contingency mission trajectories. Moving those sophisticated software tools to being onboard the vehicle can reduce the need for such complex ground support. Given that today’s spacecraft software is not yet as capable or as flexible in all circumstances as the computers depicted in movies, there is usually a trade-off between software automation cost and the flexibility of that software resulting in a trade-off between what is performed on the spacecraft and what is left to onboard crew or ground control. For missions that go beyond the Moon, software that autonomously controls nearly every aspect of a crewed mission will become a necessity given the long time delays between the spacecraft and Earth’s ground control teams. The lessons learned by Boeing and its Mission Operations team, through the design and implementation of Starliner’s hardware and software automation, will be able to inform future public and private spacecraft design. As the technologies and capabilities evolve, incorporating lessons learned in successful low Earth orbit commercial crew vehicle missions, spacecraft designs will continue to improve and be able to better enable safe execution of human missions to the Moon and beyond.

Robert C Dempsey↗

Automated Software for Manned Spacecraft - Bridging the Gap from Sci Fi to Reality

With a voice command or a few taps on the console, the spacecraft pivots on a dime at high velocity and gently docks to an orbiting space platform. This is the image most people have of the complex software computations and integrated hardware performance necessary for a spacecraft to successfully perform an automated launch, rendezvous, and docking. Today’s reality is that while computer operations are advancing rapidly, science fiction over-simplifies and over-sells current capabilities. This paper discusses the integration of spacecraft computer automation into the operation of one of the United States’ new Commercial Crew vehicles - the Boeing CST-100 Starliner. Lessons learned by the Boeing Mission Operations team, a private-public partnership with NASA, from conceptual design through real-time operation of the first test flight will be discussed. Focus will center on how operations has learned to use the automated software to their advantage while also knowing how to adjust the automation in response to spacecraft or mission anomalies. One goal of advanced spacecraft automation is the ability to reduce both the crew workload and the ground control footprint while at the same time increasing spacecraft and mission flexibility. Historically, crewed spacecraft required a large number of operators on the ground to use a plethora of tools to compute nominal and contingency mission trajectories. Moving those sophisticated software tools to being onboard the vehicle can reduce the need for such complex ground support. Given that today’s spacecraft software is not yet as capable or as flexible in all circumstances as the computers depicted in movies, there is usually a trade-off between software automation cost and the flexibility of that software resulting in a trade-off between what is performed on the spacecraft and what is left to onboard crew or ground control. For missions that go beyond the Moon, software that autonomously controls nearly every aspect of a crewed mission will become a necessity given the long time delays between the spacecraft and Earth’s ground control teams. The lessons learned by Boeing and its Mission Operations team, through the design and implementation of Starliner’s hardware and software automation, will be able to inform future public and private spacecraft design. As the technologies and capabilities evolve, incorporating lessons learned in successful low Earth orbit commercial crew vehicle missions, spacecraft designs will continue to improve and be able to better enable safe execution of human missions to the Moon and beyond.

Robert C. Dempsey↗

Automated Software for Crewed Spacecraft - Bridging the Gap from Sci Fi to Reality

With a voice command or a few taps on the console, the spacecraft pivots on a dime at high velocity and gently docks to an orbiting space platform. This is the image most people have of the complex software computations and integrated hardware performance necessary for a spacecraft to successfully perform an automated launch, rendezvous, and docking. Today’s reality is that while computer operations are advancing rapidly, science fiction over-simplifies and over-sells current capabilities. This paper discusses the integration of spacecraft computer automation into the operation of one of the United States’ new Commercial Crew vehicles - the Boeing CST-100 Starliner. Lessons learned by the Boeing Mission Operations team, a private-public partnership with NASA, from conceptual design through real-time operation of the first test flight will be discussed. Focus will center on how operations has learned to use the automated software to their advantage while also knowing how to adjust the automation in response to spacecraft or mission anomalies. One goal of advanced spacecraft automation is the ability to reduce both the crew workload and the ground control footprint while at the same time increasing spacecraft and mission flexibility. Historically, crewed spacecraft required a large number of operators on the ground to use a plethora of tools to compute nominal and contingency mission trajectories. Moving those sophisticated software tools to being onboard the vehicle can reduce the need for such complex ground support. Given that today’s spacecraft software is not yet as capable or as flexible in all circumstances as the computers depicted in movies, there is usually a trade-off between software automation cost and the flexibility of that software resulting in a trade-off between what is performed on the spacecraft and what is left to onboard crew or ground control. For missions that go beyond the Moon, software that autonomously controls nearly every aspect of a crewed mission will become a necessity given the long time delays between the spacecraft and Earth’s ground control teams. The lessons learned by Boeing and its Mission Operations team, through the design and implementation of Starliner’s hardware and software automation, will be able to inform future public and private spacecraft design. As the technologies and capabilities evolve, incorporating lessons learned in successful low Earth orbit commercial crew vehicle missions, spacecraft designs will continue to improve and be able to better enable safe execution of human missions to the Moon and beyond.

Robert C Dempsey↗

Flexible Pilot Jobs Framework for Distributed High Throughput Computing

Experimental particle physics has been at the forefront of analyzing the world’s largest datasets for decades. The high-energy physics (HEP) community was among the first to develop suitable software and computing tools for this purpose. GlideinWMS is a Glidein-based workload management system whose purpose is to provide experiments like CMS at CERN, DUNE at Fermilab, and others, a way to access and efficiently use vast amounts of computing resources. This system wants to provide a simple way to submit jobs to a set of computing resources, that will be provided to users behind the scenes. Glideins are the pilot jobs executed on the worker nodes at the grid sites, performing operations such as hardware detection, environment setup, and error handling. After all these operations, they will launch the actual user job. Many grid sites are supported, such as shared clusters, Google CE, and AWS. My internship aimed to design and code a flexible pilot jobs framework that will replace the one used by GlideinWMS, developing a modular and flexible skeleton of the Glidein and adding further functionalities. My project also focused on the application of machine learning techniques as support to this management system.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

FLASH: FPGA-Accelerated Smart Switches with GCN Case Study

Some communication switches, e.g., the Mellanox SHArP and those in the IBM BlueGene clusters, are augmented to process packets at the application level with fixed-function collectives. This approach, however, lacks flexibility, which limits their applicability in diverse and dynamic workloads. Recently, a new type of programmable packet processor, which uses high-level languages, e.g., P4, has emerged as possible candidates. P4-based switches, however, fall short in certain applications, including machine learning, where capabilities not currently supported by P4 are needed. These include more complex calculation, such as sparse computation and fused multiply-accumulate, data-intensive floating point operations, data reuse, and significant memory. The problem addressed here is that such a switch augmentation needs to support: a large amount of state, significant flexible compute capability, and ease of programming, all while maintaining full functionality, including ensuring high throughput, and demonstrating utility. In this work, we propose a programmable look-aside-type accelerator that can be embedded into, or attached to, existing communication switch pipelines and that is capable of processing packets at line-rate. The proposed in-switch accelerator is based on mixing an ISA (subset of RISC-V instructions) with dataflow graphs (found in CGRAs). To augment performance, vector instructions are also supported. To facilitate usability, we have developed a complete toolchain to compile user-provided C/C++ codes to appropriate back-end instructions for configuring the accelerator. While this approach is flexible enough to support various workloads, in this paper, we consider Graph Convolutional Networks (GCNs) as a case study. Experimental results show that this approach considerably improves the performance of distributed GCN applications.

Haghi, Pouya↗

New trends in photonic switching and optical networking architectures for data centers and computing systems [Invited]

The rapid increases in data traffic coupled with user preferences are driving the data center and computing system service providers to offer energy-efficient, intelligent, flexible, cost-effective, high-capacity, and low-latency data services without added complexity to the users. Disaggregated heterogeneous reconfigurable computing systems realized by photonic switching and interconnects can enhance throughput and energy efficiency for artificial intelligence/machine learning (AI/ML) workloads, especially when aided by the AI/ML-enhanced control plane. Photonic switching and new optical networking architectures are expected to solve many of these challenging problems. This paper discusses new trends in photonic switching and optical network architectures for future data centers and computing systems summarized as follows: (1) flat reconfigurable disaggregated computing enabled by high-radix photonic switching and interconnects in data centers; (2) chiplet-based computing architectures empowered by embedded photonics toward heterogeneous reconfigurable computing; (3) nanosecond-scale photonic switching in data centers and computing systems; (4) AI/ML in self-driving, application-aware, and situation-aware data centers; (5) the emergence of flexible networking for cloud computing, edge computing, and split computing, as well as flexible networking for 5G/6G RF-optical networks; and (6) the deployment of embedded co-designed silicon photonics being considered for future data centers.

Yoo, S. J. Ben (ORCID:0000000274201871)↗

TPSAS-NF1676L-33992-DND

The CERES Science Team integrates and fuses observations from 6 CERES instruments aboard the Terra, Aqua, S-NPP, and NOAA-20 missions with data from more than 20 other unique data sources. Following the November 2017 launch of CERES Flight Model 6 (FM6) onboard NOAA20, CERES has now amassed over 80 instrument-years of valuable Earth radiation budget data. The rapidly growing volume of CERES data coupled with the introduction of new data products alongside improvements to existing science algorithms fosters the requirement for faster, more flexible, and scalable data production and orchestration. New virtualized, cloud-centric compute hardware hosted by the NASA Langley Research Center’s (LaRC) Atmospheric Sciences Data Center (ASDC) provides an ideal environment for these ever-increasing data production demands for CERES. This poster discusses updates to the implementation of the CERES Data Management Team’s (DMT) CERES AuTomAted job Loading sYSTem (CATALYST), a custom data processing workflow engine for CERES, to use on-demand computing resources to perform automated CERES data production processing in a Linux-based container environment. Linux containers provide CERES the flexibility to build multiple production environments in containers tailored for specific workloads and allow effortless provisioning of resources based on the CERES Science Team’s data production requirements.

Thomas N. Hillyer↗

Storage media pipelining: Making good use of fine-grained media

This paper proposes a new high-performance paradigm for accessing removable media such as tapes and especially magneto-optical disks. In high-performance computing the striping of data across multiple devices is a common means of improving data transfer rates. Striping has been used very successfully for fixed magnetic disks improving overall system reliability as well as throughput. It has also been proposed as a solution for providing improved bandwidth for tape and magneto-optical subsystems. However, striping of removable media has shortcomings, particularly in the areas of latency to data and restricted system configurations, and is suitable primarily for very large I/Os. We propose that for fine-grained media, an alternative access method, media pipelining, may be used to provide high bandwidth for large requests while retaining the flexibility to support concurrent small requests and different system configurations. Its principal drawback is high buffering requirements in the host computer or file server. This paper discusses the possible organization of such a system including the hardware conditions under which it may be effective, and the flexibility of configuration. Its expected performance is discussed under varying workloads including large single I/O's and numerous smaller ones. Finally, a specific system incorporating a high-transfer-rate magneto-optical disk drive and autochanger is discussed.

Vanmeter, Rodney↗

Bridging paradigms: Designing for HPC-Quantum convergence

Here, this paper presents a comprehensive software stack architecture for integrating quantum computing (QC) capabilities with High-Performance Computing (HPC) environments. While quantum computers show promise as specialized accelerators for scientific computing, their effective integration with classical HPC systems presents significant technical challenges. We propose a hardware-agnostic software framework that supports both current noisy intermediate-scale quantum devices and future fault-tolerant quantum computers, while maintaining compatibility with existing HPC workflows. The architecture includes a quantum gateway interface, standardized APIs for resource management, and robust scheduling mechanisms to handle both simultaneous and interleaved quantum–classical workloads. Key innovations include: (1) a unified resource management system that efficiently coordinates quantum and classical resources, (2) a flexible quantum programming interface that abstracts hardware-specific details, (3) A Quantum Platform Manager API that simplifies the integration of various quantum hardware systems, and (4) a comprehensive tool chain for quantum circuit optimization and execution. We demonstrate our architecture through implementation of quantum–classical algorithms, including the variational quantum linear solver, showcasing the framework’s ability to handle complex hybrid workflows while maximizing resource utilization. This work provides a foundational blueprint for integrating QC capabilities into existing HPC infrastructures, addressing critical challenges in resource management, job scheduling, and efficient data movement between classical and quantum resources.

97 MATHEMATICS AND COMPUTING↗