Search NASASearch

SEARCH · Search NASA

Results for “service migration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

FIRE: A Failure-Adaptive RL Framework for Edge Computing Migrations

In edge computing, users' service profiles are migrated between edge servers due to user mobility. Reinforcement Learning (RL) frameworks have been proposed to do so, often trained on simulated data. However, existing RL frameworks overlook occasional server failures, which although rare, impact latency-sensitive applications like AR/VR and real- time obstacle detection. These rare failures, being not adequately represented in historical training data, pose a challenge for data-driven RL algorithms. We introduce FIRE, a framework that adapts to rare events by training a RL policy in an edge computing digital twin environment. We propose FIRE-ImRE, an importance sampling-based Q-learning algorithm, which samples rare events proportionally to their impact on the value function. FIRE considers delay, migration, failure, and backup placement costs across individual and shared service profiles. We prove FIRE-ImRE's boundedness and convergence to optimality. Next, we introduce novel deep Q-learning (FIRE-ImDQL) and actor critic (FIRE-ImACRE) versions of our algorithm to enhance scalability. Here, we extend our framework to accommodate users with varying risk tolerances of rare failure events. Through trace-driven experiments, we show that FIRE reduces edge computing costs compared to vanilla RL and the greedy baseline in the event of failures.

Edge computing

Cloud-Based Demonstration of the Eastern Interconnection Situational Awareness Monitoring System (ESAMS)

This report describes a cloud-based implementation and field demonstration of the Eastern Interconnection Situational Awareness and Monitoring System (ESAMS). ESAMS was developed to support the detection and source localization of forced oscillations using synchrophasor measurements from tie-lines connecting areas served by different reliability coordinators (RCs), so that RCs could better coordinate their response to wide-area events. A previous effort had identified deployment barriers associated with hosting shared situational awareness tools at a single RC. To address these barriers, ESAMS was migrated to Amazon Web Services and evaluated in a six-month field demonstration. ISO New England (ISO-NE) and PJM streamed data to the platform using AWS Direct Connect and a site-to-site VPN, respectively. The resulting multi-utility measurement footprint enabled regional source localization across major portions of the U.S. Eastern Interconnection and supported routine identification of oscillation events. During the final three months of the trial, 24 events above 2 MW/MVAR were detected. The largest detected oscillation approached a 25 MW peak-to-peak amplitude, and the longest persisted intermittently for more than 11 hours. The demonstration also assessed operational considerations—including data transfer volumes, end-to-end latency, and cloud computing costs—and found that network and compute requirements were modest relative to typical cloud capabilities while providing performance comparable to prior on-premises deployments. Overall, the results indicate that cloud hosting can provide a practical path to shared interconnection-wide oscillation monitoring. The cloud ESAMS demonstration establishes a foundation for broader utility participation and for building future wide-area analytics that leverage measurements across organizational boundaries.

Follum, James D.

Exploiting Kubernetes to Simplify the Deployment and Management of the Multi-purpose CMS Pilot Job Factory

GlideinWMS, a widely utilized workload management system in high-energy physics (HEP) research, serves as the backbone for efficient job provisioning across distributed computing resources. It is utilized by various experiments and organizations, including CMS, OSG, Dune, and FIFE, to create HTCondor pools as large as 600k cores. In particular, a shared factory service historically deployed at UCSD has been configured to interface with more than 500 routes to compute clusters. As part of our team’s initiative to modernize infrastructure and enhance scalability, we undertook the migration of the GlideinWMS factory service into the Kubernetes environment. Leveraging the flexibility and orchestration capabilities of Kubernetes, we successfully deployed the factory service within the OSG Tiger Kubernetes cluster. The major benefits Kubernetes gives us is it streamlines the management and monitoring of the factory infrastructure, and improves fault tolerance through its resilient deployment strategies. Through this case study, we aim to share insights, challenges, and best practices encountered during the migration process. Our experience underscores the benefits of embracing containerization and Kubernetes orchestration for HEP computing infrastructure, paving the way for scalability and resilience in distributed computing environments.

Dost, Jeffrey Michael [UC, San Diego (main)]

dCache CI/CD migration to Kubernetes

For over two decades, the dCache project has provided open-source to satisfy ever-more demanding storage requirements. More than 80 sites worldwide rely on dCache to provide services for LHC experiments, Belle-II, Eu- XFEL, and others. This can be achieved only with a well-established process from a whiteboard, where ideas are created through development, packaging, and testing. The project’s build and test infrastructure is based on Jenkins CI and a set of virtual machines. This infrastructure is maintained by dCache developers. With the introduction of the DESY-central Gitlab server, the developers have started migrating from VM-based testing to container-based deployments in the onsite Kubernetes cluster. As a result, we have packaged dCache containers and Helm charts that can be used by other sites to reproduce our test and build steps quickly or to evaluate new releases on their pre-production systems and, eventually, become a standard model of dCache deployment at the sites. This paper describes the challenges we have faced, the techniques we used to solve them, and the issues that still need to be addressed.

Mkrtchyan, Tigran [DESY]

Ab Initio Design of High-Entropy Thermal/ Environmental Barrier Coatings

Next generation thermal/environmental barrier coatings (TEBC) require carefully balancing various properties including phase stability, thermal conductivity, coefficient of thermal expansion (CTE), mechanical properties, and resistance against hot corrosion and water vapor recession. This work mainly focuses on rapid design of cost-effective high entropy rare-earth disilicates and aluminum garnets to protect SiC-based ceramic matrix composites and nickel-based superalloys in the hot section of gas turbine engines using density functional theory methods. Our calculations identify several low-cost high entropy TEBC exhibiting ultralow thermal conductivity at 1500 K and desirable CTE while maintaining good mechanical properties, including Er1/2Y3/4Yb3/4Si2O7, Gd1/4Er1/4Y3/4Yb3/4Si2O7, Eu1/4Er1/4Y3/4Yb3/4Si2O7, and (Y1/4Gd1/4Er1/4Yb1/4)3Al5O12. This work also aims to gain fundamental understanding of oxygen diffusion in model disilicates. Minimizing oxidizer (such as water vapor and oxygen) permeability through the EBC layer can significantly decrease the growth rate of thermally grown oxide and extend the service life of the coating system. Oxygen diffusion mechanisms including formation energy of defects under varying oxygen conditions and defect migration energy barriers will be presented.

coefficient of thermal expansion

Predicting runtime and resource utilization of jobs on integrated cloud and HPC systems

Recent advances in virtualization technologies used in cloud computing offer performance that closely approaches bare-metal levels. Combined with specialized instance types and high-speed networking services for cluster computing, cloud platforms have become a compelling option for high-performance computing (HPC). However, most current batch job schedulers in HPC systems are designed for homogeneous clusters and make decisions based on limited information about jobs and system status. Scientists typically submit computational jobs to these schedulers with a requested runtime that is often over- or under-estimated. More accurate runtime predictions can help schedulers make better decisions and reduce job turnaround times. Here, they can also support decisions about migrating jobs to the cloud to avoid long queue wait times in HPC systems.

97 MATHEMATICS AND COMPUTING

Hero Carbonsafe Phase 2 Project in the Columbia River Basalt Group: Technical Program Overview

The Hermiston, Oregon Basalt CarbonSAFE Phase II project (HERO CarbonSAFE) seeks to accelerate the deployment of commercial carbon dioxide (CO2) storage projects in basaltic rocks. Hermiston is located near the center of the Columbia River Basalt Group (CRBG), which is one of the largest basalt flows in the US. Basalt CO2 storage has potential advantages to conventional saline storage reservoirs including 1. The potential for rapid mineralization of CO2, 2. associated decreases in pressure and CO2 migration risks, 3. reduced long-term monitoring requirements with respect to plume tracking, 4. widespread geographic distribution and, 5. large storage potential due to thickness, porosity, and CO2 interactions with basalt. For locations such as the Pacific Northwest (PNW), Hawaii, Iceland, India and Japan, whose localities are isolated from large sedimentary basins offering conventional saline storage options, basalt may offer the only feasible option for local CO2 storage. However, mineralization/basalt storage still has many uncertainties, as there are limited field-scale assessments of CO2 storage in basalt. There are significant uncertainties hindering the effective implementation of carbon capture utilization and storage (CCUS) in basalt. These include the lack of proven storage capacities, challenges in methodologies for modeling the area of review in igneous formations, limited understanding of mineralization kinetics and timing, and uncertainties in injectivity. Additionally, the domestic availability of specialized services and drilling expertise is constrained, and existing CCUS permitting and regulatory frameworks, originally developed for conventional saline reservoirs, may not adequately address the unique requirements of basalt systems. HERO CarbonSAFE is designed to address major research gaps and uncertainties associated with basalt storage. Specifically, the project will assess the feasibility of CO2 injection in the deep layered basalts of the CRBG, long-term storage (mineralization), practical approaches for large-scale implementation (50+ million metric tons of CO2 over 30 years), lithology-specific risks, and the technoeconomic potential for CO2 storage in basalts.

58 GEOSCIENCES

Securing the Modern Grid: Federal Investments, Digitization, and Supply Chain Strategy

Across the United States (U.S.) grid expansion and modernization is underway, paving the way for accelerated load growth and intelligent resource management. Digitization of the grid is supported by several state and federal programs, providing support for utilities installing advanced metering infrastructure (AMI), AI-powered analytics systems, battery energy storage systems (BESS), and distributed energy resource management systems (DERMS) to transform the grid from a one-way power delivery system into an intelligent, responsive network that will enable faster load growth and power expansion of data centers for advanced artificial intelligence (AI) applications. The digital transformation of America's grid presents opportunity for increased efficiency and resiliency but also introduces new digital risks that require careful management. Digital equipment often contains several vulnerabilities such as unencrypted communication protocols, and persistent remote access capabilities that could be exploited to manipulate device settings, coordinate service disruptions, or inject false data into grid operations. These digital risks become particularly important as the grid must rapidly scale to support AI-driven data centers, which the administration has identified as essential for maintaining U.S. technological leadership and economic competitiveness. These vulnerabilities are compounded by supply chain realities: Chinese manufacturers currently produce 70-90% of essential grid components including inverters, batteries, and control systems, with the U.S. lacking domestic manufacturing capacity for critical assets like extra-high voltage transformers. Recent federal legislation has established Foreign Entity of Concern (FEOC) restrictions to address these risks, requiring projects to achieve escalating thresholds of non-FEOC content to receive tax credits while utilities work to expand sourcing channels for their supply chains and strengthen security measures. These restrictions arrive precisely when utilities face unprecedented electricity demand growth driven by the rapid growth in data centers, creating a considerable challenge: rapidly expanding infrastructure while navigating complex compliance requirements while lacking viable alternatives for many critical components. Idaho National Laboratory (INL) and its partners have developed practical approaches to help utilities navigate these intersecting challenges as they leverage federal investment to strengthen and grow the grid. These solutions include Cyber-Informed Engineering (CIE) principles that build resilience directly into systems, the Cirrus tool for secure cloud migration, and enhanced procurement guidance that embeds security requirements throughout equipment lifecycles. Federal initiatives, such as the Technical Assistance for Digital Assurance (TADA) project, provide direct support to utilities implementing these approaches while facilitating knowledge sharing across the industry. While these tools and frameworks cannot eliminate all risks inherent in foreign supply chain dependencies, they offer pragmatic pathways for strengthening security posture without sacrificing the deployment momentum essential to meeting surging electricity demand. Ultimately, securing America's digital energy infrastructure demands dedicated coordination across multiple fronts: building domestic supply chains, implementing robust digital assurance practices, and maintaining the aggressive modernization timeline necessary for reliability, resilience, and energy independence.

24 POWER TRANSMISSION AND DISTRIBUTION

Migrating Populus with climate change: Phenology, coppice management, cold spell susceptibility, leaf dynamics, and biomass production

Understanding the phenology and productivity of Populus species is crucial for effective management and conservation strategies amid climate change. We investigated leaf budbreak timing, susceptibility to cold damage, leaf dynamics, and biomass production of 168 Populus genotypes with diverse provenances in the southeastern United States. Our study revealed significant variation in budbreak timing across different taxa and years, with genotypes inheriting traits adapted to their parents’ local climates. Temperature emerged as a key factor triggering budbreak, while leaf development depended on other environmental cues such as photoperiod. Notably, budbreak occurred approximately 20 days earlier in 2023 compared to 2022 due to higher accumulated degree days (ADDs). Short-rotation-coppice (SRC) management delayed budbreak by five to ten days. Cold damage was significant in 2023, particularly for genotypes from northern provenances and those with P. maximowiczii parentage. Severe damage was also observed in eastern cottonwood ( Populus deltoides ​× ​ Populus deltoides (D ​× ​D)) genotypes, despite most having southeastern US parentages. Leaf dynamics, including leaf duration and leaf area index (LAI), varied across taxa and sites, with earlier budbreak correlating with extended growing seasons and increased LAI. Biomass production was intricately linked to phenological events, with earlier budbreak leading to increased biomass production and greater susceptibility to cold damage. Our findings highlight the importance of genetics, environment, and coppicing management in understanding and managing Populus phenology and biomass production. These insights provide valuable guidance for developing effective breeding, conservation, and management strategies for Populus species in the context of climate change.

Accumulated degree days (ADDs)

Fire and Drought Affect Multiple Aspects of Diversity in a Migratory Bird Stopover Community

Drought and high-severity, stand-replacing wildfires can have substantial impacts on the composition of avian communities, including stop-over communities during migration. An inextricable link exists between drought and wildfire, each operating and impacting across different timescales. Many studies have found nonlinear avian abundance trends in breeding community time series data that include pre- and post-fire observations, describing an initial decrease in abundance followed by rapid increases that can attenuate over time. Here, we use a fall bird-banding dataset to evaluate shifts in a drought-impacted avian community following wildfire from taxonomic, functional, and phylogenetic perspectives. We looked at the community as a whole and also categorized birds as residents, migrants, and breeders to assess potential varying responses at the study site. We observed post-fire shifts in functional and phylogenetic diversity that corresponded to changes in vegetation. An influx of migratory insectivores post-fire drove much of the variation between pre- and post-fire avian communities and toward a more related, less phylogenetically dispersed community. A concurrent monsoon season drought was also associated with functional and phylogenetic diversity, highlighting the intertwined pulse press effects on avian communities. Overall, our results suggest that, although bird communities are immediately impacted by fire-driven resource changes, they can rebound over time, it is unclear how long-term drought may continue to shape the composition of these avian communities.

59 BASIC BIOLOGICAL SCIENCES