Search NASA⌕ Search

SEARCH · Search NASA

Results for “distributed machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Distributed Machine Learning Workflow with PanDA and iDDS in LHC ATLAS

Machine Learning (ML) has become one of the important tools for High Energy Physics analysis. As the size of the dataset increases at the Large Hadron Collider (LHC), and at the same time the search spaces become bigger and bigger in order to exploit the physics potentials, more and more computing resources are required for processing these ML tasks. In addition, complex advanced ML workflows are developed in which one task may depend on the results of previous tasks. How to make use of vast distributed CPUs/GPUs in WLCG for these big complex ML tasks has become a popular research area. In this paper, we present our efforts enabling the execution of distributed ML workflows on the Production and Distributed Analysis (PanDA) system and intelligent Data Delivery Service (iDDS). First, we describe how PanDA and iDDS deal with large-scale ML workflows, including the implementation to process workloads on diverse and geographically distributed computing resources. Next, we report real-world use cases, such as HyperParameter Optimization, Monte Carlo Toy confidence limits calculation, and Active Learning. Finally, we conclude with future plans.

97 MATHEMATICS AND COMPUTING↗

Machine learning for modern power distribution systems: Progress and perspectives

The application of machine learning (ML) to power and energy systems (PES) is being researched at an astounding rate, resulting in a significant number of recent additions to the literature. As the infrastructure of electric power systems evolves, so does interest in deploying ML techniques to PES. However, despite growing interest, the limited number of reported real-world applications suggests that the gap between research and practice is yet to be fully bridged. To help highlight areas where this gap could be narrowed, this article discusses the challenges and opportunities in developing and adapting ML techniques for modern electric power systems, with a particular focus on power distribution systems. These systems play a crucial role in transforming the electric power sector and accommodating emerging distributed technologies to mitigate the impacts of climate change and accelerate the transition to a sustainable energy future. The objective of this article is not to provide an exhaustive overview of the state-of-the-art in the literature, but rather to make the topic accessible to readers with an engineering or computer science background and an interest in the field of ML for PES, thereby encouraging cross-disciplinary research in this rapidly developing field. To this end, the article discusses the ways in which ML can contribute to addressing the evolving operational challenges facing power distribution systems and identifies relevant application areas that exemplify the potential for ML to make near-term contributions. At the same time, key considerations for the practical implementation of ML in power distribution systems are discussed, along with suggestions for several potential future directions.

Marković, Marija (ORCID:0000000247839837)↗

Towards FAIR Workflows for Federated Experimental Sciences

A de-centralized, peer-to-peer AI metadata framework is demonstrated which can enable end-to-end metadata & lineage tracking for distributed Machine Learning pipelines spanning edge, High Performance Computing, and cloud environments. With a specific example of end-to-end microscopy algorithm and datasets, the proposed method shows how to enable reproducibility, audit trail, provenance of metadata artifacts. The emerging needs of automation in experimental sciences, ML-centric workflows, and FAIR metadata management across federated compute environments is addressed.

machine learning↗

Dynamical Sketching for Enhanced Communication Efficiency in Federated Learning

Federated learning (FL) has revolutionized distributed machine learning by enabling collaborative model training without sharing local data. However, communication efficiency and privacy guarantees remain significant challenges. This paper introduces a dynamic sketching mechanism in FL, optimizing the trade-off between communication efficiency and model accuracy. By dynamically selecting the sketch matrix size, our approach adapts to the evolving characteristics of the data and the model, ensuring optimal performance across diverse scenarios. We leverage Bayesian optimization to systematically tune the sketch parameters, achieving an effective balance between resource efficiency and model performance. Experimental results on the MNIST dataset using a convolutional neural network (CNN) architecture validate the proposed method's efficiency and scalability. Our dynamic sketching approach significantly outperforms fixed-size sketching techniques, achieving higher compression ratios (up to 62x) and providing better privacy guarantees while maintaining high model accuracy. These findings highlight the robustness and versatility of our approach and make it a valuable solution for privacy-preserving, communication-efficient federated learning.

Afrose, Sharmin [ORNL]↗

FedOSAA: Improving Federated Learning with One-Step Anderson Acceleration

Federated learning (FL) is a distributed machine learning approach that enables multiple local clients and a central server to collaboratively train a model while keeping the data on their own devices. First-order methods, particularly those incorporating variance reduction techniques, are the most widely used FL algorithms due to their simple implementation and stable performance. However, these methods tend to be slow and require a large number of communication rounds to reach the global minimizer. We propose FedOSAA, a novel approach that preserves the simplicity of first-order methods while achieving the rapid convergence typically associated with second-order methods. Our approach applies one Anderson acceleration (AA) step following classical local updates based on first-order methods with variance reduction, such as FedSVRG and SCAFFOLD, during local training. This AA step is able to leverage curvature information from the history points and gives a new update that approximates the Newton-GMRES direction, thereby significantly improving the convergence. We establish a local linear convergence rate to the global minimizer of FedOSAA for smooth and strongly convex loss functions. Numerical comparisons show that FedOSAA substantially improves the communication and computation efficiency of the original first-order methods, achieving performance comparable to second-order methods like GIANT.

Feng, Xue [University of California, Davis]↗

5G integrated edge computing platform for efficient component monitoring in coal-fired power plants

This project developed a cutting-edge 5G-integrated edge computing framework to enhance operational efficiency and reliability in coal-fired power plants through real-time component monitoring and anomaly detection. The initiative focused on leveraging distributed machine learning, federated learning, and 5G-based dynamic network slicing to support scalable, fault-tolerant monitoring environments to meet the operational requirements in industrial control systems. With a Distributed Edge Computing Service (DECS) orchestration, this project enabled federated learning at edge for condition monitoring and introduced adaptive client selection strategies to minimize communication overhead. Scalable distributed training was achieved using the Horovod framework, thus enhancing performance across edge nodes. In the realm of 5G networking, the project designed and deployed reconfigurable, QoS-aware network slicing tailored for operational technology (OT) environments, integrating software-defined networks to bolster cyber-resilience and enabling dynamic slicing for federated learning workloads. A significant milestone was the development of a virtualized ICS environment with 5G core integration—which allowed elastic and fault tolerant distributed training on real-world datasets such as NASA Bearings, Hydraulic Systems, and TEP. To broaden the impact of the project, a TRL-3 virtualized ICS testbed for research and education was designed. This project engaged several graduate and undergraduate students to conduct research on the cutting-edge technology, and it resulted in one PhD dissertation, one MS thesis, and over 14 peer-reviewed publications. With the support of this project students also participated in national cybersecurity competitions to improve their professional development skills.

20 FOSSIL-FUELED POWER PLANTS↗

Outage Cause Classification of Power Distribution Systems with Machine Learning and Real-World Data

Power distribution systems are geographically dispersed by nature. It may be affected by various factors, such as vegetation, weather, animal and human behaviors. Present response procedures to an outage event massively rely on expert experience and thus tend to be time-consuming. Automatic outage event detection and classification will help to reduce the responding and restoration time. However, this issue is less addressed with existing research done in this area. In this applied research, a set of waveform pre-processing techniques are first proposed to prepare the waveform data for being used as inputs to the classification algorithm. Further, a machine learning-based algorithm is proposed to classify the outage events according to their root causes, e.g. tree contact, animal contact, lightning, etc. Available data include three phase current & voltage waveforms and contextual information during the distribution system outages. The proposed machine learning algorithm takes the current and voltage waveforms as direct inputs in search of features that humans are unable to capture. Real data provided by a distribution company in the East Tennessee region is used to test the proposed pre-processing techniques and the classification algorithm.

Sun, Haoyuan↗

OEDI—Solar Grid Integration Data and Analytics Library

As a part of the Open Energy Data Initiative, this effort aims to develop and demonstrate novel distribution state estimation, control optimization, and transient analysis as well as provide access to data, data integration, and mapping information. More specifically, the focus of the effort will be on physics-based distribution system state estimation, hybrid (physics-based and machine learning) distribution optimal power flow, and event detection/analysis for solar integration and analytics. This work will enable reproducible, robust, replicable, and generalizable R&D in simulation and emulation of solar system integration. These test models and datasets will provide an integrated library for developing and testing power system operation technologies. To make the library user-friendly, this project will provide data curation tools such as data translators, mapping scripts and APIs, database schemas and metadata, interfaces and user dashboard, source code for the reference algorithms, description of the use-cases/scenarios, and comprehensive information on all the assumptions.

14 SOLAR ENERGY↗

Robust Decentralized Learning Using ADMM With Unreliable Agents

Many signal processing and machine learning problems can be formulated as consensus optimization problems which can be solved efficiently via a cooperative multi-agent system. However, the agents in the system can be unreliable due to a variety of reasons: noise, faults and attacks. Providing erroneous updates leads the optimization process in a wrong direction, and degrades the performance of distributed machine learning algorithms. This paper considers the problem of decentralized learning using ADMM in the presence of unreliable agents. First, we rigorously analyze the effect of erroneous updates (in ADMM learning iterations) on the convergence behavior of the multi-agent system. We show that the algorithm linearly converges to a neighborhood of the optimal solution under certain conditions and characterize the neighborhood size analytically. Next, we provide guidelines for network design to achieve a faster convergence to the neighborhood. Here, we also provide conditions on the erroneous updates for exact convergence to the optimal solution. Finally, to mitigate the influence of unreliable agents, we propose ROAD , a robust variant of ADMM, and show its resilience to unreliable agents with an exact convergence to the optimum.

97 MATHEMATICS AND COMPUTING↗

Snow distribution patterns revisited: A physics-based and machine learning hybrid approach to snow distribution mapping in the sub-Arctic Supporting Data

Snow in the Arctic and sub-Arctic is highly variable at fine scales, with deep drifts and shallow scoured areas creating a complex pattern of snow on the landscape. This fine-scale variation in snow is driven primarily by landscape and vegetation properties. Some landscape features, such as river beds, will rapidly fill in with snow during the wintertime due to high winds, while shrubs will trap blowing snow, resulting in drifts. Meanwhile, snow will blow off of exposed areas, resulting in abnormally shallow snow. These complex interactions between wind, vegetation, and terrain are difficult to represent well with physically-based models, but machine learning techniques have shown promise in the past. Here, we propose a hybrid modeling approach, where we use machine learning derived snow pattern maps to inform SnowModel, a physically-based snow process model. We develop and test this technique at the Teller 27 Seward Peninsula NGEE-Arctic study site. This dataset includes 5 *.nc files of model inputs and outputs plus one user guide (*.pdf). We present the data we use to drive SnowModel (vegetation type, digital elevation model), the snow pattern maps used to inform SnowModel (Standardized Depth Values maps, Machine Learning Snow Distribution Pattern), and our machine learning, SnowModel, and Hybrid snow depth and snow water equivalent results. The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic), a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research. The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska. Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

54 ENVIRONMENTAL SCIENCES↗

Geospatial mapping of distribution grid with machine learning and publicly-accessible multi-modal data

Abstract Detailed and location-aware distribution grid information is a prerequisite for various power system applications such as renewable energy integration, wildfire risk assessment, and infrastructure planning. However, a generalizable and scalable approach to obtain such information is still lacking. In this work, we develop a machine-learning-based framework to map both overhead and underground distribution grids using widely-available multi-modal data including street view images, road networks, and building maps. Benchmarked against the utility-owned distribution grid map in California, our framework achieves > 80% precision and recall on average in the geospatial mapping of grids. The framework developed with the California data can be transferred to Sub-Saharan Africa and maintain the same level of precision without fine-tuning, demonstrating its generalizability. Furthermore, our framework achieves a R 2 of 0.63 in measuring the fraction of underground power lines at the aggregate level for estimating grid exposure to wildfires. We offer the framework as an open tool for mapping and analyzing distribution grids solely based on publicly-accessible data to support the construction and maintenance of reliable and clean energy systems around the world.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Machine Learning-Assisted Distribution System Network Reconfiguration Problem

High penetration from volatile renewable energy resources in the grid and the varying nature of loads raise the need for frequent line switching to ensure the efficient operation of electrical distribution networks. Operators must ensure maximum load delivery, reduced losses, and the operation between voltage limits. However, computations to decide the optimal feeder configuration are often computationally expensive and intractable, making it unfavorable for real-time operations. This is mainly due to the existence of binary variables in the network reconfiguration optimization problem. To tackle this issue, we have devised an approach that leverages machine learning techniques to reshape distribution networks featuring multiple substations. This involves predicting the substation responsible for serving each part of the network. Hence, it leaves simple and more tractable Optimal Power Flow problems to be solved. This method can produce accurate results in a significantly faster time, as demonstrated using the IEEE 37-bus distribution feeder. Compared to the traditional optimization-based approaches, a feasible solution is achieved approximately ten times faster for all the tested scenarios.

deep neural networks↗

Snow Distribution Patterns Revisited: A Physics-Based and Machine Learning Hybrid Approach to Snow Distribution Mapping in the Sub-Arctic

Snowpack distribution in Arctic and alpine landscapes often occurs in repeating, year-to-year patterns due to local topographic, weather, and vegetation characteristics. Previous studies have suggested that with years of observational data, these snow distribution patterns can be statistically integrated into a snow process modeling workflow. Recent advances in snow hydrology and machine learning (ML) have increased our ability to predict snowpack distribution using in-situ observations, remote sensing data sets, and simple landscape characteristics that can be easily obtained for most environments. Here, we propose a hybrid approach to couple a ML snow distribution pattern (MLSDP) map with a physics-based, snow process model. We trained a random forest ML algorithm on tens of thousands of snow survey observations from a subarctic study area on the Seward Peninsula, Alaska, collected during peak snow water equivalent (SWE). We validated hybrid model outputs using in-situ snow depth and SWE observations, as well as a light detection and ranging data set and a distributed temperature profiling sensor data set. When the hybrid results were compared with the physics-based method, the hybrid method more accurately depicted the spatial patterns of the snowpack, areas of drifting snow, and years when no in-situ observations were used in the random forest ML training data set. The hybrid method also showed improvements in root mean squared error at 61% of locations where time-series estimations of snow depth were observed. These results can be applied to any physics-based model to improve the snow distribution patterning to reflect observed conditions in high latitude and high elevation cold region environments.

54 ENVIRONMENTAL SCIENCES↗

Knowledge-guided graph machine learning for spatially distributed prediction of daily discharge and nitrogen export dynamics

Spatially distributed prediction of streamflow and nitrogen export dynamics is essential for precision management of agricultural watersheds. While temporal deep learning models such as Long Short-Term Memory (LSTM) have shown strong performance at basin scales, their ability to generalize spatially is limited by insufficient representation of spatial dependencies and flow paths, particularly under data-scarce conditions. To address this gap, we propose HydroGraphNet, a knowledge-guided graph machine learning framework that integrates process-based knowledge and explicit spatial learning into temporal modeling. This framework incorporates directed graph topology to encode watershed connectivity and upstream inflows, with mass balance constraints to improve physical consistency. To enhance generalization in sparsely monitored regions, HydroGraphNet is pretrained on synthetic data generated by the SWAT+ (Soil and Water Assessment Tool Plus) model. We evaluated HydroGraphNet in the Upper Sangamon River Basin (44 HUC-12 subwatersheds, 2001–2020) against two LSTM baselines: a lumped basin-level model and a distributed variant. When benchmarked on SWAT+ simulations in pretraining, HydroGraphNet improved test NSEs by 8.9% (discharge) and 13.7% (NO₃–N load) in temporal extrapolation, and by 27.1% and 34.7% in spatial extrapolation, relative to the Lumped LSTM baseline. After fine-tuning with USGS monitoring data, the model achieved mean test NSE (KGE) scores of 0.768 (0.861) for discharge and 0.626 (0.664) for NO₃–N load, substantially outperforming baselines. Attribution analysis further highlighted the importance of upstream inflow representation and graph-based spatial learning in capturing cross-subwatershed dependencies. The model also reproduced seasonal hydrological and biogeochemical patterns consistent with known processes, demonstrating its robustness and process fidelity for spatially distributed prediction. Altogether, HydroGraphNet advances the integration of physical knowledge and spatially explicit learning in hydrological modeling, offering a generalizable framework for distributed modeling to support spatially targeted water quality management in data-scarce watersheds.

54 ENVIRONMENTAL SCIENCES↗

Search for New Phenomena in Two-Body Invariant Mass Distributions Using Unsupervised Machine Learning for Anomaly Detection at s = 13 TeV with the ATLAS Detector

Searches for new resonances are performed using an unsupervised anomaly-detection technique. Events with at least one electron or muon are selected from 140 fb − 1 of p p collisions at s = 13 TeV recorded by ATLAS at the Large Hadron Collider. The approach involves training an autoencoder on data, and subsequently defining anomalous regions based on the reconstruction loss of the decoder. Studies focus on nine invariant mass spectra that contain pairs of objects consisting of one light jet or b jet and either one lepton ( e , μ ) , photon, or second light jet or b jet in the anomalous regions. No significant deviations from the background hypotheses are observed. Limits on contributions from generic Gaussian signals with various widths of the resonance mass are obtained for nine invariant masses in the anomalous regions. © 2024 CERN, for the ATLAS Collaboration 2024 CERN

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Integrated analysis of X-ray diffraction patterns and pair distribution functions for machine-learned phase identification

Abstract To bolster the accuracy of existing methods for automated phase identification from X-ray diffraction (XRD) patterns, we introduce a machine learning approach that uses a dual representation whereby XRD patterns are augmented with simulated pair distribution functions (PDFs). A convolutional neural network is trained directly on XRD patterns calculated using physics-informed data augmentation, which accounts for experimental artifacts such as lattice strain and crystallographic texture. A second network is trained on PDFs generated via Fourier transform of the augmented XRD patterns. At inference, these networks classify unknown samples by aggregating their predictions in a confidence-weighted sum. We show that such an integrated approach to phase identification provides enhanced accuracy by leveraging the benefits of each model’s input representation. Whereas networks trained on XRD patterns provide a reciprocal space representation and can effectively distinguish large diffraction peaks in multi-phase samples, networks trained on PDFs provide a real space representation and perform better when peaks with low intensity become important. These findings underscore the importance of using diverse input representations for machine learning models in materials science and point to new avenues for automating multi-modal characterization.

36 MATERIALS SCIENCE↗

Generative unfolding with distribution mapping

Machine learning enables unbinned, highly-differential cross section measurements. A recent idea uses generative models to morph a starting simulation into the unfolded data. We show how to extend two morphing techniques, Schrödinger Bridges and Direct Diffusion, in order to ensure that the models learn the correct conditional probabilities. This brings distribution mapping (DM) to a similar level of accuracy as the state-of-the-art conditional generative unfolding methods. Numerical results are presented with a standard benchmark dataset of single jet substructure as well as for a new dataset describing a 22-dimensional phase space of Z+2 -jets.

Butter, Anja↗