Search NASA⌕ Search

SEARCH · Search NASA

Results for “distributed learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Business Models for Scaling Demand Flexibility Volume II – Customer relationship management strategies, challenges, and lessons learned from U.S. programs

Load growth at the grid edge is driving increased attention to the distribution system and its ability to enable customer technology adoption in an affordable and timely manner. Key industry stakeholders, including electric utilities and regulators, can benefit from strategies to manage and balance customer needs with infrastructure investments, such as demand flexibility. This report focuses on demand flexibility—the ability to reduce, shift, shed, generate, or modulate loads in response to building and grid needs—to reduce the need for costly grid upgrades by deferring investment needs and increase system reliability by shifting electricity usage during periods of high risk. Specifically, we focus on the emerging characteristics of business models for demand flexibility as a framework to understand how demand flexibility programs generate value. In this report, we focus on demand flexibility program customer relationship management strategies, which provide information on value creation and focus on ensuring customers can navigate programs smoothly. This report discusses the role of customer relationship management strategies in demand flexibility programs, characterizes customer relationship management strategies that can be considered during program design and implementation, identifies existing challenges to customer relationship management strategies, and describes lessons learned. This report is part of a series that includes reports on value propositions, stakeholder ecosystem management, and program life cycle.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Business Models for Scaling Demand Flexibility Volume I – Value proposition characteristics, challenges, and lessons learned from U.S. programs

Load growth at the grid edge is driving increased attention to the distribution system and its ability to enable customer technology adoption in an affordable and timely manner. Key industry stakeholders, including electric utilities and regulators, can benefit from strategies to manage and balance customer needs with infrastructure investments, such as demand flexibility. This report focuses on demand flexibility—the ability to reduce, shift, shed, generate, or modulate loads in response to building and grid needs—to reduce the need for costly grid upgrades by deferring investment needs and increase system reliability by shifting electricity usage during periods of high risk. Specifically, we focus on the emerging characteristics of business models for demand flexibility as a framework to understand how demand flexibility programs generate value. In this report, we focus on demand flexibility value propositions, which provide information on value creation and describe how programs deliver clear benefits that address customer and grid needs. This report discusses the role of value propositions in demand flexibility programs, provides an overview of value propositions for a range of demand flexibility stakeholders, identifies existing challenges to establishing an effective value proposition, and describes lessons learned. This report is part of a series that includes reports on customer relationship management strategies, stakeholder ecosystem management, and program life cycle.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Business Models for Scaling Demand Flexibility Volume IV – Program life cycle challenges and lessons learned from U.S. programs

Load growth at the grid edge is driving increased attention to the distribution system and its ability to enable customer technology adoption in an affordable and timely manner. Key industry stakeholders, including electric utilities and regulators, can benefit from strategies to manage and balance customer needs with infrastructure investments, such as demand flexibility. This report focuses on demand flexibility—the ability to reduce, shift, shed, generate, or modulate loads in response to building and grid needs—to reduce the need for costly grid upgrades by deferring investment needs and increase system reliability by shifting electricity usage during periods of high risk. Specifically, we focus on the emerging characteristics of business models for demand flexibility as a framework to understand how demand flexibility programs generate value. In this report, we focus on the life cycle of demand flexibility programs, which provides information on value creation and describes the various deployment phases program implementers navigate from initial program conceptualization through to program expansion and replication to new customer segments and regions. This report characterizes the key phases of the demand flexibility program life cycle, identifies existing challenges across the program deployment phases, and describes lessons learned. This report is part of a series that includes reports on customer relationship management strategies, stakeholder ecosystem management, and program life cycle.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Parallel sorting algorithm classification: is manual instrumentation necessary?

Understanding parallel algorithms is crucial for accelerating scientific simulations on complex, distributed memory, high-performance computers. Modern algorithm classification approaches learn semantics directly from source code to differentiate between algorithms, however, accessing source code is not always possible. We can learn about parallel algorithms from observing their performance, as programs running the same algorithms and using the same hardware should exhibit similar performance characteristics. We present an approach to learn algorithm classes from parallel performance data directly in order to classify algorithms without access to the source code. We extend previous work to enable classifying parallel sorting algorithms using automatic instrumentation instead of requiring manual region annotations in the source code. In this work, we design and demonstrate a study for classification of parallel sorting algorithms using parallel performance data collected from automatic instrumentation, and evaluate the performance of our new methodology on classification. We leverage Caliper to collect the performance data, Thicket for our exploratory data analysis (EDA), and PyTorch and Scikit-learn to evaluate the effectiveness of random forests, support vector machines (SVMs), decision trees, neural networks, and logistic regressions on parallel performance data. Additionally, we study noise in parallel performance data, whether the removal of noise and pre-processing of the data is necessary to accurately classify parallel sorting algorithms, and determine the effectiveness of features created from performance data. In conclusion, we demonstrate classification accuracy for these five different models of up to 97.7% across four different parallel algorithm classes.

Algorithm Classification↗

Quantifying Groundwater Response and Uncertainty in Beaver‐Influenced Mountainous Floodplains Using Machine Learning‐Based Model Calibration

Abstract Beavers ( Castor canadensis ) alter river corridor hydrology by creating ponds and inundating floodplains, and thereby improving surface water storage. However, the impact of inundation on groundwater, particularly in mountainous alluvial floodplains with permeable gravel/cobble layers overlain by a soil layer, remains uncertain. Numerical modeling across various floodplain structures considers topographic and sediment complexity and multidirectional flow, linking inundation to groundwater response. This study develops a model‐data integration workflow to address uncertainty in groundwater response to beaver‐induced inundations in a mountainous alluvial floodplain in the Upper Colorado River Basin. Uncertain factors include seasonal hydrologic dynamics, hydraulic conductivities, floodplain structures, and meteorological forcings. We employed an ensemble of groundwater models, based on geophysical and hydrologic data, with machine learning‐based calibration using a neural density estimator. This allowed us to quantify the vertical flux from the soil layer to the permeable gravel bed, the down‐valley underflow within the gravel bed, and their ratios. Results show a significant increase in the vertical flux relative to down‐valley underflow, from 2 during dry pond periods to 20 during wet periods, serving as an analogy for conditions without and with beaver ponds. The study highlights the influence of floodplain structure on groundwater storage, water balance, and water quality impacted by beaver ponds. A thick gravel bed layer, with a large down‐valley underflow, minimizes the effect of beaver‐induced inundation on water quality. We emphasize the need for field‐scale measurements of floodplain structure and improved characterization of evapotranspiration changes to reduce uncertainty in groundwater response. Plain Language Summary Beavers change the flow of water in river corridors by creating ponds, expanding wetlands, and flooding floodplains. This increases surface water area, promotes plant growth, and enhances biodiversity. However, the impact of this flooding on groundwater flow is not well understood, especially in mountainous areas with gravel layers where water moves easily beneath soil. In this study, we used numerical modeling to investigate how beaver ponds influence groundwater in a mountainous floodplain of the Upper Colorado River Basin. We adapted a machine learning method to validate our numerical models using multiple field data sets. Our findings show that beaver ponds significantly increase vertical water flow from the soil to the gravel during wet periods, compared to when the ponds are fully drained. The study also highlights the importance of floodplain structure in controlling both water flow in gravel layers along the river direction and vertical flow from the soil to the gravel with the presence of beavers. To reduce uncertainty in groundwater response, we emphasize the need for more field‐scale measurements of floodplain structure, hydraulic properties, and evapotranspiration changes. Key Points Floodplain structures and hydraulic conductivities are important for groundwater response with beaver ponds in mountainous floodplains Large down‐valley underflow in permeability‐stratified floodplains reduces beaver‐induced impacts on groundwater storage and water quality Machine learning‐based model calibration methods are effective for estimating posterior distributions of groundwater model parameters

Wang, Lijing↗

Laser aberration signatures in expelled electrons from a tenuous gas

We describe a mechanism by which the aberration content of a focused, multiterawatt laser is imprinted upon the forward angular distribution of electrons ionized within and ponderomotively expelled from the focal volume. In our experiments, the laser aberration type and magnitude are controllably varied, and the measured electron distributions are correspondingly modified in a way consistent with predictions from numerical simulations. This imprint mechanism shows potential for enabling the development of accurate focal-spot characterization of intense lasers when fired at full power and is being developed for extension to petawatt laser systems.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Convolutional Variational Autoencoder-based Unsupervised Learning for Power Systems Faults

Classification of power system event data is a growing need, particularly where non-protective relaying-based sensors are used to monitor grid performance. Given the high burden of obtaining event data with appropriate labeling, an unsupervised approach is highly valuable. This approach enables using event data without labeling, which is far easier to obtain. This paper presents an unsupervised learning method to classify and label transients observed in the distribution grid. A Convolutional Variational Autoencoder (CVAE) was developed for this purpose. We demonstrate the efficacy of our approach using the transient data generated from the simulations. The simulation data is used to train the CVAE that identifies different faults as different clusters in the latent space. The clusters are then used as the foundation model to categorize the real-world data.

Alam, Maksudul↗

Aerodynamic Sensitivities over Separable Shape Tensors

Here, we present a comprehensive aerodynamic sensitivity analysis of airfoil parameterization informed by separable shape tensors. This parameterization approach uniquely benefits the design process by isolating various well-studied shape characteristics, such as airfoil thickness, and providing a well-regulated low-dimensional parameter domain for aerodynamic designs. Exploring the aerodynamic sensitivities of this novel parameterization can provide valuable insights for more robust designs and future manufacturing efforts. We construct a data-driven parameter space of airfoils using principal geodesic analysis of separable shape tensors informed by a curated database containing almost 20,000 suitable engineering airfoils. Analyzing the shape reconstruction error and the maximum mean discrepancy between joint distributions of aerodynamic quantities, we study the dimensionality of the learned parameter space. This simple numerical experiment demonstrates a dramatic dimension reduction that retains design effectiveness and promotes regularity of the shape representations. Finally, we generate new airfoils and use the HAM2D Reynolds-averaged Navier–Stokes solver to predict lift, drag, and moment coefficients. We compute multiple sensitivity metrics to quantify and assert the consistency of parameter influence on the aerodynamic quantities. We also explore low-dimensional polynomial ridge approximations to motivate physical intuitions and offer explanations of the approximated sensitivities.

17 WIND ENERGY↗

A New Framework for Interstellar Medium Emission Line Models: Connecting Multiscale Simulations across Cosmological Volumes

The James Webb Space Telescope (JWST) and Atacama Large Millimeter/submillimeter Array have detected emission lines from the ionized interstellar medium (ISM) in some of the first galaxies at z ≳ 6. These measurements present an opportunity to better understand galaxy assembly histories and may allow important tests of state-of-the-art galaxy formation simulations. It is challenging, however, to model these lines in their proper cosmological context. In order to meet this challenge, we introduce a novel subgrid line emission modeling framework. The framework uses the high-z zoom-in simulation suite from the Feedback in Realistic Environments (FIRE) collaboration. The line emission signals from H II regions within each simulated FIRE galaxy are modeled using the semianalytic HIIL INES code. A machine learning approach is then used to determine the conditional probability distribution for the line luminosity to stellar-mass ratio from the H II regions around each simulated stellar particle. This conditional probability distribution can then be applied to predict the line luminosities around stellar particles in lower-resolution, yet larger volume cosmological simulations. As an example, we apply this approach to the IllustrisTNG simulations at z = 6. The resulting predictions for the [O II ], [O III ], and Balmer line luminosities as a function of star formation rate agree well with current observations. Our predictions differ, however, from related works in the literature, which lack detailed subgrid ISM models. This highlights the importance of our multiscale simulation modeling framework. Finally, we provide forecasts for future line luminosity function measurements from the JWST and quantify the cosmic variance in such surveys.

(ISM:) H II regions↗

Image-Based Fracture Surface Defect Characterization Methods for Additively Manufactured Ti-6Al-4V Tested in Fatigue

Abstract Fatigue initiation in additively manufactured samples/parts often occurs at processed-induced defects such as lack-of-fusion (LoF), keyhole, or other morphological/microstructural defects that have unique characteristics and measurable qualities. Attempts at identifying and minimizing such defects have utilized optimized processing conditions along with in situ and ex situ characterization that includes metallography and/or X-ray computed tomography (XCT). This paper highlights the benefits of using fracture surface analyses to detect and quantify defects that may not be detected by metallography/XCT due to sectioning and resolution limits. In addition to using manual quantification of fatigue initiating LoF and keyhole defects on fracture surfaces, image-based machine learning using convolutional neural networks such as U-Net were also used to automate the process. Statistical analyses were used to identify the extreme cases of defects that initiated and accelerated fatigue and to model the distribution of defect size and shape characteristics to distinguish the type of defect. Initial results show agreement between trained machine learning models and ground truth data in defect segmentation, and the distributions of defect characteristics are distinguishable to particular process-induced defect types.

Materials Science↗

FedEFsz: Fair Cross-Silo Federated Learning System With Error-Bounded Lossy Compression

Cross-Silo federated learning systems have been identified as an efficient approach to scaling DNN training across geographically-distributed data silos to preserve the privacy of the training data. Communication efficiency and fairness are two major issues that need to be both satisfied when federated learning systems are deployed in practice. Simultaneously guaranteeing both of them, however, is exceptionally difficult because simply combining communication reduction and fairness optimization approaches often causes non-converged training or drastic accuracy degradation. Here, to bridge this gap, we propose FedEFsz. On the one hand, it integrates the state-of-the-art error-bounded lossy compressor SZ3 into cross-silo federated learning systems to significantly reduce communication traffic during the training. On the other hand, it achieves a high fairness (i.e., rather consistent model accuracy and performance across different clients) through a carefully designed heuristic algorithm that can tune the error-bound of SZ3 for different clients during the training. Extensive experimental results based on a GPU cluster with 65 GPU cards show that FedEFsz improves the fairness across different benchmarks by up to 60.88% and meanwhile reduces the communication traffic by up to 315×.

Cross-Silo Federated Learning Systems↗

Uncertainty guided online ensemble for non-stationary data streams in fusion science

Machine Learning (ML) is poised to play a pivotal role in the development and operation of next-generation fusion devices. Fusion data shows non-stationary behavior with distribution drifts, resulted by both experimental evolution and machine wear-and-tear. ML models assume stationary distribution and fail to maintain performance when encountered with such non-stationary data streams. Online learning techniques have been leveraged in other domains, however it has been largely unexplored for fusion applications. In this paper, we investigate online learning for continuous adaptation to drifting data streams in the prediction of Toroidal Field (TF) coils deflection at the DIII-D fusion facility. We further address the short-term performance degradation inherent to standard online learning, which arises because ground truth is unavailable at prediction time. To mitigate this issue, we propose an uncertainty-guided online ensemble framework. The method leverages the Deep Gaussian Process Approximation (DGPA) for calibrated uncertainty estimation and uses these uncertainty measures to guide a meta-algorithm that aggregates predictions from learners trained over different historical horizons. Our results show that online learning reduces prediction error by 80% compared to a static model. The online ensemble and the proposed uncertainty-guided ensemble further reduce error by approximately 6%, and 10% respectively, relative to standard single-model online learning, while also providing calibrated uncertainty estimates to support operational decision-making.

AI↗

Fostering Geothermal Machine Learning Success: Elevating Big Data Accessibility and Automated Data Standardization in the Geothermal Data Repository: Preprint

The Department of Energy's (DOE) Geothermal Data Repository (GDR) has implemented improvements to both its data lakes and its data standards and automated data pipelines. The GDR data lakes have reduced storage and compute-related barriers to using large geothermal datasets, enabling these large datasets to be accessed by anyone with a modern computer and internet access. More recently, the GDR has been working to further reduce barriers through streamlining the data intake process, educating users on the process and requirements, and aiding users in accessing data from the data lakes. These improvements have augmented the quantity of datasets the GDR is able to accept into its data lakes and have enabled users who are new to cloud tools to access these datasets more easily, overall increasing the accessibility of big geothermal data for use in machine learning and other projects. In addition, the GDR now has built-in data standards and pipelines for drilling data, geospatial data, and distributed acoustic sensing (DAS) data. These standardization efforts aim to enhance the real-world applicability of geothermal machine learning outcomes by improving the quality of training data. Specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, thus allowing more time for actual research. By automating this process, the burden of standardization is lifted from the user, ultimately increasing the availability of standardized data.

accessibility↗

Fostering Geothermal Machine Learning Success: Elevating Big Data Accessibility and Automated Data Standardization in the Geothermal Data Repository

The Department of Energy's (DOE) Geothermal Data Repository (GDR) has implemented improvements to both its data lakes and its data standards and automated data pipelines. The GDR data lakes have reduced storage and compute-related barriers to using large geothermal datasets, enabling these large datasets to be accessed by anyone with a modern computer and internet access. More recently, the GDR has been working to further reduce barriers through streamlining the data intake process, educating users on the process and requirements, and aiding users in accessing data from the data lakes. These improvements have augmented the quantity of datasets the GDR is able to accept into its data lakes and have enabled users who are new to cloud tools to access these datasets more easily, overall increasing the accessibility of big geothermal data for use in machine learning and other projects. In addition, the GDR now has built-in data standards and pipelines for drilling data, geospatial data, and distributed acoustic sensing (DAS) data. These standardization efforts aim to enhance the real-world applicability of geothermal machine learning outcomes by improving the quality of training data. Specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, thus allowing more time for actual research. By automating this process, the burden of standardization is lifted from the user, ultimately increasing the availability of standardized data.

15 GEOTHERMAL ENERGY↗

Fostering Geothermal Machine Learning Success: Elevating Big Data Accessibility and Automated Data Standardization in the Geothermal Data Repository

The Department of Energy's (DOE's) Geothermal Data Repository (GDR) has implemented improvements to both its data lakes and its data standards and automated data pipelines. The GDR data lakes have reduced storage and compute-related barriers to using large geothermal datasets, enabling these large datasets to be accessed by anyone with a modern computer and internet access. More recently, the GDR has been working to further reduce barriers through streamlining the data intake process, educating users on the process and requirements, and helping users access data from the data lakes. These improvements have augmented the quantity of datasets the GDR is able to accept into its data lakes and have enabled users who are new to cloud tools to access these datasets more easily, overall increasing the accessibility of big geothermal data for use in machine learning and other projects. In addition, the GDR now has built-in data standards and pipelines for drilling data, geospatial data, and distributed acoustic sensing (DAS) data. These standardization efforts aim to enhance the real-world applicability of geothermal machine learning outcomes by improving the quality of training data. Specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, thus allowing more time for actual research. By automating this process, the burden of standardization is lifted from the user, ultimately increasing the availability of standardized data. This paper provides an update on recent improvements made to the GDR's data lakes and automated data pipelines, including: (1) streamlining the data lake intake process, (2) better educating users on the process and requirements through a new data lakes page, (3) adding data lake direct access links to GDR data lake submission pages, (4) implementing a DAS data pipeline to convert DAS data uploaded in SEG-Y format to a standardized hierarchical data format v5 (HDF5), (5) extending this pipeline to encompass data in the GDR data lake, (6) adding metadata requirements for geospatial data, (7) making user interface/user experience (UX) enhancements to the data pipelines' documentation pages, and (8) improving the GDR's data standards and pipelines pages to better guide users in ensuring that their data is standardized by the GDR's automated data pipelines. 2024 Geothermal Resources Council. All rights reserved.

accessibility↗

Machine Learning for Scalable and Optimal Load Shedding Under Power System Contingency

Prompt and effective corrective actions in response to unexpected contingencies are crucial for improving power system resilience and preventing cascading blackouts. The optimal load shedding (OLS) accounting for network limits has the potential to address the diverse system-wide impacts of contingency scenarios as compared to traditional local schemes. However, due to the fast cascading propagation of initial contingencies, real-time OLS solutions are challenging to attain in large systems with high computation and communication needs. In this paper, we propose a decentralized design that leverages offline training of a neural network (NN) model for individual load centers to autonomously construct the OLS solutions from locally available measurements. Our learning-for-OLS approach can greatly reduce the computation and communication needs during online emergency responses, thus preventing the cascading propagation of contingencies for enhanced power grid resilience. Numerical studies on both the IEEE 118-bus system and a synthetic Texas 2000-bus system have demonstrated the efficiency and effectiveness of our scalable OLS learning design for timely power system emergency operations.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Measurement of the depth of maximum of air-shower profiles with energies between 10 18.5 and 10 20 eV using the surface detector of the Pierre Auger Observatory and deep learning

We report an investigation of the mass composition of cosmic rays with energies from 3 to 100 EeV ( 1 EeV = 10 18 eV ) using the distributions of the depth of shower maximum X max . The analysis relies on ∼ 50 , 000 events recorded by the surface detector of the Pierre Auger Observatory and a deep-learning-based reconstruction algorithm. Above energies of 5 EeV, the dataset offers a 10-fold increase in statistics with respect to fluorescence measurements at the Observatory. After cross-calibration using the fluorescence detector, this enables the first measurement of the evolution of the mean and the standard deviation of the X max distributions up to 100 EeV. Our findings are threefold: (i) The evolution of the mean logarithmic mass toward a heavier composition with increasing energy can be confirmed and is extended to 100 EeV. (ii) The evolution of the fluctuations of X max toward a heavier and purer composition with increasing energy can be confirmed with high statistics. We report a rather heavy composition and small fluctuations in X max at the highest energies. (iii) We find indications for a characteristic structure beyond a constant change in the mean logarithmic mass, featuring three breaks that are observed in proximity to the ankle, instep, and suppression features in the energy spectrum. Published by the American Physical Society 2025

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗