Search NASA⌕ Search

SEARCH · Search NASA

Results for “distributed learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

Quantifying Groundwater Response and Uncertainty in Beaver‐Influenced Mountainous Floodplains Using Machine Learning‐Based Model Calibration

Abstract Beavers ( Castor canadensis ) alter river corridor hydrology by creating ponds and inundating floodplains, and thereby improving surface water storage. However, the impact of inundation on groundwater, particularly in mountainous alluvial floodplains with permeable gravel/cobble layers overlain by a soil layer, remains uncertain. Numerical modeling across various floodplain structures considers topographic and sediment complexity and multidirectional flow, linking inundation to groundwater response. This study develops a model‐data integration workflow to address uncertainty in groundwater response to beaver‐induced inundations in a mountainous alluvial floodplain in the Upper Colorado River Basin. Uncertain factors include seasonal hydrologic dynamics, hydraulic conductivities, floodplain structures, and meteorological forcings. We employed an ensemble of groundwater models, based on geophysical and hydrologic data, with machine learning‐based calibration using a neural density estimator. This allowed us to quantify the vertical flux from the soil layer to the permeable gravel bed, the down‐valley underflow within the gravel bed, and their ratios. Results show a significant increase in the vertical flux relative to down‐valley underflow, from 2 during dry pond periods to 20 during wet periods, serving as an analogy for conditions without and with beaver ponds. The study highlights the influence of floodplain structure on groundwater storage, water balance, and water quality impacted by beaver ponds. A thick gravel bed layer, with a large down‐valley underflow, minimizes the effect of beaver‐induced inundation on water quality. We emphasize the need for field‐scale measurements of floodplain structure and improved characterization of evapotranspiration changes to reduce uncertainty in groundwater response. Plain Language Summary Beavers change the flow of water in river corridors by creating ponds, expanding wetlands, and flooding floodplains. This increases surface water area, promotes plant growth, and enhances biodiversity. However, the impact of this flooding on groundwater flow is not well understood, especially in mountainous areas with gravel layers where water moves easily beneath soil. In this study, we used numerical modeling to investigate how beaver ponds influence groundwater in a mountainous floodplain of the Upper Colorado River Basin. We adapted a machine learning method to validate our numerical models using multiple field data sets. Our findings show that beaver ponds significantly increase vertical water flow from the soil to the gravel during wet periods, compared to when the ponds are fully drained. The study also highlights the importance of floodplain structure in controlling both water flow in gravel layers along the river direction and vertical flow from the soil to the gravel with the presence of beavers. To reduce uncertainty in groundwater response, we emphasize the need for more field‐scale measurements of floodplain structure, hydraulic properties, and evapotranspiration changes. Key Points Floodplain structures and hydraulic conductivities are important for groundwater response with beaver ponds in mountainous floodplains Large down‐valley underflow in permeability‐stratified floodplains reduces beaver‐induced impacts on groundwater storage and water quality Machine learning‐based model calibration methods are effective for estimating posterior distributions of groundwater model parameters

Wang, Lijing↗

Laser aberration signatures in expelled electrons from a tenuous gas

We describe a mechanism by which the aberration content of a focused, multiterawatt laser is imprinted upon the forward angular distribution of electrons ionized within and ponderomotively expelled from the focal volume. In our experiments, the laser aberration type and magnitude are controllably varied, and the measured electron distributions are correspondingly modified in a way consistent with predictions from numerical simulations. This imprint mechanism shows potential for enabling the development of accurate focal-spot characterization of intense lasers when fired at full power and is being developed for extension to petawatt laser systems.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Convolutional Variational Autoencoder-based Unsupervised Learning for Power Systems Faults

Classification of power system event data is a growing need, particularly where non-protective relaying-based sensors are used to monitor grid performance. Given the high burden of obtaining event data with appropriate labeling, an unsupervised approach is highly valuable. This approach enables using event data without labeling, which is far easier to obtain. This paper presents an unsupervised learning method to classify and label transients observed in the distribution grid. A Convolutional Variational Autoencoder (CVAE) was developed for this purpose. We demonstrate the efficacy of our approach using the transient data generated from the simulations. The simulation data is used to train the CVAE that identifies different faults as different clusters in the latent space. The clusters are then used as the foundation model to categorize the real-world data.

Alam, Maksudul↗

Aerodynamic Sensitivities over Separable Shape Tensors

Here, we present a comprehensive aerodynamic sensitivity analysis of airfoil parameterization informed by separable shape tensors. This parameterization approach uniquely benefits the design process by isolating various well-studied shape characteristics, such as airfoil thickness, and providing a well-regulated low-dimensional parameter domain for aerodynamic designs. Exploring the aerodynamic sensitivities of this novel parameterization can provide valuable insights for more robust designs and future manufacturing efforts. We construct a data-driven parameter space of airfoils using principal geodesic analysis of separable shape tensors informed by a curated database containing almost 20,000 suitable engineering airfoils. Analyzing the shape reconstruction error and the maximum mean discrepancy between joint distributions of aerodynamic quantities, we study the dimensionality of the learned parameter space. This simple numerical experiment demonstrates a dramatic dimension reduction that retains design effectiveness and promotes regularity of the shape representations. Finally, we generate new airfoils and use the HAM2D Reynolds-averaged Navier–Stokes solver to predict lift, drag, and moment coefficients. We compute multiple sensitivity metrics to quantify and assert the consistency of parameter influence on the aerodynamic quantities. We also explore low-dimensional polynomial ridge approximations to motivate physical intuitions and offer explanations of the approximated sensitivities.

17 WIND ENERGY↗

A New Framework for Interstellar Medium Emission Line Models: Connecting Multiscale Simulations across Cosmological Volumes

The James Webb Space Telescope (JWST) and Atacama Large Millimeter/submillimeter Array have detected emission lines from the ionized interstellar medium (ISM) in some of the first galaxies at z ≳ 6. These measurements present an opportunity to better understand galaxy assembly histories and may allow important tests of state-of-the-art galaxy formation simulations. It is challenging, however, to model these lines in their proper cosmological context. In order to meet this challenge, we introduce a novel subgrid line emission modeling framework. The framework uses the high-z zoom-in simulation suite from the Feedback in Realistic Environments (FIRE) collaboration. The line emission signals from H II regions within each simulated FIRE galaxy are modeled using the semianalytic HIIL INES code. A machine learning approach is then used to determine the conditional probability distribution for the line luminosity to stellar-mass ratio from the H II regions around each simulated stellar particle. This conditional probability distribution can then be applied to predict the line luminosities around stellar particles in lower-resolution, yet larger volume cosmological simulations. As an example, we apply this approach to the IllustrisTNG simulations at z = 6. The resulting predictions for the [O II ], [O III ], and Balmer line luminosities as a function of star formation rate agree well with current observations. Our predictions differ, however, from related works in the literature, which lack detailed subgrid ISM models. This highlights the importance of our multiscale simulation modeling framework. Finally, we provide forecasts for future line luminosity function measurements from the JWST and quantify the cosmic variance in such surveys.

(ISM:) H II regions↗

Managing Information On Costs

Cost Management Model, CMM, software tool for planning, tracking, and reporting costs and information related to costs. Capable of estimating costs, comparing estimated to actual costs, performing "what-if" analyses on estimates of costs, and providing mechanism to maintain data on costs in format oriented to management. Number of supportive cost methods built in: escalation rates, production-learning curves, activity/event schedules, unit production schedules, set of spread distributions, tables of rates and factors defined by user, and full arithmetic capability. Import/export capability possible with 20/20 Spreadsheet available on Data General equipment. Program requires AOS/VS operating system available on Data General MV series computers. Written mainly in FORTRAN 77 but uses SGU (Screen Generation Utility).

Taulbee, Zoe A.↗

Sensor fusion IV: Control paradigms and data structures; Proceedings of the Meeting, Boston, MA, Nov. 12-15, 1991

Various papers on control paradigms and data structures in sensor fusion are presented. The general topics addressed include: decision models and computational methods, sensor modeling and data representation, active sensing strategies, geometric planning and visualization, task-driven sensing, motion analysis, models motivated biology and psychology, decentralized detection and distributed decision, data fusion architectures, robust estimation of shapes and features, application and implementation. Some of the individual subjects considered are: the Firefly experiment on neural networks for distributed sensor data fusion, manifold traversing as a model for learning control of autonomous robots, choice of coordinate systems for multiple sensor fusion, continuous motion using task-directed stereo vision, interactive and cooperative sensing and control for advanced teleoperation, knowledge-based imaging for terrain analysis, physical and digital simulations for IVA robotics.

Schenker, Paul S.↗

Publication of science data on CD-ROM: A guide and example

CD-ROM (Compact Disk-Read Only Memory) is becoming the standard media not only in audio recording, but also in the publication of data and information accessible on many computer platforms. Little has been written about the complicated process involved in creating easy-to-use, high quality, and useful CD-ROM's containing scientific data. This document is a manual designed to aid those who are responsible for the publication of scientific data on CD-ROM. All aspects and steps of the procedure are covered, from feasibility assessment through disk design, data preparation, disc mastering, and CD-ROM distribution. General advice and actual examples are based on lessons learned from the publication of scientific data for an interdisciplinary field experiment. Appendices include actual files from a CD-ROM, a purchase request for CD-ROM mastering services, and the disk art for the first disk published for the project.

Angelici, Gary↗

Standards, Aligned Lessons, and Content of Earth Science

Never before have we been able to see so clearly how Earth breathes, moves, and lives. Global change and how we are affecting it are hot topics, and NASA is putting the necessary tools to work to collect the global data and monitor it over time. Plenty of NASA materials will be distributed free to attendees for use in their classrooms or learning centers.

Meeson, Blanche W.↗

Sampling Functions from Gaussian Processes and Structured Covariance Gaussian Networks

When learning aerodynamic models from data, it is critical to incorporate estimates of model uncertainty. This motivates the design of probabilistic aerodynamic databases which can be sampled to generate physically and statistically plausible aerodynamic models. In this talk we discuss how to sample deterministic functions from two different kinds of probabilistic models and demonstrate their use. First, Gaussian Process Regressors (GPRs) are a widely used probabilistic kernel-based model which can be thought of as Gaussian distributions over functions. GPRs are generally trained by maximizing the marginal likelihood of seeing the training data over the kernel parameter space. Sample functions are easily generated by drawing points from the Gaussian distribution at desired input points. However, when the points are not known ahead of time, the classical sampling approach is not possible since successive function samples will generate different function realizations. We present an approach for sampling consistent function evaluations from a GPR over multiple samples. Second, we describe a neural network architecture which learns a conditional Gaussian distribution by maximizing the marginal likelihood at each point in the input space. We then discuss and compare several options for generating sample functions which match this distribution. Finally, we demonstrate the use of these probabilistic aerodynamic models in an atmospheric reentry simulation.

Gaussian process regression↗

LAADS DAAC Migrates to the Cloud: Lessons Learned from Communicating About Earth Science Data on the Cloud

The Level-1 and Atmosphere Archive Distribution System (LAADS) Distributed Active Archive Center (DAAC) is migrating data to the cloud. As one of twelve DAACS supported by NASA’s Earth Science Data and Information System (ESDIS), LAADS is using moving away from on premise data storage facilities to migrating to Amazon Web Services, where the massive archive of data from the Moderate Imaging Spectroradiometer (MODIS) and the Visible Infrared Imaging Radiometer Suite (VIIRS) will be available for download and post-processing transformations online. The migration is happening in three phases and LAADS is concluding its beta testing period. This poster shows the lessons learned from communicating with a select group of users about how to effectively educate data users on using data in the cloud.

Tassia Owen↗

Image-Based Fracture Surface Defect Characterization Methods for Additively Manufactured Ti-6Al-4V Tested in Fatigue

Abstract Fatigue initiation in additively manufactured samples/parts often occurs at processed-induced defects such as lack-of-fusion (LoF), keyhole, or other morphological/microstructural defects that have unique characteristics and measurable qualities. Attempts at identifying and minimizing such defects have utilized optimized processing conditions along with in situ and ex situ characterization that includes metallography and/or X-ray computed tomography (XCT). This paper highlights the benefits of using fracture surface analyses to detect and quantify defects that may not be detected by metallography/XCT due to sectioning and resolution limits. In addition to using manual quantification of fatigue initiating LoF and keyhole defects on fracture surfaces, image-based machine learning using convolutional neural networks such as U-Net were also used to automate the process. Statistical analyses were used to identify the extreme cases of defects that initiated and accelerated fatigue and to model the distribution of defect size and shape characteristics to distinguish the type of defect. Initial results show agreement between trained machine learning models and ground truth data in defect segmentation, and the distributions of defect characteristics are distinguishable to particular process-induced defect types.

Materials Science↗

FedEFsz: Fair Cross-Silo Federated Learning System With Error-Bounded Lossy Compression

Cross-Silo federated learning systems have been identified as an efficient approach to scaling DNN training across geographically-distributed data silos to preserve the privacy of the training data. Communication efficiency and fairness are two major issues that need to be both satisfied when federated learning systems are deployed in practice. Simultaneously guaranteeing both of them, however, is exceptionally difficult because simply combining communication reduction and fairness optimization approaches often causes non-converged training or drastic accuracy degradation. Here, to bridge this gap, we propose FedEFsz. On the one hand, it integrates the state-of-the-art error-bounded lossy compressor SZ3 into cross-silo federated learning systems to significantly reduce communication traffic during the training. On the other hand, it achieves a high fairness (i.e., rather consistent model accuracy and performance across different clients) through a carefully designed heuristic algorithm that can tune the error-bound of SZ3 for different clients during the training. Extensive experimental results based on a GPU cluster with 65 GPU cards show that FedEFsz improves the fairness across different benchmarks by up to 60.88% and meanwhile reduces the communication traffic by up to 315×.

Cross-Silo Federated Learning Systems↗

Uncertainty guided online ensemble for non-stationary data streams in fusion science

Machine Learning (ML) is poised to play a pivotal role in the development and operation of next-generation fusion devices. Fusion data shows non-stationary behavior with distribution drifts, resulted by both experimental evolution and machine wear-and-tear. ML models assume stationary distribution and fail to maintain performance when encountered with such non-stationary data streams. Online learning techniques have been leveraged in other domains, however it has been largely unexplored for fusion applications. In this paper, we investigate online learning for continuous adaptation to drifting data streams in the prediction of Toroidal Field (TF) coils deflection at the DIII-D fusion facility. We further address the short-term performance degradation inherent to standard online learning, which arises because ground truth is unavailable at prediction time. To mitigate this issue, we propose an uncertainty-guided online ensemble framework. The method leverages the Deep Gaussian Process Approximation (DGPA) for calibrated uncertainty estimation and uses these uncertainty measures to guide a meta-algorithm that aggregates predictions from learners trained over different historical horizons. Our results show that online learning reduces prediction error by 80% compared to a static model. The online ensemble and the proposed uncertainty-guided ensemble further reduce error by approximately 6%, and 10% respectively, relative to standard single-model online learning, while also providing calibrated uncertainty estimates to support operational decision-making.

AI↗

Fostering Geothermal Machine Learning Success: Elevating Big Data Accessibility and Automated Data Standardization in the Geothermal Data Repository: Preprint

The Department of Energy's (DOE) Geothermal Data Repository (GDR) has implemented improvements to both its data lakes and its data standards and automated data pipelines. The GDR data lakes have reduced storage and compute-related barriers to using large geothermal datasets, enabling these large datasets to be accessed by anyone with a modern computer and internet access. More recently, the GDR has been working to further reduce barriers through streamlining the data intake process, educating users on the process and requirements, and aiding users in accessing data from the data lakes. These improvements have augmented the quantity of datasets the GDR is able to accept into its data lakes and have enabled users who are new to cloud tools to access these datasets more easily, overall increasing the accessibility of big geothermal data for use in machine learning and other projects. In addition, the GDR now has built-in data standards and pipelines for drilling data, geospatial data, and distributed acoustic sensing (DAS) data. These standardization efforts aim to enhance the real-world applicability of geothermal machine learning outcomes by improving the quality of training data. Specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, thus allowing more time for actual research. By automating this process, the burden of standardization is lifted from the user, ultimately increasing the availability of standardized data.

accessibility↗

Fostering Geothermal Machine Learning Success: Elevating Big Data Accessibility and Automated Data Standardization in the Geothermal Data Repository

The Department of Energy's (DOE) Geothermal Data Repository (GDR) has implemented improvements to both its data lakes and its data standards and automated data pipelines. The GDR data lakes have reduced storage and compute-related barriers to using large geothermal datasets, enabling these large datasets to be accessed by anyone with a modern computer and internet access. More recently, the GDR has been working to further reduce barriers through streamlining the data intake process, educating users on the process and requirements, and aiding users in accessing data from the data lakes. These improvements have augmented the quantity of datasets the GDR is able to accept into its data lakes and have enabled users who are new to cloud tools to access these datasets more easily, overall increasing the accessibility of big geothermal data for use in machine learning and other projects. In addition, the GDR now has built-in data standards and pipelines for drilling data, geospatial data, and distributed acoustic sensing (DAS) data. These standardization efforts aim to enhance the real-world applicability of geothermal machine learning outcomes by improving the quality of training data. Specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, thus allowing more time for actual research. By automating this process, the burden of standardization is lifted from the user, ultimately increasing the availability of standardized data.

15 GEOTHERMAL ENERGY↗

Fostering Geothermal Machine Learning Success: Elevating Big Data Accessibility and Automated Data Standardization in the Geothermal Data Repository

The Department of Energy's (DOE's) Geothermal Data Repository (GDR) has implemented improvements to both its data lakes and its data standards and automated data pipelines. The GDR data lakes have reduced storage and compute-related barriers to using large geothermal datasets, enabling these large datasets to be accessed by anyone with a modern computer and internet access. More recently, the GDR has been working to further reduce barriers through streamlining the data intake process, educating users on the process and requirements, and helping users access data from the data lakes. These improvements have augmented the quantity of datasets the GDR is able to accept into its data lakes and have enabled users who are new to cloud tools to access these datasets more easily, overall increasing the accessibility of big geothermal data for use in machine learning and other projects. In addition, the GDR now has built-in data standards and pipelines for drilling data, geospatial data, and distributed acoustic sensing (DAS) data. These standardization efforts aim to enhance the real-world applicability of geothermal machine learning outcomes by improving the quality of training data. Specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, thus allowing more time for actual research. By automating this process, the burden of standardization is lifted from the user, ultimately increasing the availability of standardized data. This paper provides an update on recent improvements made to the GDR's data lakes and automated data pipelines, including: (1) streamlining the data lake intake process, (2) better educating users on the process and requirements through a new data lakes page, (3) adding data lake direct access links to GDR data lake submission pages, (4) implementing a DAS data pipeline to convert DAS data uploaded in SEG-Y format to a standardized hierarchical data format v5 (HDF5), (5) extending this pipeline to encompass data in the GDR data lake, (6) adding metadata requirements for geospatial data, (7) making user interface/user experience (UX) enhancements to the data pipelines' documentation pages, and (8) improving the GDR's data standards and pipelines pages to better guide users in ensuring that their data is standardized by the GDR's automated data pipelines. 2024 Geothermal Resources Council. All rights reserved.

accessibility↗