Search NASA⌕ Search

SEARCH · Search NASA

Results for “Machine Learning for Data Analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29

Best estimate of the planetary boundary layer height from multiple remote sensing measurements

Remote sensing measurements have been widely used to estimate the planetary boundary layer height (PBLHT). Each remote sensing approach offers unique strengths and faces different limitations. In this study, we use machine learning (ML) methods to produce a best-estimate PBLHT (PBLHT-BE-ML) by integrating four PBLHT estimates derived from remote sensing measurements at the Department of Energy (DOE) Atmospheric Radiation Measurement (ARM) Southern Great Plains (SGP) observatory. Three ML models – random forest (RF) classifier, RF regressor, and light gradient-boosting machine (LightGBM) – were trained on a dataset from 2017 to 2023 that included radiosonde, various remote sensing PBLHT estimates, and atmospheric meteorological conditions. Evaluations indicated that PBLHT-BE-ML from all three models improved alignment with the PBLHT derived from radiosonde data (PBLHT-SONDE), with LightGBM demonstrating the highest accuracy under both stable and unstable boundary layer conditions. Feature analysis revealed that the most influential input features at the SGP site were the PBLHT estimates derived from (a) potential temperature profiles retrieved using Raman lidar (RL) and atmospheric emitted radiance interferometer (AERI) measurements (PBLHT-THERMO), (b) vertical velocity variance profiles from Doppler lidar (PBLHT-DL), and (c) aerosol backscatter profiles from micropulse lidar (PBLHT-MPL). The trained models were then used to predict PBLHT-BE-ML at a temporal resolution of 10 min, effectively capturing the diurnal evolution of PBLHT and its significant seasonal variations, with the largest diurnal variation observed over summer at the SGP site. We applied these trained models to data from the ARM Eastern Pacific Cloud Aerosol Precipitation Experiment (EPCAPE) field campaign (EPC), where the PBLHT-BE-ML, particularly with the LightGBM model, demonstrated improved accuracy against PBLHT-SONDE. Analyses of model performance at both the SGP and EPC sites suggest that expanding the training dataset to include various surface types, such as ocean and ice-covered areas, could further enhance ML model performance for PBLHT estimation across varied geographic regions.

Zhang, Damao [Pacific Northwest National Laborator↗

A Dynamic Hierarchical Attention Framework for Multimodal Malware Detection

The increasing use of Android in the worldwide mobile ecosystem has come along with a significant increase in advanced malware, highlighting the critical necessity for efficient, scalable, and adaptable detection systems. Despite recent advancements in machine learning improving malware detection, the majority of current solutions are limited to one, two, or three data modalities, hence neglecting the comprehensive behavioral spectrum of contemporary multi-vector threats. This thesis presents the first comprehensive multimodal framework for Android malware detection, which combines textual, time-series (temporal), graph-based (structural), and visual information using an innovative hierarchical attention mechanism and Dynamic Fusion Controller (DFC). Our methodology consistently classifies and processes modalities as either sequential or structural, facilitating content-adaptive weighting and resilient cross-modal representation learning. We advance the implementation of cutting-edge time series techniques, such as MiniRocket, for malware detection, hence creating new opportunities for temporal analysis in cybersecurity. Comprehensive experimental assessment shows that our framework performs exceptionally well, with 99.46% classification accuracy and 97.15% detection accuracy, significantly outperforming existing approaches through effective multimodal integration and hierarchical attention mechanisms.

Nazmin, Tamanna↗

Safe Physics-Informed Machine Learning for Dynamics and Control

This tutorial paper focuses on safe physics-informed machine learning in the context of dynamics and control, providing a comprehensive overview of how to integrate physical models and safety guarantees. As machine learning techniques enhance the modeling and control of complex dynamical systems, ensuring safety and stability remains a critical challenge, especially in safety-critical applications like autonomous vehicles, robotics, medical decision-making, and energy systems. We explore various approaches for embedding and ensuring safety constraints, including structural priors, Lyapunov and Control Barrier Functions, predictive control, projections, and robust optimization techniques. Additionally, we delve into methods for uncertainty quantification and safety verification, including reachability analysis and neural network verification tools, which help validate that control policies remain within safe operating bounds even in uncertain environments. The paper includes illustrative examples demonstrating the implementation aspects of safe learning frameworks that combine the strengths of data-driven approaches with the rigor of physical principles, offering a path toward the safe control of complex dynamical systems.

Drgona, Jan↗

Using Machine Learning to Predict Cloud Turbulent Entrainment–Mixing Processes

Different turbulent entrainment–mixing mechanisms between clouds and environment are essential to cloud–related processes; however, accurate representation of entrainment–mixing in weather/climate models still poses a challenge. This study exploits the use of machine learning (ML) to address this challenge. Four ML (Light Gradient Boosting Machine [LGB], eXtreme Gradient Boosting, Random Forest, and Support Vector Regression) are examined and compared. It is found that LGB performs best, and thus is selected to understand the impact of entrainment–mixing on microphysics using simulation data from Explicit Mixing Parcel Model. Compared with traditional parameterizations, the trained LGB provides more accurate microphysical properties (number concentration and cloud droplet spectral dispersion). The partial dependences of predicted microphysics on features exhibit a strong alignment with physical mechanisms and expectations, as determined by the interpreting method, thus overcoming the limitations of the “black box” scheme. The underlying mechanisms are that the smaller number concentration and larger spectral dispersion correspond to more inhomogeneous entrainment–mixing. Specifically, number concentration after entrainment–mixing is positively correlated with adiabatic number concentration and liquid water content affected by entrainment–mixing, and inversely correlated with adiabatic volume mean radius. Spectral dispersion after entrainment–mixing is negatively correlated with liquid water content affected by entrainment–mixing, turbulent dissipation rate and relative humidity of entrained air. Sensitivity analysis further suggests that number concentration is mainly determined by cloud microphysical properties whereas spectral dispersion is influenced by both cloud microphysical properties and environmental variables. The results indicate that the LGB scheme has the potential to enhance the representation of entrainment–mixing in weather/climate models.

54 ENVIRONMENTAL SCIENCES↗

Predicting Large‐Scale Systematic Missing Pipe Attributes in Water Distribution Networks

Water distribution network (WDN) models are an essential tool used by water utilities for hydraulic analysis. Unfortunately, missing data and insufficient resources often make creating and maintaining these models unfeasible. Existing methods to address missing pipe properties, like sequential imputation for missing values and reconstruction using graph metrics, are designed to accommodate random patterns of missing information and require a significant percentage of the system's attributes to be known. However, these data completeness assumptions do not always align with real‐world scenarios where large sections of the WDN model have missing data. To address this challenge, this study proposes a data‐driven approach for estimating pipe diameter when considering different spatial patterns and degrees of data completeness (i.e., 0%–90%). Using data from 16 WDNs in Kentucky, this study compares the use of machine learning (ML) using topological and geospatial features against an existing deterministic approach. Results demonstrate that WDN models with pipe diameters predicted by the proposed ML method had comparable hydraulic performance to the ground truth models. Moreover, results showed that ML method performance varies between WDNs of differing topological classification. Insights from this study help advance the ability to leverage partial data to create and maintain WDN models amid uncertainty and inadequate resources.

Poff, Jason W. [Oregon State Univ., Corvallis, OR ↗

Machine Learning Approaches to Predicting Induced Seismicity and Imaging Geothermal Reservoir Properties

This project developed machine learning (ML) methods, lab data sets, and field data to advance geothermal exploration and geothermal energy production. The work had three focus areas. One involved the development of ML methods to use microearthquakes (MEQs) for imaging geothermal reservoir properties and improving subsurface characterization – most importantly the evolution of permeability within the evolving reservoir. This part of the work included development of ML approaches for automated MEQ location, focal mechanism determination and identification of earthquake precursors. The second area focused on using MEQ signals generated by geothermal exploration and production to predict the relationship between fluid injection and seismicity. Here, we extended to reservoir scale our success in using ML to predict laboratory earthquakes and fault zone stress state. The third focus area was on lab experiments. Here, we developed new ML models for lab earthquake prediction and identification of precursors to failure to improve earthquake forecasting and early warning in geothermal settings. Major outcomes of our work include ML models that learn from MEQ signals during geothermal exploration and production to predict induced seismicity. MEQs occur naturally in connection with drilling and energy production. We developed ML methods to use the seismic waves from these events to characterize the elastic, hydraulic and poromechanical properties of reservoirs. Our work illuminated fracture geometry and the evolution of fracture permeability by incorporating seismic coda wave analysis and ML methods to relate fluid injection and seismicity. We significantly expanded laboratory earthquake prediction to include methods that use both passive measurements of microearthquakes within the lab fault zones and also active source acoustic measurements of fault zone elastic properties. These methods can now predict fault zone stress state, time to failure and the magnitude of lab earthquakes. Our work showed that repetitive stick- slip failure events during frictional sliding (the lab equivalent of earthquakes) are preceded by a cascade of micro-failure events that radiate energy in a manner that foretells unstable failure – manifest as laboratory MEQs. We documented a mapping between fracture properties and statistical attributes of elastic radiation. We extended existing works to geothermal reservoir scale and developed ML methods to determine reservoir permeability, fracture properties, and their evolution during geothermal energy production. An attractive feature of ML algorithms is their ability to handle big datasets and reveal patterns and correlations that may remain invisible to conventional analyses. Our work connected data from field, laboratory and intermediate scales to study permeability, stress, strength, fracture stiffness and geometry. At the field scale we used data from the Newberry Volcano field site, UtahFORGE, EGS Collab, and also the Bedretto underground research lab in Switzerland. These data sets are bridging the gap between the lab scale, theory, and reservoir scale. Our work produced plain language summaries to improve public understanding of DOE research. We also developed openly distributed ML and seismicity datasets for use by all researchers and we published connections between induced seismicity in geothermal areas and reservoir properties including permeability, fracture properties, and stress state. Our models are designed for the large data sets of induced seismicity typically associated with geothermal sites. We produced labeled event catalogs and used them on geothermal data to assess how ML can facilitate geothermal production and exploration. All datasets are available on the GDR Productivity: The project produced 32 publications in peer reviewed journals (two are in review). It supported the work of 6 PhD students, 40 conference presentations, 6 keynote talks at national meetings, and mentoring and professional development for 4 postdoctoral fellows.

15 GEOTHERMAL ENERGY↗

AnisONet: A deep neural operator-based anisotropic permeability upscaler from pore to Darcy scale

Directional permeability variations, which govern directional fluid flow in porous media with anisotropy, are important to accurately predict flow behavior, reactive transport, and fluid–solid interactions for various processes such as enhanced geothermal systems, energy storage devices, and biological systems. However, the intricate architecture of porous media makes it difficult to predict directional permeabilities. In this work, we present a novel machine learning (ML) framework, AnisONet, built upon an integration of a convolutional neural network, Swin transformer, and the deep operator network architecture, designed to predict anisotropic permeability and upscale predictions to larger spatial domains. First, AnisONet was evaluated with three classes of two-dimensional (2D) porous media, including synthetic circular and elliptical grains and natural sandstone grains from micro-computed tomography images. A lattice Boltzmann model (LBM) was used to calculate directional permeabilities at every 10° angle, producing 19 data points per image of porous media. AnisONet is then trained to predict permeability as a function of rotation angle. AnisONet showed strong predictive capability of directional permeability. Second, we tested our model for five upscaling cases with a large image size in the finite-element method (FEM) for 2D Darcy flow with various permeability tensor construction methods. Overall, upscaled permeability tensors in FEM simulations produce a reasonably good match with LBM results, highlighting the importance of selecting appropriate tensor formation strategies for accurate permeability upscaling. AnisONet, as a directional permeability estimator, could be further developed for more complex geometries, with the potential to develop a foundational ML model for various applications in porous media.

42 ENGINEERING↗

Leveraging differentiable programming in the inverse problem of neutron stars

Neutron stars (NSs) probe the high-density regime of the nuclear equation of state (EOS). However, inferring the EOS from observations of NSs is a computationally challenging task. Here, in this work, we efficiently solve this inverse problem by leveraging differential programming in two ways. First, we enable full Bayesian inference in under one hour of wall time on a GPU by using gradient-based samplers, without requiring pretrained machine learning emulators. Moreover, we demonstrate efficient scaling to high-dimensional parameter spaces. Second, we introduce a novel gradient-based optimization scheme that recovers the EOS of a given NS mass-radius curve. We demonstrate how our framework can reveal consistencies or tensions between nuclear physics and astrophysics. First, we show how the breakdown density of a metamodel description of the EOS can be determined from NS observations. Second, we demonstrate how degeneracies in EOS modeling using nuclear empirical parameters can influence the inverse problem during gradient-based optimization. Looking ahead, our approach opens up new theoretical studies of the relation between NS properties and the EOS, while effectively tackling the data analysis challenges brought by future detectors.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Updimensioning strategy derived from synthetic equiaxed grain structures for approximating 3D grain size distributions from 2D visualizations with 1D parameters

We generated synthetic equiaxed grain structures using computer graphics software to explore the relationship between various grain size determination methods and true three-dimensional (3D) grain diameters. Mirroring grain measurement techniques, the synthetic 3D grain structures are imaged as 2D micrographs which are measured to yield 1D grain size parameters. Synthetic grain structures provide data at a mass scale and permit exploration of both polished and fractured surface micrographs, revealing one-to-one correspondence between exposed 2D grain cross-sections and individual 3D grains. Analysis of this correspondence yielded a procedure to approximate 3D equiaxed grain size and volume distributions based on the mode of the 2D fractograph grain size distribution. The 3D approximation procedure is shown to be less susceptible to different imaging conditions that affect small, undiscernible grains compared to the standard planimetric and linear intercept methods, which by design also tend to underestimate the 3D grain diameter. The procedure requires larger sample sizes to lower variance and a deeper analysis which could become more practical with machine learning (ML) models for grain boundary segmentation, which synthetic grain structures can help train. This work lays the foundation for analyzing other grain distributions such as columnar and composite grains in similar depth.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Remote Instrumentation and Data Acquisition: An Internship Research Report

This report outlines the development and implementation of a remote data acquisition system for waveform analysis using a Rohde & Schwarz oscilloscope. The project involved capturing waveform data, and transferring it to a local machine for visualization and analysis. The core logic was developed in C++ with a focus on object oriented programming and the use of polymorphism so the main application can interact with any instrument without knowing its exact type, simplifying the overall logic and making it easier to add or swap out components without changing the rest of the codebase.. The system issues Standard Commands for Programmable Instruments (SCPI) via a socket connection and parses the oscilloscope’s ASCII waveform data. The C++ application was containerized using Docker for ease of portability, and reproducibility. Emphasis was placed on secure networking practices, error handling, and effective data capture. The report describes the technical steps taken, challenges encountered, and lessons learned, providing insight into the practical integration of hardware interfacing with remote computational environments.

Parikh, Jaymil [Fermilab]↗

Novel CHI3L1 ‐Associated Angiogenic Phenotypes Define Glioma Microenvironments: Insights From Multi‐Omics Integration

ABSTRACT The CHI3L1 signaling pathway significantly influences glioma angiogenesis, but its role in the tumor microenvironment (TME) remains elusive. We propose a novelCHI3L1‐associated vascular phenotype classification for glioma through integrative analyses of multiple datasets with bulk and single‐cell transcriptome, genomics, digital pathology, and clinical data. We investigated the biological characteristics, genomic alterations, therapeutic vulnerabilities, and immune profiles within these phenotypes through a comprehensive multi‐omics approach. We constructed the vascular‐related risk (VR) score based onCHI3L1‐associated vascular signatures (CAVS) identified by machine learning algorithms. Utilizing unsupervised consensus clustering, gliomas were stratified into three distinct vascular phenotypes: Cluster A, marked by high vascularization and stromal activation with a relatively low levels of tumor‐infiltrating lymphocytes (TILs); Cluster B, characterized by moderate vascularization and stromal activity, coupled with a high density of TILs; and Cluster C, defined by low vascularization and sparse immune cell infiltration. We observed that the CAVS effectively indicated glioma‐associated angiogenesis and immune suppression by single‐cell RNA‐seq analysis. Moreover, the high‐VR‐score group exhibited enhanced angiogenic activity, reduced immune response, resistance to immunotherapy, and poorer clinical outcomes. The VR score independently predicted glioma prognosis and, combined with a nomogram, provided a robust clinical decision‐making tool. Potential drug prediction based on transcription factors for high‐risk patients was also performed. Our study reveals thatCHI3L1‐associated vascular phenotypes shape distinct immune landscapes in gliomas, offering insights for optimizing therapeutic strategies to improve patient outcomes.

Oncology↗

Machine learning of factors for improving oyster hatchery production

Oyster aquaculture and restoration in the Chesapeake Bay are vital, yet hatcheries frequently struggle with inconsistent larval growth and sudden mass mortality events. Unpredictable disruptions in larval production cause large economic losses, represent a perceived risk to growers, and impede industry expansion. To better understand associations between production yield and its potential predictors, we applied machine learning (random forest, and neural network) and statistical (generalized additive model) models to a comprehensive dataset of environmental, water quality, and operational parameters from a Maryland oyster hatchery, aiming to identify key yield predictors and develop a robust forecasting tool. We used recursive Boruta algorithm for variable selection, pinpointing critical predictors, and employed cross-validation to fine-tune model settings. Shapley value analysis offered crucial insights into model interpretations, highlighting week number, Normalized Difference Vegetation Index, salinity, turbidity, and fecundity as primary drivers of yield variability. For low-yield cases, salinity-related variables were particularly important. Our findings provide an early warning system for potential production downturns, empowering hatchery operators to make data-driven decisions for optimizing water conditions, feeding schedules, and broodstock management. By boosting predictability and efficiency, this research directly supports economic stability of the oyster industry and ecological health of the Chesapeake Bay.

Vishwakarma, Srishti [Oak Ridge National Laborator↗

Enhancing Data Quality Monitoring at CMS with Interactive Visualization Tools and Automated Reference Run Selection

Current data quality monitoring (DQM) tools at CMS offer granularity limited to per-run analysis. Consequently, issues manifesting at the per-lumisection level can go unnoticed or, even if detectable, often lead to the classification of the whole run as bad, resulting in unnecessary data loss. Additionally, shifters have to evaluate a large set of monitoring elements during their long shifts, increasing the probability of human errors or overlooked problems. In this contribution, we present ongoing work on the development of tools that will provide shifters with an accessible, granularity-enhanced view of DQM data through interactive and dynamic visualizations. Furthermore, we introduce a reference run selection tool currently under development, which will automate the selection based on data-taking conditions and will offer a curated set of training data for machine learning models that will be used for the partial automation of the offline data certification process. These endeavors will be integrated into the DIALS website, enabling enhancements in data certification accuracy and improving the accessibility of DQM at CMS.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Challenges for monitoring and data analytics in a leadership public data repository

The availability and disposition of data has assumed increasing importance in large-scale computational science. Data repositories are evolving to meet new classes of requirements: compliance with government access guidelines, support for reproducibility of experimental results, and long-term availability of data products. The Constellation public data repository at the Oak Ridge Leadership Computing Facility faces these issues while being situated in one of the most productive data centers in the world. While monitoring and operational data analysis are ingrained in the operation of the OLCF’s large-scale high performance computing platforms, data repositories do not have this history of support. Problems faced by Constellation range from data size (over 7 petabytes in current holdings) to analytic complexity (detailed curation is both absolutely necessary for many data sets and absolutely impossible for humans to accomplish in any practical manner) to deployment environment (OLCF storage resources are oriented toward the needs of the compute platforms). In this paper we describe some of the challenges for collecting monitoring and analytic data from a leadership public data repository. We also discuss various strategies we are pursuing in order to address these challenges, from manual data collection to plans for introducing machine learning-based curatorial techniques.

Widener, Patrick [ORNL] (ORCID:0000000258820816)↗

ICAT: The Interactive Corpus Analysis Tool

The Interactive Corpus Analysis Tool (ICAT) is a Python library for creating dashboards to explore textual datasets and build simple binary classification models to help filter through them and focus on entries of interest. This tool uses a form of interactive machine learning (IML), a paradigm of “machine teaching” (Simard et al., 2017) that sits at the intersection of the fields of human computer interaction (HCI), visual analytics, and machine learning. The intent of ICAT is to allow subject matter experts (SME) with limited to no experience in machine learning to benefit from an iterative human-in-the-loop (HITL) approach to building their own model without needing to understand the details of the underlying algorithm. This interactivity is achieved by allowing the user to create features, label data points, and visually manipulate a representation of the features to manually cluster and investigate data, while a model is trained on the fly based on these actions. ICAT is built on top of the Panel (Holoviz, 2018) library, using a combination of Vega, a custom IPyWidget using D3, and ipyvuetify, and is intended to be used inside of a Jupyter environment.

Martindale, Nathan [Oak Ridge National Laboratory ↗

Super-resolution model for overlapping peak detection and improved spatial resolution in high-energy diffraction microscopy

Reconstruction quality in Far-field High-Energy Diffraction Microscopy (FF-HEDM) is limited by the spatial resolution of area detectors and the frequent occurrence of overlapping diffraction spots. To address these challenges, we developed a super-resolution (SR) framework using convolutional neural networks (CNNs) to recreate 2D diffraction peaks at up to ×8 resolution from raw detector data. A specialized simulation tool was created to generate synthetic training datasets with varying degrees of peak overlap. Integrated into the Microstructural Imaging using Diffraction Analysis Software (MIDAS), the SR model improves the spatial accuracy and precision of 3D grain reconstruction by an order of magnitude. This approach provides a robust solution for investigation of complex micromechanical states and material classes where the analysis is limited by the presence of overlapping peaks. Furthermore, the methodology developed here can potentially be extended to other techniques that require sub-pixel accuracy for high-fidelity data analysis.

High-energy diffraction microscopy↗

MapsTorch : automatic differentiation for X-ray fluorescence data analysis

X-ray fluorescence (XRF) is a popular spectroscopy technique for elemental analysis. Spectrum fitting and parameter tuning are at the core of XRF analysis and are conventionally manually intensive, especially for synchrotron experiments involving large amounts of diverse samples. This work introduces the automatic differentiation (AD) technique to XRF and an open-source package called MapsTorch. By transforming an analytical model of the XRF spectrum into a differentiable computation graph with AD, MapsTorch enables robust optimization of parameters and elemental intensities. We evaluate MapsTorch by conducting computational experiments on a large number of historical synchrotron XRF datasets and compare its performance with the currently practiced fitting tool NLopt. The results show that MapsTorch consistently achieves high-quality fits and often leads to better fitting quality than NLopt, particularly in tasks such as initial spectrum fitting and elemental intensity refinement. The robust performance of MapsTorch paves the way for developing automated and high-throughput XRF data analysis workflows to handle the increasing data volumes expected from next-generation synchrotron facilities.

X-ray fluorescence↗

Electrical Load Forecasting Over Multihop Smart Metering Networks With Federated Learning

Electric load forecasting is essential for power management and stability in smart grids. This is mainly achieved via advanced metering infrastructure, where smart meters (SMs) record household energy data. Traditional machine learning (ML) methods are often employed for load forecasting, but require data sharing, which raises data privacy concerns. Federated learning (FL) can address this issue by running distributed ML models at local SMs without data exchange. However, current FL-based approaches struggle to achieve efficient load forecasting due to imbalanced data distribution across heterogeneous SMs. Here, this article presents a novel personalized FL (PFL) method for high-quality load forecasting in metering networks. A meta-learning-based strategy is developed to address data heterogeneity at local SMs in the collaborative training of local load forecasting models. Moreover, to minimize the load forecasting delays in our PFL model, we study a new latency optimization problem based on optimal resource allocation at SMs. A theoretical convergence analysis is also conducted to provide insights into FL design for federated load forecasting. Extensive simulations from real-world datasets show that our method outperforms existing approaches regarding better load forecasting and reduced operational latency costs.

Rahman, Ratun [Univ. of Alabama, Huntsville, AL (U↗