Search NASA⌕ Search

SEARCH · Search NASA

Results for “Machine Learning for Data Analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20

Predictive analytics of selections of russet potatoes

We explore the application of machine learning algorithms specifically to enhance the selection process of Russet potato (Solanum tuberosum L.) clones in breeding trials by predicting their suitability for advancement. This study addresses the challenge of efficiently identifying high-yield, disease-resistant, and climate-resilient potato varieties that meet processing industry standards. Leveraging manually collected data from trials in the state of Oregon, we investigate the potential of a wide variety of state-of-the-art binary classification models. The dataset includes 1086 clones, with data on 38 attributes recorded for each clone, focusing on yield, size, appearance, and frying characteristics, with several control varieties planted consistently across four Oregon regions from 2013 to 2021. We conduct a comprehensive analysis of the dataset that includes preprocessing, feature engineering, and imputation to address missing values. We focus on several key metrics such as accuracy, F1-score, and Matthews correlation coefficient (MCC) for model evaluation. The top-performing models, namely a feedforward neural network classifier (Neural Net), a histogram-based gradient boosting classifier (HGBC), and a support vector machine classifier (SVM), demonstrate consistent and significant results. To further validate our findings, we conducted a simulation study using the aims, data-generating mechanisms, estimands, methods, and performance measures (ADEMP) framework, simulating different data-generating scenarios to assess model robustness and performance through true positive, true negative, false positive, and false negative distributions, area under the receiver operating characteristic curve (AUC-ROC) and MCC. The simulation results highlight that non-linear models like SVM and HGBC consistently show higher AUC-ROC and MCC than logistic regression, thus outperforming the traditional linear model across various distributions, and emphasizing the importance of model selection and tuning in agricultural trials. Variable selection further enhances model performance and identifies influential features in predicting trial outcomes. The findings emphasize the potential of machine learning in streamlining the selection process for potato varieties, offering benefits such as increased efficiency, substantial cost savings, and judicious resource utilization. Our study contributes insights into precision agriculture and showcases the relevance of advanced technologies for informed decision-making in breeding programs.

60 APPLIED LIFE SCIENCES↗

SRF Cavity Instability Detection with Machine Learning at CEBAF

During the operation of the Continuous Electron Beam Accelerator Facility (CEBAF), one or more unstable superconducting radio-frequency (SRF) cavities often cause beam loss trips while the unstable cavities themselves do not necessarily trip off. The present RF controls for the legacy cavities report at only 1 Hz, which is too slow to detect fast transient instabilities during these trip events. These challenges make the identification of an unstable cavity out of the hundreds installed at CEBAF a difficult and time-consuming task. To tackle these issues, a fast data acquisition system (DAQ) for the legacy SRF cavities has been developed, which records the sample at 5 kHz. A Principal Component Analysis (PCA) approach is being developed to identify anomalous SRF cavity behavior. We will discuss the present status of the DAQ system and PCA model, along with initial performance metrics. Overall, our method offers a practical solution for identifying unstable SRF cavities, contributing to increased beam availability and machine reliability.

Carpenter, A.↗

SRF Cavity Instability Detection with Machine Learning at CEBAF

During the operation of the Continuous Electron Beam Accelerator Facility (CEBAF), one or more unstable superconducting radio-frequency (SRF) cavities often cause beam loss trips while the unstable cavities themselves do not necessarily trip off. The present RF controls for the legacy cavities report at only 1 Hz, which is too slow to detect fast transient instabilities during these trip events. These challenges make the identification of an unstable cavity out of the hundreds installed at CEBAF a difficult and time-consuming task. To tackle these issues, a fast data acquisition system (DAQ) for the legacy SRF cavities has been developed, which records the sample at 5 kHz. A Principal Component Analysis (PCA) approach is being developed to identify anomalous SRF cavity behavior. We will discuss the present status of the DAQ system and PCA model, along with initial performance metrics. Overall, our method offers a practical solution for identifying unstable SRF cavities, contributing to increased beam availability and machine reliability.

Carpenter, A.↗

Machine learning inversion of interatomic force constants from single-crystal inelastic neutron scattering

Atomic vibrations govern many macroscopic properties of materials, but experiments to comprehensively probe them remain challenging. Inelastic neutron scattering (INS) is a powerful technique to map phonon dispersions in crystals, especially when leveraging modern time-of-flight (ToF) spectrometers with large detectors. However, efficiently and robustly extracting interatomic force constants (FCs) parameterizing phonon dynamics from experimental spectra remains a bottleneck due to the complexity and high dimensionality of ToF INS datasets. Here, we present a machine learning approach for the direct inversion of FCs from single-crystal INS measurements. The framework leverages synthetic training data generated using universal machine-learned force fields and an efficient physics-based forward model. We benchmark two neural architectures–one emphasizing structured latent representation learning and the other direct, supervised spectral regression–across simulated datasets for two materials under idealized and noisy conditions. The latent-representation model is subsequently applied to experimental single-crystal INS data on germanium. The model is shown to reproduce FCs derived from both first-principles simulations and from iterative optimization, and furthermore achieves reliable inference even from sparse, single-orientation measurements representing short data acquisitions. Analysis of the learned latent space reveals semantically continuous and physically interpretable encodings that support strong cross-domain generalization. By bridging theoretical and experimental domains, we establish a path toward rapid inversion of experimental spectra and data-driven interpretation of temperature-dependent lattice dynamics.

42 ENGINEERING↗

Self-Supervised and Interpretable Anomaly Detection Using Network Transformers

Machine learning and deep neural networks (DNNs) have been proposed as a tool to identify anomalies in computer network communications. However, due the obfuscated nature of off-the-shelf machine learning models, their output often does not provide enough information to isolate the source of the anomaly to take corrective measures. In this article, we introduce the network transformer (NeT), a DNN model for anomaly detection that incorporates the graph structure of the communication network in order to improve interpretability. Further, the presented approach has the following advantages: first, enhanced interpretability by incorporating the graph structure of computer networks; second, provides a hierarchical set of features that enables analysis at different levels of granularity; second, self-supervised training that does not require labeled data. The NeT model was evaluated on a set of anomalous scenarios executed in a real industrial control system. The presented approach successfully identified the anomalies, the devices affected, and the specific connections causing the anomalies, providing a data-driven hierarchical approach to analyze the behavior of a cyber network.

97 MATHEMATICS AND COMPUTING↗

Scalable edge clustering of dynamic graphs via weighted line graphs

Timestamped relational datasets consisting of records (or connections) between pairs of entities are ubiquitous in network science. For applications like peer-to-peer communication, email, various social network interactions, and computer network security, it is useful to organize these records into groups based on how and when they are occurring. Weighted line graphs offer a natural way to model how records are related in such datasets but for large real-world graph topologies, building and utilizing the line graph is prohibitively expensive. Here, we present the framework to cluster the edges of a dynamic graph via the associated line graph that contains two major contributions. The first is a method to work with the line graph implicitly and the second is a distributed scale implementation of an agglomerative hierarchical graph clustering algorithm. We outline a novel hierarchical dynamic graph edge clustering approach that efficiently breaks massive relational datasets into small sets of edges containing events at various timescales. This is in stark contrast to traditional graph clustering algorithms that prioritize highly connected (clique-like) community structures. Our approach relies on constructing a sufficient subgraph of a weighted line graph and applying a hierarchical agglomerative clustering. This approach is related to scalable techniques from spatial clustering, nonlinear-dimension reduction, topological data analysis, and draws particular inspiration from HDBSCAN. As an edge clustering, this method yields an overlapping node clustering. Our algorithm is parallelizable and we demonstrate efficient clustering of a billion-scale, real-world dynamic graph into small edge sets that correlate in topology and time. The entire clustering process for a graph with tens of billions of edges takes just a few minutes of run time on 256 nodes of a distributed compute environment. We argue how the output of the edge clustering is useful for a multitude of data visualization and powerful machine learning tasks, both involving the original massive dynamic graph data and metadata associated with the nodes and edges. Finally, we describe how this approach can be extended to dynamic hypergraphs and dynamic graphs/hypergraphs with unstructured data living on vertices and edges.

Data Analysis↗

Observational evidence for groundwater influence on crop yields in the United States

As climate change shifts crop exposure to dry and wet extremes, a better understanding of factors governing crop response is needed. Recent studies identified shallow groundwater—groundwater within or near the crop rooting zone—as influential, yet existing evidence is largely based on theoretical crop model simulations, indirect or static groundwater data, or small-scale field studies. Here, we use observational satellite yield data and dynamic water table simulations from 1999 to 2018 to provide field-scale evidence for shallow groundwater effects on maize yields across the United States Corn Belt. We identify three lines of evidence supporting groundwater influence: 1) crop model simulations better match observed yields after improvements in groundwater representation; 2) machine learning analysis of observed yields and modeled groundwater levels reveals a subsidy zone between 1.1 and 2.5 m depths, with yield penalties at shallower depths and no effect at deeper depths; and 3) locations with groundwater typically in the subsidy zone display higher yield stability across time. We estimate an average 3.4% yield increase when groundwater levels are at optimum depth, and this effect roughly doubles in dry conditions. Groundwater yield subsidies occur ~35% of years on average across locations, with 75% of the region benefitting in at least 10% of years. Overall, we estimate that groundwater-yield interactions had a net monetary contribution of approximately $10 billion from 1999 to 2018. This study provides empirical evidence for region-wide groundwater yield impacts and further underlines the need for better quantification of groundwater levels and their dynamic responses to short- and long-term weather conditions.

60 APPLIED LIFE SCIENCES↗

Utilizing machine learning to predict tensile ductility and yield strength of CoNiV-based multi-principal elements alloys

This study explores the use of machine learning (ML) as a computational tool to accelerate the design of multi-principal element alloys (MPEAs) with improved tensile elongation. An ML model was trained using available experimental data from the literature along with theoretically derived features to predict yield strength (YS) and ductility. A subset of ML-predicted compositions—CoNiVFe, CoNiVTi, CoNiVTiFe, and CoCrNiVTi—was synthesized and evaluated through tensile testing. The ML model underpredicted YS by approximately 20–30 % and overpredicted ductility by 60–70 % for Ti-containing alloys. Microstructural analysis revealed that Ti segregation at interdendritic regions contributed to early fracture, leading to discrepancies in ductility predictions. Ti segregation at these regions likely drives the increased YS due to segregation strengthening. In contrast, the CoNiVFe alloy showed good agreement with both experimental YS and elongation, with prediction errors of ∼10.2 % and ∼20.7 %, respectively. Microstructural characterization revealed minimal segregation in this alloy, suggesting that the ML model can reliably predict the properties of alloys with little to no segregation. These findings highlight the capability of ML in predicting YS with good accuracy but underscore its limitations in capturing defect-driven failure mechanisms such as segregation-induced embrittlement.

36 MATERIALS SCIENCE↗

Search for a heavy resonance decaying into a Z and a Higgs boson in events with an energetic jet and two electrons, two muons, or missing transverse momentum in proton-proton collisions at $\sqrt{s}$ = 13 TeV

A search is presented for a heavy resonance decaying into a Z boson and a Higgs (H) boson. The analysis is based on data from proton-proton collisions at a centre-of-mass energy of 13 TeV corresponding to an integrated luminosity of 138 fb$^{−1}$, recorded with the CMS experiment in the years 2016–2018. Resonance masses between 1.4 and 5 TeV are considered, resulting in large transverse momenta of the Z and H bosons. Final states that result from Z boson decays to pairs of electrons, muons, or neutrinos are considered. The H boson is reconstructed as a single large-radius jet, recoiling against the Z boson. Machine-learning flavour-tagging techniques are employed to identify decays of a Lorentz-boosted H boson into pairs of charm or bottom quarks, or into four quarks via the intermediate H → WW$^{*}$ and ZZ$^{*}$ decays. The analysis targets H boson decays that were not generally included in previous searches using the H → $ \textrm{b}\overline{\textrm{b}} $ channel. Compared with previous analyses, the sensitivity for high resonance masses is improved significantly in the channel where at most one b quark is tagged.[graphic not available: see fulltext]

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Considerations for Distributed Edge Data Centers and Use of Building Loads to Support Large Interconnections

The rapid expansion of artificial intelligence (AI) and machine learning is driving unprecedented electricity demand from data centers. It is predicted that by 2030, 90% of AI workloads will be inference-based, requiring interconnection of multiple low-latency edge data centers (<20 MW) sited closer to end users - often on already constrained distribution feeders. Although individually small, these loads can aggregate to large loads per feeder, straining infrastructure, creating multi-year interconnection delays, and driving up customer costs. This paper proposes a data center-focused grid-integration framework that combines feeder hosting capacity analysis with building energy efficiency, building load flexibility, and waste heat reuse to expand effective feeder and substation headroom. Such approaches can reduce interconnection delays, lower costs for ratepayers, and accelerate AI-ready infrastructure deployment.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Decoding the proton’s gluonic density with lattice QCD-informed machine learning

We present a first machine learning-based decoding of the gluonic structure of the proton from lattice QCD using a variational autoencoder inverse mapper (VAIM). Harnessing the power of generative AI, we predict the parton distribution function (PDF) of the gluon given information on the reduced pseudo-Ioffe-time distributions (RpITDs) as calculated from an ensemble with lattice spacing a ≈ 0.09 fm and a pion mass of M π ≈ 310 MeV. The resulting gluon PDF is consistent with phenomenological global fits within uncertainties, particularly in the intermediate-to-high-x region where lattice data are most constraining. A subsequent correlation analysis confirms that the VAIM learns a meaningful latent representation, highlighting the potential of generative AI to bridge lattice QCD and phenomenological extractions within a unified analysis framework.

Gluon parton distribution function↗

Machine learning models of intermittent operation of RO wellhead water treatment for salinity reduction and nitrate removal

Machine learning models were developed for intermittent multi-mode operation of a wellhead reverse osmosis water purification and desalination system to predict salt passage, nitrate passage, and permeate flux. The models, based on long short-term memory (LSTM) recurrent neural network (RNN) architecture, included an attention mechanism to increase model performance in proximity of the regulatory limit for nitrate. Training and testing of the models for the Startup, Production, Shutdown and Flushing operational modes were based on operational data (consisting of 22 process variables per data sample) acquired every 2–5 s over a six-month period. The significant sets of model input attributes for the different operational modes were assessed via Spearman ranking correlation, Self-Organizing Map (SOM) analysis and feed forward feature selection (FFFS). Although the variability of nitrate passage, salt passage and permeate flux was significant over the four operational modes, prediction performance for the three outcomes were with R2 and Average Absolute Relative Error (AARE) of 0.78–0.95 and 2.96–6.16 %, respectively. Model updates post membrane elements replacement demonstrated similar levels of prediction accuracy. The study results suggest that there is merit in exploring the utility of multi-mode models for sensor fault detection, data imputation, and for potential use in model-predictive control.

Intermittent RO operation↗

Transforming Agricultural Productivity with AI-Driven Forecasting: Innovations in Food Security and Supply Chain Optimization

Global food security is under significant threat from climate change, population growth, and resource scarcity. This review examines how advanced AI-driven forecasting models, including machine learning (ML), deep learning (DL), and time-series forecasting models like SARIMA/ARIMA, are transforming regional agricultural practices and food supply chains. Through the integration of Internet of Things (IoT), remote sensing, and blockchain technologies, these models facilitate the real-time monitoring of crop growth, resource allocation, and market dynamics, enhancing decision making and sustainability. The study adopts a mixed-methods approach, including systematic literature analysis and regional case studies. Highlights include AI-driven yield forecasting in European hydroponic systems and resource optimization in southeast Asian aquaponics, showcasing localized efficiency gains. Furthermore, AI applications in food processing, such as plasma, ozone and Pulsed Electric Field (PEF) treatments, are shown to improve food preservation and reduce spoilage. Key challenges—such as data quality, model scalability, and prediction accuracy—are discussed, particularly in the context of data-poor environments, limiting broader model applicability. The paper concludes by outlining future directions, emphasizing context-specific AI implementations, the need for public–private collaboration, and policy interventions to enhance scalability and adoption in food security contexts.

99 GENERAL AND MISCELLANEOUS↗

In-Silico Analysis of High Refractive Index Materials Through Principles of Materials Design

The intent of the paper is to use specific principles of Materials Design that were developed and applied in the electronics industry for enabling understanding and design of improved high refractive index materials. Further, by combining first-principle based ab-initio, semiempirical interatomic potential methods, and machine learning approaches in conjunction with experimental data, we identified specific determinants of high refractive index materials, which can be critically applied for informing materials design and accelerating discovery. Specifically, it was demonstrated that chalcogenides and perovskites as bulk materials can exhibit higher refractive indices with appropriate engineering of specific aspects of the materials.

36 MATERIALS SCIENCE↗

Advanced Data Science Model for Detecting Intelligent Malware

This study focused on developing a robust artificial intelligence (AI) model capable of detecting and characterizing advanced malware in Internet of Things (IoT) devices using network data. By analyzing network traffic with various machine learning (ML) models, our AI model can identify and characterize malicious activities to significantly improve malware detection accuracy and reliability as compared to traditional methods. The developed AI/ML model was trained using network data from IoT devices, leveraging classifiers such as Random Forest, Gradient Boosting, AdaBoost, and others to optimize detection performance. This project demonstrates a scalable framework for real-time malware detection and characterization in IoT networks, capable of identifying infected devices and facilitating the necessary steps to remove or isolate them, thereby preventing further infections. Although digital twin (DT) integration is not yet implemented in the current model, it represents a promising future enhancement. By creating a virtual replica of physical IoT devices, DT technology would allow for real-time monitoring and analysis without directly accessing operational technology, thus reducing the risk of compromising or reducing the performance of actual devices. This integration would further enhance the security of IoT ecosystems, combining AI technology to better flag and detect indications of malware-infected devices within a nuclear system environment.

42 ENGINEERING↗

Side-by-Side Comparison of Subhourly Clipping Models

Over the past several years there have been numerous attempts at quantifying the inherent power clipping of inverters due to subhourly irradiance variability that is not captured in hourly PV performance models. Different models have been proposed to correct for these clipping losses in PV performance estimates, including matrix lookup models, distribution modeling of the PV power performance within a given hour, and machine learning methods. To date, there have been few comprehensive quantitative comparisons of these inverter clipping correction modeling approaches to evaluate the effectiveness of these approaches in predicting the actual behavior of PV system inverter clipping. In this study, we perform such a comparison, evaluating the Allen and Walker correction loss modeling approaches recently implemented in the System Advisor Model (SAM) against clipping losses modeled with 1-minute climate data. These comparisons were performed across a variety of climate locations and inverter loading ratios to thoroughly analyze the effectiveness of these modeling approaches relative to each other. Results from this analysis reveal that both clipping correction approaches improve annual energy accuracy to within 2% of 1-minute modeled energy yield. The two models predict annual clipping loss more accurately than simple hourly power limit clipping, with the Allen method typically being slightly more accurate at typical ILR values and the Walker method often being slightly more accurate at high ILR values The models can improve accuracy over the status quo clipping approach up to 3 percentage points in systems with ILR of 2.0, showing the importance of this modeling factor in energy yield estimates.

accuracy↗

“Understanding Robustness Lottery”: A Geometric Visual Comparative Analysis of Neural Network Pruning Approaches

Deep learning approaches have provided state-of-the-art performance in many applications by relying on large and overparameterized neural networks. However, such networks are very brittle and are difficult to deploy on resource-limited platforms. Model pruning, i.e., reducing the size of the network, is a widely adopted strategy that can lead to a more robust and compact model. Many heuristics exist for model pruning, but our understanding of the pruning process remains limited due to the black-box nature of a neural network model. Empirical studies show that some heuristics improve performance whereas others can make models more brittle. Here, this work aims to shed light on how different pruning methods alter the network’s internal feature representation and the corresponding impact on model performance. To facilitate a comprehensive comparison and characterization of the high-dimensional model feature space, we introduce a visual geometric analysis of feature representations. We evaluated a set of critical geometric concepts decomposed from the commonly adopted classification loss and used them to design a visualization system to compare and highlight the impact of pruning on model performance and feature representation. The proposed tool provides an environment for an in-depth comparison of pruning methods and a comprehensive understanding of how the model responds to common data corruption. By leveraging the proposed visualization, machine learning researchers can reveal the similarities between pruning methods and redundancy in robustness evaluation benchmarks, obtain geometric insights about the differences between pruned models that achieve superior robustness performance, and identify samples that are robust or fragile to model pruning and common data corruption.

Li, Zhimin [Univ. of Utah, Salt Lake City, UT (Uni↗

Scattering-based structural inversion of soft materials via Kolmogorov–Arnold networks

Small-angle scattering techniques are indispensable tools for probing the structure of soft materials. However, traditional analytical models often face limitations in structural inversion for complex systems, primarily due to the absence of closed-form expressions of scattering functions. To address these challenges, we present a machine learning framework based on the Kolmogorov–Arnold Network (KAN) for directly extracting real-space structural information from scattering spectra in reciprocal space. This model-independent, data-driven approach provides a versatile solution for analyzing intricate configurations in soft matter. By applying the KAN to lyotropic lamellar phases and colloidal suspensions—two representative soft matter systems—we demonstrate its ability to accurately and efficiently resolve structural collectivity and complexity. Here, our findings highlight the transformative potential of machine learning in enhancing the quantitative analysis of soft materials, paving the way for robust structural inversion across diverse systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗