Search NASA⌕ Search

SEARCH · Search NASA

Results for “Machine Learning for Data Analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29

Self-Supervised and Interpretable Anomaly Detection Using Network Transformers

Machine learning and deep neural networks (DNNs) have been proposed as a tool to identify anomalies in computer network communications. However, due the obfuscated nature of off-the-shelf machine learning models, their output often does not provide enough information to isolate the source of the anomaly to take corrective measures. In this article, we introduce the network transformer (NeT), a DNN model for anomaly detection that incorporates the graph structure of the communication network in order to improve interpretability. Further, the presented approach has the following advantages: first, enhanced interpretability by incorporating the graph structure of computer networks; second, provides a hierarchical set of features that enables analysis at different levels of granularity; second, self-supervised training that does not require labeled data. The NeT model was evaluated on a set of anomalous scenarios executed in a real industrial control system. The presented approach successfully identified the anomalies, the devices affected, and the specific connections causing the anomalies, providing a data-driven hierarchical approach to analyze the behavior of a cyber network.

97 MATHEMATICS AND COMPUTING↗

Scalable edge clustering of dynamic graphs via weighted line graphs

Timestamped relational datasets consisting of records (or connections) between pairs of entities are ubiquitous in network science. For applications like peer-to-peer communication, email, various social network interactions, and computer network security, it is useful to organize these records into groups based on how and when they are occurring. Weighted line graphs offer a natural way to model how records are related in such datasets but for large real-world graph topologies, building and utilizing the line graph is prohibitively expensive. Here, we present the framework to cluster the edges of a dynamic graph via the associated line graph that contains two major contributions. The first is a method to work with the line graph implicitly and the second is a distributed scale implementation of an agglomerative hierarchical graph clustering algorithm. We outline a novel hierarchical dynamic graph edge clustering approach that efficiently breaks massive relational datasets into small sets of edges containing events at various timescales. This is in stark contrast to traditional graph clustering algorithms that prioritize highly connected (clique-like) community structures. Our approach relies on constructing a sufficient subgraph of a weighted line graph and applying a hierarchical agglomerative clustering. This approach is related to scalable techniques from spatial clustering, nonlinear-dimension reduction, topological data analysis, and draws particular inspiration from HDBSCAN. As an edge clustering, this method yields an overlapping node clustering. Our algorithm is parallelizable and we demonstrate efficient clustering of a billion-scale, real-world dynamic graph into small edge sets that correlate in topology and time. The entire clustering process for a graph with tens of billions of edges takes just a few minutes of run time on 256 nodes of a distributed compute environment. We argue how the output of the edge clustering is useful for a multitude of data visualization and powerful machine learning tasks, both involving the original massive dynamic graph data and metadata associated with the nodes and edges. Finally, we describe how this approach can be extended to dynamic hypergraphs and dynamic graphs/hypergraphs with unstructured data living on vertices and edges.

Data Analysis↗

Observational evidence for groundwater influence on crop yields in the United States

As climate change shifts crop exposure to dry and wet extremes, a better understanding of factors governing crop response is needed. Recent studies identified shallow groundwater—groundwater within or near the crop rooting zone—as influential, yet existing evidence is largely based on theoretical crop model simulations, indirect or static groundwater data, or small-scale field studies. Here, we use observational satellite yield data and dynamic water table simulations from 1999 to 2018 to provide field-scale evidence for shallow groundwater effects on maize yields across the United States Corn Belt. We identify three lines of evidence supporting groundwater influence: 1) crop model simulations better match observed yields after improvements in groundwater representation; 2) machine learning analysis of observed yields and modeled groundwater levels reveals a subsidy zone between 1.1 and 2.5 m depths, with yield penalties at shallower depths and no effect at deeper depths; and 3) locations with groundwater typically in the subsidy zone display higher yield stability across time. We estimate an average 3.4% yield increase when groundwater levels are at optimum depth, and this effect roughly doubles in dry conditions. Groundwater yield subsidies occur ~35% of years on average across locations, with 75% of the region benefitting in at least 10% of years. Overall, we estimate that groundwater-yield interactions had a net monetary contribution of approximately $10 billion from 1999 to 2018. This study provides empirical evidence for region-wide groundwater yield impacts and further underlines the need for better quantification of groundwater levels and their dynamic responses to short- and long-term weather conditions.

60 APPLIED LIFE SCIENCES↗

Advanced Analytics and Big Earth Data

NASA's Earth Science Data Systems process, archive and distribute petabytes of Earth Observation data to a variety of end users. These end users will face dramatically increased data size in the near future, bringing about new challenges and opportunities in analyzing those data. One area of particular ferment currently is Machine Learning. Many Machine Learning methods are black boxes, limiting direct insight into the data's properties. However, they can be used for a variety of data enhancement purposes, such as parameter retrieval, data fusion and image classification and segmentation. The Earth Observing System Data and Information System is also evolving to host large data volumes in the cloud, enabling data proximal analysis. As part of this effort, an Analytics framework is being developed to support and enhance user analysis of the data. By using standards based services in the framework, diverse user communities can be served, while also allowing inter-system collaboration in the analysis process.

Cloud Computing↗

Utilizing machine learning to predict tensile ductility and yield strength of CoNiV-based multi-principal elements alloys

This study explores the use of machine learning (ML) as a computational tool to accelerate the design of multi-principal element alloys (MPEAs) with improved tensile elongation. An ML model was trained using available experimental data from the literature along with theoretically derived features to predict yield strength (YS) and ductility. A subset of ML-predicted compositions—CoNiVFe, CoNiVTi, CoNiVTiFe, and CoCrNiVTi—was synthesized and evaluated through tensile testing. The ML model underpredicted YS by approximately 20–30 % and overpredicted ductility by 60–70 % for Ti-containing alloys. Microstructural analysis revealed that Ti segregation at interdendritic regions contributed to early fracture, leading to discrepancies in ductility predictions. Ti segregation at these regions likely drives the increased YS due to segregation strengthening. In contrast, the CoNiVFe alloy showed good agreement with both experimental YS and elongation, with prediction errors of ∼10.2 % and ∼20.7 %, respectively. Microstructural characterization revealed minimal segregation in this alloy, suggesting that the ML model can reliably predict the properties of alloys with little to no segregation. These findings highlight the capability of ML in predicting YS with good accuracy but underscore its limitations in capturing defect-driven failure mechanisms such as segregation-induced embrittlement.

36 MATERIALS SCIENCE↗

Human Factors in Accidents Involving Remotely Piloted Aircraft

This presentation examines human factors that contribute to RPA mishaps and provides analysis of lessons learned. RPA accident data from U.S. military and government agencies were reviewed and analyzed to identify human factors issues. Common contributors to RPA mishaps fell into several major categories: cognitive factors (pilot workload), physiological factors (fatigue and stress), environmental factors (situational awareness), staffing factors (training and crew coordination), and design factors (human machine interface).

Merlin, Peter William↗

SkICAT: A cataloging and analysis tool for wide field imaging surveys

We describe an integrated system, SkICAT (Sky Image Cataloging and Analysis Tool), for the automated reduction and analysis of the Palomar Observatory-ST ScI Digitized Sky Survey. The Survey will consist of the complete digitization of the photographic Second Palomar Observatory Sky Survey (POSS-II) in three bands, comprising nearly three Terabytes of pixel data. SkICAT applies a combination of existing packages, including FOCAS for basic image detection and measurement and SAS for database management, as well as custom software, to the task of managing this wealth of data. One of the most novel aspects of the system is its method of object classification. Using state-of-theart machine learning classification techniques (GID3* and O-BTree), we have developed a powerful method for automatically distinguishing point sources from non-point sources and artifacts, achieving comparably accurate discrimination a full magnitude fainter than in previous Schmidt plate surveys. The learning algorithms produce decision trees for classification by examining instances of objects classified by eye on both plate and higher quality CCD data. The same techniques will be applied to perform higher-level object classification (e.g., of galaxy morphology) in the near future. Another key feature of the system is the facility to integrate the catalogs from multiple plates (and portions thereof) to construct a single catalog of uniform calibration and quality down to the faintest limits of the survey. SkICAT also provides a variety of data analysis and exploration tools for the scientific utilization of the resulting catalogs. We include initial results of applying this system to measure the counts and distribution of galaxies in two bands down to Bj is approximately 21 mag over an approximate 70 square degree multi-plate field from POSS-II. SkICAT is constructed in a modular and general fashion and should be readily adaptable to other large-scale imaging surveys.

Weir, N.↗

Search for a heavy resonance decaying into a Z and a Higgs boson in events with an energetic jet and two electrons, two muons, or missing transverse momentum in proton-proton collisions at $\sqrt{s}$ = 13 TeV

A search is presented for a heavy resonance decaying into a Z boson and a Higgs (H) boson. The analysis is based on data from proton-proton collisions at a centre-of-mass energy of 13 TeV corresponding to an integrated luminosity of 138 fb$^{−1}$, recorded with the CMS experiment in the years 2016–2018. Resonance masses between 1.4 and 5 TeV are considered, resulting in large transverse momenta of the Z and H bosons. Final states that result from Z boson decays to pairs of electrons, muons, or neutrinos are considered. The H boson is reconstructed as a single large-radius jet, recoiling against the Z boson. Machine-learning flavour-tagging techniques are employed to identify decays of a Lorentz-boosted H boson into pairs of charm or bottom quarks, or into four quarks via the intermediate H → WW$^{*}$ and ZZ$^{*}$ decays. The analysis targets H boson decays that were not generally included in previous searches using the H → $ \textrm{b}\overline{\textrm{b}} $ channel. Compared with previous analyses, the sensitivity for high resonance masses is improved significantly in the channel where at most one b quark is tagged.[graphic not available: see fulltext]

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Considerations for Distributed Edge Data Centers and Use of Building Loads to Support Large Interconnections

The rapid expansion of artificial intelligence (AI) and machine learning is driving unprecedented electricity demand from data centers. It is predicted that by 2030, 90% of AI workloads will be inference-based, requiring interconnection of multiple low-latency edge data centers (<20 MW) sited closer to end users - often on already constrained distribution feeders. Although individually small, these loads can aggregate to large loads per feeder, straining infrastructure, creating multi-year interconnection delays, and driving up customer costs. This paper proposes a data center-focused grid-integration framework that combines feeder hosting capacity analysis with building energy efficiency, building load flexibility, and waste heat reuse to expand effective feeder and substation headroom. Such approaches can reduce interconnection delays, lower costs for ratepayers, and accelerate AI-ready infrastructure deployment.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Exploring Anomalous PM 2.5 from Wildfires and Dust Storms using Data and Services at NASA GES DISC

The presence of fine particles in the atmosphere with a diameter of less than 2.5 µm, called particulate matter 2.5 (PM 2.5 ), poses a significant threat to human health as a criteria air pollutant. Fortunately, NASA's Goddard Earth Sciences Data and Information Services Center (GES DISC) provides easy access to several PM 2.5 concentration products. These datasets include the reanalysis of global hourly and monthly aerosol components including PM 2.5 data from the Modern-Era Retrospective analysis for Research and Applications, version 2 (MERRA-2), as well as 3-hourly real-time ensemble forecasts of PM 2.5 from the Hazardous Air Quality Ensemble System (HAQES). The HAQES products are developed by the George Mason University Air Quality Laboratory as part of NASA's Health Air Quality Applied Science Team (HAQAST). The GES DISC is actively collaborating with scientists in the HAQAST program to further expand air quality data collections. Two new datasets are currently being archived: one is the machine learning-based global hourly PM 2.5 derived from MERRA-2; the other is the localized data (NO 2 , O 3 , and PM 2.5 ) time series derived from NASA's GEOS Composition Forecasting (GEOS-CF) system. In this presentation, we will explore the spatial patterns and long-distance transport characteristics of elevated PM 2.5 during extreme pollution events, such as the June 2023 Canadian wildfires, which are still active at the time of writing; and severe spring dust storms in 2023 over Asia. To gain comprehensive insights, we will utilize various PM 2.5 data in conjunction with satellite-observed aerosol data from TROPOspheric Monitoring Instrument (TROPOMI) on Sentinel-5P. The primary focus of this presentation will be to demonstrate effective use of data tools and services to visualize and explore extreme air pollution phenomena. Additionally, we will provide guidance on how users can download specific data of interest, facilitating further analysis and research in this critical area.

air quality↗

Decoding the proton’s gluonic density with lattice QCD-informed machine learning

We present a first machine learning-based decoding of the gluonic structure of the proton from lattice QCD using a variational autoencoder inverse mapper (VAIM). Harnessing the power of generative AI, we predict the parton distribution function (PDF) of the gluon given information on the reduced pseudo-Ioffe-time distributions (RpITDs) as calculated from an ensemble with lattice spacing a ≈ 0.09 fm and a pion mass of M π ≈ 310 MeV. The resulting gluon PDF is consistent with phenomenological global fits within uncertainties, particularly in the intermediate-to-high-x region where lattice data are most constraining. A subsequent correlation analysis confirms that the VAIM learns a meaningful latent representation, highlighting the potential of generative AI to bridge lattice QCD and phenomenological extractions within a unified analysis framework.

Gluon parton distribution function↗

Applications of artificial intelligence 1993: Knowledge-based systems in aerospace and industry; Proceedings of the Meeting, Orlando, FL, Apr. 13-15, 1993

The present volume on applications of artificial intelligence with regard to knowledge-based systems in aerospace and industry discusses machine learning and clustering, expert systems and optimization techniques, monitoring and diagnosis, and automated design and expert systems. Attention is given to the integration of AI reasoning systems and hardware description languages, care-based reasoning, knowledge, retrieval, and training systems, and scheduling and planning. Topics addressed include the preprocessing of remotely sensed data for efficient analysis and classification, autonomous agents as air combat simulation adversaries, intelligent data presentation for real-time spacecraft monitoring, and an integrated reasoner for diagnosis in satellite control. Also discussed are a knowledge-based system for the design of heat exchangers, reuse of design information for model-based diagnosis, automatic compilation of expert systems, and a case-based approach to handling aircraft malfunctions.

Fayyad, Usama M.↗

Collaborative Clustering for Sensor Networks

Traditionally, nodes in a sensor network simply collect data and then pass it on to a centralized node that archives, distributes, and possibly analyzes the data. However, analysis at the individual nodes could enable faster detection of anomalies or other interesting events, as well as faster responses such as sending out alerts or increasing the data collection rate. There is an additional opportunity for increased performance if individual nodes can communicate directly with their neighbors. Previously, a method was developed by which machine learning classification algorithms could collaborate to achieve high performance autonomously (without requiring human intervention). This method worked for supervised learning algorithms, in which labeled data is used to train models. The learners collaborated by exchanging labels describing the data. The new advance enables clustering algorithms, which do not use labeled data, to also collaborate. This is achieved by defining a new language for collaboration that uses pair-wise constraints to encode useful information for other learners. These constraints specify that two items must, or cannot, be placed into the same cluster. Previous work has shown that clustering with these constraints (in isolation) already improves performance. In the problem formulation, each learner resides at a different node in the sensor network and makes observations (collects data) independently of the other learners. Each learner clusters its data and then selects a pair of items about which it is uncertain and uses them to query its neighbors. The resulting feedback (a must and cannot constraint from each neighbor) is combined by the learner into a consensus constraint, and it then reclusters its data while incorporating the new constraint. A strategy was also proposed for cleaning the resulting constraint sets, which may contain conflicting constraints; this improves performance significantly. This approach has been applied to collaborative clustering of seismic and infrasonic data collected by the Mount Erebus Volcano Observatory in Antarctica. Previous approaches to distributed clustering cannot readily be applied in a sensor network setting, because they assume that each node has the same view of the data set. A view is the set of features used to represent each object. When a single data set is partitioned across several computational nodes, distributed clustering works; all objects have the same view. But when the data is collected from different locations, using different sensors, a more flexible approach is needed. This approach instead operates in situations where the data collected at each node has a different view (e.g., seismic vs. infrasonic sensors), but they observe the same events. This enables them to exchange information about the likely cluster membership relations between objects, even if they do not use the same features to represent the objects.

Wagstaff. Loro :/↗

Machine learning models of intermittent operation of RO wellhead water treatment for salinity reduction and nitrate removal

Machine learning models were developed for intermittent multi-mode operation of a wellhead reverse osmosis water purification and desalination system to predict salt passage, nitrate passage, and permeate flux. The models, based on long short-term memory (LSTM) recurrent neural network (RNN) architecture, included an attention mechanism to increase model performance in proximity of the regulatory limit for nitrate. Training and testing of the models for the Startup, Production, Shutdown and Flushing operational modes were based on operational data (consisting of 22 process variables per data sample) acquired every 2–5 s over a six-month period. The significant sets of model input attributes for the different operational modes were assessed via Spearman ranking correlation, Self-Organizing Map (SOM) analysis and feed forward feature selection (FFFS). Although the variability of nitrate passage, salt passage and permeate flux was significant over the four operational modes, prediction performance for the three outcomes were with R2 and Average Absolute Relative Error (AARE) of 0.78–0.95 and 2.96–6.16 %, respectively. Model updates post membrane elements replacement demonstrated similar levels of prediction accuracy. The study results suggest that there is merit in exploring the utility of multi-mode models for sensor fault detection, data imputation, and for potential use in model-predictive control.

Intermittent RO operation↗

Transforming Agricultural Productivity with AI-Driven Forecasting: Innovations in Food Security and Supply Chain Optimization

Global food security is under significant threat from climate change, population growth, and resource scarcity. This review examines how advanced AI-driven forecasting models, including machine learning (ML), deep learning (DL), and time-series forecasting models like SARIMA/ARIMA, are transforming regional agricultural practices and food supply chains. Through the integration of Internet of Things (IoT), remote sensing, and blockchain technologies, these models facilitate the real-time monitoring of crop growth, resource allocation, and market dynamics, enhancing decision making and sustainability. The study adopts a mixed-methods approach, including systematic literature analysis and regional case studies. Highlights include AI-driven yield forecasting in European hydroponic systems and resource optimization in southeast Asian aquaponics, showcasing localized efficiency gains. Furthermore, AI applications in food processing, such as plasma, ozone and Pulsed Electric Field (PEF) treatments, are shown to improve food preservation and reduce spoilage. Key challenges—such as data quality, model scalability, and prediction accuracy—are discussed, particularly in the context of data-poor environments, limiting broader model applicability. The paper concludes by outlining future directions, emphasizing context-specific AI implementations, the need for public–private collaboration, and policy interventions to enhance scalability and adoption in food security contexts.

99 GENERAL AND MISCELLANEOUS↗

In-Silico Analysis of High Refractive Index Materials Through Principles of Materials Design

The intent of the paper is to use specific principles of Materials Design that were developed and applied in the electronics industry for enabling understanding and design of improved high refractive index materials. Further, by combining first-principle based ab-initio, semiempirical interatomic potential methods, and machine learning approaches in conjunction with experimental data, we identified specific determinants of high refractive index materials, which can be critically applied for informing materials design and accelerating discovery. Specifically, it was demonstrated that chalcogenides and perovskites as bulk materials can exhibit higher refractive indices with appropriate engineering of specific aspects of the materials.

36 MATERIALS SCIENCE↗

Advanced Data Science Model for Detecting Intelligent Malware

This study focused on developing a robust artificial intelligence (AI) model capable of detecting and characterizing advanced malware in Internet of Things (IoT) devices using network data. By analyzing network traffic with various machine learning (ML) models, our AI model can identify and characterize malicious activities to significantly improve malware detection accuracy and reliability as compared to traditional methods. The developed AI/ML model was trained using network data from IoT devices, leveraging classifiers such as Random Forest, Gradient Boosting, AdaBoost, and others to optimize detection performance. This project demonstrates a scalable framework for real-time malware detection and characterization in IoT networks, capable of identifying infected devices and facilitating the necessary steps to remove or isolate them, thereby preventing further infections. Although digital twin (DT) integration is not yet implemented in the current model, it represents a promising future enhancement. By creating a virtual replica of physical IoT devices, DT technology would allow for real-time monitoring and analysis without directly accessing operational technology, thus reducing the risk of compromising or reducing the performance of actual devices. This integration would further enhance the security of IoT ecosystems, combining AI technology to better flag and detect indications of malware-infected devices within a nuclear system environment.

42 ENGINEERING↗

Side-by-Side Comparison of Subhourly Clipping Models

Over the past several years there have been numerous attempts at quantifying the inherent power clipping of inverters due to subhourly irradiance variability that is not captured in hourly PV performance models. Different models have been proposed to correct for these clipping losses in PV performance estimates, including matrix lookup models, distribution modeling of the PV power performance within a given hour, and machine learning methods. To date, there have been few comprehensive quantitative comparisons of these inverter clipping correction modeling approaches to evaluate the effectiveness of these approaches in predicting the actual behavior of PV system inverter clipping. In this study, we perform such a comparison, evaluating the Allen and Walker correction loss modeling approaches recently implemented in the System Advisor Model (SAM) against clipping losses modeled with 1-minute climate data. These comparisons were performed across a variety of climate locations and inverter loading ratios to thoroughly analyze the effectiveness of these modeling approaches relative to each other. Results from this analysis reveal that both clipping correction approaches improve annual energy accuracy to within 2% of 1-minute modeled energy yield. The two models predict annual clipping loss more accurately than simple hourly power limit clipping, with the Allen method typically being slightly more accurate at typical ILR values and the Walker method often being slightly more accurate at high ILR values The models can improve accuracy over the status quo clipping approach up to 3 percentage points in systems with ILR of 2.0, showing the importance of this modeling factor in energy yield estimates.

accuracy↗