Search NASA⌕ Search

SEARCH · Search NASA

Results for “machine data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29

Monitoring and Modeling Performance of Communications in Computational Grids

Computational grids may include many machines located in a number of sites. For efficient use of the grid we need to have an ability to estimate the time it takes to communicate data between the machines. For dynamic distributed grids it is unrealistic to know exact parameters of the communication hardware and the current communication traffic and we should rely on a model of the network performance to estimate the message delivery time. Our approach to a construction of such a model is based on observation of the messages delivery time with various message sizes and time scales. We record these observations in a database and use them to build a model of the message delivery time. Our experiments show presence of multiple bands in the logarithm of the message delivery times. These multiple bands represent multiple paths messages travel between the grid machines and are incorporated in our multiband model.

Frumkin, Michael A.↗

Inductive System Health Monitoring

The Inductive Monitoring System (IMS) software was developed to provide a technique to automatically produce health monitoring knowledge bases for systems that are either difficult to model (simulate) with a computer or which require computer models that are too complex to use for real time monitoring. IMS uses nominal data sets collected either directly from the system or from simulations to build a knowledge base that can be used to detect anomalous behavior in the system. Machine learning and data mining techniques are used to characterize typical system behavior by extracting general classes of nominal data from archived data sets. IMS is able to monitor the system by comparing real time operational data with these classes. We present a description of learning and monitoring method used by IMS and summarize some recent IMS results.

Iverson, David L.↗

INDUCTIVE SYSTEM HEALTH MONITORING WITH STATISTICAL METRICS

Model-based reasoning is a powerful method for performing system monitoring and diagnosis. Building models for model-based reasoning is often a difficult and time consuming process. The Inductive Monitoring System (IMS) software was developed to provide a technique to automatically produce health monitoring knowledge bases for systems that are either difficult to model (simulate) with a computer or which require computer models that are too complex to use for real time monitoring. IMS processes nominal data sets collected either directly from the system or from simulations to build a knowledge base that can be used to detect anomalous behavior in the system. Machine learning and data mining techniques are used to characterize typical system behavior by extracting general classes of nominal data from archived data sets. In particular, a clustering algorithm forms groups of nominal values for sets of related parameters. This establishes constraints on those parameter values that should hold during nominal operation. During monitoring, IMS provides a statistically weighted measure of the deviation of current system behavior from the established normal baseline. If the deviation increases beyond the expected level, an anomaly is suspected, prompting further investigation by an operator or automated system. IMS has shown potential to be an effective, low cost technique to produce system monitoring capability for a variety of applications. We describe the training and system health monitoring techniques of IMS. We also present the application of IMS to a data set from the Space Shuttle Columbia STS-107 flight. IMS was able to detect an anomaly in the launch telemetry shortly after a foam impact damaged Columbia's thermal protection system.

Iverson, David L.↗

Architectures Toward Reusable Science Data Systems

Science Data Systems (SDS) comprise an important class of data processing systems that support product generation from remote sensors and in-situ observations. These systems enable research into new science data products, replication of experiments and verification of results. NASA has been building systems for satellite data processing since the first Earth observing satellites launched and is continuing development of systems to support NASA science research and NOAAs Earth observing satellite operations. The basic data processing workflows and scenarios continue to be valid for remote sensor observations research as well as for the complex multi-instrument operational satellite data systems being built today. System functions such as ingest, product generation and distribution need to be configured and performed in a consistent and repeatable way with an emphasis on scalability. This paper will examine the key architectural elements of several NASA satellite data processing systems currently in operation and under development that make them suitable for scaling and reuse. Examples of architectural elements that have become attractive include virtual machine environments, standard data product formats, metadata content and file naming, workflow and job management frameworks, data acquisition, search, and distribution protocols. By highlighting key elements and implementation experience we expect to find architectures that will outlast their original application and be readily adaptable for new applications. Concepts and principles are explored that lead to sound guidance for SDS developers and strategists.

Data Processing↗

Global 3D Data Visualization and Analysis Platform with Advanced Machine Learning Capabilities in Support of Lunar Exploration

Introduction: The science goals for NASA’s Artemis program include: a) Understanding the character and origin of lunar polar volatiles, b) Conducting experimental science in the lunar environment and c) Investigating and mitigating exploration risks. The permanently shadowed regions (PSRs) on the Lunar south pole are expected to host large quantities of water-ice and volatiles that are important for sustainable Lunar exploration. There are several missions such as onboard Korea Pathfinder Lunar Orbiter (KPLO: Korean name Danuri) with onboard ShadowCam camera, Astrobotic Peregrine Mission One [4], and other efforts underway to obtain high resolution topographic, minerals, volatiles and other information on the moon. We envision a need in immediate future for platforms to integrate these data sets, provide rendering and visualization capabilities in the context of a 3D Lunar globe for easier information access and analysis. NASA's Celestial Mapping System (CMS) is developed to address the need for 3D tools for planetary science investigations, mission planning, in-situ operations, in a 3D-first design constructed around a unified view of a planetary globe. At present CMS provides many critical functionalities that include: 1) equipment planning and optimized placement on Lunar surface 2) line-of-sight (LOS) analysis 3) powerful measurement tools based on 3D terrain with realistic 3D models to represent rovers, astronauts and equipment 4) visualization of derived mapping products (e.g. resource maps), and 5) a data engine for hosting new observations that are not available in other contemporary lunar data tools. Planetary Data Ingestion: CMS can consume and analyze data from locally hosted and external third party sources. It is compatible with Open Geospatial Consortium (OGC) data and file standards and currently integrates datasets from the Astrogeology Science Center of USGS. This includes global and local data acquired from NASA (LRO, Clementine, Lunar Orbiter) and JAXA (SELENE/Kaguya), with capability of integrating more datasets. In addition, users can specify other WMS-hosted data endpoints, which CMS can then query and stream data from automatically. To set-up an automated process for ingestion and accurate rendering, visualization and analysis of external 3rd party planetary datasets within CMS, we initiated the process of ingesting unique dataset of super-enhanced images of the permanently shadowed regions (PSRs) at the lunar poles which were produced by the Hyper-effective nOise Removal U-net Software (HORUS) tool. This tool was developed to enhance the extremely low-light images of the interior of PSRs and provide the ability to see within these regions at and discern surface features (i.e. boulders and craters) down to 3 meters in size. We focused on the Nobile region on the Lunar south pole, selected site for VIPER mission and stitched several images to create a high-resolution map within one of the PSR of Nobile crater. Figure 1 shows the dark PSR zone form the original NAC layer of LRO as the base layer (left image) and the illuminated areas within that crater (center) which was created by ingesting and merging several of HORUS generated images. At present we employ a semi-automated process to ensure spatial accuracy and merger of several overlapping zones. However, we are in the process of completely automating this process by employing AI based techniques that would rank, sort, and stack the images based on their information density. The georectification of the images would employ selected features. Analysis on Ingested Planetary Datasets: Once an external planetary data-set is successfully ingested, georectified and merged seamlessly as a data-layer; CMS’ numerous analysis tools can be used on this data. A Line of Sight (LOS) tool has been developed for CMS which analyzes terrain profiles and obstructions to determine visibility for remote observers. Figure 1 (right image) shows the viewshed analysis on the same PSR in the Nobile region. The yellow pin shows the observer location outside the PSR. The yellow area shows the visible part of PSR. The obstructed area with no visibility for the observer is shown in red. The Measurements tool allows the user to take area and distance measurements of features on the terrain using various shapes. Measurement type can be specified in a number of ways: Line, Path, Polygon, Circle, Ellipse, Square, Rectangle or Freehand. Once the shape is specified, elevation information can then be extracted along each of these shapes. Figure 2 (left) shows the measurements performed on a crater n illuminated PSR in Nobile region. The equipment placement tool allows the user to place a 3D equipment model at a desired location and analyze its coverage area. The equipment placement tool is coupled with LOS to determine the coverage. Figure 2 (right) shows an equipment placed on the Lunar terrain and it’s coverage area. The red rays are blocked sight lines and the green rays are non-obstructed sight lines with the cyan lines showing the point of intersection with the terrain. More details are provided in the video demonstrations in Reference 5. Overcoming Polar Distortions: 3D geospatial applications exhibit significant distortions in polar imagery due to several reasons: 1) distortions in the source imagery, 2) incompatible tessellation algorithms at the poles, and 3) map projections. We are leveraging new tessellation algorithms and reprojecting data using projections that are better suited for Lunar poles. The goal is to seamlessly switch to polar projections while maintaining 3D view and navigation.

Maps↗

Global 3D Data Visualization and Analysis Platform With Advanced Machine Learning Capabilities in Support of Lunar Exploration

Introduction: The science goals for NASA’s Artemis program include: a) Understanding the character and origin of lunar polar volatiles, b) Conducting experimental science in the lunar environment and c) Investigating and mitigating exploration risks [1]. The permanently shadowed regions (PSRs) on the Lunar south pole are expected to host large quantities of water-ice and volatiles that are important for sustainable Lunar exploration [2]. There are several missions such as onboard Korea Pathfinder Lunar Orbiter (KPLO: Korean name Danuri) with onboard ShadowCam camera [3], Astrobotic Peregrine Mission One [4], and other efforts underway to obtain high resolution topographic, minerals, volatiles and other information on the moon. We envision a need in immediate future for platforms to integrate these data sets, provide rendering and visualization capabilities in the context of a 3D Lunar globe for easier information access and analysis. NASA's Celestial Mapping System (CMS) [5] is developed to address the need for 3D tools for planetary science investigations, mission planning, in-situ operations, in a 3D-first design constructed around a unified view of a planetary globe. At present CMS provides many critical functionalities that include: 1) equipment planning and optimized placement on Lunar surface 2) line-of-sight (LOS) analysis 3) powerful measurement tools based on 3D terrain with realistic 3D models to represent rovers, astronauts and equipment 4) visualization of derived mapping products (e.g. resource maps), and 5) a data engine for hosting new observations that are not available in other contemporary lunar data tools [5, 7]. Planetary Data Ingestion: CMS can consume and analyze data from locally hosted and external third party sources. It is compatible with Open Geospatial Consortium (OGC) data and file standards and currently integrates datasets from the Astrogeology Science Center of USGS. This includes global and local data acquired from NASA (LRO, Clementine, Lunar Orbiter) and JAXA (SELENE/Kaguya), with capability of integrating more datasets. In addition, users can specify other WMS-hosted data endpoints, which CMS can then query and stream data from automatically. To set-up an automated process for ingestion and accurate rendering, visualization and analysis of external 3rd party planetary datasets within CMS, we initiated the process of ingesting unique dataset of super-enhanced images of the permanently shadowed regions (PSRs) at the lunar poles which were produced by the Hyper-effective nOise Removal U-net Software (HORUS) tool [8]. This tool was developed to enhance the extremely low-light images of the interior of PSRs and provide the ability to see within these regions at and discern surface features (i.e. boulders and craters) down to 3 meters in size. We focused on the Nobile region on the Lunar south pole, selected site for VIPER mission and stitched several images to create a high-resolution map within one of the PSR of Nobile crater. Figure 1 shows the dark PSR zone form the original NAC layer of LRO as the base layer (left image) and the illuminated areas within that crater (center) which was created by ingesting and merging several of HORUS generated images. At present we employ a semi-automated process to ensure spatial accuracy and merger of several overlapping zones. However, we are in the process of completely automating this process by employing AI based techniques that would rank, sort, and stack the images based on their information density. The georectification of the images would employ selected features. Analysis on Ingested Planetary Datasets: Once an external planetary data-set is successfully ingested, georectified and merged seamlessly as a data-layer; CMS’ numerous analysis tools can be used on this data. A Line of Sight (LOS) tool has been developed for CMS which analyzes terrain profiles and obstructions to determine visibility for remote observers [5,6]. Figure 1 (right image) shows the viewshed analysis on the same PSR in the Nobile region. The yellow pin shows the observer location outside the PSR. The yellow area shows the visible part of PSR. The obstructed area with no visibility for the observer is shown in red. The Measurements tool allows the user to take area and distance measurements of features on the terrain using various shapes. Measurement type can be specified in a number of ways: Line, Path, Polygon, Circle, Ellipse, Square, Rectangle or Freehand. Once the shape is specified, elevation information can then be extracted along each of these shapes. Figure 2 (left) shows the measurements performed on a crater n illuminated PSR in Nobile region. The equipment placement tool allows the user to place a 3D equipment model at a desired location and analyze its coverage area. The equipment placement tool is coupled with LOS to determine the coverage. Figure 2 (right) shows an equipment placed on the Lunar terrain and it’s coverage area. The red rays are blocked sight lines and the green rays are non-obstructed sight lines with the cyan lines showing the point of intersection with the terrain. More details are provided in the video demonstrations in Reference 5. Overcoming Polar Distortions: 3D geospatial applications exhibit significant distortions in polar imagery due to several reasons: 1) distortions in the source imagery, 2) incompatible tessellation algorithms at the poles, and 3) map projections. We are leveraging new tessellation algorithms and reprojecting data using projections that are better suited for Lunar poles. The goal is to seamlessly switch to polar projections while maintaining 3D view and navigation.

Maps↗

Application of Machine Learning to Rotorcraft Health Monitoring

Machine learning is a powerful tool for data exploration and model building with large data sets. This project aimed to use machine learning techniques to explore the inherent structure of data from rotorcraft gear tests, relationships between features and damage states, and to build a system for predicting gear health for future rotorcraft transmission applications. Classical machine learning techniques are difficult, if not irresponsible to apply to time series data because many make the assumption of independence between samples. To overcome this, Hidden Markov Models were used to create a binary classifier for identifying scuffing transitions and Recurrent Neural Networks were used to leverage long distance relationships in predicting discrete damage states. When combined in a workflow, where the binary classifier acted as a filter for the fatigue monitor, the system was able to demonstrate accuracy in damage state prediction and scuffing identification. The time dependent nature of the data restricted data exploration to collecting and analyzing data from the model selection process. The limited amount of available data was unable to give useful information, and the division of training and testing sets tended to heavily influence the scores of the models across combinations of features and hyper-parameters. This work built a framework for tracking scuffing and fatigue on streaming data and demonstrates that machine learning has much to offer rotorcraft health monitoring by using Bayesian learning and deep learning methods to capture the time dependent nature of the data. Suggested future work is to implement the framework developed in this project using a larger variety of data sets to test the generalization capabilities of the models and allow for data exploration.

machine learning↗

A Machine Learning Approach to Improve Air Traffic Management Initiatives

Collaborating closely with commercial air carriers and related organizations, the Federal Aviation Administration(FAA) regulates air traffic and ensures the safety and efficiency of air operations. Air traffic controllers make strategic decisions, such as delaying, rerouting, or canceling flights, partly based on guidance provided by the FAA’s Air TrafficControl System Command Center (ATCSCC). The guidance includes, among other things, control measures known asTraffic Management Initiatives (TMIs) designed to enhance safety and improve operational efficiency. TMIs play a crucial role in managing the demand and capacity within the U.S. National Airspace System (NAS). Two major TMIs that are routinely used (primarily to mitigate the adverse effects of bad weather) are Ground Delay Programs (GDPs) andGround Stops (GSs). In a GDP, flights destined for airports facing thunderstorm activity experience delays at their origin airports. This proactive approach minimizes the risk of routing aircraft through hazardous weather conditions and also replaces (fuel burning) airborne delays with ground delays. In a GS, a temporary restriction is imposed on the departure or arrival of aircraft at a specific airport or within a designated airspace. Although other TMIs (e.g., miles-in-trail) are also implemented as part of (air) traffic flow management in the NAS, the focus of this work is on GDPs and GSs. Since TMIs, by design, lead to flight delays or cancellations, it is crucial to put in place the right set of parameters(e.g., scope and duration of the GDP). For example, when the end time of a GDP extends beyond what is necessary, it imposes unnecessary delays on departing flights. This situation could occur as a result of inaccurate prediction of the(required) duration of the GDP based on the weather forecast. On the other hand, if a GDP ends prematurely before the underlying capacity constraints are resolved at the destination airport, it may result in airborne holding. The delicate balance lies in matching the termination of the GDP precisely with the resolution of capacity constraints, avoiding both the imposition of unnecessary ground delays and the need for airborne holding due to premature program termination.Failing to specify the right parameters for TMIs also leads to flight delays, creating a significant obstacle in managing the increasing traffic volumes causing increased work load for the controllers. To address this issue, we propose the integration of Machine Learning (ML) models in the traffic flow management(TFM) pipeline. In current operations, decisions are made by human experts based on extensive training, historical patterns, available traffic and weather data. Since we have an abundance of data from past events that tell us the likely impact of various TMIs, by ingesting historical data, properly trained ML models can offer valuable insights and aid human decision-making. With the FAA increasingly exploring advanced analytics, ML emerges as a focal point for enhancing TFM within the National Airspace System (NAS). As a first step, this study aims to provide traffic controllers with decision-making support for the issuance and adjustment of TMIs. Data analytics and machine learning have been previously employed to address some of the challenges associated with TMIs. Numerous studies have concentrated on various facets of TMI issuance, exploring factors influencing TMI parameters, including arrival rate, airport capacity, and delay prediction. For example, using weather forecasts, several statistical methods were used to produce probabilistic capacity profiles which in conjunction with deterministic models provided insights into the GDP planning process [1–4]. The downside of using deterministic models is that they rely on fixed inputs and predetermined rules, which lack the ability to account for the inherent uncertainty and variability present in real-world scenarios. In a separate series of studies, researchers aimed to predict the occurrences of GDPs and GSs. The majority of these studies utilized various supervised learning methods, including Decision Trees, Naive Bayes, Support VectorMachines, and Random Forests to analyze the influence of weather conditions and arrival demand on TMI incidents[5–8]. However, these studies primarily focused on predicting the incidence of TMIs without explicitly addressing the scope of TMIs, including their duration and their geographical coverage. Furthermore, the emphasis of these studies was largely on GDPs, given their higher frequency and longer duration when compared to GSs. A limited number of studies focused on predicting the parameters of TMIs, specifically addressing their duration and extent. In one such study focusing on optimizing the TMI parameters at San Francisco International Airport (SFO),the authors utilized a probabilistic forecast of fog [9]. They simulated various capacity scenarios based on the (fog)burn-off forecasts, selecting GDP parameters that minimized airborne and overall ground delays. However, this approach exclusively emphasizes stratus (fog) burn-off as the primary determinant of GDP and GS, neglecting other influential factors like severe weather events, runway closures, lower capacity than traffic demand, and other important variables. Given the complexity of predicting the TMI and determining its scope, we seek a more holistic approach. We aim to consider all significant factors that could impact TMIs and their parameters. What sets this research apart is the fusion of all data sources relevant to the issuance and adjustment of TMIs and it represents the first comprehensive attempt to optimize TMIs in this manner. Since this comprehensive solution involves various aspects, we break down the problem into smaller components and input all parameters into a unified model called the “TMI Adjuster”. Figure 1 shows the overall framework and the list of datasets used in each model. The objective of the TMI Adjuster module is to deliver reliable, consistent and expedited recommendations for the progression, adjustment, and termination of TMIs. The ML solution entails developing a pipeline capable of predicting the necessity of a TMI (e.g., GS or GDP) along with its various parameters. For example, in the case of a GS, this includes the scope of the GS either in terms of distance from the destination airport or based on pre-defined airspace sectors. Here, scope refers to those regions and departing airports that are subject to the GS. In this paper, we concentrate on the issuance of GSs in the three major airports in the New York area — LaGuardia(LGA), John F. Kennedy International (JFK), and Newark Liberty International (EWR). We fuse traffic, weather and other relevant aviation data from years 2017 to 2019 to train and validate the ML models. In particular, we use the following datasets: •Terminal Aerodrome Forecast (TAF): meteorological forecasts specific to each airport, issued four times a day, covering predefined time periods. •TMI data: includes all GSs and GDPs along with their respective parameters. •Aviation System Performance Metrics (ASPM): includes traffic related data such as aircraft delays, arrival, and departure rates. •Notices to Airmen (NOTAMs): utilized to extract runway closure data and manage interdependencies between terminals in close proximity. •Flight cancellation data •Airspace Flow Programs (AFP): includes information on flight airborne holdings caused by TMIs. The data preprocessing entails transforming ASPM, TMI, AFP, NOTAMs, and weather data into an hourly format and consolidating all datasets by merging them based on date and time as the primary key. The TMI Adjuster framework comprises two parallel models: one dedicated to GS and a second model focused on GDP. As previously mentioned, our specific focus is on the GS model as a multi-classification problem. In this framework, each data point of the GS model input summarizes ten hours of data. Specifically, the data loader for the GS model generates the input and output of the model as follows: at a given time step, the input includes the actual traffic, weather, and TMI data from the two-hour window before the time step, alongside the weather forecast and scheduled traffic for the next 8 hours starting from the time step. Based on this information, the output of the GS model for each time interval consists of three dimensions. The first dimension represents a binary decision on whether there should be a GS in place for the next hour or not. The second dimension is related to the scope of the GS in the United States, and the third dimension is related to the scope of the GS in Canada (i.e., to determine if the GS impacts airports in Canada).One of the challenges with TMI modeling is the sparsity of TMI events, particularly regarding its scope. To address this challenge in the scope of the GS model output, we implement grouping. The GS scope for the US region is defined based on a list of centers that should be included when the GS is in place. With 20 centers in the US, we utilized historical data to group them into 4 categories. In particular, we summarized our historical data in a graph format where nodes represent centers, and link weights are defined based on the co-occurrence of centers in the scope parameter ofTMIs. By identified strongly connected components in this graph, we were able to partition the centers into four groups. We consider two model structures for the GS Model. Firstly, a hierarchical classification model [10], where the human decision-making for a GS is of hierarchical nature. The decision-maker first decides whether there is a need fora GS, and if the answer is yes, determines the scope. A hierarchical classification model organizes the problem into a class hierarchy, typically a tree or a Directed Acyclic Graph (DAG) structure, and considers the dependency of the decision in the previous step to the next component [10]. Here, we employ the local classifier per level approach, which involves training one multi-class classifier for each level of the class hierarchy. The second structure is the independent structure. In this setting, as the name suggests, we do not consider the dependency of the decisions in the different dimensions of the output of the model. Instead, for each dimension, we train a multi-class classifier independently. Table 1 summarizes GS model statistics for training, validation and testing. The table documents the effect of limiting data to the time steps when there was actually a TMI in place or when a TMI had just terminated. This resulted in a more balanced distribution of the GS class(GS positive class)versus “No GS”(GS negative class), which might help the training process. While JFK and LGA follow very similar distributions, with 40% and 42% GS positive class respectively, EWR has proportionally fewer GS incidents at 28%. Our subsequent phase involves evaluating the performance of both hierarchical structure and independent structure using different state-of-the-art multi-class classifier models such as Random Forest, Decision Trees, K-nearest Neighbors, and Logistic Regression and forecast the duration and scope of the GSs.

Farzan Masrour Shalmani↗

Run-time scheduling and execution of loops on message passing machines

Sparse system solvers and general purpose codes for solving partial differential equations are examples of the many types of problems whose irregularity can result in poor performance on distributed memory machines. Often, the data structures used in these problems are very flexible. Crucial details concerning loop dependences are encoded in these structures rather than being explicitly represented in the program. Good methods for parallelizing and partitioning these types of problems require assignment of computations in rather arbitrary ways. Naive implementations of programs on distributed memory machines requiring general loop partitions can be extremely inefficient. Instead, the scheduling mechanism needs to capture the data reference patterns of the loops in order to partition the problem. First, the indices assigned to each processor must be locally numbered. Next, it is necessary to precompute what information is needed by each processor at various points in the computation. The precomputed information is then used to generate an execution template designed to carry out the computation, communication, and partitioning of data, in an optimized manner. The design is presented for a general preprocessor and schedule executer, the structures of which do not vary, even though the details of the computation and of the type of information are problem dependent.

Crowley, Kay↗

Run-time scheduling and execution of loops on message passing machines

Sparse system solvers and general purpose codes for solving partial differential equations are examples of the many types of problems whose irregularity can result in poor performance on distributed memory machines. Often, the data structures used in these problems are very flexible. Crucial details concerning loop dependences are encoded in these structures rather than being explicitly represented in the program. Good methods for parallelizing and partitioning these types of problems require assignment of computations in rather arbitrary ways. Naive implementations of programs on distributed memory machines requiring general loop partitions can be extremely inefficient. Instead, the scheduling mechanism needs to capture the data reference patterns of the loops in order to partition the problem. First, the indices assigned to each processor must be locally numbered. Next, it is necessary to precompute what information is needed by each processor at various points in the computation. The precomputed information is then used to generate an execution template designed to carry out the computation, communication, and partitioning of data, in an optimized manner. The design is presented for a general preprocessor and schedule executer, the structures of which do not vary, even though the details of the computation and of the type of information are problem dependent.

Saltz, Joel↗

Characterization and Quantification of Radiation-Induced Clusters/Precipitates in RPV Steels Using STEM-EDS and Machine Learning

Over the operational lifespan of a nuclear reactor, reactor pressure vessel (RPV) steels are subjected to significant neutron irradiation, resulting in complex microstructural changes and the consequent degradation of mechanical properties. Various physically motivated correlation models have been developed to predict neutron irradiation-induced embrittlement of RPVs under different irradiation conditions. However, the efficient and accurate characterizations and quantification of radiation-induced clusters in RPVs are still challenging, which will affect the precision of the predictive models for embrittlement of RPV components. In the DOE Visiting Faculty Program (VFP) research work at Oak Ridge National Lab (ORNL), I integrate machine learning to aid Scanning Transmission Electron Microscopy – Energy Dispersive X-ray Spectroscopy (STEM-EDS) analyses, which improve the characterization and quantification of radiation-induced clusters in RPV steels, thereby enabling more accurate predictions of material behavior under irradiation. The surveillance base- and welded- RPV steels were annealed at various temperatures of 340 °C, 450 °C and 500 °C for up to 168 hours, respectively. Afterwards, I have characterized radiation-induced clusters using advanced STEM-EDS techniques and subsequently applying machine learning algorithms to analyze and refine STEM-EDS datasets, enhancing the quantification of clusters compositions and distributions. In the end, an efficient workflow for integrating STEM-EDS data analysis with machine learning to address challenges including noise reduction has been developed. The completion of this VFP work will support bridge critical gaps in the accurate quantification of radiation-induced clusters in RPV steels using STEM-EDS and support the development of more precise models for predicting RPV embrittlement in the Light Water Reactor Sustainability program supported by Department of Energy and enhancing the collaboration between ORNL and Alred University. The outcome of the VFP project will leverage a few research papers submission to peer-reviewed journals in the relevant scientific field and a few oral presentations at national and international conferences.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Driver Identification Midyear Report

First, we create a profile for each authorized driver based on their existing driving data. We then train a machine learning model on the driving data from this profile, yielding an individualized model for each driver. Finally during a drive, we pass the Controller Area Network (CAN) data to the model and authen ticate the driver’s identity in real-time. This verification or lack thereof could be used to alert supervisors of threats to their drivers or transported materials. Deviations from their normal driving behavior could indicate high-risk situations, medical events, or even insider threats.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Infrared astronomical data base and catalog of infrared observations

The NASA/Goddard Space Flight Center has developed a computer data base of infrared astronomical observations. The data base represents a machine-readable library of infrared observational data published in the relevant literature since 1960 for celestial sources outside the solar system. It likewise includes the contents of infrared surveys and catalogs. A catalog of infrared observations has been developed in both printed and magnetic-tape formats. The data base will be accessed through a bibliographic guide and an atlas of infrared source names and positions. Future plans also include two-dimensional graphical displays of infrared data and a user-interactive data terminal.

Schmitz, M.↗

Verification and Validation Data for Computational Unsteady Aerodynamics

Computational Unsteady Aerodynamics computer codes are being increasingly used. In order to validate their results they must be tested against valid experimental data. The present report aims at collecting reliable experimental data on unsteady aerodynamics and presenting them in a form which permits use for verification of codes. For ease of handling, the data are also presented in machine readable form (CD-ROM). Data on increasingly complex generic forms were selected and the following categories are covered: flutter, buffet, stability and control, dynamic stall, cavity flows, store separation. Computational solutions are included in order to permit evaluation of codes and analysis of solutions which differ from experimental data.

Applied Vehicle Technology Panel (AVT)Task Group↗

Enhancing Electron Microscopy Image Classification Using Data Augmentation

Manual labeling for machine learning tasks such as image classification is tedious and labor-intensive; as a result, scientific datasets suitable for deep learning applications are scarce and limited. While data augmentation techniques have shown promise for extending image datasets, very little work has been done to understand the impact of combining multiple augmentation methods sequentially or the limits of their effectiveness when combined. Our work addresses this gap by examining how standard and combinatorial data augmentation affects the performance of machine learning models when trained on small datasets for label classification tasks. For our analysis, we generate single, double and quadruple-augmented datasets for a microscopy image classification task using six standard augmentation methods, and compare the resultant improvements observed in binary classification accuracy with three standard image classification models (DenseNet169, MobileNetV2, ResNet101V2). Our experiments show a non-monotonic relationship between the number of simultaneous augmentation methods and classification accuracy, indicating that there is a trade-off between the degree of augmentation and the model performance. These findings suggest that the optimal number of augmentation methods will vary by domain and use case. We also find that the order in which augmentation methods are applied to a limited dataset matters when combining augmentation schemes, with our use case showing performance differences up to 2.6% when the augmentation order is reversed for double-augmented datasets. Our work offers insights to the limits of data augmentation when working on image classification tasks with limited datasets.

Welsman, Jordan A↗

Unsupervised Anomaly Detection in High-Dimensional Flight Data Using Convolutional Variational Auto-Encoder

The modern National Airspace System (NAS) is an extremely safe system. The industry has experienced a steady decrease in fatalities over the years. This can be contributed to both improved flight critical systems with redundant hardware and software protections as well as an increased focus on active monitoring and response to real time and historically identified vulnerabilities by implementing more resilient procedures and protocols. The main practice for identifying vulnerabilities in operations leverages domain expertise using knowledge about how the system should behave with the expected tolerances to known safety margins. This approach works well when the system has a well-defined operating condition. However, the operations in the NAS can be highly complex with various nuances that render it difficult to clearly pre-define all known safety vulnerabilities. With the advancement of data science and machine learning techniques, the potential to automatically identify emerging vulnerabilities in the observed operations has become more practical in recent years. The state-of-the-art anomaly detection approaches in aerospace data usually rely on supervised or semi-supervised learning. However, in many real-world problems such as flight safety creating labels for the data requires huge amount of efforts and is largely expensive. As a result, in this article, we develop a Convolutional Variational Auto-Encoder (CVAE), an unsupervised learning approach for anomaly detection in high-dimensional heterogeneous time-series data. We validate performance of CVAE compared to the state-of-the-art supervised learning approach (as an upper bound) as well as an supervised clustering based on K-Means (as a lower bound) on Yahoo!'s benchmark time series anomaly detection data. Finally, we showcase performance of CVAE on a case study of identifying anomalies in the first 60 seconds of commercial flights' take-offs using Flight Operational Quality Assurance (FOQA) data.

Milad Memarzadeh↗

Documentation for the machine-readable version of the Fourth Cambridge Radio Survey Catalogue (4C) (Pilkington, Gower, Scott and Wills 1965, 1967)

The machine readable catalogue contains survey data from the papers of Pilkington and Scott and Gower, Scott and Wills. These data result from a survey of radio sources between declinations -07 deg and +80 deg using the large Cambridge interferometer at 178 MHz. The computerized catalog contains for each source the 4C number, 1950 position, measured flux density, and accuracy class. For some sources miscellaneous brief comments such as cross identifications to the 3C catalog or remarks on contamination from nearby sources are given at the ends of the data records. A detailed description of the machine readable catalog as it is currently being distributed by the Astronomical Data Center is given to enable users to read and process the data.

Warren, W. H., Jr.↗

Evaluation of Machine Learning and Deep Learning Algorithms for Fire Prediction in Southeast Asia

Vegetation fires are prevalent in South/Southeast Asian countries, making fire prediction crucial due to their potential environmental, economic, and social impacts. Accurate predictions of fires facilitate timely interventions, helping to mitigate uncontrolled fires that can lead to biodiversity loss and air quality issues. In this study, we utilize VIIRS satellite-derived fire data alongside six machine learning and deep learning models—Simple Persistence, Multi-Layer Perceptron (MLP), Convolutional Neural Network (CNN), Long Short-Term Memory (LSTM), CNN-LSTM, and ConvLSTM—to determine the most effective fire prediction model, using Root Mean Square Error (RMSE) as the metric. Our results indicate that the CNN model is the most reliable in regions with spatial dependencies, such as Brunei, Indonesia, Malaysia, the Philippines, Timor-Leste, and Thailand. Conversely, the ConvLSTM model excels in countries with complex spatiotemporal dynamics like Laos, Myanmar, and Vietnam. The CNN-LSTM hybrid model also performed well in Cambodia, suggesting a need for a balanced approach in areas requiring both spatial and temporal feature extraction. Furthermore, simpler models like Persistence and MLP showed limitations in capturing dynamic patterns and temporal dependencies. Our findings highlight the importance of evaluating models before implementing any decision support systems (DSS) in fire management. By tailoring models to specific regional fire data, we can enhance prediction accuracy and responsiveness, ultimately improving fire risk management in Southeast Asia and beyond.

Deep learning↗