Search NASA⌕ Search

SEARCH · Search NASA

Results for “training data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Analysis of the Westland Data Set

The "Westland" set of empirical accelerometer helicopter data with seeded and labeled faults is analyzed with the aim of condition monitoring. The autoregressive (AR) coefficients from a simple linear model encapsulate a great deal of information in a relatively few measurements; and it has also been found that augmentation of these by harmonic and other parameters call improve classification significantly. Several techniques have been explored, among these restricted Coulomb energy (RCE) networks, learning vector quantization (LVQ), Gaussian mixture classifiers and decision trees. A problem with these approaches, and in common with many classification paradigms, is that augmentation of the feature dimension can degrade classification ability. Thus, we also introduce the Bayesian data reduction algorithm (BDRA), which imposes a Dirichlet prior oil training data and is thus able to quantify probability of error in all exact manner, such that features may be discarded or coarsened appropriately.

Wen, Fang↗

Condition Monitoring for Helicopter Data

In this paper the classical "Westland" set of empirical accelerometer helicopter data is analyzed with the aim of condition monitoring for diagnostic purposes. The goal is to determine features for failure events from these data, via a proprietary signal processing toolbox, and to weigh these according to a variety of classification algorithms. As regards signal processing, it appears that the autoregressive (AR) coefficients from a simple linear model encapsulate a great deal of information in a relatively few measurements; it has also been found that augmentation of these by harmonic and other parameters can improve classification significantly. As regards classification, several techniques have been explored, among these restricted Coulomb energy (RCE) networks, learning vector quantization (LVQ), Gaussian mixture classifiers and decision trees. A problem with these approaches, and in common with many classification paradigms, is that augmentation of the feature dimension can degrade classification ability. Thus, we also introduce the Bayesian data reduction algorithm (BDRA), which imposes a Dirichlet prior on training data and is thus able to quantify probability of error in an exact manner, such that features may be discarded or coarsened appropriately.

Wen, Fang↗

TPSAS-NF1676L-35322-DND

This talk discusses emerging methods that seek to fuse and integrate physics-based modeling with machine learning. With the recent rise of machine learning and artificial intelligence, there has been a huge surge in data-driven approaches to solve computational science and engineering problems. However, neglecting a priori knowledge of established physical laws and relying solely on data-driven methods can yield unreliable, less interpretable, and/or non-physical results, especially when data is sparse or predictions are required outside of the training data domain. This two part talk presents two distinct approaches for accelerating predictions with machine learning that are grounded and constrained by relevant physics and their application to problems at NASA.

Julian Cuevas Paniagua↗

Machine Learning Methods for Estimating Propeller Source Noise Spheres

In this work, several neural network function approximations are compared for inter- polating, storing, and sampling acoustic source spheres with applications to propeller noise estimation. These methods are compared using an acoustic model of the three bladed GL-10 propeller at different flight conditions, with training data generated using NASA’s ANOPP-PAS module. The source spheres used to train the networks capture the tonal propeller noise due to both the blade thickness and loading. This tonal noise prediction method allows the vehicle noise to be estimated for auralization and acoustic control. Three radial basis function neural network architectures are compared in this work. The first two networks directly estimate the parameters of the source sphere at different flight conditions but differ in the number of layers used. The third network estimates the parameters of the source sphere using a weighted combination of spherical basis functions. These networks are trained on numerically generated source spheres, with operating points given in terms of the propeller rotation rate, freestream speed, and propeller angle of attack. The performance of the neural network is determined using a validation dataset of withheld data points. This performance is quantified in terms of the approximation error, training time, and sample time. The third network, which estimates the weights of the spherical basis functions, performs the best in both average and maximum approximation errors in all cases. This network’s worst case performance is 5.6 % relative dif- ference of a model parameter associated with acoustic pressure. The direct estimation network with a single layer has the worst approximation error in all cases. Additionally, the spherically defined network has the slowest sample time at 0.05 seconds per thousand points. Both direct estimation methods produce a thousand sample points in approximately 0.001 seconds.

Acoustics↗

Active Learning with Irrelevant Examples

An improved active learning method has been devised for training data classifiers. One example of a data classifier is the algorithm used by the United States Postal Service since the 1960s to recognize scans of handwritten digits for processing zip codes. Active learning algorithms enable rapid training with minimal investment of time on the part of human experts to provide training examples consisting of correctly classified (labeled) input data. They function by identifying which examples would be most profitable for a human expert to label. The goal is to maximize classifier accuracy while minimizing the number of examples the expert must label. Although there are several well-established methods for active learning, they may not operate well when irrelevant examples are present in the data set. That is, they may select an item for labeling that the expert simply cannot assign to any of the valid classes. In the context of classifying handwritten digits, the irrelevant items may include stray marks, smudges, and mis-scans. Querying the expert about these items results in wasted time or erroneous labels, if the expert is forced to assign the item to one of the valid classes. In contrast, the new algorithm provides a specific mechanism for avoiding querying the irrelevant items. This algorithm has two components: an active learner (which could be a conventional active learning algorithm) and a relevance classifier. The combination of these components yields a method, denoted Relevance Bias, that enables the active learner to avoid querying irrelevant data so as to increase its learning rate and efficiency when irrelevant items are present. The algorithm collects irrelevant data in a set of rejected examples, then trains the relevance classifier to distinguish between labeled (relevant) training examples and the rejected ones. The active learner combines its ranking of the items with the probability that they are relevant to yield a final decision about which item to present to the expert for labeling. Experiments on several data sets have demonstrated that the Relevance Bias approach significantly decreases the number of irrelevant items queried and also accelerates learning speed.

Wagstaff, Kiri↗

Machine learning for Deep Space Network antenna motions detection

Highly stable frequency and timing standards are essential for deep-space missions and radio science. At the NASA Deep Space Network (DSN), these standards are distributed through a network of underground fiber cables to support several Goldstone antennas. Independently developed frequency-measuring instruments generate tremendous quantities of data to monitor and validate the antennas’ stringent frequency requirements. In this paper, we propose a lightweight processing tool capable of detecting disturbances on the frequency signal caused by DSN antenna motions. Our training data is sampled from the movement log of the antenna of interest and the generated data from the fiber optic metrology instrument linked to the antenna. We demonstrate that a convolutional neural network (CNN) model can achieve high accuracies on classifying instances of antenna movements and is an effective predictor when used iteratively on longer, variable stretches of metrology data. The simplicity, low training cost, and high accuracies of our model strongly suggest its efficacy in identifying and troubleshooting frequency disturbances caused by the antenna.

Yi, Lin↗

Neural Network and Regression Approximations in High Speed Civil Transport Aircraft Design Optimization

Nonlinear mathematical-programming-based design optimization can be an elegant method. However, the calculations required to generate the merit function, constraints, and their gradients, which are frequently required, can make the process computational intensive. The computational burden can be greatly reduced by using approximating analyzers derived from an original analyzer utilizing neural networks and linear regression methods. The experience gained from using both of these approximation methods in the design optimization of a high speed civil transport aircraft is the subject of this paper. The Langley Research Center's Flight Optimization System was selected for the aircraft analysis. This software was exercised to generate a set of training data with which a neural network and a regression method were trained, thereby producing the two approximating analyzers. The derived analyzers were coupled to the Lewis Research Center's CometBoards test bed to provide the optimization capability. With the combined software, both approximation methods were examined for use in aircraft design optimization, and both performed satisfactorily. The CPU time for solution of the problem, which had been measured in hours, was reduced to minutes with the neural network approximation and to seconds with the regression method. Instability encountered in the aircraft analysis software at certain design points was also eliminated. On the other hand, there were costs and difficulties associated with training the approximating analyzers. The CPU time required to generate the input-output pairs and to train the approximating analyzers was seven times that required for solution of the problem.

Patniak, Surya N.↗

A Machine Learning Approach to Objective Identification of Dust in Satellite Imagery

Airborne dust has broad adverse effects on human activity, including aviation, human health, and agriculture. Remote sensing observations are used to detect dust and aerosols in the atmosphere using long established techniques. False color Red-Green-Blue (RGB) imagery using band differences sensitive to dust absorption (Dust RGB) is currently used operationally to assist forecasters and decision-makers in identifying dust at night, but there are still limitations, subjectivity, and nuances to image interpretation making night-time dust identification difficult even for experts. This study applies machine learning to the problem of night-time dust detection with a simple random forest (RF) model using Geostationary Operational Environmental Satellite-16 (GOES-16) Advanced Baseline Imager (ABI) infrared imagery, band differences sensitive to dust absorption, and Dust RGB color components as inputs to the model. The RF model achieves an Area-Under-Curve (AUC) of 0.97 with a standard deviation of 0.04 for dust cases. For images with dust present, the model correctly labels 85% of dust pixels and 99.96% of no-dust pixels for all dust images in the validation data set. The addition of a single null case to the training data set drastically reduces error in labeling no-dust pixels as dust from 45% to 14.5%. Application of the machine learning model to the April 13–14, 2019 dust event demonstrates the ability of the model to identify dust during night-time hours when visual dust detection is limited by the cooling ground surface characteristics.

dust↗

Parameterization of Vertical Cloud Distribution from C3M and MERRA Data Using ML Method

Clouds play a key role in regulating the hydrological cycle and the Earth's radiative energy budget. However, global climate models (GCMs) with a horizontal grid spacing on the order of 100 km have limitations in representing sub-grid cloud dynamics with spatial scales on the order of 1 km, leading to potential uncertainties in cloud radiative feedback on the global scale. In our research, we will leverage the capabilities of Deep Machine Learning (DML) methods to construct parameterizations of sub-grid volumetric cloud fraction (VCF), which is the frequency of occurrence on a grid volume accumulated in the horizontal and vertical directions. Our investigation delves into the intricate relationship between VCF obtained from the NASA CALIPSO-CloudSat-CERES-MODIS (CCCM) satellite observation data and 3-D MERRA-2 reanalysis meteorological profiling data (e.g., wind, relative humidity, temperature). Through a comprehensive one-year data training utilizing the Sequence to Sequence DML method, we have successfully disentangled the complicated cloud formation dynamics across diverse meteorological conditions through a day-to-day analysis framework. Preliminary findings reveal promising statistical agreements in geographical and vertical distributions and seasonal variations of volumetric cloud fraction between ML prediction and satellite measurements. These results underscore the aptitude of our DML model to discern underlying cloud physical processes and accurately represent sub-grid cloud formation dynamics. Additionally, we have also employed trained neural network to analyze uncertainties arising from errors in meteorological data, further enhancing the robustness of our VCF parameterization.

Shan Zeng↗

NASA Pilot-Engaged Expert Response Using IBM Watson Technology: Prototype Evaluation of Knowledge Retrieval System

NASA Langley Research Center and IBM have been investigating the use of IBM Watson technology in aerospace research and development. One application of Watson technology is the Pilot-Engaged Expert Response (PEER) use case. The PEER system is envisioned as an in-cockpit advisor that will act as a source of situationally-relevant information for pilots and other flight crew members to assist in decision making about real-time events and situations that arise in the course of aircraft operations. PEER will make available vast stores of knowledge and information quickly and directly, putting important informational resources where they are needed most. IBM has worked with NASA to develop an architecture and articulate a roadmap for the development of the PEER system. That vision is built around Watson Discovery Advisor (WDA) software solution, derived from IBM's Jeopardy!-winning automatic question answering system. PEER makes use of WDA's sophisticated question-answering capabilities as its core, adding important User Interface components and other customizations for the cockpit environment, including communication with flight systems and other external data sources. The development plan for PEER includes four development stages, with the current project constituting the first phase. In this project, a prototype instance of PEER was successfully adapted to the aviation domain, enabling users to ask questions about aviation topics and receive useful and accurate answers to these questions. Major tasks accomplished include the development of procedures for domain adaptation through automatic lexicon extraction from domain glossaries; generation of question-answer training data which was used to train the system; and assessment of the effectiveness of domain adaptation, which showed a dramatic improvement in the ability of the PEER system to answer domain-relevant questions. In addition, the vision for the PEER system was pushed forward by the articulation of a plan for the automatic enhancement of question-answering with contextual information. This initial phase focused on two main goals: 1) the targeted domain adaptation of the underlying WDA system to the aviation domain; and, 2) the design of the software systems needed to leverage flight-contextual data. Domain adaptation of the WDA system proceeds via three main activities: Domain data ingestion, lexical customization and model training. A textual corpus consisting of 1,147 individual documents with more than 7.5 million words of text was ingested into the system and this served as the basis of all further development. A domain lexicon of over 3,500 aviation-domain terms was semi-automatically generated from domain documents and used to train the system. In addition, a set of over 500 question-answer (QA) pairs relevant to the PEER use case was developed; these were used to train and assess the system. These important first steps established the basis for the PEER system. In addition, steps were taken towards the integration of the PEER system into the cockpit environment with the development of a functional design for the Contextual Data Augmentation (CDA) subsystem. This subsystem brings to bear contextual data to improve system responses. It has three main submodules: the Contextual Data Collection module, the Contextual Data Selection module, and the Contextual QA Augmentation module. These modules form a processing pipeline that addresses the problems associated with automatically integrating information from external resources into the knowledge-retrieval mechanism.

Machine learning↗

Transfer-AE: A novel autoencoder-based impact detection model for structural digital twin

Accurately detecting the location and intensity of impacts is crucial for ensuring structural safety. Currently, AI-based structural impact detection methods are widely used for their excellent detection accuracy. However, their generalization capability is limited by the scenarios present in the training data. Many complex and dangerous impact scenarios are difficult to conduct real-world experiments on to collect sufficient samples. To capture all impact scenarios and fully leverage the advantages of AI-based detection technologies, advanced methods involve combining real-world structural monitoring data with corresponding numerical models to construct digital twins. These methods continuously refine the created numerical models with limited real-world data and provide diverse impact scenarios through numerical model simulations. However, there are inevitable differences between digital models and physical models that are challenging to correct through mechanical means. This discrepancy in data distribution between the two models significantly hinders the application of digital twin technology in impact/event identification tasks. To address this challenge, this study proposes a novel model based on autoencoders, named Transfer-AE. Transfer-AE encodes the common features of digital twins in the latent space to bridge the uncertainty gap at a macro scale between numerical models and physical models and synchronously fits the magnitude and location of the impact load in the decoder. This enables consistent detection results for the same impact event, whether the sample comes from the numerical model or the physical model. Transfer-AE includes two operating modes: Mode 1 has a fixed computational complexity with stable inference speed, but the training cost and difficulty increase with data distribution. Mode 2's computational complexity increases with data distribution, but it has a fixed training cost and speed. In both cases involving the geodesic dome structure simulating a deep space habitat and the IASC-ASCE benchmark structure, Transfer-AE demonstrated the best performance in impact localization and quantification tasks compared to mainstream domain-adaptive transfer models.

Chengjia Han↗

Needs assessment final report

The stated purposes of the Management Science Faculty Fellowship Project were to: (1) provide a comprehensive analysis of KSC management training for engineers and other management professionals from project/program lead through executive levels; and (2) development of evaluation methodologies which can be used to perform ongoing program-wide course-to-course assessments. This report will focus primarily in the first stated purpose for the project. Ideally, the analysis of KSC management training will build in the current system and efficiently propose improvements to achieve existing goals and objectives while helping to identify new visions and new outcomes for the Center's Management Training Mission. Section 2 describes the objectives, approach, and specific tasks used to analyze KSC's Management training System. Section 3 discusses the main conclusions derived from an analysis of the available training data. Section 4 discusses the characteristics and benefits envisioned for a Management Training System. Section 5 proposes a Training System as identified by the results of a Needs Assessment exercise conducted at KSC this summer. Section 6 presents a number of recommendations for future work.

Sepulveda, Jose A.↗

Automatic Speech Recognition for Launch Control Center Communication Using Recurrent Neural Networks with Data Augmentation and Custom Language Model

Transcribing voice communications in NASA’s launch control center is important for information utilization. However, automatic speech recognition in this environment is particularly challenging due to the lack of training data, unfamiliar words in acronyms, multiple different speakers and accents, and conversational characteristics of speaking. We used bidirectional deep recurrent neural networks to train and test speech recognition performance. We showed that data augmentation and custom language models can improve speech recognition accuracy. Transcribing communications from the launch control center will help the machine analyze information and accelerate knowledge generation.

Chow, Edward↗

Evaluation of a segment-based LANDSAT full-frame approach to corp area estimation

As the registration of LANDSAT full frames enters the realm of current technology, sampling methods should be examined which utilize other than the segment data used for LACIE. The effect of separating the functions of sampling for training and sampling for area estimation. The frame selected for analysis was acquired over north central Iowa on August 9, 1978. A stratification of he full-frame was defined. Training data came from segments within the frame. Two classification and estimation procedures were compared: statistics developed on one segment were used to classify that segment, and pooled statistics from the segments were used to classify a systematic sample of pixels. Comparisons to USDA/ESCS estimates illustrate that the full-frame sampling approach can provide accurate and precise area estimates.

Bauer, M. E.↗

Vector-Ordering Filter Procedure for Data Reduction

The vector-ordering filter (VOF) technique involves a procedure for sampling a large population of data vectors to select a subset of data vectors that fully characterize the state space of the large population. The VOF technique enables a large reduction of the volume of data that must be handled in the automated monitoring system and method discussed in the two immediately preceding articles. In so doing, the VOF technique enables the development of data-driven mathematical models of a monitored asset from sets of data that would otherwise exceed the memory capacities of conventional engineering computers. Data-driven mathematical models have been shown to offer high fidelity for purposes of control and monitoring of assets. In practice, a collection of asset-operating observations is acquired with the intention that the collection contain observations characteristic of the full dynamic range of operation of the asset. Often, such a collection contains an extremely large number of observations, many of which are redundant. The VOF technique fills the need for a means to extract, from the original collection of observational data, a reduced data matrix that excludes redundant data while maintaining the full statistical character and dynamic range of the original data. The reduced data matrix can then be used as the input data for development of a mathematical model of the monitored asset, or as training data for a neural-network substitute for an explicit mathematical model of the asset. Alternatively, the reduced data matrix can, itself, be used directly as a mathematical model of the monitored asset, as is commonly done in multivariate state-estimation techniques. The original data are collected from the asset over a range of operating states and are put in matrix form. Each column vector in the original data matrix represents the signal values acquired at a particular operational state of the asset. Thus, the number of columns of the original data matrix equals the number of observed states and the number of rows in this matrix equals the number of signals acquired at each observation. In the VOF technique, one extracts the reduced data matrix from the original data matrix through the selection of a representative subset of the column (state) vectors.

Bickford, Randall L.↗

BEAM Technology Flight Demonstration

As technologies advance, their growing complexity makes them harder to maintain. Detection methods for isolating and identifying impending problems are needed to balance this complexity. Through comparison of signal pairs from onboard sensors, the Beacon-based Exception Analysis For Multimissions (BEAM) algorithm can identify and help classify deviations in system operation from a data-trained statistical model. The goal of this task is to mature BEAM and validate its performance on a flying test bed. A series of F-18 flight demonstrations with BEAM monitoring engine parameters in real time was used to demonstrate in-the-field readiness. Captured F-18 and simulated F-18 engine data were used in model creation and training. The algorithm was then ported to the embedded system with a data buffering, file writing, and data-time-stamp monitoring shell to reduce the impact of embedded system faults on BEAM'S ability to correctly identify engine faults. Embedded system testing identified hardware related restrictions and contributed to iterative improvements in the code's runtime performance. The system was flown with forced engine flameouts and other pilot induced faults to simulate operation out of the norm. Successful detection of these faults, confirmed through post-flight data analysis, helped BEAM achieve TRL6.

Wang, David↗

GL4U: Training the next generation of bioinformaticians, one omics datatype at a time

Spaceflight modifies gene expression in every organism examined to date, including humans. Understanding how these gene expression changes affect physiology is crucial for the development of countermeasures to enable long-duration manned missions. NASA’s GeneLab project provides researchers open access to multi-omics data, including genetic and gene expression data, from spaceflight experiments that can be mined to understand the effects of spaceflight on biological systems. To ensure new knowledge generation through data re-use, it is important to maximize the number of scientists who utilize GeneLab data. Training students on the GeneLab platform is the best way to create long-term adopters of this NASA database and its tools. Turning students into future instructors and advocates will also accelerate the dissemination of these data and tools to the broader scientific community. Therefore, in collaboration with the GeneLab Educational Working Group (EWG), GeneLab has created GeneLab for Colleges and Universities (GL4U). GL4U provides space biology-relevant training in bioinformatics to the next generation of scientists through direct and indirect approaches. The GeneLab team plans to host two annual data processing bootcamps, one for college-level students (direct) and one for college educators (indirect – training of trainers), in which participants learn to analyze GeneLab’s space-relevant omics data. During the bootcamp, educators will receive materials and training to enable them to run the bootcamp at their home institutions or alternatively to adapt the content to implement within existing courses, thereby extending the reach of this initiative. The GL4U direct training pilot program was conducted in June 2021 in collaboration with USRA and San Jose State University (SJSU). During the pilot, SJSU students participated in a week-long bootcamp consisting of space biology-specific lectures and hands-on instruction using Jupyter Notebooks to analyze RNA sequence data. This pilot demonstrates the capacity of GL4U for training young scientists and encouraging data re-use.

Jonathan Matthew Galazka↗