Search NASA⌕ Search

SEARCH · Search NASA

Results for “training data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Using Decision Trees to Detect and Isolate Simulated Leaks in the J-2X Rocket Engine

The goal of this work was to use data-driven methods to automatically detect and isolate faults in the J-2X rocket engine. It was decided to use decision trees, since they tend to be easier to interpret than other data-driven methods. The decision tree algorithm automatically "learns" a decision tree by performing a search through the space of possible decision trees to find one that fits the training data. The particular decision tree algorithm used is known as C4.5. Simulated J-2X data from a high-fidelity simulator developed at Pratt & Whitney Rocketdyne and known as the Detailed Real-Time Model (DRTM) was used to "train" and test the decision tree. Fifty-six DRTM simulations were performed for this purpose, with different leak sizes, different leak locations, and different times of leak onset. To make the simulations as realistic as possible, they included simulated sensor noise, and included a gradual degradation in both fuel and oxidizer turbine efficiency. A decision tree was trained using 11 of these simulations, and tested using the remaining 45 simulations. In the training phase, the C4.5 algorithm was provided with labeled examples of data from nominal operation and data including leaks in each leak location. From the data, it "learned" a decision tree that can classify unseen data as having no leak or having a leak in one of the five leak locations. In the test phase, the decision tree produced very low false alarm rates and low missed detection rates on the unseen data. It had very good fault isolation rates for three of the five simulated leak locations, but it tended to confuse the remaining two locations, perhaps because a large leak at one of these two locations can look very similar to a small leak at the other location.

Schwabacher, Mark A.↗

From Simulation to Reality With Random Noise

The challenging environment of autonomous vehicle (AV) navigation necessitates certain functions be performed by deep neural networks. Optimizing these models involves collecting vast quantities of domain-specific training data and ensuring that the dataset is representative of expected conditions. High-fidelity simulation plays a vital role in making this process feasible, allowing a wide range of scenarios to be explored at low cost. However, learning from simulation introduces subtle biases into models, which can degrade real-world performance in unpredictable ways. This effect can be mitigated with learning schemes specialized to bridge distributional shifts (transfer learning). Given the complex nature of these methods, the underlying models, and their environments, meaningfully evaluating performance is notstraight forward. Many unrelated factors can effect an improvement in generalization accuracy, but a full ablation analysis is often difficult. To tease out signal from noise, it is necessary to understand how transfer learning performance is affected by noise itself. The goals of this paper are (i) to establish a domain randomization baseline for a simple classification transfer learning task and (ii) to validate the RRAV testbed as a platform for further research in sim-to-real learning. We generate imagery from a simulation of NASA Ames Research Center and train a small convolutional neural network (ConvNet) to classify position relative to a centerline. Further models are trained with different types of noise progressively added to the data. The models are deployed aboard the on-site test vehicle to test real-world performance. In our experiments, we find that such naive domain randomization raises sim-to-real accuracy from 64% to 79%, while training directly on real data yields an 89% accuracy ceiling. These results suggest that the isolated mechanism of domain randomization can significantly improve generalization.

simulation↗

A New Machine Learning Based Analysis for Improving Satellite Retrieved Atmospheric Composition Data: OMI SO2 as an Example

Despite recent progress, satellite retrievals of anthropogenic SO2 still suffer from relatively low signal-tonoise ratios. In this study, we demonstrate a new machine learning data analysis method to improve the quality of satellite SO2 products. In the absence of large ground-truth datasets for SO2, we start from SO2 slant column densities (SCDs) retrieved from the Ozone Monitoring Instrument (OMI) using a data-driven, physically based algorithm and calculate the ratio between the SCD and the root mean square (rms) of the fitting residuals for each pixel. To build the training data, we select presumably clean pixels with small SCD / rms ratios (SRRs) and set their target SCDs to zero. For polluted pixels with relatively large SRRs, we set the target to the original retrieved SCDs. We then train neural networks (NNs) to reproduce the target SCDs using predictors including SRRs for individual pixels, solar zenith, viewing zenith and phase angles, scene reflectivity, and O3 column amounts, as well as the monthly mean SRRs. For data analysis, we employ two NNs: (1) one trained daily to produce analyzed SO2 SCDs for polluted pixels each day and (2) the other trained once every month to produce analyzed SCDs for less polluted pixels for the entire month. Test results for 2005 show that our method can significantly reduce noise and artifacts over background regions. Over polluted areas, the monthly mean NN-analyzed and original SCDs generally agree to within ±15 %, indicating that our method can retain SO2 signals in the original retrievals except for large volcanic eruptions. This is further confirmed by running both the NN-analyzed and original SCDs through a topdown emission algorithm to estimate the annual SO2 emissions for ∼ 500 anthropogenic sources, with the two datasets yielding similar results. We also explore two alternative approaches to the NN-based analysis method. In one, we employ a simple linear interpolation model to analyze the original SCD retrievals. In the other, we develop a PCA–NN algorithm that uses OMI measured radiances, transformed and dimension-reduced with a principal component analysis (PCA) technique, as inputs to NNs for SO2 SCD retrievals. While the linear model and the PCA–NN algorithm can reduce retrieval noise, they both underestimate SO2 over polluted areas. Overall, the results presented here demonstrate that our new data analysis method can significantly improve the quality of existing OMI SO2 retrievals. The method can potentially be adapted for other sensors and/or species and enhance the value of satellite data in air quality research and applications.

Can Li↗

Influence of soils on Landsat spectral signatures of corn

Landsat data have been investigated extensively to determine crop types and acreage. However, confounding site factors have been found to reduce accuracy. Soils data in a small, contiguous area in southeast South Dakota were used to stratify Landsat data. A June 5 and July 29 CCT were used in a statistical analysis of corn training data. Significant soil parameters causing differences in study area soils were slope and parent material. Implication of the results is that, in this region, stratification of CCT data along parent material boundaries would improve corn classification accuracy. Research expanding on the interaction of soils and crops is both in progress and scheduled for additional studies in east central South Dakota.

Dalsted, K. J.↗

Transfer of training on manual control systems differing in short period frequency and damping characteristics

Each of four groups of 16 subjects was trained on one of four compensatory tracking tasks that differed with regard to short period natural frequency and damping characteristics. After completion of the training sessions, the members of each group either transferred to a task on which they had not been trained or continued with their original task. Analysis of the training data indicated that relative task difficulty was largely determined by system damping which, however, had little effect on the amount of transfer during the transfer trials. The effect of system frequency was essentially reversed, and a marked interaction between training and transfer frequencies was observed in the transfer data. Similar results were obtained both with relative error scores and transinformation scores. Positive transfer was exhibited by most of the groups when they transferred to tasks on which they had not been trained.

Lincoln, R. S.↗

Tracking Historical NASA EVA Training: Lifetime Surveillance of Astronaut Health (LSAH) Development of the EVA Suit Exposure Tracker (EVA SET)

During a spacewalk, designated as extravehicular activity (EVA), an astronaut ventures from the protective environment of the spacecraft into the vacuum of space. EVAs are among the most challenging tasks during a mission, as they are complex and place the astronaut in a highly stressful environment dependent on the spacesuit for survival. Due to the complexity of EVA, NASA has conducted various training programs on Earth to mimic the environment of space and to practice maneuvers in a more controlled and forgiving environment. However, rewards offset the risks of EVA, as some of the greatest accomplishments in the space program were accomplished during EVA, such as the Apollo moonwalks and the Hubble Space Telescope repair missions. Water has become the environment of choice for EVA training on Earth, using neutral buoyancy as a substitute for microgravity. During EVA training, an astronaut wears a modified version of the spacesuit adapted for working in water. This high fidelity suit allows the astronaut to move in the water while performing tasks on full-sized mockups of space vehicles, telescopes, and satellites. During the early Gemini missions, several EVA objectives were much more difficult than planned and required additional time. Later missions demonstrated that "complex (EVA) tasks were feasible when restraints maintained body position and underwater simulation training ensured a high success probability".1,2 EVA training has evolved from controlling body positioning to perform basic tasks to complex maintenance of the Hubble Space Telescope and construction of the International Space Station (ISS). Today, preparation is centered at special facilities built specifically for EVA training, such as the Neutral Buoyancy Laboratory (NBL) at NASA's Johnson Space Center ([JSC], Houston) and the Hydrolab at the Gagarin Cosmonaut Training Centre ([GCTC], Star City, outside Moscow). Underwater training for an EVA is also considered hazardous duty for NASA astronauts. This activity places astronauts at risk for decompression sickness and barotrauma as well as various musculoskeletal disorders from working in the spacesuit. The medical, operational and research communities over the years have requested access to EVA training data to better understand the risks. As a result of these requests, epidemiologists within the Lifetime Surveillance of Astronaut Health (LSAH) team have compiled records from numerous EVA training venues to quantify the exposure to EVA training. The EVA Suit Exposure Tracker (EVA SET) dataset is a compilation of ground-based training activities using the extravehicular mobility unit (EMU) in neutrally buoyant pools to enhance EVA performance on orbit. These data can be used by the current ISS program and future exploration missions by informing physicians, researchers, and operational personnel on the risks of EVA training in order that future suit and mission designs incorporate greater safety. The purpose of this technical report is to document briefly the various facilities where NASA astronauts have performed EVA training while describing in detail the EVA training records used to generate the EVA SET dataset.

Laughlin, Mitzi S.↗

Integrated Cloud-Aerosol-Radiation Product using CERES, MODIS, CALIPSO and CloudSat Data

This paper documents the development of the first integrated data set of global vertical profiles of clouds, aerosols, and radiation using the combined NASA A-Train data from the Aqua Clouds and Earth's Radiant Energy System (CERES) and Moderate Resolution Imaging Spectroradiometer (MODIS), Cloud-Aerosol Lidar and Infrared Pathfinder Satellite Observations (CALIPSO), and CloudSat. As part of this effort, cloud data from the CALIPSO lidar and the CloudSat radar are merged with the integrated column cloud properties from the CERES-MODIS analyses. The active and passive datasets are compared to determine commonalities and differences in order to facilitate the development of a 3- dimensional cloud and aerosol dataset that will then be integrated into the CERES broadband radiance footprint. Preliminary results from the comparisons for April 2007 reveal that the CERES-MODIS global cloud amounts are, on average, 0.14 less and 0.15 greater than those from CALIPSO and CloudSat, respectively. These new data will provide unprecedented ability to test and improve global cloud and aerosol models, to investigate aerosol direct and indirect radiative forcing, and to validate the accuracy of global aerosol, cloud, and radiation data sets especially in polar regions and for multi-layered cloud conditions.

Sun-Mack, Sunny↗

Towards Reliable Evaluation of Anomaly-Based Intrusion Detection Performance

This report describes the results of research into the effects of environment-induced noise on the evaluation process for anomaly detectors in the cyber security domain. This research was conducted during a 10-week summer internship program from the 19th of August, 2012 to the 23rd of August, 2012 at the Jet Propulsion Laboratory in Pasadena, California. The research performed lies within the larger context of the Los Angeles Department of Water and Power (LADWP) Smart Grid cyber security project, a Department of Energy (DoE) funded effort involving the Jet Propulsion Laboratory, California Institute of Technology and the University of Southern California/ Information Sciences Institute. The results of the present effort constitute an important contribution towards building more rigorous evaluation paradigms for anomaly-based intrusion detectors in complex cyber physical systems such as the Smart Grid. Anomaly detection is a key strategy for cyber intrusion detection and operates by identifying deviations from profiles of nominal behavior and are thus conceptually appealing for detecting "novel" attacks. Evaluating the performance of such a detector requires assessing: (a) how well it captures the model of nominal behavior, and (b) how well it detects attacks (deviations from normality). Current evaluation methods produce results that give insufficient insight into the operation of a detector, inevitably resulting in a significantly poor characterization of a detectors performance. In this work, we first describe a preliminary taxonomy of key evaluation constructs that are necessary for establishing rigor in the evaluation regime of an anomaly detector. We then focus on clarifying the impact of the operational environment on the manifestation of attacks in monitored data. We show how dynamic and evolving environments can introduce high variability into the data stream perturbing detector performance. Prior research has focused on understanding the impact of this variability in training data for anomaly detectors, but has ignored variability in the attack signal that will necessarily affect the evaluation results for such detectors. We posit that current evaluation strategies implicitly assume that attacks always manifest in a stable manner; we show that this assumption is wrong. We describe a simple experiment to demonstrate the effects of environmental noise on the manifestation of attacks in data and introduce the notion of attack manifestation stability. Finally, we argue that conclusions about detector performance will be unreliable and incomplete if the stability of attack manifestation is not accounted for in the evaluation strategy.

cyber defense↗

Development of TEMPO Products and Tools to Support Air Quality Management Decisions

The TEMPO mission has been observing air pollutants every hour during the daytime across its Field of Regard (FoR) covering greater North America since First Light on August 2, 2023. The highly anticipated public release of TEMPO data occurred on May 20, 2024, consisting of level 2 and level 3 trace gas data products of nitrogen dioxide, formaldehyde, and ozone. Our project at the NASA SPoRT Center is developing value-added products and tools to support the TEMPO mission and Early Adopters program with special attention on the air quality management community. The initial focus of this project is evaluating the TEMPO products over stakeholder target areas using Pandora and surface monitor observations. Methods for oversampling TEMPO data to 1 km resolution are being applied over the target areas to resolve fine-scale emission sources and pollutant gradients. Machine learning techniques using TEMPO, surface monitor, and model data to estimate surface-level nitrogen dioxide concentrations are being developed over the target areas. Our SPoRT viewer has been updated to include visualizations of the TEMPO products and an ArcGIS dashboard is being designed for enabling air quality management stakeholders to efficiently analyze TEMPO data. Training materials including user guides are being developed to ensure the effective and sustained use of TEMPO data in air quality management applications. The major outcome of this project is to support the inclusion of TEMPO data in exceptional event demonstrations by active engagement with stakeholders and ultimately enable more informed air quality management decisions in the future. This talk will provide an update on our project activities and showcase use cases of TEMPO data for monitoring different emission sources including wildland fire smoke.

air quality↗

Transitioning TEMPO Data for Air Quality Management Applications at the NASA SPoRT Center

The TEMPO mission has been observing air pollutants every hour during the daytime across its Field of Regard (FoR) covering greater North America since First Light on August 2, 2023. The highly anticipated public release of TEMPO data occurred on May 20, 2024, consisting of level 2 and level 3 trace gas data products of nitrogen dioxide, formaldehyde, and ozone. The NASA SPoRT Center is developing value-added products and tools to support the TEMPO mission and Early Adopters program with special attention on stakeholder from air agencies. One component of our work is focused on evaluating the TEMPO products over stakeholder target areas using Pandora and surface monitor observations to characterize the uncertainties and develop best practices for processing and analyzing TEMPO data. Methods for oversampling TEMPO data to 1 km resolution are being applied over the target areas to resolve fine-scale emission sources and pollutant gradients. Machine learning techniques using TEMPO, surface monitor, and model data to estimate surface-level nitrogen dioxide concentrations are being developed over the target areas. A SPoRT viewer for TEMPO has been launched for providing visualizations of the TEMPO products and an ArcGIS dashboard is being designed for enabling air agency stakeholders to efficiently analyze TEMPO data. Training materials including user guides are being developed to ensure the effective and sustained use of TEMPO data at our stakeholder agencies. One major goal of our TEMPO initiatives at SPoRT is to better enable the inclusion of TEMPO data in air quality management applications such as exceptional event demonstrations through our close engagement with stakeholders. This talk will provide an update on our SPoRT activities and showcase use cases of TEMPO data for monitoring different emission sources including wildland fire smoke.

Air Quality↗

A function approximation approach to anomaly detection in propulsion system test data

Ground test data from propulsion systems such as the Space Shuttle Main Engine (SSME) can be automatically screened for anomalies by a neural network. The neural network screens data after being trained with nominal data only. Given the values of 14 measurements reflecting external influences on the SSME at a given time, the neural network predicts the expected nominal value of a desired engine parameter at that time. We compared the ability of three different function-approximation techniques to perform this nominal value prediction: a novel neural network architecture based on Gaussian bar basis functions, a conventional back propagation neural network, and linear regression. These three techniques were tested with real data from six SSME ground tests containing two anomalies. The basis function network trained more rapidly than back propagation. It yielded nominal predictions with, a tight enough confidence interval to distinguish anomalous deviations from the nominal fluctuations in an engine parameter. Since the function-approximation approach requires nominal training data only, it is capable of detecting unknown classes of anomalies for which training data is not available.

Whitehead, Bruce A.↗

Teaching artificial neural systems to drive: Manual training techniques for autonomous systems

A methodology was developed for manually training autonomous control systems based on artificial neural systems (ANS). In applications where the rule set governing an expert's decisions is difficult to formulate, ANS can be used to extract rules by associating the information an expert receives with the actions taken. Properly constructed networks imitate rules of behavior that permits them to function autonomously when they are trained on the spanning set of possible situations. This training can be provided manually, either under the direct supervision of a system trainer, or indirectly using a background mode where the networks assimilates training data as the expert performs its day-to-day tasks. To demonstrate these methods, an ANS network was trained to drive a vehicle through simulated freeway traffic.

Shepanski, J. F.↗

Estimating Helicopter Noise Abatement Information with Machine Learning

Machine learning techniques are applied to the NASA Langley Research Center's expansive database of helicopter noise measurements containing over 1500 steady flight conditions for ten different helicopters. These techniques are then used to develop models capable of predicting the operating conditions under which significant Blade-Vortex Interaction noise will be generated for any conventional helicopter. A measure for quantifying the overall ground noise exposure of a particular helicopter operating condition is developed. This measure is then used to classify the measured flight conditions as noisy or not-noisy. These data are then parameterized on a nondimensional basis that defines the main rotor operating condition and are then scaled to remove bias. Several machine learning methods are then applied to these data. The developed models show good accuracy in identifying the noisy operating region for helicopters not included in the training data set. Noisy regions are accurately identified for a variety of different helicopters. One of these models is applied to estimate changes in the noisy operating region as vehicle drag and ambient atmospheric conditions are varied.

Greenwood, Eric↗

Mass Inferencing Model Creation and Deployment to the RASSOR Lunar Excavation Robot

The Regolith Advanced Surface Systems Operations Robot (RASSOR) Excavator is a teleoperated mobile robotic platform with a unique space regolith excavation capability. The Intelligent Capabilities Enhanced RASSOR research project developed functionality for inferencing regolith mass ingested during RASSOR operation, enhancing RASSOR’s ability to successfully complete ISRU missions. To teleoperate or run autonomously, it is crucial for the quantity of regolith mass ingested by RASSOR to be available as a system state for efficient operation. For example, during autonomous operation, RASSOR should navigate and move to a processing plant to offload the collected regolith when the drums are full; without knowledge of how much mass is in the drums, this type of high-level planning is not possible. Four distinct modeling approaches were employed in developing a mass inferencing approach that could work on RASSOR. All take in system states, such as arm/drum positions, velocities, currents, voltages, and robot pose, and output a mass prediction for each set of the robot’s bucket drums.1) A neural network model that takes a vector of normalized system states; 2) A model that uses the integrated power consumption of an arm-raise (normalized by velocity); 3) A model that uses average drum current over a variable length interval of the drum disengaged from the surface; and 4) A real-time estimation model that aggregates excavation drum current. The developed models run in real time, outputting predictions for the front and rear drums, timestamp of the last prediction, and total mass in RASSOR’s drums. Further testing is required to validate the arm-raise model (2), though initial tests indicate reasonable performance (<10% mean error) on the hardware. The linear fit of average drum-current model (3) had a front value of r^2=0.99 and a rear value of r^2=0.98 on the validation dataset. This model currently has the best performance on unseen data. The real time model (4) is still in development, though initial results on a small subset of the training data show that it has high accuracy in predicting the increase in mass during excavation. Though work remains to be done with deploying a high-fidelity model to the physical system that makes predictions with error below the desired threshold, the modular architecture for model development allows quick adjustment of parameters to increase model fidelity. This architecture can also be adapted to use lunar excavation data to create models that are reflective of RASSOR’s dynamics when operating on the lunar surface. The results are promising as it has been shown that models can be developed that accurately estimate excavated regolith mass.

rassor↗

Metric Learning for Hyperspectral Image Segmentation

We present a metric learning approach to improve the performance of unsupervised hyperspectral image segmentation. Unsupervised spatial segmentation can assist both user visualization and automatic recognition of surface features. Analysts can use spatially-continuous segments to decrease noise levels and/or localize feature boundaries. However, existing segmentation methods use tasks-agnostic measures of similarity. Here we learn task-specific similarity measures from training data, improving segment fidelity to classes of interest. Multiclass Linear Discriminate Analysis produces a linear transform that optimally separates a labeled set of training classes. The defines a distance metric that generalized to a new scenes, enabling graph-based segmentation that emphasizes key spectral features. We describe tests based on data from the Compact Reconnaissance Imaging Spectrometer (CRISM) in which learned metrics improve segment homogeneity with respect to mineralogical classes.

Compact Reconnaissance Imaging Spectrometer (CRISM↗

Variability of wetland reflectance and its effect on automatic categorization of satellite imagery

The author has identified the following significant results. Land cover categorization of data from the same overpass in four test wetland areas was carried out using a four category classification system. The tests indicate that training data based on in situ reflectance measurements and atmospheric correction of LANDSAT data can produce comparable accuracy of categorization to that achieved using more than four wetlands cover categories (salt marsh cordgrass, salt hay, unvegetated, and water tidal flat) produced overall classification accuracies of 85% by conventional and relative radiance training and 81% by use of in situ measurements. Overall mapping accuracies were 76% and 72% respectively.

Klemas, V.↗

Classification results using spacially correlated Landsat data

Tubbs and Coberly (1978) demonstrated that Landsat multispectral scanner data are not independent random observations, but, are in fact highly correlated. They also demonstrated that the correlation structure for the data is similar to that of a stationary autoregressive process of order one. This paper investigates the effect that serially correlated training data have upon both the estimation of parameters and the classification problem. Results are included for both the Bayesian and maximum likelihood classification procedures.

Tubbs, J. D.↗

Automated Cardiovascular Pathology Assessment using Semantic Segmentation and Ensemble Learning

Cardiac magnetic resonance imaging provides high spatial resolution, enabling improved extraction of important functional and morphological features for cardiovascular disease staging. Segmentation of ventricular cavities and myocardium in cardiac cine sequencing provides a basis to quantify cardiac measures such as ejection fraction. A method is presented that curtails the expense and observer bias of manual cardiac evaluation by combining semantic segmentation and disease classification into a fully automatic processing pipeline. The initial processing element consists of a robust dilated convolutional neural network architecture for voxel-wise segmentation of the myocardium and ventricular cavities. The resulting comprehensive volumetric feature matrix captures diagnostic clinical procedure data and is utilized by the final processing element to model a cardiac pathology classifier. Our approach evaluated anonymized cardiac images from a training data set of 100 patients (4 pathology groups, 1 healthy group, 20 patients per group) examined at the University Hospital of Dijon. The top average Dice index scores achieved were 0.940, 0.886, 0.849 for structure segmentation of the left ventricle (LV), myocardium and right ventricle (RV) respectively. A 5-ary pathology classification accuracy of 90% was recorded on an independent test set using the trained model. Performance results demonstrate potential for advanced machine learning methods to deliver accurate, efficient and reproducible cardiac pathological assessment.

Semantic Segmentation↗