Search NASA⌕ Search

SEARCH · Search NASA

Results for “machine data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

A New Machine Learning Based Analysis for Improving Satellite Retrieved Atmospheric Composition Data: OMI SO2 as an Example

Despite recent progress, satellite retrievals of anthropogenic SO2 still suffer from relatively low signal-tonoise ratios. In this study, we demonstrate a new machine learning data analysis method to improve the quality of satellite SO2 products. In the absence of large ground-truth datasets for SO2, we start from SO2 slant column densities (SCDs) retrieved from the Ozone Monitoring Instrument (OMI) using a data-driven, physically based algorithm and calculate the ratio between the SCD and the root mean square (rms) of the fitting residuals for each pixel. To build the training data, we select presumably clean pixels with small SCD / rms ratios (SRRs) and set their target SCDs to zero. For polluted pixels with relatively large SRRs, we set the target to the original retrieved SCDs. We then train neural networks (NNs) to reproduce the target SCDs using predictors including SRRs for individual pixels, solar zenith, viewing zenith and phase angles, scene reflectivity, and O3 column amounts, as well as the monthly mean SRRs. For data analysis, we employ two NNs: (1) one trained daily to produce analyzed SO2 SCDs for polluted pixels each day and (2) the other trained once every month to produce analyzed SCDs for less polluted pixels for the entire month. Test results for 2005 show that our method can significantly reduce noise and artifacts over background regions. Over polluted areas, the monthly mean NN-analyzed and original SCDs generally agree to within ±15 %, indicating that our method can retain SO2 signals in the original retrievals except for large volcanic eruptions. This is further confirmed by running both the NN-analyzed and original SCDs through a topdown emission algorithm to estimate the annual SO2 emissions for ∼ 500 anthropogenic sources, with the two datasets yielding similar results. We also explore two alternative approaches to the NN-based analysis method. In one, we employ a simple linear interpolation model to analyze the original SCD retrievals. In the other, we develop a PCA–NN algorithm that uses OMI measured radiances, transformed and dimension-reduced with a principal component analysis (PCA) technique, as inputs to NNs for SO2 SCD retrievals. While the linear model and the PCA–NN algorithm can reduce retrieval noise, they both underestimate SO2 over polluted areas. Overall, the results presented here demonstrate that our new data analysis method can significantly improve the quality of existing OMI SO2 retrievals. The method can potentially be adapted for other sensors and/or species and enhance the value of satellite data in air quality research and applications.

Can Li↗

Flux Improvement based on Machine Learning for the CERES FluxByCldTyp Data Product

The NASA Clouds and the Earth's Radiant Energy System (CERES) product provides over 20 years of accurately observed top-of-the-atmosphere (TOA) and surface flux data record for climate monitoring and diagnostic studies. The interaction between clouds and radiation interaction is a key factor that dominate climate feedbacks but is not well understood. To further advance our understanding of the cloud-radiation interaction, a new CERES FluxByCldTyp (FBCT) product has been developed that contains radiative fluxes by cloud-type, which can provide more stringent constraints when validating models and reveal more insight into the interactions between clouds and climate. For CERES partly cloudy and multiple cloud-type footprints, the FBCT product utilizes Moderate Resolution Imaging Spectroradiometer (MODIS) narrow-band (NB) imager channel radiances partitioned by cloud-type within a CERES footprint to estimate the cloud-type broadband fluxes. The MODIS multi-channel derived broadband fluxes were compared with the CERES observed footprint fluxes and were found to be within 1% and 2.5% for LW and SW, respectively, as well as being mostly free of cloud property dependencies. The FBCT all-sky and clear-sky monthly averaged fluxes were found to be consistent with the CERES SSF1deg product. This study takes advantage of recent progress in machine learning (ML) field by applying deep neural network algorithm to improve fluxes based on MODIS NB radiances. The preliminary study shows ML produce are an improvement over the current FBCT Edition 4 NB2BB algorithm. Furthermore, unlike Ed4 NB2BB, the new ML method convert NB radiances directly to broadband fluxes. For future Ed5, new NB radiances are proposed and used by ML to improve fluxes calculation. Prelimary results show significant LW improvement.

Moguo Sun↗

Mapping soils, crops, and rangelands by machine analysis of multitemporal ERTS-1 data

ERTS-1 data, obtained during the period 25 August 1972 to 5 September 1973 over a range of test sites in the Central United States, have been used for identifying and mapping differences in soil patterns, species and conditions of cultivated crops, and conditions of rangelands. Multispectral scanner data from multiple ERTS passes over certain test sites have provided the opportunity to study temporal changes in the scene. Multispectral classifications delineating soils boundaries in different test sites compared well with existing soil association maps prepared by conventional means. Spectral analysis of ERTS data was used to identify, maps, and make areal measurements of wheat in western Kansas. Multispectral analysis of ERTS-1 data provided patterns in rangelands which can be related to soils differences, range management practices, and the extent of infestation of grasslands by mesquite (prosopis fuliflora) and juniper (juniperus spp.).

Baumgardner, M. F.↗

Snow cover monitoring by machine processing of multitemporal LANDSAT MSS data

LANDSAT frames were geometrically corrected and data sets from six different dates were overlaid to produce a 24 channel (six dates and four wavelength bands) data tape. Changes in the extent of the snowpack could be accurately and easily determined using a change detection technique on data which had previously been classified by the LARSYS software system. A second phase of the analysis involved determination of the relationship between spatial resolution or data sampling frequency and accuracy of measuring the area of the snowpack.

Luther, S. G.↗

Automatic Data Traffic Control on DSM Architecture

We study data traffic on distributed shared memory machines and conclude that data placement and grouping improve performance of scientific codes. We present several methods which user can employ to improve data traffic in his code. We report on implementation of a tool which detects the code fragments causing data congestions and advises user on improvements of data routing in these fragments. The capabilities of the tool include deduction of data alignment and affinity from the source code; detection of the code constructs having abnormally high cache or TLB misses; generation of data placement constructs. We demonstrate the capabilities of the tool on experiments with NAS parallel benchmarks and with a simple computational fluid dynamics application ARC3D.

Frumkin, Michael↗

Confidence-Based Feature Acquisition

Confidence-based Feature Acquisition (CFA) is a novel, supervised learning method for acquiring missing feature values when there is missing data at both training (learning) and test (deployment) time. To train a machine learning classifier, data is encoded with a series of input features describing each item. In some applications, the training data may have missing values for some of the features, which can be acquired at a given cost. A relevant JPL example is that of the Mars rover exploration in which the features are obtained from a variety of different instruments, with different power consumption and integration time costs. The challenge is to decide which features will lead to increased classification performance and are therefore worth acquiring (paying the cost). To solve this problem, CFA, which is made up of two algorithms (CFA-train and CFA-predict), has been designed to greedily minimize total acquisition cost (during training and testing) while aiming for a specific accuracy level (specified as a confidence threshold). With this method, it is assumed that there is a nonempty subset of features that are free; that is, every instance in the data set includes these features initially for zero cost. It is also assumed that the feature acquisition (FA) cost associated with each feature is known in advance, and that the FA cost for a given feature is the same for all instances. Finally, CFA requires that the base-level classifiers produce not only a classification, but also a confidence (or posterior probability).

Wagstaff, Kiri L.↗

Exploring the Utility of Machine Learning-Based Passive Microwave Brightness Temperature Data Assimilation over Terrestrial Snow in High Mountain Asia

This study explores the use of a support vector machine (SVM) as the observation operator within a passive microwave brightness temperature data assimilation framework (herein SVM-DA) to enhance the characterization of snow water equivalent (SWE) over High Mountain Asia (HMA). A series of synthetic twin experiments were conducted with the NASA Land Information System (LIS) at a number of locations across HMA. Overall, the SVM-DA framework is effective at improving SWE estimates (~70% reduction in RMSE relative to the Open Loop) for SWE depths less than 200 mm during dry snowpack conditions. The SVM-DA framework also improves SWE estimates in deep, wet snow (~45% reduction in RMSE) when snow liquid water is well estimated by the land surface model, but can lead to model degradation when snow liquid water estimates diverge from values used during SVM training. In particular, two key challenges of using the SVM-DA framework were observed over deep, wet snowpacks. First, variations in snow liquid water content dominate the brightness temperature spectral difference (TB) signal associated with emission from a wet snowpack, which can lead to abrupt changes in SWE during the analysis update. Second, the ensemble of SVM-based predictions can collapse (i.e., yield a near-zero standard deviation across the ensemble) when prior estimates of snow are outside the range of snow inputs used during the SVM training procedure. Such a scenario can lead to the presence of spurious error correlations between SWE and TB, and as a consequence, can result in degraded SWE estimates from the analysis update. These degraded analysis updates can be largely mitigated by applying rule-based approaches. For example, restricting the SWE update when the standard deviation of the predicted TB is greater than 0.05 K helps prevent the occurrence of filter divergence. Similarly, adding a thin layer (i.e., 5 mm) of SWE when the synthetic TB is larger than 5 K can improve SVM-DA performance in the presence of a precipitation dry bias. The study demonstrates that a carefully constructed SVM-DA framework cognizant of the inherent limitations of passive microwave-based SWE estimation holds promise for snow mass data assimilation.

Kwon, Yonghwan↗

CyberGAN: Generating High-fidelity Cybersecurity Data With Generative Adversarial Networks

Machine learning for cyber defense offers the promise of detecting adversarial activity against the ground data systems managing critical space assets. A fundamental challenge facing machine learning research in cybersecurity is the lack of high-fidelity, shareable datasets for robust evaluation and testing of machine learning-based solutions. High-fidelity, real-world datasets are necessary for reliable benchmarking of nominal system behavior and malicious activity. Unfortunately, such realistic datasets of both nominal and adversarial activity are rarely shared publicly by data owners due to security and privacy concerns. Besides, the available adversarial data is sparse, which makes training models on malicious activity much harder. This situation has impeded and continues to impede the research and successful adoption of machine learning methods for cyber defense. Researchers have dealt with this problem by generating data within a low-fidelity lab environment, using classified and thus unshareable datasets, or downloading low-fidelity public datasets made available by others. We propose an innovative solution to the problem by employing machine learning methods to generate high-fidelity data. Specifically, we propose the use of Generative Adversarial Networks (GANs) to generate high-fidelity data for cybersecurity purposes. GANs have found successful image processing and natural language applications, but have not yet been investigated for cyber data generation. Our proposed approach first involves training the `discriminator' network of the GAN with a sample of real-world data consisting of malicious and nominal samples. We then use the `generator' network to generate new high-fidelity data samples consisting of an appropriate mix of malicious and nominal activity. We demonstrate applications of our architecture by generating high-fidelity cybersecurity data containing both malicious and nominal samples. We thoroughly evaluate the fidelity of our generated data using heuristics and evaluate its usefulness for machine learning applications using three different datasets. Overall, our approach results in high-fidelity, shareable datasets.

Zhang, Yuening↗

Evaluation of human exposure to the noise from large wind turbine generators

The human perception of a nuisance level of noise was quantified in tests and attempts were made to define criteria for acceptable sound levels from wind turbines. Comparisons were made between the sound necessary to cause building vibration, which occurred near the Mod-1 wind turbine, and human perception thresholds for building noise and building vibration. Thresholds were measured for both broadband and impulsive noise, with the finding that noise in the 500-2000 Hz region, and impulses with a 1 Hz fundamental, were most noticeable. Curves were developed for matching a receiver location with expected acoustic output from a machine to determine if the sound levels were offensive. In any case, further data from operating machines are required before definitive criteria can be established.

Shepherd, K. P.↗

Implementing nested conditional statements in SIMD machines

Single instruction, multiple data (SIMD) computers consist of a very large number of processors executing a common sequence of instructions. Maintaining the full speedup potential of such machines is most sensitive to conditional execution in their programs, regions of code where some processing elements (PEs) perform no useful work. Techniques are presented for efficiently implementing nested conditional statements, specifically if and case statements, in SIMD machines, while adding minimal specialized hardware.

Middleton, David↗

High-Performance Wireless Telemetry

Prior technology for machinery data acquisition used slip rings, FM radio communication, or non-real-time digital communication. Slip rings are often noisy, require much space that may not be available, and require access to the shaft, which may not be possible. FM radio is not accurate or stable, and is limited in the number of channels, often with channel crosstalk, and intermittent as the shaft rotates. Non-real-time digital communication is very popular, but complex, with long development time, and objections from users who need continuous waveforms from many channels. This innovation extends the amount of information conveyed from a rotating machine to a data acquisition system while keeping the development time short and keeping the rotating electronics simple, compact, stable, and rugged. The data are all real time. The product of the number of channels, times the bit resolution, times the update rate, gives a data rate higher than available by older methods. The telemetry system consists of a data-receiving rack that supplies magnetically coupled power to a rotating instrument amplifier ring in the machine being monitored. The ring digitizes the data and magnetically couples the data back to the rack, where it is made available. The transformer is generally a ring positioned around the axis of rotation with one side of the transformer free to rotate and the other side held stationary. The windings are laid in the ring; this gives the data immunity to any rotation that may occur. A medium-frequency sine-wave power source in a rack supplies power through a cable to a rotating ring transformer that passes the power on to a rotating set of electronics. The electronics power a set of up to 40 sensors and provides instrument amplifiers for the sensors. The outputs from the amplifiers are filtered and multiplexed into a serial ADC. The output from the ADC is connected to another rotating ring transformer that conveys the serial data from the rotating section to the stationary section. From there, a cable conveys the serial data to the remote rack, where it is reconditioned to logic level specifications, de-serialized, and converted back to analog. In the rotating electronics are code generators to indicate the beginning of files for data synchronization.

Griebeler, Elmer↗