Search NASA⌕ Search

SEARCH · Search NASA

Results for “machine data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Quantum Leap: Evaluating the Feasibility of Quantum Machine Learning Using NASA Earth Observational Data

This study explores the feasibility of leveraging quantum machine learning (QML) to analyze NASA Earth Observational (EO) data for climate change research, with a particular focus on the phenomenon of ”crop frosting” which has become more prevalent due to climate change. We implemented and evaluated two QML models, the Variational Quantum Classifier (VQC) and Quantum Support Vector Classifier (QSVC), in both simulated and real quantum computing environments using a 127 qubit IBM quantum processor. Our study emphasizes the scientific rigor in comparing these quantum models with a classical Support Vector Machine (SVM) classifier, highlighting their performance in processing climate data. The results offer valuable insights into the potential scientific advantages, limitations, and scalability of QML for analyzing EO datasets, thus paving the way for more advanced climate modeling and predictive analytics using quantum computing. We showcased how Environmental Interaction Knowledge Graphs (EIKGs) and Digital Twins (DTs) can be integrated into this study. This research underscores the transformative potential of Classical and QML leveraging KGs and DT to address the multifaceted challenges posed by climate change.

Quantum Computing↗

Machine processing of remotely sensed data; Proceedings of the Fifth Annual Symposium, Purdue University, West Lafayette, Ind., June 27-29, 1979

Papers are presented on techniques and applications for the machine processing of remotely sensed data. Specific topics include the Landsat-D mission and thematic mapper, data preprocessing to account for atmospheric and solar illumination effects, sampling in crop area estimation, the LACIE program, the assessment of revegetation on surface mine land using color infrared aerial photography, the identification of surface-disturbed features through a nonparametric analysis of Landsat MSS data, the extraction of soil data in vegetated areas, and the transfer of remote sensing computer technology to developing nations. Attention is also given to the classification of multispectral remote sensing data using context, the use of guided clustering techniques for Landsat data analysis in forest land cover mapping, crop classification using an interactive color display, and future trends in image processing software and hardware.

Tendam, I. M.↗

Conjugate point determination for multitemporal data overlay

The machine processing is discussed of spatially variant multitemporal data such as imagery obtained at different times which requires that these data be in geometrical registration so that the analysis processor may obtain the datum for a specified ground resolution element in each of the sets of imagery being utilized for analysis. Misregistration between corresponding subsets of imagery contains both a displacement and a geometrical distortion component. Search techniques utilizing the moduli of the Fourier Transforms are developed for estimating the coefficients of geometrical distortion components of this model. Following the correction of these distortion components, the displacement is located by the crosscorrelation of a template obtained from one set of data, termed the reference, with the second, or background data. This template, derived for the optimum discrimination of the reference data embedded in the background, is determined by the solution of a system of equations involving the reference data and the covariance matrix of these data.

Emmert, R. A.↗

Communication Studies of DMP and SMP Machines

Understanding the interplay between machines and problems is key to obtaining high performance on parallel machines. This paper investigates the interplay between programming paradigms and communication capabilities of parallel machines. In particular, we explicate the communication capabilities of the IBM SP-2 distributed-memory multiprocessor and the SGI PowerCHALLENGEarray symmetric multiprocessor. Two benchmark problems of bitonic sorting and Fast Fourier Transform are selected for experiments. Communication-efficient algorithms are developed to exploit the overlapping capabilities of the machines. Programs are written in Message-Passing Interface for portability and identical codes are used for both machines. Various data sizes and message sizes are used to test the machines' communication capabilities. Experimental results indicate that the communication performance of the multiprocessors are consistent with the size of messages. The SP-2 is sensitive to message size but yields a much higher communication overlapping because of the communication co-processor. The PowerCHALLENGEarray is not highly sensitive to message size and yields a low communication overlapping. Bitonic sorting yields lower performance compared to FFT due to a smaller computation-to-communication ratio.

Sohn, Andrew↗

Machine-Learned Committor Functions for Reactive Molecular Dynamics

Reactive molecular dynamics (MD) is a powerful tool for atomistic-scale modeling of a diverse range of chemical processes. However, scaling these simulations to large systems and long times scales remains a challenge because of the complexity of the potential energy function required. The authors previously developed a heuristic approach, called REACTER, that incorporates reactivity in MD simulations in a less general but much more computationally efficient manner. REACTER uses standard, fixed valence force fields as the underlying potentialenergy surface for describing all interatomic interactions but adds a procedure for enforcing user-defined reactions that occur when certain geometric constraints on relative atomic positions are satisfied. Further, these bonding changes can be accepted or rejected with a probability related tothe local thermal energy. This work seeks to generalize this approach by replacing the set of user defined geometric constraints and energetic criteria with a committor function that specifies the probability of a reaction occurring on the basis of the local atomic configuration. The committor function is a useful mathematical tool for modeling rare events but, unfortunately, is very difficult to compute for realistic systems in a general way. This work describes a method for approximating the committor function using a machine learning approach, specifically a deep neural network trained with data from reactive MD and DFT-based dynamics simulations. This network is coupled to the existing REACTER protocol, as implemented in the LAMMPS MD package, and used to make on-the-fly predictions of reaction probabilities without the more extensive user input previously required. The new method is demonstrated using the polymerization of polystyrene as a case study. Although very dependent on the quality and quantity of training data, machine-learned committor functions show promise as a method for incorporating reaction probability from higher level calculations into highly scalable MD simulations.

polymer simulations↗

Machine processing of remotely sensed data; Proceedings of the Conference, Purdue University, West Lafayette, Ind., October 16-18, 1973

Topics discussed include the management and processing of earth resources information, special-purpose processors for the machine processing of remotely sensed data, digital image registration by a mathematical programming technique, the use of remote-sensor data in land classification (in particular, the use of ERTS-1 multispectral scanning data), the use of remote-sensor data in geometrical transformations and mapping, earth resource measurement with the aid of ERTS-1 multispectral scanning data, the use of remote-sensor data in the classification of turbidity levels in coastal zones and in the identification of ecological anomalies, the problem of feature selection and the classification of objects in multispectral images, the estimation of proportions of certain categories of objects, and a number of special systems and techniques. Individual items are announced in this issue.

Source record↗

Towards a program of record of inland water quality: Exploiting present and heritage multispectral sensors for maximum information extraction

Degradation of Earth’s inland water resources due to anthropogenic perturbations and climate anomalies at both local and global scales continues to place human health at substantial risk. There is now a growing necessity to develop pragmatic approaches that allow timely and effective extrapolation of local processes, to spatially resolved global products, and to promote operational and sustainable resource policy management. This research exploits recent advancements in bio-optical modeling, cloud computing, and machine learning to enhance our capacity to leverage present and heritage satellite data. Recent research suggests that sensors with low spectral resolution, such as Sentinel 2 and Landsat missions, contain enough hidden spectral variation which can be exploited using data-driven approaches. The availability of three decades of archival imagery will open doors to discover global trends of eutrophication and increased cyanobacteria dominance and provide valuable insight to the development of predictive methodologies. Preliminary efforts in synthetic emulation of global natural inland waters will be discussed and contextualized against satellite radiometric measurement uncertainty, satellite data product uncertainties and causal signal ambiguity over the visible wavelength range, supported by high quality field and image data for selected inland aquatic sites. Insights on water quality estimation via data-driven machine learning models versus matrix inversions will be discussed, and how we can exploit spectral-spatial relationships in high spatial resolution data. A cross-sensor synergistic approach with detailed uncertainty analysis based on optical water types, will allow for unprecedented global snapshots of fine scale ecological dynamics of inland waters.

Inland↗

Advancing NASA SatCORPS Global Data Products with Cloud Computing and Machine Learning

Operational satellite imager radiances are valuable for deriving many different physical parameters that can be used for a variety of weather, aviation, and energy applications. The NASA Satellite ClOud and Radiation Property retrieval System (SatCORPS) applies a suite of algorithms to meteorological satellite data to provide cloud properties, radiative fluxes and other parameters on a global scale. The use of cloud computing has enabled recent enhancements to process a constellation of geostationary satellites at higher spatiotemporal resolutions than previously possible that meet low latency and near real-time needs. Data taken from Meteosat-8 and -11, Himawari, GOES-16, and -17 are processed and combined with operational polar orbiting satellite data and composited on a 3-km grid to provide global coverage. To improve the utility of the data products, machine learning and other innovative methods are applied in various ways to help minimize data product uncertainties under the most challenging conditions and to improve their consistency at all times of day. An update on recent SatCORPS enhancements is presented, highlighting the community benefits achieved with the use of cloud computing and machine learning.

William L Smith↗

Concurrent Image Processing Executive (CIPE). Volume 1: Design overview

The design and implementation of a Concurrent Image Processing Executive (CIPE), which is intended to become the support system software for a prototype high performance science analysis workstation are described. The target machine for this software is a JPL/Caltech Mark 3fp Hypercube hosted by either a MASSCOMP 5600 or a Sun-3, Sun-4 workstation; however, the design will accommodate other concurrent machines of similar architecture, i.e., local memory, multiple-instruction-multiple-data (MIMD) machines. The CIPE system provides both a multimode user interface and an applications programmer interface, and has been designed around four loosely coupled modules: user interface, host-resident executive, hypercube-resident executive, and application functions. The loose coupling between modules allows modification of a particular module without significantly affecting the other modules in the system. In order to enhance hypercube memory utilization and to allow expansion of image processing capabilities, a specialized program management method, incremental loading, was devised. To minimize data transfer between host and hypercube, a data management method which distributes, redistributes, and tracks data set information was implemented. The data management also allows data sharing among application programs. The CIPE software architecture provides a flexible environment for scientific analysis of complex remote sensing image data, such as planetary data and imaging spectrometry, utilizing state-of-the-art concurrent computation capabilities.

Lee, Meemong↗

Concurrent Image Processing Executive (CIPE)

The design and implementation of a Concurrent Image Processing Executive (CIPE), which is intended to become the support system software for a prototype high performance science analysis workstation are discussed. The target machine for this software is a JPL/Caltech Mark IIIfp Hypercube hosted by either a MASSCOMP 5600 or a Sun-3, Sun-4 workstation; however, the design will accommodate other concurrent machines of similar architecture, i.e., local memory, multiple-instruction-multiple-data (MIMD) machines. The CIPE system provides both a multimode user interface and an applications programmer interface, and has been designed around four loosely coupled modules; (1) user interface, (2) host-resident executive, (3) hypercube-resident executive, and (4) application functions. The loose coupling between modules allows modification of a particular module without significantly affecting the other modules in the system. In order to enhance hypercube memory utilization and to allow expansion of image processing capabilities, a specialized program management method, incremental loading, was devised. To minimize data transfer between host and hypercube a data management method which distributes, redistributes, and tracks data set information was implemented.

Lee, Meemong↗

Detecting Abnormal Machine Characteristics in Cloud Infrastructures

In the cloud computing environment resources are accessed as services rather than as a product. Monitoring this system for performance is crucial because of typical pay-peruse packages bought by the users for their jobs. With the huge number of machines currently in the cloud system, it is often extremely difficult for system administrators to keep track of all machines using distributed monitoring programs such as Ganglia1 which lacks system health assessment and summarization capabilities. To overcome this problem, we propose a technique for automated anomaly detection using machine performance data in the cloud. Our algorithm is entirely distributed and runs locally on each computing machine on the cloud in order to rank the machines in order of their anomalous behavior for given jobs. There is no need to centralize any of the performance data for the analysis and at the end of the analysis, our algorithm generates error reports, thereby allowing the system administrators to take corrective actions. Experiments performed on real data sets collected for different jobs validate the fact that our algorithm has a low overhead for tracking anomalous machines in a cloud infrastructure.

Bhaduri, Kanishka↗

The Social Life of a Data Base

This paper presents the complex social life of a large data base. The topics include: 1) Social Construction of Mechanisms of Memory; 2) Data Bases: The Invisible Memory Mechanism; 3) The Human in the Machine; 4) Data of the Study: A Large-Scale Problem Reporting Data Base; 5) The PRACA Study; 6) Description of PRACA; 7) PRACA and Paper; 8) Multiple Uses of PRACA; 9) The Work of PRACA; 10) Multiple Forms of Invisibility; 11) Such Systems are Everywhere; and 12) Two Morals to the Story. This paper is in viewgraph form.

Linde, Charlotte↗

Probabilistic Modeling of Heavy Machinery Using Machine Learning

NASA Glenn Research Center’s facility operations seeks to leverage its extensive instrumentation and historical data with machine learning to increase system efficiencies. The ultimate goal of this effort is to probabilistically model the behavior of Glenn’s central air service compressors for optimal decision making and planning. This project is a first step in that direction. We propose a multimodel approach that uses high-dimensional models to ask simple questions about complex dependent structures, and low-dimensional models to ask complex questions about simple dependent structures. While the low-dimensional models make strong assumptions, they can be visualized and they can be insightful. We show good fits for univariate models of compressor sensors, and preliminary work on high-dimensional multivariate models.

Machine learning↗

Validating a large geophysical data set: Experiences with satellite-derived cloud parameters

We are validating the global cloud parameters derived from the satellite-borne HIRS2 and MSU atmospheric sounding instrument measurements, and are using the analysis of these data as one prototype for studying large geophysical data sets in general. The HIRS2/MSU data set contains a total of 40 physical parameters, filling 25 MB/day; raw HIRS2/MSU data are available for a period exceeding 10 years. Validation involves developing a quantitative sense for the physical meaning of the derived parameters over the range of environmental conditions sampled. This is accomplished by comparing the spatial and temporal distributions of the derived quantities with similar measurements made using other techniques, and with model results. The data handling needed for this work is possible only with the help of a suite of interactive graphical and numerical analysis tools. Level 3 (gridded) data is the common form in which large data sets of this type are distributed for scientific analysis. We find that Level 3 data is inadequate for the data comparisons required for validation. Level 2 data (individual measurements in geophysical units) is needed. A sampling problem arises when individual measurements, which are not uniformly distributed in space or time, are used for the comparisons. Standard 'interpolation' methods involve fitting the measurements for each data set to surfaces, which are then compared. We are experimenting with formal criteria for selecting geographical regions, based upon the spatial frequency and variability of measurements, that allow us to quantify the uncertainty due to sampling. As part of this project, we are also dealing with ways to keep track of constraints placed on the output by assumptions made in the computer code. The need to work with Level 2 data introduces a number of other data handling issues, such as accessing data files across machine types, meeting large data storage requirements, accessing other validated data sets, processing speed and throughput for interactive graphical work, and problems relating to graphical interfaces.

Kahn, Ralph↗

Mobile and replicated alignment of arrays in data-parallel programs

When a data-parallel language like FORTRAN 90 is compiled for a distributed-memory machine, aggregate data objects (such as arrays) are distributed across the processor memories. The mapping determines the amount of residual communication needed to bring operands of parallel operations into alignment with each other. A common approach is to break the mapping into two stages: first, an alignment that maps all the objects to an abstract template, and then a distribution that maps the template to the processors. We solve two facets of the problem of finding alignments that reduce residual communication: we determine alignments that vary in loops, and objects that should have replicated alignments. We show that loop-dependent mobile alignment is sometimes necessary for optimum performance, and we provide algorithms with which a compiler can determine good mobile alignments for objects within do loops. We also identify situations in which replicated alignment is either required by the program itself (via spread operations) or can be used to improve performance. We propose an algorithm based on network flow that determines which objects to replicate so as to minimize the total amount of broadcast communication in replication. This work on mobile and replicated alignment extends our earlier work on determining static alignment.

Chatterjee, Siddhartha↗

A Machine Learning-Based Approach to Time-Series Wave Identification in the Solar Wind

The Wind spacecraft has yielded several decades of high-resolution magnetic field data, a large fraction of which displays small-scale structures. In particular, the solar wind is full of wavelike fluctuations that appear in both the field magnitude and its components. The nature of these fluctuations can be tied to the properties of other structures in the solar wind, such as shocks, that have implications for the time evolution of the solar wind. As such, having a large collection of wave events would facilitate further study of the effects that these fluctuations have on solar wind evolution. Given the large volume of magnetic field data available, machine learning is the most practical approach to classifying the myriad small-scale structures observed. To this end, a subset of Wind data is labeled and used as a training set for a multi-branch 1D convolutional neural network aimed at classifying circularly polarized wave modes. Using this algorithm, a preliminary statistical study of one year of data is performed, yielding about 300,000 wave intervals out of about 5,000,000 solar wind intervals. The wave intervals come about more often in the fast solar wind and at higher temperatures, and the number of waves per day is highly periodic. This machine learning-based approach to wave detection has the potential to be a powerful, inexpensive way to catalog waves throughout decades of spacecraft data.

Samuel Fordin↗