Search NASA⌕ Search

SEARCH · Search NASA

Results for “big data applications”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Research Data Alliance: Understanding Big Data Analytics Applications in Earth Science

The Research Data Alliance (RDA) enables data to be shared across barriers through focused working groups and interest groups, formed of experts from around the world - from academia, industry and government. Its Big Data Analytics (BDA) interest groups seeks to develop community based recommendations on feasible data analytics approaches to address scientific community needs of utilizing large quantities of data. BDA seeks to analyze different scientific domain applications (e.g. earth science use cases) and their potential use of various big data analytics techniques. These techniques reach from hardware deployment models up to various different algorithms (e.g. machine learning algorithms such as support vector machines for classification). A systematic classification of feasible combinations of analysis algorithms, analytical tools, data and resource characteristics and scientific queries will be covered in these recommendations. This contribution will outline initial parts of such a classification and recommendations in the specific context of the field of Earth Sciences. Given lessons learned and experiences are based on a survey of use cases and also providing insights in a few use cases in detail.

Riedel, Morris↗

Introduction to Big Earth Data Applications

Climate and weather modeling generate enormous volumes that make iterative analysis challenging, spurring the development of new ways to work with the data. A theme going across applications is the need to identify and highlight "interesting" data for the scientist to focus on. Operational applications often scale up from small, local studies to larger spatial scales with more analysis targets.

parallel processing (computers)↗

Introduction to Big Earth Data Applications

Climate and weather modeling generate enormous volumes that make iterative analysis challenging, spurring the development of new ways to work with the data. At the same time in the Earth Observation area, technology advances are enabling new sensors and satellites that will increase data volume, velocity and application variety. Scaling up can also be seen when operational applications expand from small, local studies to larger spatial scales with more analysis targets.

Christopher Lynnes↗

Efficient First-Order Algorithms for Large-Scale, Non-Smooth Maximum Entropy Models with Application to Wildfire Science

Maximum entropy (MaxEnt) models are a class of statistical models that use the maximum entropy principle to estimate probability distributions from data. Due to the size of modern data sets, MaxEnt models need efficient optimization algorithms to scale well for big data applications. State-of-the-art algorithms for MaxEnt models, however, were not originally designed to handle big data sets; these algorithms either rely on technical devices that may yield unreliable numerical results, scale poorly, or require smoothness assumptions that many practical MaxEnt models lack. In this paper, we present novel optimization algorithms that overcome the shortcomings of state-of-the-art algorithms for training large-scale, non-smooth MaxEnt models. Our proposed first-order algorithms leverage the Kullback–Leibler divergence to train large-scale and non-smooth MaxEnt models efficiently. For MaxEnt models with discrete probability distribution of n elements built from samples, each containing m features, the stepsize parameter estimation and iterations in our algorithms scale on the order of O(mn) operations and can be trivially parallelized. Moreover, the strong ℓ1 convexity of the Kullback–Leibler divergence allows for larger stepsize parameters, thereby speeding up the convergence rate of our algorithms. To illustrate the efficiency of our novel algorithms, we consider the problem of estimating probabilities of fire occurrences as a function of ecological features in the Western US MTBS-Interagency wildfire data set. Our numerical results show that our algorithms outperform the state of the art by one order of magnitude and yield results that agree with physical models of wildfire occurrence and previous statistical analyses of wildfire drivers.

Physics↗

An overview of visualization and visual analytics applications in water resources management

Recent advances in information, communication, and environmental monitoring technologies have increased the availability, spatiotemporal resolution, and quality of water-related data, thereby leading to the emergence of many innovative big data applications. Among these applications, visualization and visual analytics, also known as the visual computing techniques, empower the synergy of computational methods (e.g., machine learning and statistical models) with human reasoning to improve the understanding and solution toward complex science and engineering problems. These approaches are frequently integrated with geographic information systems and cyberinfrastructure to provide new opportunities and methods for enhancing water resources management. Here, we present a comprehensive review of recent hydroinformatics applications that employ visual computing techniques to (1) support complex data-driven research problems, and (2) support the communication and decision-makings in the water resources management sector. Then, we conduct a technical review of the state-of-the-art web-based visualization technologies and libraries to share our experiences on developing shareable, adaptive, and interactive visualizations and visual interfaces for water resources management applications. We close with a vision that applies the emerging visual computing technologies and paradigms to develop the next generation of hydroinformatics applications.

54 ENVIRONMENTAL SCIENCES↗

SNPP and N20 VIIRS Thermal Emissive Bands Calibration Comparison Using the GEO-LEO Double Difference Method

The VIIRS instruments onboard the SNPP and NOAA-20 satellites have identical spatial resolutions and the same spectral bands. Similar prelaunch tests and identical on-orbit calibration algorithms established the foundation for their consistent Earth measurements. Calibration assessment and consistency comparisons are useful to maintain their performance and measurement accuracy. Simultaneous nadir overpasses (SNO) between two satellites are commonly used for a direct calibration comparison between sensors. However, there are no SNO between SNPP and NOAA20. Hence, a reference sensor or Earth measurements are normally used to bridge the comparison. As a reference, we focus on the Advanced Baseline Imager (ABI) onboard the GOES-R series spacecraft and its application to the SNPP and NOAA-20 VIIRS comparison. GOES16 and GOES17are the first two satellites of the GOES-R series and were launched on November 19, 2016, and March 12, 2018, respectively. Their operational positions are on the equator with longitudes of 75.2° West over land and 137.2° West over ocean, respectively. The ABI is the primary imaging instrument of these spacecrafts for the Earth’s weather, oceans, and environment, with observations (every 10 minutes) that provide vast data for GEO-Low Earth orbit (LEO)and LEO-LEO comparisons utilizing it as an intermediate reference sensor. VIIRS and ABI have spectrally matched bands and can have simultaneous measurements over any selected site every day. The simultaneous measurements over the same site also have various scan angles. These features provide advantages for a VIIRS-to-ABI comparison. The spectral response function difference between instruments, sites selected, and view angles will have effects on the instrument measurements. Their impacts on the calibration comparison, including the use of double differences, will be discussed. By collecting VIIRS measurements over a large range of view angles, the view angle effect will also be investigated. The collection of an ample amount of data provides an advantage for statistical analyses and potential big data applications to sensor calibration assessments. This method can also be applied to other sensor calibration comparison and performance assessments, such as GOES16 and GOES17 ABI, and Terra and Aqua MODIS.

Tiejun Chang↗

CMOS-Based Single-Cycle in-Memory XOR/XNOR

Big data applications are on the rise, and so is the number of data centers. The ever-increasing massive data pool needs to be periodically backed up in a secure environment. Moreover, a massive amount of securely backed-up data is required for training binary convolutional neural networks for image classification. XOR and XNOR operations are essential for large-scale data copy verification, encryption, and classification algorithms. The disproportionate speed of existing compute and memory units makes the von Neumann architecture inefficient to perform these Boolean operations. Compute-in-memory (CiM) has proved to be an optimum approach for such bulk computations. The existing CiM-based XOR/XNOR techniques either require multiple cycles for computing or add to the complexity of the fabrication process. Here, we propose a CMOS-based hardware topology for single-cycle in-memory XOR/XNOR operations. Our design provides at least 2× improvement in the latency compared with other existing CMOS-compatible solutions. We verify the proposed system through circuit/system-level simulations and evaluate its robustness using a 5000-point Monte Carlo variation analysis. This all-CMOS design paves the way for practical implementation of CiM XOR/XNOR at scaled technology nodes.

97 MATHEMATICS AND COMPUTING↗

Next Generation Big Data Storage for Long Space Missions

This paper presents the results of the HELIOS (Hardened Extremely Long Life In-formation Optical Storage) mission on the International Space Station (ISS) which tested a unique solution for the long-term storage and retrieval of data in space. For this mission Creative Technology (CTech) developed test media—termed WORF (Write Once, Read Forever)—to validate whether this patented technology will survive all critical parameters for harsh space-based environments including microgravity and ionizing radiation. The HELIOS experiment confirmed that the WORF media is impervious to ionizing radiation, microgravity, solar (plasma) eruptions, and the stress from 8 Gs of the launch including extreme temperature expo-sure. The principal results indicate that there has been no discernible degradation of the media after 8 months on the ISS as compared to a control set of media stored on the ground. This data validated the media’s survivability for harsh space environments for long-term and deep space missions. In addition to the space environment, we are confident that WORF technology can be used for data storage where space-related and other long-term or archival integrity is critical such as: geospatial collections from satellites; space weather archives; past, ongoing, and future space mission media and documentation files; the deep space Gate-way program; as well as Big Data applications such as the Vera C. Rubin astronomical observatory (formerly the LSST). WORF technology for the HELIOS experiment uses a proven archival media, redesigned, re-purposed and patented by CTech to store digital data for long periods, measured in decades and possibly centuries. The media stores standing waves embedded in a substrate that capture the precise col-ors or wavelengths projected onto the media. The colors represent numerical data, with each data location storing multiple superimposed wavelengths, which facilitate the storage of multiple data bytes (rather than just zeros and ones); advanced mathematical permutations allow for extremely large data density equal to or greater than contemporary data storage de-vices. These colors cannot fade or degrade over time since the standing waves are physically stabilized (fully oxidized) metallic silver; no dyes are embedded for this storage system, and silver ions resist micro-bacterial and fungal contamination

Rodney Grubbs↗

Cyber Security: Big Data Think II Working Group Meeting

This presentation focuses on approaches that could be used by a data computation center to identify attacks and ensure malicious code and backdoors are identified if planted in system. The goal is to identify actionable security information from the mountain of data that flows into and out of an organization. The approaches are applicable to big data computational center and some must also use big data techniques to extract the actionable security information from the mountain of data that flows into and out of a data computational center. The briefing covers the detection of malicious delivery sites and techniques for reducing the mountain of data so that intrusion detection information can be useful, and not hidden in a plethora of false alerts. It also looks at the identification of possible unauthorized data exfiltration.

computer security↗

Introduction to NASA Goddard Workshop on Artificial Intelligence

Artificial Intelligence (AI) is a collection of advanced technologies that allows machines to think and act, both humanly and rationally, through sensing, comprehending, acting and learning. AI's foundations lie at the intersection of several traditional fields Philosophy, Mathematics, Economics, Neuroscience, Psychology and Computer Science. Although the inception of AI started in the 1950's, it has recently made a strong comeback in all aspects of society and all over the world; this is mainly due to the timely combination of increased data volumes, advanced and mature algorithms, and improvements in computing power and storage. Current AI applications include big data analytics, robotics, intelligent sensing, assisted decision making, and speech recognition just to name a few.This workshop will be investigating how AI technologies can be adapted or developed to address the following challenges: Discover events of interest and correlations in large amounts of science data; improve the outcomes of science modeling and data assimilation using improved data processing, integration, and analysis. Design advisors for mission planning and operations, including anomaly detection and spacecraft health monitoring. Develop tools for engineering support, including advanced manufacturing, orbit determination, new component design and system engineering. Customize intelligent user interfaces, including visual analytics and natural language processing.

Le Moigne, Jacqueline↗

Overview of Artificial Intelligence (AI) at NASA Goddard

Artificial Intelligence (AI) is a collection of advanced technologies that allows machines to think and act, both humanly and rationally, through sensing, comprehending, acting and learning. AI's foundations lie at the intersection of several traditional fields Philosophy, Mathematics, Economics, Neuroscience, Psychology and Computer Science. Although the inception of AI started in the 1950's, it has recently made a strong comeback in all aspects of society and all over the world; this is mainly due to the timely combination of increased data volumes, advanced and mature algorithms, and improvements in computing power and storage. Current AI applications include big data analytics, robotics, intelligent sensing, assisted decision making, and speech recognition just to name a few. During the Tour, we will show a few examples of the current AI activities at NASA Goddard.

Le Moigne, Jacqueline↗

Nonnegative canonical tensor decomposition with linear constraints: nnCANDELINC

Abstract There is an emerging interest for tensor factorization applications in big‐data analytics and machine learning. To speed up the factorization of extra‐large datasets, organized in multidimensional arrays (also known as tensors), easy to compute compression‐based tensor representations, such as, Tucker and tensor train formats, are used to approximate the initial large‐tensor. Further, tensor factorization is used to extract latent features that can facilitate discoveries of new mechanisms and signatures hidden in the data, where the explainability of the latent features is of principal importance. Nonnegative tensor factorization extracts latent features that are naturally sparse and parts of the data, which makes them easily interpretable. However, to take into account available domain knowledge and subject matter expertise, often additional constraints need to be imposed, which lead us to canonical decomposition with linear constraints (CANDELINC), a canonical polyadic decomposition with rank deficient factors. In CANDELINC, Tucker compression is used as a preprocessing step, which lead to a larger residual error but to more explainable latent features. Here, we propose a nonnegative CANDELINC (nnCANDELINC) accomplished via a specific nonnegative Tucker decomposition; we refer to as minimal or canonical nonnegative Tucker. We derive several results required to understand the specificity of nnCANDELINC, focusing on the difficulties of preserving the nonnegative rank of a tensor to its Tucker core and comparing the real valued to nonnegative case. Finally, we demonstrate nnCANDELINC performance on synthetic and real‐world examples.

97 MATHEMATICS AND COMPUTING↗

Learning Global Proliferation Expertise Evolution Using AI-Driven Analytics and Public Information

Detecting and anticipating global proliferation expertise and capability evolution from unstructured, noisy, and incomplete public data streams is a highly desired, but extremely challenging task. Here, in this article, we present our pioneering data-driven approach to support the non-proliferation mission to detect and explain the evolution of proliferation expertise and capability development globally from terabytes of publicly available information (PAI), focusing on our knowledge extraction pipeline and descriptive analytics. We first discuss how we fuse nine open-source data streams, including multilingual data, to convert 4 TB of unstructured data to structured knowledge and encode dynamically evolving proliferation expertise representations—content and context graphs. For this, we rely on natural language processing (NLP) and deep learning (DL) models to perform information extraction, topic modeling, and distributed text representation (aka embedding) learning. We then present interactive, usable, and explainable descriptive analytics to refine domain knowledge and present it in a human-understandable form. Finally, we introduce future work avenues that will leverage our dynamic knowledge representations and descriptive analytics to enable predictive and prescriptive inferences to achieve real-time domain understanding and contextual reasoning about global proliferation expertise and capability evolution.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Fault Detection Utilizing Convolution Neural Network on Timeseries Synchrophasor Data From Phasor Measurement Units

An end-to-end supervised learning method is proposed for fault detection in the electric grid using Big Data from multiple Phasor Measurement Units (PMUs). The approach consists of preprocessing steps aimed at reducing data noise and dimensionality, followed by utilization of six classification models considered for detecting faults. Three of the models were variants of Convolutional Neural Network (CNN) architectures that consider a single type of measurement (voltage, current or frequency) at all PMUs or all types together also at all PMUs. CNN based models were compared to traditional methods of Logistic Regression (LR), Multi-layer Perceptron (MLP) and Support Vector Machine (SVM). Evaluation was conducted on two-year data measured by PMUs at 37 locations in a large electric grid. Here, the response variable for classification were extracted from the grid-wide outage event log. Experiments show that CNN-based models outperformed traditional methods on one year out-of-sample outage detection over the entire grid.

42 ENGINEERING↗

Roadmap on artificial intelligence and big data techniques for superconductivity

This paper presents a roadmap to the application of AI techniques and big data (BD) for different modelling, design, monitoring, manufacturing and operation purposes of different superconducting applications. To help superconductivity researchers, engineers, and manufacturers understand the viability of using AI and BD techniques as future solutions for challenges in superconductivity, a series of short articles are presented to outline some of the potential applications and solutions. These potential futuristic routes and their materials/technologies are considered for a 10–20 yr time-frame.

machine learning, neural network↗

NASA GES DISC Giovanni: Current and Future

Giovanni (Geospatial Interactive Online Visualization and Analysis Infrastructure), developed by the NASA Goddard Earth Sciences Data and Information Services Center (GES DISC), has established a reputation among NASA users for easy access, analysis, and visualization of NASA Earth science data. Currently, Giovanni supports over 1900 variables in eight disciplinary areas. Like any other enterprise application, Giovanni faces big data challenges, such as servicing increasingly large data volumes and more complex data types, while at the same time addressing the demands of a more diverse user community, e.g., placing requests for long-term time series from multiple spatially and temporally dense data records. I will present how Giovanni has been evolving from an on-premises, monolithic software application towards a cloud-enabled implementation to address these challenges.

Analytics↗