Search NASA⌕ Search

SEARCH · Search NASA

Results for “big data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Enabling Analytics in the Cloud for Earth Science Data

The purpose of this workshop was to hold interactive discussions where providers, users, and other stakeholders could explore the convergence of three main elements in the rapidly developing world of technology: Big Data, Cloud Computing, and Analytics, [for earth science data].

Analytics↗

Interactive Multi-Instrument Database of Solar Flares

The fundamental motivation of the project is that the scientific output of solar research can be greatly enhanced by better exploitation of the existing solar/heliosphere space-data products jointly with ground-based observations. Our primary focus is on developing a specific innovative methodology based on recent advances in "big data" intelligent databases applied to the growing amount of high-spatial and multi-wavelength resolution, high-cadence data from NASA's missions and supporting ground-based observatories. Our flare database is not simply a manually searchable time-based catalog of events or list of web links pointing to data. It is a preprocessed metadata repository enabling fast search and automatic identification of all recorded flares sharing a specifiable set of characteristics, features, and parameters. The result is a new and unique database of solar flares and data search and classification tools for the Heliophysics community, enabling multi-instrument/multi-wavelength investigations of flare physics and supporting further development of flare-prediction methodologies.

Heliophysics↗

Pyroscopegridding: Geo-Leo Aerosol Data Fusion (an Open-Source Package)

The retrieval of aerosol optical depths (AODs) from sun-synchronous polar orbiting satellites, such as MODISs, VIIRSs, OMI, TROPOMI, etc., has been widely adopted as a method for obtaining information regarding particulate matter (PM) and related atmospheric processes. However, the advent of recently launched geostationary satellites, such as GOES-16/17/18, Himawari-8/9, and Meteosat Third Generation (MTG), has led to an increased temporal resolution of AOD observations (order of 10 minutes), resulting in typically one or more images per hour during daylight hours, compared to the once-per-day observations obtained from LEO satellites. By integrating these observations, the diurnal cycle of global AOD can be characterized at local, regional, and global scales. The scientific community is still examining the novel data from geostationary satellite observations and evaluating methods for effectively merging these observations with differing spatial and temporal resolutions. This presents a significant ""Big Data"" challenge, encompassing not only data storage, but also data discoverability, accessibility, and migration within cloud computing environments. This study presents our attempts at fusing Level 2 aerosol data from six satellites, three of which are geostationary (GOES-16/17 and Himawari-8) and three of which are polar orbiting (TERRA/MODIS, AQUA/MODIS, and SNPP-VIIRS), using the Dark Target aerosol retrieval algorithm. The ability to fuse remote sensing products on demand into desired temporal and spatial domains empowers researchers and practitioners to more efficiently work with satellite and sensor data. It is our hope that through making our open-source package and accompanying functionality available, the scientific community will have improved access to aerosol data processing resources.

Jennifer Wei↗

An Advanced Open-Source Platform for Air Quality Analysis, Visualization, and Prediction

Ambient air pollution is the largest environmental health risk factor, leading to several million premature deaths globally per year. The challenge of combating poor air quality is exacerbated by growing urban populations, changing emissions, and a warming climate. While there have been many advances monitoring and modeling of atmospheric composition, reflected in the dramatic increase in archived Earth Observations, there is no single measurement or method that alone can provide an accurate depiction of the entire atmosphere. The rapidly growing collections of observational and modeling data require us to be smarter about what data to include, and how such data is used. In recent years, NASA has invested significantly in advancing the concepts for Analytics Collaborative Framework (ACF) [5] and New Observing Strategies (NOS) [4] to tackle our software infrastructure need for harmonized data management and dynamic acquisition of diverse measurements for on-demand, interactive, multivariate analysis, and access [3]. It is not enough to have a big data, standalone analytics solution; it is critical that we start integrating data from remote sensing, modeling, and in-situ networks in a harmonized manner that enables timely and data-driven decision-making for air quality management. This work presents the design and development of an Air Quality Analytics Collaborative Framework (AQ ACF), as part of NASA’s Advanced Information Systems Technology (AIST) effort, to establish a data, machine-learning, and numerically driven platform for air quality analysis, visualization, and prediction.

Liu, Qian↗

Battery Health Quantification for TDRS Spacecraft by Using Signature Discriminability Measurement

The NASA/GSFC Space Network Project Office (SN) currently operates a constellation of ten geosynchronous TDRS spacecraft launched over the past 30 years. The SN project collects up to 16.5 Gigabytes of telemetry every month. Generally, the spacecraft health and functionality are obtained by the use of real-time telemetry data for the multiple spacecraft subsystems, which are transmitted to the main ground station at the White Sands Complex in Las Cruces, NM. Recently, the SN has instituted a program of Big Data to analyze the large amounts of data using a variety of tools including Machine Learning, Artificial Intelligence, development of training sets, and a variety of mathematical modeling tools. The goal is to improve spacecraft management and obtain a more accurate prediction of the spacecraft end of life. The combination of these efforts with those of the Aerospace Corporation, which has a contract with the SN to produce yearly reliability estimates for the TDRS fleet, will be performed. This paper presents a new concept called telemetry quality quantification (TQQ) and discusses the progress that has been made in battery performance estimation for the second-generation TDRS spacecraft using a signature discriminability measures (SDM) algorithm combined with the Aerospace Corp. battery life estimation models. This activity is important because many of the TDRS fleet of spacecraft have exceeded their on-orbit design lifetime and, therefore, NASA must carefully manage the spacecraft to continue operations while avoiding an end-of-mission scenario that leaves a non-functioning spacecraft in geosynchronous orbit.

Ma, Kenneth Y.↗

Fusion Approach for Remotely-Sensed Mapping of Agriculture (FARMA): A Scalable Open Source Method for Land Cover Monitoring Using Data Fusion

The increasing availability of very-high resolution (VHR; <2 m) imagery has the potential to enable agricultural monitoring at increased resolution and cadence, particularly when used in combination with widely available moderate-resolution imagery. However, scaling limitations exist at the regional level due to big data volumes and processing constraints. Here, we demonstrate the Fusion Approach for Remotely-Sensed Mapping of Agriculture (FARMA), using a suite of open source software capable of efficiently characterizing time-series field-scale statistics across large geographical areas at VHR resolution. We provide distinct implementation examples in Vietnam and Senegal to demonstrate the approach using WorldView VHR optical, Sentinel-1 Synthetic Aperture Radar, and Sentinel-2 and Sentinel-3 optical imagery. This distributed software is open source and entirely scalable, enabling large area mapping even with modest computing power. FARMA provides the ability to extract and monitor sub-hectare fields with multisensor raster signals, which previously could only be achieved at scale with large computational resources. Implementing FARMA could enhance predictive yield models by delineating boundaries and tracking productivity of smallholder fields, enabling more precise food security observations in low and lower-middle income countries.

fusion↗

Air Traffic Management TestBed Data Exchange Model

The Air Traffic Management (ATM) TestBed is a Platform as a Service that is being developed by the National Aeronautics and Space Administration (NASA) to help design, configure, integrate, run, and monitor air traffic simulations. The platform is designed to provide cloud services including back-end, big-data analytics tools, on-demand computing resource management, data storage, and communication middleware. The ATM TestBed reduces the time to test concepts and technologies, supports interactions among various methods such as human-in-the-loop and automation-in-the-loop simulations, and enables collaborative simulations by sharing technologies and tools in the ATM community. In order to allow easier access to simulation components, TestBed provides a messaging support layer for connectivity using a consistent set of input/output interfaces. In addition, a standard data format is introduced to facilitate communication between the components. The data exchange model, supported in the messaging support layer, standardizes the format of the information to be exchanged among the components. This document describes the messaging data model currently developed in TestBed and provides data dictionaries for references to component developers as well as simulation engineers.

air traffic simulation↗

Enabling Real-time Multi-messenger Astrophysics Discoveries with Deep Learning

Multi-messenger astrophysics is a fast-growing, interdisciplinary field that combines data, which vary in volume and speed of data processing, from many different instruments that probe the Universe using different cosmic messengers: electromagnetic waves, cosmic rays, gravitational waves and neutrinos. In this Expert Recommendation, we review the key challenges of real-time observations of gravitational wave sources and their electromagnetic and astroparticle counterparts, and make a number of recommendations to maximize their potential for scientific discovery. These recommendations refer to the design of scalable and computationally efficient machine learning algorithms; the cyber-infrastructure to numerically simulate astrophysical sources, and to process and interpret multi-messenger astrophysics data; the management of gravitational wave detections to trigger real-time alerts for electromagnetic and astroparticle follow-ups; a vision to harness future developments of machine learning and cyber-infrastructure resources to cope with the big-data requirements; and the need to build a community of experts to realize the goals of multi-messenger astrophysics.

E A Huerta↗

Applications of Anomaly Detection and Precursor Identification in Airspace Operations

As we continue to advance the U.S. National Airspace into the next generation of air traffic, we face challenges in both increase in complexity, as well as, a significant growth in traffic volume. Addressing these challenges, while maintaining the same level of safety is an important application of data mining. Because of these significant shifts in airspace design and usage there is a need to identify current and emergent safety risks along with their potential precursors. In recent years NASA has made advancements in developing scalable methods to address this effort in the Big Data paradigm. Multiple kernel anomaly detection approaches have been employed on both surveillance radar data and flight operational quality assurance data to identify operationally significant safety risks. Additionally, events have been explored with a recently developed precursor identification tool to discover states that reveal an increased probability of a safety event. These tools can be used to discover emerging safety risks that may not be currently monitored, which allows for mitigation tactics to be employed and ultimately make the overall airspace safer. This talk will discuss an overview of these methods and a discussion of the findings.

anomaly detection↗

Investigation and Evaluation of Advanced Spectrum Management Concepts for Aeronautical Communications

With the emergence of new aerial vehicles into the airspace and the continued growth of aviation operations, there will be an increasing demand for voice and data communications within the National Airspace System (NAS). The continued use of existing VHF and UHF frequency allocations is not a sustainable approach, and as a result, a new spectrum management solution is required to support future mission needs. The proposed spectrum management concepts leverage modern advancements such as artificial intelligence (AI) and big data to dynamically optimize the spectrum utilization based on the predicted communications demand throughout the airspace. This technical investigation considers both air-ground and air-air communications networks, and can be applied to both existing applications such as the air traffic control (ATC) system, as well as future applications, such as the emerging Advanced Air Mobility (AAM). To support the evaluation of the proposed concepts and technologies, a modeling and simulation capability is currently under development and will continue to evolve to support new and advanced airspace applications. It is anticipated that this proposed spectrum concept will better serve the spectrum needs of future NAS applications.

Eric J Knoblock↗

Investigation and Evaluation of Advanced Spectrum Management Concepts for Aeronautical Communications

With the emergence of new aerial vehicles into the airspace and the continued growth of aviation operations, there will be an increasing demand for voice and data communications within the National Airspace System (NAS). The continued use of existing VHF and UHF frequency allocations is not a sustainable approach, and as a result, a new spectrum management solution is required to support future mission needs. The proposed spectrum management concepts leverage modern advancements such as artificial intelligence (AI) and big data to dynamically optimize the spectrum utilization based on the predicted communications demand throughout the airspace. This technical investigation considers both air-ground and air-air communications networks, and can be applied to both existing applications such as the air traffic control (ATC) system, as well as future applications, such as the emerging Advanced Air Mobility (AAM). To support the evaluation of the proposed concepts and technologies, a modeling and simulation capability is currently under development and will continue to evolve to support new and advanced airspace applications. It is anticipated that this proposed spectrum concept will better serve the spectrum needs of future NAS applications.

Eric J Knoblock↗

The Importance of Data Visualization: Incorporating Storytelling into the Scientific Presentation

From its inception in 2000, one of the primary tasks of the Biomedical Data Reduction Analysis (BDRA) group has been translation of large amounts of data into information that is relevant to the audience receiving it. BDRA helps translate data into an integrated model that supports both operational and research activities. This data integrated model and subsequent visual data presentations have contributed to BDRA's success in delivering the message (i.e., the story) that its customers have needed to communicate. This success has led to additional collaborations among groups that had previously not felt they had much in common until they worked together to develop solutions in an integrated fashion. As more emphasis is placed on working with "big data" and on showing how NASA's efforts contribute to the greater good of the American people and of the world, it becomes imperative to visualize the story of our data to communicate the greater message we need to share. METHODS To create and expand its data integrated model, BDRA has incorporated data from many different collaborating partner labs and other sources. Data are compiled from the repositories of the Lifetime Surveillance of Astronaut Health and the Life Sciences Data Archive, and from the individual laboratories at Johnson Space Center that support collection of data from medical testing, environmental monitoring, and countermeasures, as designated in the Medical Requirements Integration Documents. Ongoing communication with the participating collaborators is maintained to ensure that the message and story of the data are retained as data are translated into information and visual data presentations are delivered in different venues and to different audiences. RESULTS We will describe the importance of storytelling through an integrated model and of subsequent data visualizations in today's scientific presentations and discuss the collaborative methods used. We will illustrate the discussion with examples of graphs from BDRA's past work supporting operations and/or research efforts.

Babiak-Vazquez, A.↗

A Heuristic Approach to Correlating ERAM Flight Data from Twenty Centers

Among its many other functions, the Federal Aviation Administration’s En Route Automation Modernization (ERAM) provides external systems with real-time air traffic data for flights in enroute airspace in the National Airspace System. It replaced the En Route Host computer and backup system used at 20 FAA Air Route Traffic Control Centers (Centers) nationwide. Among the new features of ERAM, its output data stream of flight plan and track data includes a unique identifier for a flight originating in any one of the 20 ERAM Centers. The unique identifier, called the Global Unique Flight Identifier (GUFI), is persistent across all the Centers that track the flight. However, certain factors make it difficult to correlate data using the GUFI. First, the value of the GUFI is only unique within a time window of seven days. Second, the GUFI is attached only to flight-plan related data messages. Finally, track positions reported by ERAM do not reference the GUFI. In order to correlate historical as well as real time flight-plan and position related ERAM data, an efficient, heuristic approach was developed, and a prototype was developed. The approach showed that the processing speed, through parallel processing, is sufficient to correlate ERAM data in real-time. As described in this paper, when there are multiple track positions reported from multiple Centers within a few seconds, each position is assigned with a weighted score to indicate the quality of the position relative to its last know position. The weighted score can be used to eliminate potentially duplicate track positions. The approach is database-agnostic, and can be implemented in a Big Data system such as an Apache Hadoop system, as well as in traditional database systems.

Correlating ERAM Flight Data↗

Spatial validation reveals poor predictive performance of large-scale ecological mapping models

Mapping aboveground forest biomass is central for assessing the global carbon balance. However, current large-scale maps show strong disparities, despite good validation statistics of their underlying models. Here, we attribute this contradiction to a flaw in the validation methods, which ignore spatial autocorrelation (SAC) in data, leading to overoptimistic assessment of model predictive power. To illustrate this issue, we reproduce the approach of large-scale mapping studies using a massive forest inventory dataset of 11.8 million trees in central Africa to train and validate a random forest model based on multispectral and environmental variables. A standard nonspatial validation method suggests that the model predicts more than half of the forest biomass variation, while spatial validation methods accounting for SAC reveal quasi-null predictive power. This study underscores how a common practice in big data mapping studies shows an apparent high predictive power, even when predictors have poor relationships with the ecological variable of interest, thus possibly leading to erroneous maps and interpretations.

biomass↗

Leveraging the Usage of GPUs in SAR Processing for the NISAR Mission

The NASA ISRO Synthetic Aperture Radar (NISAR) mission will redefine the future of earth science in terms of both the quality as well as the quantity of data that will be downlinked daily. The current software architecture used to process this data is the InSAR Scientific Computing Environment (ISCE), a powerful and modular platform that applies a combination of novel and legacy processing modules to many sources of SAR data. Until recently, this architecture could process most images in a reasonable amount of time; however in the case of the NISAR mission (where the daily influx as well as the size of the images themselves are significantly larger) the current architecture can take hours to process even a single image. This paper explores new efforts to use a Graphics Processing Unit (GPU) to accelerate one of the processing modules to achieve unprecedented runtimes with no loss in precision, potentially setting a new standard in radar processing in the world of “Big Data”.

Cohen, Joshua↗

Time and Measurement Days

Questions in data analysis involving the concepts of time and measurement are often pushed into the background or reserved for a philosophical discussion. Some examples are: a) Is causality a consequence of the laws of physics, or can the arrow of time be reversed? b) Can we determine the arrow of time of an event? c) Do we need the continuum hypothesis for the underlying function in any measurement process? d) Can we say anything about the analyticity of the underlying process of an event? e) Would it be valid to model a non-analytical process as function of time? f) What are the implications of all these questions for classical Fourier techniques? However, in the age of big data gathered either from space missions supplying ultra-precise long time series, or e.g. LIGO data from the ground, the moment to bring these questions to the foreground seems arrived. The limitations of our understanding of some fundamental processes is emphasized by the lack of solution for problems open for more than 2 decades, such as the non-detection of solar g-modes, or the modal identification of main sequence stellar pulsators like delta Scuti stars. Flicker noise or 1/f noise, for example, attributed in solar-like stars to granulation, is analyzed mostly only to apply noise reduction techniques, neither considering the classical problem of 1/f noise that was introduced a 100 years ago, nor taking into account ergodic or non-ergodic solutions that make inapplicable spectral analysis techniques in practice. This topic was discussed by Nicholas W. Watkins during the ITISE meeting held in Granada in 2016. There he presented preliminary results of his research on Mandelbrot's related work. We reproduce here his quotation of Mandelbrot (1999) "There is a sharp contrast between a highly anomalous ("non-white") noise that proceeds in ordinary clock time and a noise whose principal anomaly is that it is restricted to fractal time", suggesting a connection with the above proposed topics that could be phrased as the following additional questions:a) Is self-organized criticality (SOC) frequent in astrophysical phenomena? b) Could all fractals in nature be considered stochastic? c) Could we establish mathematical/physical relationships between chaotic and fractal behaviors in time series? d) Could the differences between fractals and chaos in terms of analyticity be used to understand the residuals of the fitting of stellar light curves? In this meeting we would like to approximate these problems from a holistic and multidisciplinary perspective, taking into account not only technical issues but also the deeper implications. In particular the concept of connectivity (introduced in Pascual-Granado et al. A&A, 2015) could be used to implement, within the framework of ARMA processes, an "arrow of time" (see attached document), and so studying the possible implications in the concept of time as envisaged by Watkins.

data analysis↗

Implementing Journaling in a Linux Shared Disk File System

In computer systems today, speed and responsiveness is often determined by network and storage subsystem performance. Faster, more scalable networking interfaces like Fibre Channel and Gigabit Ethernet provide the scaffolding from which higher performance computer systems implementations may be constructed, but new thinking is required about how machines interact with network-enabled storage devices. In this paper we describe how we implemented journaling in the Global File System (GFS), a shared-disk, cluster file system for Linux. Our previous three papers on GFS at the Mass Storage Symposium discussed our first three GFS implementations, their performance, and the lessons learned. Our fourth paper describes, appropriately enough, the evolution of GFS version 3 to version 4, which supports journaling and recovery from client failures. In addition, GFS scalability tests extending to 8 machines accessing 8 4-disk enclosures were conducted: these tests showed good scaling. We describe the GFS cluster infrastructure, which is necessary for proper recovery from machine and disk failures in a collection of machines sharing disks using GFS. Finally, we discuss the suitability of Linux for handling the big data requirements of supercomputing centers.

Preslan, Kenneth W.↗

Demonstrations at Supercomputing 1998

Contents of the papers include: Visualization techniques- The next generation, Small-machine visualizing very big data sets, Public display of Lunar Prospector data, Steering molecules via Feel, High-dimensional data browser, Multi-source visualization, Interactive surface flow visualization, and the Virtual Windtunnel.

Bryson, Steve↗