Search NASA⌕ Search

SEARCH · Search NASA

Results for “big data applications”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Research Data Alliance: Understanding Big Data Analytics Applications in Earth Science

The Research Data Alliance (RDA) enables data to be shared across barriers through focused working groups and interest groups, formed of experts from around the world - from academia, industry and government. Its Big Data Analytics (BDA) interest groups seeks to develop community based recommendations on feasible data analytics approaches to address scientific community needs of utilizing large quantities of data. BDA seeks to analyze different scientific domain applications (e.g. earth science use cases) and their potential use of various big data analytics techniques. These techniques reach from hardware deployment models up to various different algorithms (e.g. machine learning algorithms such as support vector machines for classification). A systematic classification of feasible combinations of analysis algorithms, analytical tools, data and resource characteristics and scientific queries will be covered in these recommendations. This contribution will outline initial parts of such a classification and recommendations in the specific context of the field of Earth Sciences. Given lessons learned and experiences are based on a survey of use cases and also providing insights in a few use cases in detail.

Riedel, Morris↗

Introduction to Big Earth Data Applications

Climate and weather modeling generate enormous volumes that make iterative analysis challenging, spurring the development of new ways to work with the data. A theme going across applications is the need to identify and highlight "interesting" data for the scientist to focus on. Operational applications often scale up from small, local studies to larger spatial scales with more analysis targets.

parallel processing (computers)↗

Introduction to Big Earth Data Applications

Climate and weather modeling generate enormous volumes that make iterative analysis challenging, spurring the development of new ways to work with the data. At the same time in the Earth Observation area, technology advances are enabling new sensors and satellites that will increase data volume, velocity and application variety. Scaling up can also be seen when operational applications expand from small, local studies to larger spatial scales with more analysis targets.

Christopher Lynnes↗

SNPP and N20 VIIRS Thermal Emissive Bands Calibration Comparison Using the GEO-LEO Double Difference Method

The VIIRS instruments onboard the SNPP and NOAA-20 satellites have identical spatial resolutions and the same spectral bands. Similar prelaunch tests and identical on-orbit calibration algorithms established the foundation for their consistent Earth measurements. Calibration assessment and consistency comparisons are useful to maintain their performance and measurement accuracy. Simultaneous nadir overpasses (SNO) between two satellites are commonly used for a direct calibration comparison between sensors. However, there are no SNO between SNPP and NOAA20. Hence, a reference sensor or Earth measurements are normally used to bridge the comparison. As a reference, we focus on the Advanced Baseline Imager (ABI) onboard the GOES-R series spacecraft and its application to the SNPP and NOAA-20 VIIRS comparison. GOES16 and GOES17are the first two satellites of the GOES-R series and were launched on November 19, 2016, and March 12, 2018, respectively. Their operational positions are on the equator with longitudes of 75.2° West over land and 137.2° West over ocean, respectively. The ABI is the primary imaging instrument of these spacecrafts for the Earth’s weather, oceans, and environment, with observations (every 10 minutes) that provide vast data for GEO-Low Earth orbit (LEO)and LEO-LEO comparisons utilizing it as an intermediate reference sensor. VIIRS and ABI have spectrally matched bands and can have simultaneous measurements over any selected site every day. The simultaneous measurements over the same site also have various scan angles. These features provide advantages for a VIIRS-to-ABI comparison. The spectral response function difference between instruments, sites selected, and view angles will have effects on the instrument measurements. Their impacts on the calibration comparison, including the use of double differences, will be discussed. By collecting VIIRS measurements over a large range of view angles, the view angle effect will also be investigated. The collection of an ample amount of data provides an advantage for statistical analyses and potential big data applications to sensor calibration assessments. This method can also be applied to other sensor calibration comparison and performance assessments, such as GOES16 and GOES17 ABI, and Terra and Aqua MODIS.

Tiejun Chang↗

Next Generation Big Data Storage for Long Space Missions

This paper presents the results of the HELIOS (Hardened Extremely Long Life In-formation Optical Storage) mission on the International Space Station (ISS) which tested a unique solution for the long-term storage and retrieval of data in space. For this mission Creative Technology (CTech) developed test media—termed WORF (Write Once, Read Forever)—to validate whether this patented technology will survive all critical parameters for harsh space-based environments including microgravity and ionizing radiation. The HELIOS experiment confirmed that the WORF media is impervious to ionizing radiation, microgravity, solar (plasma) eruptions, and the stress from 8 Gs of the launch including extreme temperature expo-sure. The principal results indicate that there has been no discernible degradation of the media after 8 months on the ISS as compared to a control set of media stored on the ground. This data validated the media’s survivability for harsh space environments for long-term and deep space missions. In addition to the space environment, we are confident that WORF technology can be used for data storage where space-related and other long-term or archival integrity is critical such as: geospatial collections from satellites; space weather archives; past, ongoing, and future space mission media and documentation files; the deep space Gate-way program; as well as Big Data applications such as the Vera C. Rubin astronomical observatory (formerly the LSST). WORF technology for the HELIOS experiment uses a proven archival media, redesigned, re-purposed and patented by CTech to store digital data for long periods, measured in decades and possibly centuries. The media stores standing waves embedded in a substrate that capture the precise col-ors or wavelengths projected onto the media. The colors represent numerical data, with each data location storing multiple superimposed wavelengths, which facilitate the storage of multiple data bytes (rather than just zeros and ones); advanced mathematical permutations allow for extremely large data density equal to or greater than contemporary data storage de-vices. These colors cannot fade or degrade over time since the standing waves are physically stabilized (fully oxidized) metallic silver; no dyes are embedded for this storage system, and silver ions resist micro-bacterial and fungal contamination

Rodney Grubbs↗

Cyber Security: Big Data Think II Working Group Meeting

This presentation focuses on approaches that could be used by a data computation center to identify attacks and ensure malicious code and backdoors are identified if planted in system. The goal is to identify actionable security information from the mountain of data that flows into and out of an organization. The approaches are applicable to big data computational center and some must also use big data techniques to extract the actionable security information from the mountain of data that flows into and out of a data computational center. The briefing covers the detection of malicious delivery sites and techniques for reducing the mountain of data so that intrusion detection information can be useful, and not hidden in a plethora of false alerts. It also looks at the identification of possible unauthorized data exfiltration.

computer security↗

Introduction to NASA Goddard Workshop on Artificial Intelligence

Artificial Intelligence (AI) is a collection of advanced technologies that allows machines to think and act, both humanly and rationally, through sensing, comprehending, acting and learning. AI's foundations lie at the intersection of several traditional fields Philosophy, Mathematics, Economics, Neuroscience, Psychology and Computer Science. Although the inception of AI started in the 1950's, it has recently made a strong comeback in all aspects of society and all over the world; this is mainly due to the timely combination of increased data volumes, advanced and mature algorithms, and improvements in computing power and storage. Current AI applications include big data analytics, robotics, intelligent sensing, assisted decision making, and speech recognition just to name a few.This workshop will be investigating how AI technologies can be adapted or developed to address the following challenges: Discover events of interest and correlations in large amounts of science data; improve the outcomes of science modeling and data assimilation using improved data processing, integration, and analysis. Design advisors for mission planning and operations, including anomaly detection and spacecraft health monitoring. Develop tools for engineering support, including advanced manufacturing, orbit determination, new component design and system engineering. Customize intelligent user interfaces, including visual analytics and natural language processing.

Le Moigne, Jacqueline↗

Overview of Artificial Intelligence (AI) at NASA Goddard

Artificial Intelligence (AI) is a collection of advanced technologies that allows machines to think and act, both humanly and rationally, through sensing, comprehending, acting and learning. AI's foundations lie at the intersection of several traditional fields Philosophy, Mathematics, Economics, Neuroscience, Psychology and Computer Science. Although the inception of AI started in the 1950's, it has recently made a strong comeback in all aspects of society and all over the world; this is mainly due to the timely combination of increased data volumes, advanced and mature algorithms, and improvements in computing power and storage. Current AI applications include big data analytics, robotics, intelligent sensing, assisted decision making, and speech recognition just to name a few. During the Tour, we will show a few examples of the current AI activities at NASA Goddard.

Le Moigne, Jacqueline↗

NASA GES DISC Giovanni: Current and Future

Giovanni (Geospatial Interactive Online Visualization and Analysis Infrastructure), developed by the NASA Goddard Earth Sciences Data and Information Services Center (GES DISC), has established a reputation among NASA users for easy access, analysis, and visualization of NASA Earth science data. Currently, Giovanni supports over 1900 variables in eight disciplinary areas. Like any other enterprise application, Giovanni faces big data challenges, such as servicing increasingly large data volumes and more complex data types, while at the same time addressing the demands of a more diverse user community, e.g., placing requests for long-term time series from multiple spatially and temporally dense data records. I will present how Giovanni has been evolving from an on-premises, monolithic software application towards a cloud-enabled implementation to address these challenges.

Analytics↗

Use of Schema on Read in Earth Science Data Archives

Traditionally, NASA Earth Science data archives have file-based storage using proprietary data file formats, such as HDF and HDF-EOS, which are optimized to support fast and efficient storage of spaceborne and model data as they are generated. The use of file-based storage essentially imposes an indexing strategy based on data dimensions. In most cases, NASA Earth Science data uses time as the primary index, leading to poor performance in accessing data in spatial dimensions. For example, producing a time series for a single spatial grid cell involves accessing a large number of data files. With exponential growth in data volume due to the ever-increasing spatial and temporal resolution of the data, using file-based archives poses significant performance and cost barriers to data discovery and access. Storing and disseminating data in proprietary data formats imposes an additional access barrier for users outside the mainstream research community. At the NASA Goddard Earth Sciences Data Information Services Center (GES DISC), we have evaluated applying the schema-on-read principle to data access and distribution. We used Apache Parquet to store geospatial data, and have exposed data through Amazon Web Services (AWS) Athena, AWS Simple Storage Service (S3), and Apache Spark. Using the schema-on-read approach allows customization of indexing spatially or temporally to suit the data access pattern. The storage of data in open formats such as Apache Parquet has widespread support in popular programming languages. A wide range of solutions for handling big data lowers the access barrier for all users. This presentation will discuss formats used for data storage, frameworks with This presentation will discuss formats used for data storage, frameworks with support for schema-on-read used for data access, and common use cases covering data usage patterns seen in a geospatial data archive.

cloud applications↗

NASA’s Prototype Spectral Water Inversion Processor and Emulator (SWIPE): Towards Global Coastal and Inland Water Quality and Algal Biodiversity Monitoring

Degradation of Earth’s inland water resources due to anthropogenic perturbations and climate anomalies at both local and global scales continues to place human health at substantial risk. There is now a growing necessity to develop pragmatic approaches that allow timely and effective extrapolation of local processes, to spatially resolved global products, and to promote operational and sustainable resource policy management. This presentation will provide updates on NASA’s prototype open-source aquatic modeling platform, Spectral Water Inversion Processor and Emulator (SWIPE), which is a comprehensive, multi-faceted modeling platform for both forward and inverse modeling of diverse aquatic ecosystems from the benthos to top-of-atmosphere (TOA). SWIPE provides a cohesive application which leverages recent advancements in particle modeling, Big Data analytics, and machine learning to develop a high-fidelity synthetic training ground for sensitivity studies and algorithm development for multispectral or upcoming hyperspectral missions. Some of the prominent features of SWIPE to be discussed include: 1. Advanced hyperspectral modeling of globally diverse algal and non-algal particles using a novel two-layer coated sphere scattering model and radiative transfer modeling, 2. Massive, highly detailed synthetic spectral libraries of Analysis-Ready-Data (ARD) which include spectral libraries of particle microphysics, water biogeophysical and optical properties, as well as surface and TOA reflectances at 1 nm resolution, 3. An ensemble of pre-built analytic, machine learning, and deep learning inversion algorithms for various water quality and biodiversity related retrieval parameters and uncertainty quantification, 4. Sensor-agnostic water quality inversion at wide ranging spatial and spectral resolutions including a codebase for seamless application in the Google Earth Engine and NASA Earth Exchange (NEX) for planetary scale analysis. SWIPE will be a fully open-source platform based in python with comprehensive documentation, tutorials, and options for distributed computing on high performance computing clusters or on single, local machines. Further, we will discuss how we envision SWIPE contributing towards a global analysis of coastal and inland water quality dynamics.

top-of-atmosphere (TOA)↗

Restructuring Big Data to Improve Data Access and Performance in Analytic Services Making Research More Efficient for the Study of Extreme Weather Events and Application User Communities

By developing and enhancing various services and tools, the GES DISC provides users with the capability to access and visualize data, and to make comparisons of data from multiple sensor and models via a number of cross-discipline projects. Discovering Data via Faceted Web Interface Web interface to data products and services Search and Download mechanisms Dataset Landing Pages Accessing Data through Interoperable Services: GDS – GrADS Data Server OPeNDAP - Open-source Project for a Network Data Access Protocol WMS – OGC service GIS connector – allowing IS tools to access data easier (coming soon) HTTPS -- direct online access Downloading Data Basics: Subset and egridding Service – Parameter, Spatial, Time, Vertical, Mean averaging, format conversion, and regridding for L3/L4 gridded data Swath Data Subsetter – Parameter, spatial subset of L2 /L1 data. Visualizing Data Online: Giovanni –Visualization and Analysis L3/L4 gridded data AIRS NRT Viewer – AIRS near-real-time DQVis – L2 data quality visualization

data cube↗

High Resolution Nature Runs and the Big Data Challenge

NASA's Global Modeling and Assimilation Office at Goddard Space Flight Center is undertaking a series of very computationally intensive Nature Runs and a downscaled reanalysis. The nature runs use the GEOS-5 as an Atmospheric General Circulation Model (AGCM) while the reanalysis uses the GEOS-5 in Data Assimilation mode. This paper will present computational challenges from three runs, two of which are AGCM and one is downscaled reanalysis using the full DAS. The nature runs will be completed at two surface grid resolutions, 7 and 3 kilometers and 72 vertical levels. The 7 km run spanned 2 years (2005-2006) and produced 4 PB of data while the 3 km run will span one year and generate 4 BP of data. The downscaled reanalysis (MERRA-II Modern-Era Reanalysis for Research and Applications) will cover 15 years and generate 1 PB of data. Our efforts to address the big data challenges of climate science, we are moving toward a notion of Climate Analytics-as-a-Service (CAaaS), a specialization of the concept of business process-as-a-service that is an evolving extension of IaaS, PaaS, and SaaS enabled by cloud computing. In this presentation, we will describe two projects that demonstrate this shift. MERRA Analytic Services (MERRA/AS) is an example of cloud-enabled CAaaS. MERRA/AS enables MapReduce analytics over MERRA reanalysis data collection by bringing together the high-performance computing, scalable data management, and a domain-specific climate data services API. NASA's High-Performance Science Cloud (HPSC) is an example of the type of compute-storage fabric required to support CAaaS. The HPSC comprises a high speed Infinib and network, high performance file systems and object storage, and a virtual system environments specific for data intensive, science applications. These technologies are providing a new tier in the data and analytic services stack that helps connect earthbound, enterprise-level data and computational resources to new customers and new mobility-driven applications and modes of work. In our experience, CAaaS lowers the barriers and risk to organizational change, fosters innovation and experimentation, and provides the agility required to meet our customers' increasing and changing needs

big data analysis↗

DSD Characteristics of a Mid-Winter Tornadic Storm Using C-Band Polarimetric Radar and Two 2D-Video Disdrometers

Drop size distributions in an evolving tornadic storm are examined using C-band polarimetric radar observations and two 2D-video disdrometers. The E-F2 storm occurred in mid-winter (21 January 2010) in northern Alabama, USA, and caused widespread damage. The evolution of the storm occurred within the C-band radar coverage and moreover, several minutes prior to touch down, the storm passed over a site where several disdrometers including two 2D video disdrometers (2DVD) had been installed. One of the 2DVDs is a low profile unit and the other is a new next generation compact unit currently undergoing performance evaluation. Analyses of the radar data indicate that the main region of precipitation should be treated as a "big-drop" regime case. Even the measured differential reflectivity values (i.e. without attenuation correction) were as high as 6-7 dB within regions of high reflectivity. Standard attenuation-correction methods using differential propagation phase have been "fine tuned" to be applicable to the "big drop" regime. The corrected reflectivity and differential reflectivity data are combined with the co-polar correlation coefficient and specific differential phase to determine the mass-weighted mean diameter, Dm, and the width of the mass spectrum, (sigma)M, as well as the intercept parameter , Nw. Significant areas of high Dm (3-4 mm) were retrieved within the main precipitation areas of the tornadic storm. The "big drop" regime assumption is substantiated by the two sets of 2DVD measurements. The Dm values calculated from 1-minute drop size distributions reached nearly 4 mm, whilst the maximum drop diameters were over 6 mm. The fall velocity measurements from the 2DVD indicate almost all hydrometeors to be fully melted at ground level. Drop shapes for this event are also being investigated from the 2DVD camera data.

Thurai, M.↗

Big-data Efficient and Automated Science Transfer (BEAST): An Open-Source Software Architecture for Arc Jet Data Management, Modeling, and Automation

Big-data Efficient and Automated Science Transfer (BEAST) is a facility data management application developed for the NASA Ames arc jet facilities. The current decentralized data management practices limit statistical tracking, synchronization between video/time series, search capability, data throughput, and data processing speed/efficiency. Consequently, BEAST was developed to provide a new data infrastructure with streamlined data collection, processing, transfer, and analysis. This new framework also seeks to implement the FAIR principles of data stewardship: Findable, Accessible, Interoperable, and Reusable. The BEAST framework is based on a combination of the Python Django web framework and the Python data stack to provide a monolithic, open-source platform for data management, automation, and machine learning. This architecture was chosen for maintainability and scalability for a small, in-house development team. This paper will describe the application framework, deployment, and discuss the benefits and future plans for the system.

Data management↗

Bundle Data Approach at GES DISC Targeting Natural Hazards

Severe natural phenomena such as hurricane, volcano, blizzard, flood and drought have the potential to cause immeasurable property damages, great socioeconomic impact, and tragic loss of human life. From searching to assessing the Big, i.e., massive and heterogeneous scientific data (particularly, satellite and model products) in order to investigate those natural hazards, it has, however, become a daunting task for Earth scientists and applications researchers, especially during recent decades. The NASA Goddard Earth Sciences Data and Information Service Center (GES DISC) has served Big Earth science data, and the pertinent valuable information and services to the aforementioned users of diverse communities for years. In order to help and guide our users to online readily (i.e., with a minimum effort) acquire their requested data from our enormous resource at GES DISC for studying their targeted hazard event, we have thus initiated a Bundle Data approach in 2014, first targeting the hurricane event topic. We have recently worked on new topics such as volcano and blizzard. The bundle data of a specific hazard event is basically a sophisticated integrated data package consisting of a series of proper datasets containing a group of relevant (knowledge--based) data variables readily accessible to users via a system-prearranged table linking those data variables to the proper datasets (URLs). This online approach has been developed by utilizing a few existing data services such as Mirador as search engine; Giovanni for visualization; and OPeNDAP for data access, etc. The online Data Cookbook site at GES DISC is the current host for the bundle data. We are now also planning on developing an Automated Virtual Collection Framework that shall eventually accommodate the bundle data, as well as further improve our management in Big Data.

GES DISC↗

MERRA Analytic Services: Meeting the Big Data Challenges of Climate Science Through Cloud-enabled Climate Analytics-as-a-service

Climate science is a Big Data domain that is experiencing unprecedented growth. In our efforts to address the Big Data challenges of climate science, we are moving toward a notion of Climate Analytics-as-a-Service (CAaaS). We focus on analytics, because it is the knowledge gained from our interactions with Big Data that ultimately produce societal benefits. We focus on CAaaS because we believe it provides a useful way of thinking about the problem: a specialization of the concept of business process-as-a-service, which is an evolving extension of IaaS, PaaS, and SaaS enabled by Cloud Computing. Within this framework, Cloud Computing plays an important role; however, we it see it as only one element in a constellation of capabilities that are essential to delivering climate analytics as a service. These elements are essential because in the aggregate they lead to generativity, a capacity for self-assembly that we feel is the key to solving many of the Big Data challenges in this domain. MERRA Analytic Services (MERRAAS) is an example of cloud-enabled CAaaS built on this principle. MERRAAS enables MapReduce analytics over NASAs Modern-Era Retrospective Analysis for Research and Applications (MERRA) data collection. The MERRA reanalysis integrates observational data with numerical models to produce a global temporally and spatially consistent synthesis of 26 key climate variables. It represents a type of data product that is of growing importance to scientists doing climate change research and a wide range of decision support applications. MERRAAS brings together the following generative elements in a full, end-to-end demonstration of CAaaS capabilities: (1) high-performance, data proximal analytics, (2) scalable data management, (3) software appliance virtualization, (4) adaptive analytics, and (5) a domain-harmonized API. The effectiveness of MERRAAS has been demonstrated in several applications. In our experience, Cloud Computing lowers the barriers and risk to organizational change, fosters innovation and experimentation, facilitates technology transfer, and provides the agility required to meet our customers' increasing and changing needs. Cloud Computing is providing a new tier in the data services stack that helps connect earthbound, enterprise-level data and computational resources to new customers and new mobility-driven applications and modes of work. For climate science, Cloud Computing's capacity to engage communities in the construction of new capabilies is perhaps the most important link between Cloud Computing and Big Data.

Data Analytics↗

Machine Learning Lifecycle for Earth Science Application: A Practical Insight into Production Deployment

Earth science domain presents unique sets of problems that are increasingly being solved using data driven approaches. The availability of big Earth science data offers immense potential for Machine learning (ML) as evident from numerous research publications lately. However, many of these publications are not ending up as production applications mainly because the data scientists who develop the ML models are now expected to complete the ML lifecycle by deploying and scaling the models in production. We introduce ML lifecycle to the Earth science community including the opportunities and challenges that lie ahead in each phase of the lifecycle. We demonstrate the lifecycle using an Earth science problem that we used ML to address and transitioned to production.

Maskey, Manil↗