Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data Challenge”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

An Overview of the Mock LISA Data Challenges

The LISA International Science Team Working Group on Data Analysis (LIST-WG1B) is sponsoring several rounds of mock data challenges, with the purpose of fostering the development of LISA data-analysis capabilities, and of demonstrating technical readiness for the maximum science exploitation of the LISA data. The first round of challenge data sets were released at this Symposium. We describe the objectives, structure, and timeline of this program.

black holes↗

The Mock LISA Data Challenge Round 3: New and Improved Sources

The Mock LISA Data Challenges are a program to demonstrate and encourage the development of data-analysis capabilities for LISA. Each round of challenges consists of several data sets containing simulated instrument noise and gravitational waves from sources of undisclosed parameters. Participants are asked to analyze the data sets and report the maximum information they can infer about the source parameters. The challenges are being released in rounds of increasing complexity and realism. Challenge 3. currently in progress, brings new source classes, now including cosmic-string cusps and primordial stochastic backgrounds, and more realistic signal models for supermassive black-hole inspirals and galactic double white dwarf binaries.

Baker, John↗

MERRA Analytic Services: Meeting the Big Data Challenges of Climate Science Through Cloud-enabled Climate Analytics-as-a-service

Climate science is a Big Data domain that is experiencing unprecedented growth. In our efforts to address the Big Data challenges of climate science, we are moving toward a notion of Climate Analytics-as-a-Service (CAaaS). We focus on analytics, because it is the knowledge gained from our interactions with Big Data that ultimately produce societal benefits. We focus on CAaaS because we believe it provides a useful way of thinking about the problem: a specialization of the concept of business process-as-a-service, which is an evolving extension of IaaS, PaaS, and SaaS enabled by Cloud Computing. Within this framework, Cloud Computing plays an important role; however, we it see it as only one element in a constellation of capabilities that are essential to delivering climate analytics as a service. These elements are essential because in the aggregate they lead to generativity, a capacity for self-assembly that we feel is the key to solving many of the Big Data challenges in this domain. MERRA Analytic Services (MERRAAS) is an example of cloud-enabled CAaaS built on this principle. MERRAAS enables MapReduce analytics over NASAs Modern-Era Retrospective Analysis for Research and Applications (MERRA) data collection. The MERRA reanalysis integrates observational data with numerical models to produce a global temporally and spatially consistent synthesis of 26 key climate variables. It represents a type of data product that is of growing importance to scientists doing climate change research and a wide range of decision support applications. MERRAAS brings together the following generative elements in a full, end-to-end demonstration of CAaaS capabilities: (1) high-performance, data proximal analytics, (2) scalable data management, (3) software appliance virtualization, (4) adaptive analytics, and (5) a domain-harmonized API. The effectiveness of MERRAAS has been demonstrated in several applications. In our experience, Cloud Computing lowers the barriers and risk to organizational change, fosters innovation and experimentation, facilitates technology transfer, and provides the agility required to meet our customers' increasing and changing needs. Cloud Computing is providing a new tier in the data services stack that helps connect earthbound, enterprise-level data and computational resources to new customers and new mobility-driven applications and modes of work. For climate science, Cloud Computing's capacity to engage communities in the construction of new capabilies is perhaps the most important link between Cloud Computing and Big Data.

Data Analytics↗

New Era, New Opportunity, Is GES DISC Ready for Big Data Challenge?

The new era of Big Data has opened doors for many new opportunities, as well as new challenges, for both Earth science research/application and data communities. As one of the twelve NASA data centers - Goddard Earth Sciences Data and Information Services Center (GES DISC), one of our great challenges has been how to help research/application community efficiently (quickly and properly) accessing, visualizing and analyzing the massive and diverse data in natural hazard research, management, or even prediction. GES DISC has archived over 2000 TB data on premises and distributed over 23,000 TB of data since 2010. Our data has been widely used in every phase of natural hazard management and research, i.e. long term risk assessment and reduction, forecasting and predicting, monitoring and detection, early warning, damage assessment and response. The big data challenge is not just about data storage, but also about data discoverability and accessibility, and even more, about data migration/mirroring in the cloud. This paper is going to demonstrate GES DISC’s efforts and approaches of evolving our overall Web services and powerful Giovanni (Geospatial Interactive Online Visualization ANd aNalysis Infrastructure) tool into further improving data discoverability and accessibility. Prototype works will also be presented.

Li, A.↗

The Mock LISA Data Challenges: History, Status, Prospects

This slide presentation reviews the importance for the Mock LISA Data Challenges (MLDC). Laser Interferometer Space Antenna (LISA) is a gravitational wave (GW) observatory that will return data such that data analysis is integral to the measurement concept. Further rationale of the MLDC are to kickstart the development of a LISA data-analysis computational infrastructure, and to encourage, track, and compare progress in LISA data-analysis development in the open community. The MLDCs is a coordinated, voluntary effort in GW community, that will periodically issue datasets with synthetic noise and GW signals from sources of undisclosed parameters; increasing difficulty. The challenge participants return parameter estimates and descriptions of search methods. Some of the challenges and the resultant entries are reviewed. The aim is to show that LISA data analysis is possible, and to develop new techniques, using multiple international teams for the development of LISA core analysis tools

data analysis↗

Big Data Challenges at CCMC

Like other research centers, the Community Coordinated Modeling Center (CCMC, https://ccmc.gsfc.nasa.gov) at NASA Goddard Space Flight Center (GSFC) is also experiencing the big data challenges. CCMC hosts over 80 space weather models for Runs On Request (ROR), Continuous Runs and Instant Runs simulation services for the research community. In addition, CCMC has started to support simulation output onboarding in response to the Open Science initiative. Overall, we have accumulated over petabytes of simulation output data and are rapidly growing. In this presentation, we will discuss our data and storage challenges. We will present our attempts to address our challenges and any associated lessons learned. CCMC uses Apache Airflow to ensure data transfer is consistent. We will give a brief overview on how we leverage Apache Airflow to enhance our environment.

space weather↗

Big Data Challenges at the Community Coordinated Modeling Center (CCMC)

Like other research centers, the Community Coordinated Modeling Center (CCMC, https://ccmc.gsfc.nasa.gov) at NASA Goddard Space Flight Center (GSFC) is also experiencing the big data challenges. CCMC hosts over 80 space weather models for Runs On Request (ROR), Continuous Runs and Instant Runs simulation services for the research community. In addition, CCMC has started to support simulation output onboarding in response to the Open Science initiative. Overall, we have accumulated over petabytes of simulation output data and are rapidly growing. In this presentation, we will discuss our data and storage challenges. We will present our attempts to address our challenges and any associated lessons learned. CCMC uses Apache Airflow to ensure data transfer is consistent. We will give a brief overview on how we leverage Apache Airflow to enhance our environment.

space weather↗

The Status of the Mock LISA Data Challenges

For the last four years, many gravitational-wave researchers around the world have participated in the Mock LISA Data Challenges (MLDCs), a program to demonstrate and encourage the development of LISA data-analysis capabilities, tools and techniques. In this poster, we present a summary of the results of MLDC 3, which was completed in 2009. During MLDC 3, 27 participants from 15 institutions successfully analyzed data sets that included Galactic binaries, coalescing spinning massive black holes, extreme-mass-ratio inspirals, cosmic-string cusp bursts and a stochastic gravitational-wave background. We also describe the technical and scientific challenges that will be addressed by future MLI)Cs, starting with MLDC 4, which is currently in progress.

Baker, John↗

GLOBE Observer and the GO on a Trail Data Challenge: A Citizen Science Approach to Generating a Global Land Cover Land Use Reference Dataset

Land cover and land use are highly visible indicators of climate change and human disruption to natural processes. While land cover is frequently monitored over a large area using satellite data, ground-based reference data is valuable as a comparison point. The NASA-funded GLOBE Observer (GO) program provides volunteer-collected land cover photos tagged with location, date and time, and, in some cases, land cover type. When making a full land cover observation, volunteers take six photos of the site, one facing north, south, east, and west (N-S-E-W) respectively, one pointing straight up to capture canopy and sky, and one pointing down to document ground cover. Together, the photos document a 100 meter square of land. Volunteers may then optionally tag each N-S-E-W photo with the land cover types present. Volunteers collect the data through a smartphone app, also called GLOBE Observer, resulting in consistent data. While land cover data collected through GLOBE Observer is ongoing, this paper presents the results of a data challenge held between June 1 and October 15, 2019. Called “GO on a Trail,” the challenge focused on collecting land cover observations at designated points along the Lewis and Clark National Historic Trail (LCNHT), which transects approximately 7,900 kilometers of the United States from east to west. The challenge also included an international component in which participants were recognized for the most data collected during the period. The challenge resulted in more than 2800 land cover data points from around the world with just over 900 land cover data points along the LCNHT. The data collected along the LCNHT is largely within a 1,000 meter corridor along the historic route. Internationally, the challenge had strong participation in Australia, where a Scouts Australia team-based competition was conducted in conjunction with GO on a Trail. GLOBE Observer collections can serve as reference data, ground truthing satellite imagery for the improvement and verification of broad land cover maps. Continued collection using this protocol will build a database documenting climate-related land cover and land use change into the future.

land cover↗

High Resolution Nature Runs and the Big Data Challenge

NASA's Global Modeling and Assimilation Office at Goddard Space Flight Center is undertaking a series of very computationally intensive Nature Runs and a downscaled reanalysis. The nature runs use the GEOS-5 as an Atmospheric General Circulation Model (AGCM) while the reanalysis uses the GEOS-5 in Data Assimilation mode. This paper will present computational challenges from three runs, two of which are AGCM and one is downscaled reanalysis using the full DAS. The nature runs will be completed at two surface grid resolutions, 7 and 3 kilometers and 72 vertical levels. The 7 km run spanned 2 years (2005-2006) and produced 4 PB of data while the 3 km run will span one year and generate 4 BP of data. The downscaled reanalysis (MERRA-II Modern-Era Reanalysis for Research and Applications) will cover 15 years and generate 1 PB of data. Our efforts to address the big data challenges of climate science, we are moving toward a notion of Climate Analytics-as-a-Service (CAaaS), a specialization of the concept of business process-as-a-service that is an evolving extension of IaaS, PaaS, and SaaS enabled by cloud computing. In this presentation, we will describe two projects that demonstrate this shift. MERRA Analytic Services (MERRA/AS) is an example of cloud-enabled CAaaS. MERRA/AS enables MapReduce analytics over MERRA reanalysis data collection by bringing together the high-performance computing, scalable data management, and a domain-specific climate data services API. NASA's High-Performance Science Cloud (HPSC) is an example of the type of compute-storage fabric required to support CAaaS. The HPSC comprises a high speed Infinib and network, high performance file systems and object storage, and a virtual system environments specific for data intensive, science applications. These technologies are providing a new tier in the data and analytic services stack that helps connect earthbound, enterprise-level data and computational resources to new customers and new mobility-driven applications and modes of work. In our experience, CAaaS lowers the barriers and risk to organizational change, fosters innovation and experimentation, and provides the agility required to meet our customers' increasing and changing needs

big data analysis↗

Big Data Challenges for Large Radio Arrays

Future large radio astronomy arrays, particularly the Square Kilometre Array (SKA), will be able to generate data at rates far higher than can be analyzed or stored affordably with current practices. This is, by definition, a "big data" problem, and requires an end-to-end solution if future radio arrays are to reach their full scientific potential. Similar data processing, transport, storage, and management challenges face next-generation facilities in many other fields.

Combining↗

Clouds Around the World: How a Simple Citizen Science Data Challenge Became a Worldwide Success

Citizen science is often recognized for its potential to directly engage the public in science, and is uniquely positioned to support and extend participants’ learning in science. In March 2018, the Global Learning and Observations to Benefit the Environment (GLOBE) Program, NASA’s largest and longest-lasting citizen science program about Earth, organized a month-long event that asked people around the world to contribute daily cloud observations and photographs of the sky (15 March–15 April 2018). What was considered a simple engagement activity turned into an unprecedented worldwide event that garnered major public interest and media recognition, collecting over 55,000 observations from 99 different countries, in more than 15,000 locations, on every continent including Antarctica. The event was called the “Spring Cloud Challenge” and was created to 1) engage the general public in the scientific process and promote the use of the GLOBE Observer app, 2) collect ground-based visual observations of varying cloud types during boreal spring, and 3) increase the number and locations of ground-based visual cloud observations collocated with cloud-observing satellites. The event resulted in roughly 3 times more observations than during the historic and highly publicized 2017 North American total solar eclipse. The dataset also includes observations over the Drake Passage in Antarctica and reports from intense Saharan dust events. This article describes how the challenge was crafted, outreach to volunteer scientists around the world, details of the data collected, and impact of the data.

GLOBE↗

Data Quality Challenges for Analysis Ready Data (ARD)

Data quality plays a critical role in research and applications. The Earth Science Information Partners (ESIP) Information Quality Cluster (IQC) defines four aspects of information quality: Science, Product, Stewardship, and Services. The ESIP IQC has become internationally recognized as an authoritative and responsive resource of information and guidance to data producers and distributors on how to implement data quality standards and best practices for their science data systems, datasets, and data/metadata dissemination services. In recent years, cloud computing environments have provided scale-up capabilities such as data archives and services, enabling interdisciplinary science and applications. More value-added products are expected from data service providers, including Analysis Ready Data (ARD). ARD refers to data that has been preprocessed into a form that allows immediate analysis by the end user, processed to a minimum set of requirements and provides interoperability over time and across multiple datasets. Once a dataset has been developed from its original form to produce ARD, what quality characteristics should the derived dataset or ARD possess? Also, is it safe to assume that the quality of the ARD is consistent with the quality of the source data, or are there special attributes to an ARD that would warrant a secondary, independent quality assessment? What provenance (also called “data lineage”) information needs to be included in ARD? It is important to answer these questions, especially given the ease of use of ARD, and the consequent temptation by users to trust ARD without understanding the limitations or possible variations in quality compared to the source data. In this presentation, we will discuss data quality challenges for ARD products and services and introduce IQC for participation.

data quality↗

Using global aerosol models and satellite data for air quality studies: Challenges and data needs

Aerosol particles, also known as PM2.5 (particle diameter less than 2.5 pm) and PM10 (particle diameter less than 10 pm), are one of the key atmospheric components that determines air quality. Yet, air quality forecasts for PM are still in their infancy and remain a challenging task. It is difficult to simply relate PM levels to local meteorological conditions, and large uncertainties exist in regional air quality model emission inventories and initial and boundary conditions. Especially challenging are periods when a significant amount of aerosol comes from outside the regional modeling domain through long-range transport. In the past few years, NASA has launched several satellites with global aerosol measurement capabilities, providing large-scale chemical weather pictures. NASA has also supported development of global models which simulate atmospheric transport and transformation processes of important atmospheric gas and aerosol species. I will present the current modeling and satellite capabilities for PM2.5 studies, the possibilities and challenges in using satellite data for PM2.5 forecasts, and the needs of future remote sensing data for improving air quality monitoring and modeling.

Chin, Mian↗

Data Democratization: Challenges and Opportunities

Democratizing Earth data is one of the challenges many organizations around the world face in order to maximize the use of their Earth data for research, applications, education, and societal benefits. For example, at the NASA Goddard Earth Sciences (GES) Data and Information Services Center (DISC), over 1600 global and regional datasets in several NASA Earth science focus areas, including atmospheric composition, water and energy cycles, and climate variability, are archived and distributed to the public. Giovanni, the Geospatial Interactive Online Visualization and Analysis Infrastructure, was developed by GES DISC to facilitate data access and exploration, especially for novice users of Earth science. With Giovanni, users can analyze and visualize over 2000 Earth science variables (e.g., precipitation, aerosol, surface wind) without downloading data, software, the expert understanding of data formats and structures, and coding skills, lowering the barrier to data analysis/comparison by preprocessing and accessing to the data. Results of data analysis and visualization can be accessed in several popular formats (e.g., NetCDF, CSV). As a result of Giovanni's efforts, more than 3000 referral papers have been published in various fields. In spite of this, Giovanni is still difficult to use for some users. For instance, if one searches for "precipitation," it will return over 150 related variables. The question is, which one to use? Furthermore, variables from different data providers (e.g., satellites and models) are named differently with different units, further confusing users, especially those outside the communities. Data democratization is complex and multifaceted. Challenges include service and data discovery, user experiences, visualization, data quality, trustworthiness, and more. In this presentation, we will examine Giovanni as an example of challenges and opportunities in developing data democratization services.

data democratization↗