Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data Challenge”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

A Parameter-masked Mock Data Challenge for Beyond-two-point Galaxy Clustering Statistics

The past few years have seen the emergence of a wide array of novel techniques for analyzing high-precision data from upcoming galaxy surveys, which aim to extend the statistical analysis of galaxy clustering data beyond the linear regime and the canonical two-point (2pt) statistics. We test and benchmark some of these new techniques in a community data challenge named “Beyond-2pt,” initiated during the Aspen 2022 Summer Program “Large-Scale Structure Cosmology beyond 2-Point Statistics,” whose first round of results we present here. The challenge data set consists of high-precision mock galaxy catalogs for clustering in real space, in redshift space, and on a light cone. Participants in the challenge have developed end-to-end pipelines to analyze mock catalogs and extract unknown (“masked”) cosmological parameters of the underlying ΛCDM models with their methods. The methods represented are density-split clustering, nearest neighbor statistics, BACCO power spectrum emulator, void statistics, LEFTfield field-level inference using effective field theory (EFT), and joint power spectrum and bispectrum analyses using both EFT and simulation-based inference. In this work, we review the results of the challenge, focusing on problems solved, lessons learned, and future research needed to perfect the emerging beyond-2pt approaches. The unbiased parameter recovery demonstrated in this challenge by multiple statistics and the associated modeling and inference frameworks supports the credibility of cosmology constraints from these methods. The challenge data set is publicly available, and we welcome future submissions from methods that are not yet represented.

Krause, Elisabeth [Univ. of Arizona, Tucson, AZ (U↗

An Overview of the Mock LISA Data Challenges

The LISA International Science Team Working Group on Data Analysis (LIST-WG1B) is sponsoring several rounds of mock data challenges, with the purpose of fostering the development of LISA data-analysis capabilities, and of demonstrating technical readiness for the maximum science exploitation of the LISA data. The first round of challenge data sets were released at this Symposium. We describe the objectives, structure, and timeline of this program.

black holes↗

The Mock LISA Data Challenge Round 3: New and Improved Sources

The Mock LISA Data Challenges are a program to demonstrate and encourage the development of data-analysis capabilities for LISA. Each round of challenges consists of several data sets containing simulated instrument noise and gravitational waves from sources of undisclosed parameters. Participants are asked to analyze the data sets and report the maximum information they can infer about the source parameters. The challenges are being released in rounds of increasing complexity and realism. Challenge 3. currently in progress, brings new source classes, now including cosmic-string cusps and primordial stochastic backgrounds, and more realistic signal models for supermassive black-hole inspirals and galactic double white dwarf binaries.

Baker, John↗

The LSST AGN Data Challenge: Selection Methods

Abstract Development of the Rubin Observatory Legacy Survey of Space and Time (LSST) includes a series of Data Challenges (DCs) arranged by various LSST Scientific Collaborations that are taking place during the project's preoperational phase. The AGN Science Collaboration Data Challenge (AGNSC-DC) is a partial prototype of the expected LSST data on active galactic nuclei (AGNs), aimed at validating machine learning approaches for AGN selection and characterization in large surveys like LSST. The AGNSC-DC took place in 2021, focusing on accuracy, robustness, and scalability. The training and the blinded data sets were constructed to mimic the future LSST release catalogs using the data from the Sloan Digital Sky Survey Stripe 82 region and the XMM-Newton Large Scale Structure Survey region. Data features were divided into astrometry, photometry, color, morphology, redshift, and class label with the addition of variability features and images. We present the results of four submitted solutions to DCs using both classical and machine learning methods. We systematically test the performance of supervised models (support vector machine, random forest, extreme gradient boosting, artificial neural network, convolutional neural network) and unsupervised ones (deep embedding clustering) when applied to the problem of classifying/clustering sources as stars, galaxies, or AGNs. We obtained classification accuracy of 97.5% for supervised models and clustering accuracy of 96.0% for unsupervised ones and 95.0% with a classic approach for a blinded data set. We find that variability features significantly improve the accuracy of the trained models, and correlation analysis among different bands enables a fast and inexpensive first-order selection of quasar candidates.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

MERRA Analytic Services: Meeting the Big Data Challenges of Climate Science Through Cloud-enabled Climate Analytics-as-a-service

Climate science is a Big Data domain that is experiencing unprecedented growth. In our efforts to address the Big Data challenges of climate science, we are moving toward a notion of Climate Analytics-as-a-Service (CAaaS). We focus on analytics, because it is the knowledge gained from our interactions with Big Data that ultimately produce societal benefits. We focus on CAaaS because we believe it provides a useful way of thinking about the problem: a specialization of the concept of business process-as-a-service, which is an evolving extension of IaaS, PaaS, and SaaS enabled by Cloud Computing. Within this framework, Cloud Computing plays an important role; however, we it see it as only one element in a constellation of capabilities that are essential to delivering climate analytics as a service. These elements are essential because in the aggregate they lead to generativity, a capacity for self-assembly that we feel is the key to solving many of the Big Data challenges in this domain. MERRA Analytic Services (MERRAAS) is an example of cloud-enabled CAaaS built on this principle. MERRAAS enables MapReduce analytics over NASAs Modern-Era Retrospective Analysis for Research and Applications (MERRA) data collection. The MERRA reanalysis integrates observational data with numerical models to produce a global temporally and spatially consistent synthesis of 26 key climate variables. It represents a type of data product that is of growing importance to scientists doing climate change research and a wide range of decision support applications. MERRAAS brings together the following generative elements in a full, end-to-end demonstration of CAaaS capabilities: (1) high-performance, data proximal analytics, (2) scalable data management, (3) software appliance virtualization, (4) adaptive analytics, and (5) a domain-harmonized API. The effectiveness of MERRAAS has been demonstrated in several applications. In our experience, Cloud Computing lowers the barriers and risk to organizational change, fosters innovation and experimentation, facilitates technology transfer, and provides the agility required to meet our customers' increasing and changing needs. Cloud Computing is providing a new tier in the data services stack that helps connect earthbound, enterprise-level data and computational resources to new customers and new mobility-driven applications and modes of work. For climate science, Cloud Computing's capacity to engage communities in the construction of new capabilies is perhaps the most important link between Cloud Computing and Big Data.

Data Analytics↗

New Era, New Opportunity, Is GES DISC Ready for Big Data Challenge?

The new era of Big Data has opened doors for many new opportunities, as well as new challenges, for both Earth science research/application and data communities. As one of the twelve NASA data centers - Goddard Earth Sciences Data and Information Services Center (GES DISC), one of our great challenges has been how to help research/application community efficiently (quickly and properly) accessing, visualizing and analyzing the massive and diverse data in natural hazard research, management, or even prediction. GES DISC has archived over 2000 TB data on premises and distributed over 23,000 TB of data since 2010. Our data has been widely used in every phase of natural hazard management and research, i.e. long term risk assessment and reduction, forecasting and predicting, monitoring and detection, early warning, damage assessment and response. The big data challenge is not just about data storage, but also about data discoverability and accessibility, and even more, about data migration/mirroring in the cloud. This paper is going to demonstrate GES DISC’s efforts and approaches of evolving our overall Web services and powerful Giovanni (Geospatial Interactive Online Visualization ANd aNalysis Infrastructure) tool into further improving data discoverability and accessibility. Prototype works will also be presented.

Li, A.↗

The Mock LISA Data Challenges: History, Status, Prospects

This slide presentation reviews the importance for the Mock LISA Data Challenges (MLDC). Laser Interferometer Space Antenna (LISA) is a gravitational wave (GW) observatory that will return data such that data analysis is integral to the measurement concept. Further rationale of the MLDC are to kickstart the development of a LISA data-analysis computational infrastructure, and to encourage, track, and compare progress in LISA data-analysis development in the open community. The MLDCs is a coordinated, voluntary effort in GW community, that will periodically issue datasets with synthetic noise and GW signals from sources of undisclosed parameters; increasing difficulty. The challenge participants return parameter estimates and descriptions of search methods. Some of the challenges and the resultant entries are reviewed. The aim is to show that LISA data analysis is possible, and to develop new techniques, using multiple international teams for the development of LISA core analysis tools

data analysis↗

RTN-107: The Rubin Observatory Target-of-Opportunity Mock Data Challenge

We describe the activities of the Target-of-Opportunity mock data challenge, taking place from Sep 22 2025 - Oct 18 2025. We center this activity in four questions that are critical for maximizing the scientific output of the ToO system: (1) How quickly can Rubin Observatory start observing after a ToO alert is received? (2) How efficient is Rubin Observatory at recovering the host of a ToO event? (3) How accurate are the observing strategies that the community has created for the ToO program? (4) How can expert ToO scientists interact effectively with the ToO system, where many processes are fully automated? In this challenge, and the report summarized herein, we aim to answer the aforementioned questions to better the Rubin ToO program.

79 ASTRONOMY AND ASTROPHYSICS↗

Big Data Challenges at CCMC

Like other research centers, the Community Coordinated Modeling Center (CCMC, https://ccmc.gsfc.nasa.gov) at NASA Goddard Space Flight Center (GSFC) is also experiencing the big data challenges. CCMC hosts over 80 space weather models for Runs On Request (ROR), Continuous Runs and Instant Runs simulation services for the research community. In addition, CCMC has started to support simulation output onboarding in response to the Open Science initiative. Overall, we have accumulated over petabytes of simulation output data and are rapidly growing. In this presentation, we will discuss our data and storage challenges. We will present our attempts to address our challenges and any associated lessons learned. CCMC uses Apache Airflow to ensure data transfer is consistent. We will give a brief overview on how we leverage Apache Airflow to enhance our environment.

space weather↗

Big Data Challenges at the Community Coordinated Modeling Center (CCMC)

Like other research centers, the Community Coordinated Modeling Center (CCMC, https://ccmc.gsfc.nasa.gov) at NASA Goddard Space Flight Center (GSFC) is also experiencing the big data challenges. CCMC hosts over 80 space weather models for Runs On Request (ROR), Continuous Runs and Instant Runs simulation services for the research community. In addition, CCMC has started to support simulation output onboarding in response to the Open Science initiative. Overall, we have accumulated over petabytes of simulation output data and are rapidly growing. In this presentation, we will discuss our data and storage challenges. We will present our attempts to address our challenges and any associated lessons learned. CCMC uses Apache Airflow to ensure data transfer is consistent. We will give a brief overview on how we leverage Apache Airflow to enhance our environment.

space weather↗

The Status of the Mock LISA Data Challenges

For the last four years, many gravitational-wave researchers around the world have participated in the Mock LISA Data Challenges (MLDCs), a program to demonstrate and encourage the development of LISA data-analysis capabilities, tools and techniques. In this poster, we present a summary of the results of MLDC 3, which was completed in 2009. During MLDC 3, 27 participants from 15 institutions successfully analyzed data sets that included Galactic binaries, coalescing spinning massive black holes, extreme-mass-ratio inspirals, cosmic-string cusp bursts and a stochastic gravitational-wave background. We also describe the technical and scientific challenges that will be addressed by future MLI)Cs, starting with MLDC 4, which is currently in progress.

Baker, John↗

GLOBE Observer and the GO on a Trail Data Challenge: A Citizen Science Approach to Generating a Global Land Cover Land Use Reference Dataset

Land cover and land use are highly visible indicators of climate change and human disruption to natural processes. While land cover is frequently monitored over a large area using satellite data, ground-based reference data is valuable as a comparison point. The NASA-funded GLOBE Observer (GO) program provides volunteer-collected land cover photos tagged with location, date and time, and, in some cases, land cover type. When making a full land cover observation, volunteers take six photos of the site, one facing north, south, east, and west (N-S-E-W) respectively, one pointing straight up to capture canopy and sky, and one pointing down to document ground cover. Together, the photos document a 100 meter square of land. Volunteers may then optionally tag each N-S-E-W photo with the land cover types present. Volunteers collect the data through a smartphone app, also called GLOBE Observer, resulting in consistent data. While land cover data collected through GLOBE Observer is ongoing, this paper presents the results of a data challenge held between June 1 and October 15, 2019. Called “GO on a Trail,” the challenge focused on collecting land cover observations at designated points along the Lewis and Clark National Historic Trail (LCNHT), which transects approximately 7,900 kilometers of the United States from east to west. The challenge also included an international component in which participants were recognized for the most data collected during the period. The challenge resulted in more than 2800 land cover data points from around the world with just over 900 land cover data points along the LCNHT. The data collected along the LCNHT is largely within a 1,000 meter corridor along the historic route. Internationally, the challenge had strong participation in Australia, where a Scouts Australia team-based competition was conducted in conjunction with GO on a Trail. GLOBE Observer collections can serve as reference data, ground truthing satellite imagery for the improvement and verification of broad land cover maps. Continued collection using this protocol will build a database documenting climate-related land cover and land use change into the future.

land cover↗

MSD CoP Webinar: AI and Extreme Events - Overcoming Data Challenges for Improved Characterization of Climate Extremes

Context: This webinar was hosted by the MultiSector Dynamics Community of Practice (MSD CoP; https://multisectordynamics.org). Abstract: Artificial Intelligence (AI) models require large volumes of data for training and testing. Data requirements present challenges for using AI to explore extreme events with limited observational data. This webinar will showcase two innovative methods developed by part of the European Climate Intelligence (CLINT) project to overcome data challenges and harness AI to improve our understanding of climate extremes. Dr. Ascenso will present his research on data augmentation methods to improve estimates of tropical cyclones using satellite data. His presentation will review established methods for data augmentation and explore opportunities and challenges for using generative AI to generate images of extreme, life-threatening tropical cyclones. Next, Dr. Plesiat will present his research on deep learning techniques to overcome limited observational data sets. His presentation will illustrate deep learning methods to develop AI reconstructions of four climate indices across Europe. Presenters : Dr. Guido Ascenso (post-doctoral researcher, Politecnico di Milano); Dr. Étienne Plésiat (German Climate Computing Centre - DKRZ) Moderator(s): Stefano Galelli (MSD CoP WG Co-Lead), David Gold (MSD CoP WG Co-Lead), Jillian Sturtevant (MSD CoP WG Communications Officer), Matteo Giuliani (Politecnico di Milano, MSD CoP WG Member, Moderator and Organizer) This webinar was held on: October 11, 2024 from 11AM - 1PM ET

AI↗

High Resolution Nature Runs and the Big Data Challenge

NASA's Global Modeling and Assimilation Office at Goddard Space Flight Center is undertaking a series of very computationally intensive Nature Runs and a downscaled reanalysis. The nature runs use the GEOS-5 as an Atmospheric General Circulation Model (AGCM) while the reanalysis uses the GEOS-5 in Data Assimilation mode. This paper will present computational challenges from three runs, two of which are AGCM and one is downscaled reanalysis using the full DAS. The nature runs will be completed at two surface grid resolutions, 7 and 3 kilometers and 72 vertical levels. The 7 km run spanned 2 years (2005-2006) and produced 4 PB of data while the 3 km run will span one year and generate 4 BP of data. The downscaled reanalysis (MERRA-II Modern-Era Reanalysis for Research and Applications) will cover 15 years and generate 1 PB of data. Our efforts to address the big data challenges of climate science, we are moving toward a notion of Climate Analytics-as-a-Service (CAaaS), a specialization of the concept of business process-as-a-service that is an evolving extension of IaaS, PaaS, and SaaS enabled by cloud computing. In this presentation, we will describe two projects that demonstrate this shift. MERRA Analytic Services (MERRA/AS) is an example of cloud-enabled CAaaS. MERRA/AS enables MapReduce analytics over MERRA reanalysis data collection by bringing together the high-performance computing, scalable data management, and a domain-specific climate data services API. NASA's High-Performance Science Cloud (HPSC) is an example of the type of compute-storage fabric required to support CAaaS. The HPSC comprises a high speed Infinib and network, high performance file systems and object storage, and a virtual system environments specific for data intensive, science applications. These technologies are providing a new tier in the data and analytic services stack that helps connect earthbound, enterprise-level data and computational resources to new customers and new mobility-driven applications and modes of work. In our experience, CAaaS lowers the barriers and risk to organizational change, fosters innovation and experimentation, and provides the agility required to meet our customers' increasing and changing needs

big data analysis↗

Big Data Challenges for Large Radio Arrays

Future large radio astronomy arrays, particularly the Square Kilometre Array (SKA), will be able to generate data at rates far higher than can be analyzed or stored affordably with current practices. This is, by definition, a "big data" problem, and requires an end-to-end solution if future radio arrays are to reach their full scientific potential. Similar data processing, transport, storage, and management challenges face next-generation facilities in many other fields.

Combining↗

Clouds Around the World: How a Simple Citizen Science Data Challenge Became a Worldwide Success

Citizen science is often recognized for its potential to directly engage the public in science, and is uniquely positioned to support and extend participants’ learning in science. In March 2018, the Global Learning and Observations to Benefit the Environment (GLOBE) Program, NASA’s largest and longest-lasting citizen science program about Earth, organized a month-long event that asked people around the world to contribute daily cloud observations and photographs of the sky (15 March–15 April 2018). What was considered a simple engagement activity turned into an unprecedented worldwide event that garnered major public interest and media recognition, collecting over 55,000 observations from 99 different countries, in more than 15,000 locations, on every continent including Antarctica. The event was called the “Spring Cloud Challenge” and was created to 1) engage the general public in the scientific process and promote the use of the GLOBE Observer app, 2) collect ground-based visual observations of varying cloud types during boreal spring, and 3) increase the number and locations of ground-based visual cloud observations collocated with cloud-observing satellites. The event resulted in roughly 3 times more observations than during the historic and highly publicized 2017 North American total solar eclipse. The dataset also includes observations over the Drake Passage in Antarctica and reports from intense Saharan dust events. This article describes how the challenge was crafted, outreach to volunteer scientists around the world, details of the data collected, and impact of the data.

GLOBE↗