Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data Challenge”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

A Parameter-masked Mock Data Challenge for Beyond-two-point Galaxy Clustering Statistics

The past few years have seen the emergence of a wide array of novel techniques for analyzing high-precision data from upcoming galaxy surveys, which aim to extend the statistical analysis of galaxy clustering data beyond the linear regime and the canonical two-point (2pt) statistics. We test and benchmark some of these new techniques in a community data challenge named “Beyond-2pt,” initiated during the Aspen 2022 Summer Program “Large-Scale Structure Cosmology beyond 2-Point Statistics,” whose first round of results we present here. The challenge data set consists of high-precision mock galaxy catalogs for clustering in real space, in redshift space, and on a light cone. Participants in the challenge have developed end-to-end pipelines to analyze mock catalogs and extract unknown (“masked”) cosmological parameters of the underlying ΛCDM models with their methods. The methods represented are density-split clustering, nearest neighbor statistics, BACCO power spectrum emulator, void statistics, LEFTfield field-level inference using effective field theory (EFT), and joint power spectrum and bispectrum analyses using both EFT and simulation-based inference. In this work, we review the results of the challenge, focusing on problems solved, lessons learned, and future research needed to perfect the emerging beyond-2pt approaches. The unbiased parameter recovery demonstrated in this challenge by multiple statistics and the associated modeling and inference frameworks supports the credibility of cosmology constraints from these methods. The challenge data set is publicly available, and we welcome future submissions from methods that are not yet represented.

Krause, Elisabeth [Univ. of Arizona, Tucson, AZ (U↗

The LSST AGN Data Challenge: Selection Methods

Abstract Development of the Rubin Observatory Legacy Survey of Space and Time (LSST) includes a series of Data Challenges (DCs) arranged by various LSST Scientific Collaborations that are taking place during the project's preoperational phase. The AGN Science Collaboration Data Challenge (AGNSC-DC) is a partial prototype of the expected LSST data on active galactic nuclei (AGNs), aimed at validating machine learning approaches for AGN selection and characterization in large surveys like LSST. The AGNSC-DC took place in 2021, focusing on accuracy, robustness, and scalability. The training and the blinded data sets were constructed to mimic the future LSST release catalogs using the data from the Sloan Digital Sky Survey Stripe 82 region and the XMM-Newton Large Scale Structure Survey region. Data features were divided into astrometry, photometry, color, morphology, redshift, and class label with the addition of variability features and images. We present the results of four submitted solutions to DCs using both classical and machine learning methods. We systematically test the performance of supervised models (support vector machine, random forest, extreme gradient boosting, artificial neural network, convolutional neural network) and unsupervised ones (deep embedding clustering) when applied to the problem of classifying/clustering sources as stars, galaxies, or AGNs. We obtained classification accuracy of 97.5% for supervised models and clustering accuracy of 96.0% for unsupervised ones and 95.0% with a classic approach for a blinded data set. We find that variability features significantly improve the accuracy of the trained models, and correlation analysis among different bands enables a fast and inexpensive first-order selection of quasar candidates.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

RTN-107: The Rubin Observatory Target-of-Opportunity Mock Data Challenge

We describe the activities of the Target-of-Opportunity mock data challenge, taking place from Sep 22 2025 - Oct 18 2025. We center this activity in four questions that are critical for maximizing the scientific output of the ToO system: (1) How quickly can Rubin Observatory start observing after a ToO alert is received? (2) How efficient is Rubin Observatory at recovering the host of a ToO event? (3) How accurate are the observing strategies that the community has created for the ToO program? (4) How can expert ToO scientists interact effectively with the ToO system, where many processes are fully automated? In this challenge, and the report summarized herein, we aim to answer the aforementioned questions to better the Rubin ToO program.

79 ASTRONOMY AND ASTROPHYSICS↗

MSD CoP Webinar: AI and Extreme Events - Overcoming Data Challenges for Improved Characterization of Climate Extremes

Context: This webinar was hosted by the MultiSector Dynamics Community of Practice (MSD CoP; https://multisectordynamics.org). Abstract: Artificial Intelligence (AI) models require large volumes of data for training and testing. Data requirements present challenges for using AI to explore extreme events with limited observational data. This webinar will showcase two innovative methods developed by part of the European Climate Intelligence (CLINT) project to overcome data challenges and harness AI to improve our understanding of climate extremes. Dr. Ascenso will present his research on data augmentation methods to improve estimates of tropical cyclones using satellite data. His presentation will review established methods for data augmentation and explore opportunities and challenges for using generative AI to generate images of extreme, life-threatening tropical cyclones. Next, Dr. Plesiat will present his research on deep learning techniques to overcome limited observational data sets. His presentation will illustrate deep learning methods to develop AI reconstructions of four climate indices across Europe. Presenters : Dr. Guido Ascenso (post-doctoral researcher, Politecnico di Milano); Dr. Étienne Plésiat (German Climate Computing Centre - DKRZ) Moderator(s): Stefano Galelli (MSD CoP WG Co-Lead), David Gold (MSD CoP WG Co-Lead), Jillian Sturtevant (MSD CoP WG Communications Officer), Matteo Giuliani (Politecnico di Milano, MSD CoP WG Member, Moderator and Organizer) This webinar was held on: October 11, 2024 from 11AM - 1PM ET

AI↗

Challenges of open data in aquatic sciences: issues faced by data users and data providers

Free use and redistribution of data (i.e., Open Data) increases the reproducibility, transparency, and pace of aquatic sciences research. However, barriers to both data users and data providers may limit the adoption of Open Data practices. Here, we describe common Open Data challenges faced by data users and data providers within the aquatic sciences community (i.e., oceanography, limnology, hydrology, and others). These challenges were synthesized from literature, authors’ experiences, and a broad survey of 174 data users and data providers across academia, government agencies, industry, and other sectors. Through this work, we identified seven main challenges: 1) metadata shortcomings, 2) variable data quality and reusability, 3) open data inaccessibility, 4) lack of standardization, 5) authorship and acknowledgement issues 6) lack of funding, and 7) unequal barriers around the globe. Our key recommendation is to improve resources to advance Open Data practices. This includes dedicated funds for capacity building, hiring and maintaining of skilled personnel, and robust digital infrastructures for preparation, storage, and long-term maintenance of Open Data. Further, to incentivize data sharing we reinforce the need for standardized best practices to handle data acknowledgement and citations for both data users and data providers. We also highlight and discuss regional disparities in resources and research practices within a global perspective.

54 ENVIRONMENTAL SCIENCES↗

Revolutionizing thermal Management in Next-Generation AI data centers: Challenges and breakthrough innovations

Data centers (DCs) serve as critical infrastructure for powering the growth and evolution of AI. Next-generation AI DCs present unique challenges in thermal management driven by unprecedented computational demands. This paper provides a comprehensive summary of key stakeholder perspectives on technology gaps, infrastructure requirements, test bed needs, emerging opportunities, and preliminary solutions related to thermal management for AI DCs. It establishes six strategic pillars of thermal management for next generation AI DC: reliability, deployability, efficiency, resilience, measurability, and valorization. The discussion spans a range of critical topics, including advanced cooling technologies, thermal strategies for emerging modular and edge DCs, system-level optimization and control frameworks, infrastructure planning and grid integration designs, benchmarking approaches, and pathways for waste heat recovery and reuse. The proposed research, development, and demonstration efforts are aimed at accelerating the deployment of AI DCs while ensuring energy efficiency, reliability, safety, and regulatory compliance.

Wang, Pengtao [ORNL] (ORCID:0000000214713429)↗

Application of a Prize Mechanism to Address Data Utilization Challenges at Utilities

The electric industry sector is facing an “explosion” of data from a variety of sources. Electric sector stakeholders need to define how to capitalize on large datasets, both those they create and those from other sources (like data on weather, buildings, electric vehicles, etc.), to improve reliability and resilience and meet the changing system dynamics from renewable integration. For the electricity sector to fully utilize these vast new datasets, it must undergo a transformation in how it manages data quality, storage, and processing. The U.S. Department of Energy (DOE) Office of Electricity (OE) is committed to accelerating research, development, and demonstration of new technologies and tools within the electricity sector to advance reliability, resilience, and affordable operation of the power system. Through the prize mechanism, OE identified two widespread data-related challenges for utilities—load modeling and data analysis automation—and offered an opportunity for utilities and teams of software engineers to identify additional challenges faced by utilities. After completing one round of the American-Made Digitizing Utilities Prize, OE, the National Renewable Energy Laboratory (NREL) as the prize administrator, and Pacific Northwest National Laboratory (PNNL) as the domain experts have compiled the results and lessons learned to feed into the second round of the prize.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

ECON-T and ECON-D: Endcap Concentrator ASICs for the CMS HGCAL

With over 6 million channels, the High Granularity Calorimeter (HGCAL) for the CMS HL-LHC Upgrade presents a unique data challenge. The ECON ASICs provide critical on-detector data reduction for the 40 MHz trigger path (ECON-T) and 750 kHz data acquisition path (ECON-D) of the HGCAL. The ASICs, fabricated in 65 nm CMOS, are rad-tolerant (600 Mrad) with low power consumption (<2.5 mW/channel). This presentation is the first comprehensive description of the ECON designs, first functionality and radiation tests for the ECON-T ASIC, and first results from the full production of 75k ECON-D and ECON-T ASICs.

Bergamin, G. [CERN] (ORCID:0000000285758704)↗

Navigating Integration: Key Challenges for Data Centers, Nuclear Stakeholders, and Utility Operators

The rapid expansion of data centers, driven by the exponential growth in data-processing and storage needs, presents significant challenges and opportunities for various stakeholders, including data center developers, nuclear energy providers, and utility companies. Data centers are projected to consume 6.7–12% of United States (U.S.) electricity by 2028, driven by artificial intelligence (AI) and cloud-computing demands. Nuclear energy offers reliability and dispatchable baseload power, but data centers need power now while nuclear still needs time to address siting, fast power ramping, and regulatory hurdles. Utilities must keep pace with the unprecedented acceleration of large load interconnection requests and urgently adapt to high-density loads while maintaining grid stability, reliability, and accelerating interconnection timelines. This report dives into these challenges and proposes key collaboration strategies to streamline data center integration that aligns with recent federal initiatives like America’s AI Action Plan and related executive orders that emphasize the importance of data center growth, nuclear energy expansion, and maintaining a competitive edge in the global AI race.

21 - SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLAN↗

The Persistent Challenge of Data Locality in the Post-Exascale Era

The era of exascale computing, exemplified by systems like Frontier achieving exaflop-level performance, marks a milestone. However, the quest for sheer compute power leads to strong imbalance in system design. Hence, scaling advancements in memory, network bandwidth, and storage are also necessary and pose challenges, with a crucial need to address data locality issues. This article underscores the fundamental importance of data locality as a key abstraction for optimizing application performance. Despite notable software solutions, the growing complexity of parallelism and memory hierarchy demands performance-portable data locality solutions across diverse computing platforms. Additionally, the article revisits data locality aspects, covering hardware considerations, application perspectives, software stack abstractions, and tool support. It concludes with insights into data locality challenges and opportunities, emphasizing the ongoing significance of collaborative research for progress in this critical issue.

Unat, Didem [Koc University, Istanbul (Turkey)] (O↗

GPS-supported smartphone app-based integrated travel diary and time-use data collection: challenges and lessons learned

Travel behaviour and time-use data are two vital data sources for travel demand modelling. Travel behaviour is traditionally collected through household travel surveys, enhanced by using GPS-supported smartphone apps for passive location data collection. However, recruiting individuals willing to install these apps with sustained motivation to continue participation has been a critical challenge. This paper shares insights from a travel and time-use data collection procedure in Chicago and Sydney using the Fourstep app. Social media platforms were utilised as a solution to recruit participants in Chicago, where an international market research company failed to accomplish the task. This paper also discusses the challenges we faced and suggests ways to overcome them, offering valuable guidance to researchers in recruiting participants for smartphone application-based data collection. It also offers an analysis of travel, time-use, and travel-based multitasking behaviours based on the data collected from the Chicago and Sydney samples.

GPS-supported smartphone apps↗

Navigating Integration: Key Challenges for Data Centers, Nuclear Stakeholders, and Utility Operators

he exponential growth of data centers—driven by artificial intelligence and cloud computing—is reshaping the U.S. energy landscape, presenting urgent challenges and transformative opportunities for data center developers, nuclear energy providers, and utility operators. As data centers are projected to consume up to 12% of U.S. electricity by 2028, stakeholders must address rapid deployment needs, grid congestion, and the demand for reliable, high-quality power. This presentation explores the multifaceted barriers to integrating data centers with nuclear and utility infrastructure, including land use constraints, public perception, regulatory complexity, and workforce alignment. It highlights the distinct priorities and operational cultures of each sector, and the friction that arises from misaligned planning horizons and risk tolerances. We examine collaborative strategies such as co-siting, hybrid power-purchase agreements, unified community engagement, and innovative financing models.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

LCLS Big Data Handling – How I Learned to Stop Worrying and Love the Data Deluge

Advanced data and computing systems are vital to Linac Coherent Light Source (LCLS) operations, data interpretation and overall scientific productivity. The transition to MHz-era operation marks a fundamental change in scale that requires new infrastructure and architectures to link LCLS to the required scale of computing needed for scientific interpretation. The LCLS-II Data System meets big data challenges by implementing configurable data reduction that can adapt to multiple science areas, real-time analysis frameworks to provide visualization and fast feedback, and the ability to transfer data to local and remote computational facilities for near real time analysis at the appropriate scale. Feature extracted information generated in the data analysis pipeline - at the edge, local compute, or remote High-Performance Computing (HPC) resources - can be used to steer experiments and inform user decisions during beam time. Artificial Intelligence and Machine Learning (AI/ML) techniques present new opportunities to rapidly analyse large datasets and direct experiments, but create new challenges in scaling, adaptability, complexity, and trustworthiness. We describe how the LCLS-II Data System architecture addresses its data-driven challenges in the areas of data acquisition, data processing, data management, and workflow orchestration to decrease the overall time-to-science and provide a vision for future developments.

artificial intelligence↗

Measurement of the Neutron Electromagnetic Form Factor Ratio at High Momentum Transfer

The inner structure of the nucleon (proton and neutron) remains a topic of great interest in nuclear and particle physics, after many decades of study. For example, understanding the quark-gluon dynamics inside the nucleon would shed light on how 99% of the nucleon mass is created. The neutron electromagnetic form factors, Gn E and Gn M , give important insights into the neutron structure. The Super BigBite Spectrometer (SBS) program at Jefferson Lab (JLab) seeks to extend the form factor measurements for both the proton and the neutron. The neutron electric form actor, Gn E , has been historically difficult to measure due to the short lifetime of the free neutron and the small value of Gn E . The GEn-II experiment is part of the SBS program and seeks to measure Gn E , significantly increasing the high momentum transfer coverage. A newly designed polarized 3He target increased the figure of merit by three times compared to previous measurements. The analysis of this data is especially challenging due to the unprecedented high-rate environment caused by the open nature of the spectrometer with a direct line of sight to the target. This required developing new Gas Electron Multiplier (GEM) particle trackers which can cover large areas demanded by this setup and handle particle rates up to 500 kHz/cm2. Rates this high over a large area is unprecedented in particle tracking systems and came with a number of challenges. Data taken in the SBS program was critical to understanding hardware and software solutions that improved the track reconstruction efficiency to be >97% with a position resolution of 70 ?m. In previous experiments the proton electromagnetic form factors, Gp E and Gp M were measured up to Q2 = 8.5 GeV2 and Q2 = 30 GeV2, respectively, while Gn E has only been measured up to Q2 = 3.4 GeV2. The GEn-II experiment has measured the neutron form factor ratio, Gn E/Gn M, at Q2 values of 2.90, 6.50, and 9.47 GeV2 by scattering a polarized electron beam with a polarized 3He target, used here as an effective polarized neutron target, and measuring the double spin asymmetry of the cross section. Previous Gn E measurements do not extend above Q2 = 3.4 GeV2, and therefore this analysis has extended the world data by almost three times. The background correction is especially difficult at the higher Q2 settings leading to large systematic errors. As very exploratory results from this early analysis of the data, we find for Q2 = 2.90 GeV2, Gn E = 0.0157 ±stat 0.0016 ±sys 0.0011, for Q2 = 6.50 GeV2, Gn E = 0.0067 ±stat 0.0019 ±sys 0.0005, and for Q2 = 9.46 GeV2, Gn E = 0.0046 ±stat 0.0023 ±sys 0.0005. These results are compared to predictions from the Dyson-Schwinger Equations (DSE) model and a Relativistic Constituent Quark Model (RCQM).

Jeffas, Sean↗

Exascale Computing and Data Handling: Challenges and Opportunities for Weather and Climate Prediction

The emergence of exascale computing and artificial intelligence offer tremendous potential to significantly advance Earth system prediction capabilities. However, enormous challenges must be overcome to adapt models and prediction systems to use these new technologies effectively. A 2022 WMO report on exascale computing recommends “urgency in dedicating efforts and attention to disruptions associated with evolving computing technologies that will be increasingly difficult to overcome, threatening continued advancements in weather and climate prediction capabilities.” Further, the explosive growth in data from observations, model and ensemble output, and postprocessing threatens to overwhelm the ability to deliver timely, accurate, and precise information needed for decision-making. Artificial intelligence (AI) offers untapped opportunities to alter how models are developed, observations are processed, and predictions are analyzed and extracted for decision-making. Given the extraordinarily high cost of computing, growing complexity of prediction systems, and increasingly unmanageable amount of data being produced and consumed, these challenges are rapidly becoming too large for any single institution or country to handle. This paper describes key technical and budgetary challenges, identifies gaps and ways to address them, and makes a number of recommendations.

Atmosphere↗

Modern chemical graph theory

Abstract Graph theory has a long history in chemistry. Yet as the breadth and variety of chemical data is rapidly changing, so too do graph encoding methods and analyses that yield qualitative and quantitative insights. Using illustrative cases within a basic mathematical framework, we showcase modern chemical graph theory's utility in Chemists' analysis and model development toolkit. The encoding of both experimental and simulation data is discussed at various levels of granularity of information. This is followed by a discussion of the two major classes of graph theoretical analyses: identifying connectivity patterns and partitioning methods. Measures, metrics, descriptors, and topological indices are then introduced with an emphasis upon enhancing interpretability and incorporation into physical models. Challenging data cases are described that include strategies for studying time dependence. Throughout, we incorporate recent advancements in computer science and applied mathematics that are propelling chemical graph theory into new domains of chemical study. This article is categorized under: Molecular and Statistical Mechanics > Molecular Dynamics and Monte‐Carlo Methods Structure and Mechanism > Computational Materials Science Structure and Mechanism > Molecular Structures

Leite, Leonardo S. G.↗