Search NASA⌕ Search

SEARCH · Search NASA

Results for “Benchmark data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25

Investigation of Benchmark $k$ eff Sensitivity and Uncertainty for 239 Pu fission in Specific Energy Ranges

Nuclear data at intermediate energies (from 1 to 100s of keV) are evaluated based on scarce differential data and theory unable to capture physics’ expected structure. There is also a lack of integral data. This is a known deficiency and is challenging to address. Calculated effective multiplication factor, k eff , values for intermediate energy experiments are ~25× further from experiment than for fast energies and are often well outside the experimental uncertainties. The goal of the PARADIGM (PARallel Approach of Differential and InteGral Measurements) project is to significantly re duce the uncertainties of intermediate energy nuclear data for 239 Pu. To this end, PARADIGM simultaneously optimizes experiments at both the Los Alamos Neutron Science Center (LANSCE) and National Criticality Experiments Research Center (NCERC). The combined set of data will inform new intermediate-energy nuclear data. By execution of differential and integral experiments, establishment of new theory, and undertaking nuclear data evaluation in parallel, the timeline to deliver improved nuclear data to users will be reduced significantly that is to three years. For the PARADIGM project, it was decided to optimize an integral experiment for two neutron energy ranges, within the full intermediate energy range. The low energy range goes from 1 to 30 keV, while the higher energy range goes from 30 to 600 keV. This work focuses on nuclear data sensitivities and uncertainties for 239 Pu fission for existing experiments in the International Criticality Safety Benchmark Evaluation Project (ICSBEP). When designing new experiments, it is important to understand what benchmarks currently exist. For a more traditional experiment design (in which a specific application model(s) exists), comparisons would be made between the application model(s) and existing benchmarks. For PARADIGM, there is no specific application model, but instead the specific nuclear data reaction and energy ranges of interest can be explored for existing benchmarks.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

IFAR Liner Benchmark Challenge #1 - DLR Impedance Eduction of Uniform and Axially Segmented Liners and Comparison with NASA Results

This paper presents the contribution from the German Aerospace Center (DLR) to the first liner benchmark challenge under the framework of the International Forum for Aviation Research (IFAR).Therefore, two sets of acoustically damping wall treatment, called ’liner samples’, have been produced by additive manufacturing based on the design data provided by NASA coordinating this benchmark. These liner samples have been integrated and acoustically characterized in the liner flow test facility DUCT-R at DLR Berlin as well as in the liner flow test facility GFIT at NASA Langley. Besides the dissipation coefficients and the axial pressure profiles, the liner wall impedance was educed by first determining the axial wave numbers and then applying a straightforward method based on the one-dimensional Convected Helmholtz Equation. Finally, the comparison of the liner impedance values to the NASA results show a fairly good agreement.

liner characterization↗

Instrumentation for the In-Core Real-Time Mechanical Testing of Structural Materials (INCREASE) Project

Idaho National Laboratory (INL), in collaboration with the Electric Power Research Institute (EPRI), the Nuclear Regulatory Commission (NRC), the French Atomic and Alternative Energies Commission (CEA), the Joint Research Center (JRC), the Nuclear Research and Consultancy Group (NRG), and the Research Center Rez (CVR), started a Joint Experimental Program (JEEP) project that operates within the Nuclear Energy Agency’s Framework for Irradiation Experiments (FIDES II) program in order to develop capabilities for the in-core real-time mechanical testing of structural materials. This effort will focus on designing a shared capsule capable of housing a variety of in-core mechanical testing instrumentation allowing enhanced experiments for the material science community. The outcome of the project would be high-priority, stress relaxation data for stainless-steel-based materials provided by EPRI and CEA. Stress relaxation is a major phenomenon that contributes to material degradation in nuclear reactor components. Currently, nuclear material stress relaxation is assessed both before and after irradiation, using complex and costly post-irradiation examination (PIE) activities. In-situ data would support the development of precision modeling and simulation of this degradation phenomena and would provide validation and benchmarking for existing models using the PIE data. As part of the U.S. Department of Energy Advanced Sensor and Instrumentation (ASI) program, INL has fabricated and tested out-of-core mechanical test instrumentation. This instrumentation was designed for easy adaptation to the irradiation capsule proposed under this JEEP, and can be deployed to measure real-time stress relaxation under pressurized-water reactor (PWR) conditions. This initial effort will partially serve to replace the testing capabilities lost because of shutting down the Halden Boiling Water Reactor (HBWR). The project would provide these capabilities to the international community via a shared capsule design that is easily adaptable to additional material test reactors. The design features will incorporate expansion to PWR and non-light-water reactor (LWR) environments that will be developed in future work. The capsule and instrumentation will be demonstrated in the Massachusetts Institute of Technology Reactor (MITR) for phase I and the Petten High Flux Reactor (HFR) for phase II irradiations to deliver real-time stress relaxation data on high priority stainless steel structural materials.

36 MATERIALS SCIENCE↗

Mic-hackathon 2024: hackathon on machine learning for electron and scanning probe microscopy

Microscopy is one of the primary sources of information on materials structure and functionality at the nanometer and atomic scales. The data generated through microscopy is often contained in well-structured datasets, enriched with extensive metadata and sample histories, although not always with the same level of detail or storage format. The broad incorporation of data management plans by major funding agencies ensures the preservation and accessibility of this data. However, deriving insights from these rich datasets remains challenging due to the lack of established code ecosystems, standardized benchmarks, and integration strategies. Correspondingly, the efficiency of data usage is very low, and time expenditures at the analysis stage are enormous. In addition to post-acquisition data analysis, the emergence of application programming interfaces by major microscope manufacturers now creates opportunities for real-time ML-based data analytics to enable automated decision making, and particularly ML-agent controlled real-time microscope operation. Despite these opportunities, there is a significant gap in integrating the ML community with the broader microscopy community, limiting the value that these methods bring to physics and materials discovery and materials optimization. Hackathons address these challenges by fostering collaboration between ML experts and microscopy professionals, encouraging the development of innovative solutions that leverage ML for microscopy and preparing the workforce of the future both for microscopy-intensive domains areas, instrument manufacturers, and ML scientists interested in real world applications for fundamental research, materials optimization, and manufacturing. The hackathon generated benchmark datasets and digital twins of microscopes that further contribute to the development of the field and establish data analysis ecosystems. All the codes can be found at GitHub(https://github.com/KalininGroup/Mic-hackathon-2024-codes-publication/tree/1.0.0.1) and Zenodo (https://zenodo.org/records/15579940).

97 MATHEMATICS AND COMPUTING↗

Statistical Estimation of Orbital Debris Populations with a Spectrum of Object Size

Orbital debris is a real concern for the safe operations of satellites. In general, the hazard of debris impact is a function of the size and spatial distributions of the debris populations. To describe and characterize the debris environment as reliably as possible, the current NASA Orbital Debris Engineering Model (ORDEM2000) is being upgraded to a new version based on new and better quality data. The data-driven ORDEM model covers a wide range of object sizes from 10 microns to greater than 1 meter. This paper reviews the statistical process for the estimation of the debris populations in the new ORDEM upgrade, and discusses the representation of large-size (greater than or equal to 1 m and greater than or equal to 10 cm) populations by SSN catalog objects and the validation of the statistical approach. Also, it presents results for the populations with sizes of greater than or equal to 3.3 cm, greater than or equal to 1 cm, greater than or equal to 100 micrometers, and greater than or equal to 10 micrometers. The orbital debris populations used in the new version of ORDEM are inferred from data based upon appropriate reference (or benchmark) populations instead of the binning of the multi-dimensional orbital-element space. This paper describes all of the major steps used in the population-inference procedure for each size-range. Detailed discussions on data analysis, parameter definition, the correlation between parameters and data, and uncertainty assessment are included.

Xu, Y. -l↗

Determination of Aerothermal Environment and Ablator Material Response Using Inverse Methods

The Mars Science Laboratory (MSL) was protected during entry into the Martian atmosphere by a thermal protection system that used NASA’s Phenolic Impregnated Carbon Ablator (PICA). The heat shield of the probe was instrumented with the Mars Entry Descent and Landing Instrument (MEDLI) suite of sensors. MEDLI Integrated Sensor Plugs (MISP) included thermocouples that measured in-depth temperatures at various locations on the heatshield. The flight data has been used as a benchmark for validating ablation codes within NASA. This work seeks to refine the estimate of the material properties for the MSL heat shield and the aerothermal environment during Mars entry using estimation methods in DAKOTA on the temperature data obtained from MEDLI.

Thornton, John M.↗

Investigation of Error Patterns in Geographical Databases

The objective of the research conducted in this project is to develop a methodology to investigate the accuracy of Airport Safety Modeling Data (ASMD) using statistical, visualization, and Artificial Neural Network (ANN) techniques. Such a methodology can contribute to answering the following research questions: Over a representative sampling of ASMD databases, can statistical error analysis techniques be accurately learned and replicated by ANN modeling techniques? This representative ASMD sample should include numerous airports and a variety of terrain characterizations. Is it possible to identify and automate the recognition of patterns of error related to geographical features? Do such patterns of error relate to specific geographical features, such as elevation or terrain slope? Is it possible to combine the errors in small regions into an error prediction for a larger region? What are the data density reduction implications of this work? ASMD may be used as the source of terrain data for a synthetic visual system to be used in the cockpit of aircraft when visual reference to ground features is not possible during conditions of marginal weather or reduced visibility. In this research, United States Geologic Survey (USGS) digital elevation model (DEM) data has been selected as the benchmark. Artificial Neural Networks (ANNS) have been used and tested as alternate methods in place of the statistical methods in similar problems. They often perform better in pattern recognition, prediction and classification and categorization problems. Many studies show that when the data is complex and noisy, the accuracy of ANN models is generally higher than those of comparable traditional methods.

Dryer, David↗

LandScan Global 2023: Silver Edition

For a quarter of a century, the LandScan Global (LSG) project has annually released a global, high-resolution gridded population dataset representing the ambient or unwarned population at a 30 arcsecond resolution. LSG supports a range of applications such as emergency management, disaster response, and human health and security for understanding populations at risk. The 2023 release of LSG, the LandScan Silver Edition, represents a major methodological leap forward while also leveraging previous knowledge—the previous year was the baseline for the current annual update carrying forward valuable knowledge of the built environment for the past quarter century—to train the machine learning models. Compared with annual releases over the past 24years, multiple advancements were made to different aspects of the methodology to achieve reproducibility, transparency, and consistent global propagation of solutions to modeling or population distribution issues identified during the review process. These novel changes include incorporation of the latest available geospatial inputs across the globe, machine learning models instead of manual modifications, population feature importance analysis, open-source solutions vs. proprietary software, generation of multiple global versions, analytic validations, and human-in-the-loop revisions to produce the final version. Additionally, algorithms—such as anomaly detection—were introduced to quickly identify areas of focus to develop a new and robust systematic review. Significant changes in modeled population distributions were observed between the 2022 and 2023 releases, largely attributable to improvements in data and methods and discussed thoroughly within this report. In summation, the LandScan Silver Edition leverages the best of the past quarter century of LSG legacy knowledge and continues a tradition of applying cutting-edge enhancements to serve as a new benchmark for accurate, actionable gridded population data

Lebakula, Viswadeep↗

Impacts of benchmarking choices on inferred model skill of the Arctic–Boreal terrestrial carbon cycle

Abstract Land surface models require continuous validation against observations to improve and reduce simulation uncertainty. However, inferred model performance can be heavily influenced by subjective choices made in the selection and application of observational data products. A key area often misrepresented by models is the Arctic–Boreal region, which is a potential tipping point region in Earth’s climate system due to large permafrost carbon stocks that are vulnerable to release with climate warming. We use the International Land Model Benchmarking (ILAMB) framework to evaluate how the model skill of TRENDY-v9 models varies based on the choice of observational-based benchmark and how benchmarks are applied in model evaluation. This analysis uses global datasets integrated into ILAMB and new, regionally-specific observational products from the Arctic–Boreal Vulnerability Experiment. Our results cover the overall time period of 1979–2019 and show that model scores can vary substantially depending on the data product applied, with higher model scores indicating better model performance against observations. The lowest model scores occur when benchmarked against regional, compared to global, datasets. We also evaluate observed and modeled functional relationships between ecosystem respiration and air temperature and between gross primary production and precipitation. Here, we find that the magnitude and shape of the responses are strongly impacted by the choice of observational dataset and the approach used to construct the functional relationship benchmark. These results suggest that model evaluation studies could conclude a false sense of model skill if only using a single benchmark data product or if not applying regional data products when performing a regional model analysis. Collectively, our findings highlight the influence of benchmarking choices on model evaluation and point to the need for benchmarking guidelines when assessing model skill.

Poe, Jeralyn (ORCID:0000000318495278)↗

Evaluating and Quantifying the Climate-Driven Interannual Variability in Global Inventory Modeling and Mapping Studies (GIMMS) Normalized Difference Vegetation Index (NDVI3g) at Global Scales

Satellite observations of surface reflected solar radiation contain informationabout variability in the absorption of solar radiation by vegetation. Understanding thecauses of variability is important for models that use these data to drive land surface fluxesor for benchmarking prognostic vegetation models. Here we evaluated the interannualvariability in the new 30.5-year long global satellite-derived surface reflectance index data,Global Inventory Modeling and Mapping Studies normalized difference vegetation index(GIMMS NDVI3g). Pearsons correlation and multiple linear stepwise regression analyseswere applied to quantify the NDVI interannual variability driven by climate anomalies, andto evaluate the effects of potential interference (snow, aerosols and clouds) on the NDVIsignal. We found ecologically plausible strong controls on NDVI variability by antecedent precipitation and current monthly temperature with distinct spatial patterns. Precipitation correlations were strongest for temperate to tropical water limited herbaceous systemswhere in some regions and seasons 40 of the NDVI variance could be explained byprecipitation anomalies. Temperature correlations were strongest in northern mid- to-high-latitudes in the spring and early summer where up to 70 of the NDVI variance was explained by temperature anomalies. We find that, in western and central North America,winter-spring precipitation determines early summer growth while more recent precipitation controls NDVI variability in late summer. In contrast, current or prior wetseason precipitation anomalies were correlated with all months of NDVI in sub-tropical herbaceous vegetation. Snow, aerosols and clouds as well as unexplained phenomena still account for part of the NDVI variance despite corrections. Nevertheless, this study demonstrates that GIMMS NDVI3g represents real responses of vegetation to climate variability that are useful for global models.

interference↗

Comparison of the Effects of using Tygon Tubing in Rocket Propulsion Ground Test Pressure Transducer Measurements

This paper documents acoustics environments data collected during liquid oxygen- ethanol hot-fire rocket testing at NASA Marshall Space Flight Center in November- December 2003. The test program was conducted during development testing of the RS-88 development engine thrust chamber assembly in support of the Orbital Space Plane Crew Escape System Propulsion Program Pad Abort Demonstrator. In addition to induced environments analysis support, coincident data collected using other sensors and methods has allowed benchmarking of specific acoustics test measurement methodologies during propulsion tests. Qualitative effects on data characteristics caused by using tygon sense lines of various lengths in pressure transducer measurements is discussed here.

Farr, Rebecca A.↗

Infrared Sensing Aeroheating Flight Experiment: STS-96 Flight Results

Major elements of an experiment called the Infrared Sensing Aeroheating Flight Experiment are discussed. The primary experiment goal is to provide reentry global temperature images from infrared measurements to define the characteristics of hypersonic boundary-layer transition during flight. Specifically, the experiment is to identify, monitor, and quantity hypersonic boundary layer windward surface transition of the X-33 vehicle during flight. In addition, the flight data will serve as a calibration and validation of current boundary layer transition prediction techniques, provide benchmark laminar, transitional, and fully turbulent global aeroheating data in order to validate existing wind tunnel and computational results, and to advance aeroheating technology. Shuttle Orbiter data from STS-96 used to validate the data acquisition and data reduction to global temperatures, in order to mitigate the experiment risks prior to the maiden flight of the X-33, is discussed. STS-96 reentry midwave (3-5 micron) infrared data were collected at the Ballistic Missile Defense Organization/Innovative Sciences and Technology Experimentation Facility site at NASA-Kennedy Space Center and subsequently mapped into global temperature contours using ground calibrations only. A series of image mapping techniques have been developed in order to compare each frame of infrared data with thermocouple data collected during the flight. Comparisons of the ground calibrated global temperature images with the corresponding thermocouple data are discussed. The differences are shown to be generally less than about 5%, which is comparable to the expected accuracy of both types of aeroheating measurements.

Blanchard, Robert C.↗

Infrared Sensing Aeroheating Flight Experiement: STS-96 Flight Results

Major elements of an experiment called the Infrared Sensing Aeroheating Flight Experiment are discussed. The primary experiment goal is to provide reentry global temperature images from infrared measurements to define the characteristics of hypersonic boundary-layer transition during flight. Specifically, the experiment is to identify, monitor, and quantify hypersonic boundary layer windward surface transition of the X-33 vehicle during flight. In addition, the flight data will serve as a calibration and validation of current boundary layer transition prediction techniques, provide benchmark laminar, transitional, and fully turbulent global aeroheating data in order to validate existing wind tunnel and computational results, and to advance aeroheating technology. Shuttle Orbiter data from STS-96 used to validate the data acquisition and data reduction to global temperatures, in order to mitigate the experiment risks prior to the maiden flight of the X-33, is discussed. STS-96 reentry mid-wave (3-5 Pm) infrared data were collected at the Ballistic Missile Defense Organization/Innovative Sciences and Technology Experimentation Facility site at NASA-Kennedy Space Center and subsequently mapped into global temperature contours using ground calibrations only. A series of image mapping techniques have been developed in order to compare each frame of infrared data with thermocouple data collected during the flight. Comparisons of the ground calibrated global temperature images with the corresponding thermocouple data are discussed. The differences are shown to be generally less than about 5%, which is comparable to the expected accuracy of both types of aeroheating measurements.

Blanchard, Robert C.↗

AI Model Benchmarking for Nonproliferation Applications: Steel Thread Benchmarking Task Force Technical Report (Rev. 2)

Steel Thread is a NA-22 venture that seeks to build trustworthy, reliable AI models that can be used in a wide variety of nonproliferation tasks. A key aspect of building these models is developing appropriate benchmarks and evaluation methods, which will enable the venture to identify and adapt models to provide the most value in the nonproliferation domain. Benchmarks must be relevant to key tasks in this domain, such as question answering, information retrieval, document summarization and classification, consensus analysis, and image and data analysis. This report 1) provides an overview of benchmark design, evaluation, and challenges; 2) reviews a variety of open benchmarks, with a focus on language models and tasks; and 3) identifies benchmarks that are most relevant to Steel Thread. This report is intended to serve as a basis for further efforts to classify and evaluate benchmarks and their correlation with success on nonproliferation-specific tasks. The Steel Thread venture has defined benchmarks to be a particular combination of a dataset (or datasets) and a metric (or metrics) conceptualized as representing one or more specific tasks or sets of abilities for a specific modality. It is adopted by a research community as a shared framework for comparing methods.1 It includes 1) Data: Labeled (a designated subset not used for training, which could be all the data), 2) Metric: A way to quantify performance, 3) Task/Ability: What the benchmark is testing, 4) Protocol: A structured and repeatable evaluation process, 5) Baseline/Reference Model: For comparison; could be statistical, rule-based, SME-derived, or another model, and 6) Maintenance Plan: to update with new information over time; important for long-term utility. For further clarity, the definition includes what a benchmark, in this context, is not. It is not a corpus of training data, specific to a model (it is intended to apply to a range of models), a universal evaluation of performance, a guarantee that the ‘top’ model on the leaderboard will be the best fit for every specific use case, an all-encompassing proof of a model’s universal quality, nor is it a one-size-fits-all measure of success. It does not cover every real-world constraint (like operational, ethical, or cost considerations), a systems integration test, or a unit test. This definition was inspired by and resulted from discussions within the Steel Thread Benchmarking Task Force. This group was formed to define what we would mean as a benchmark within Steel Thread but persisted as the need to develop a thorough understanding of the large and expanding existing benchmarking space. This technical report is a result of the group’s divide and conquer approach to exploring this space. The release of benchmarks might not be progressing as quickly as model development, but it is moving very fast, as many benchmarks quickly become saturated, when state-of-the-art models score so close to the benchmark’s ceiling that their results are virtually indistinguishable. At that point, the test no longer differentiates between new systems, so researchers usually stop reporting scores as the benchmark no longer informs about improvements from the next generation of models. In the OpenAI announcement of GPT-5, they reported results on six flagship public benchmarks (AIME 2025, SWE-bench Verified, Aider Polyglot, MMMU, HealthBench Hard, GPQA) but the full system-card covers roughly thirty-five separate evaluations, comprising hundreds of test task items in total. There have been some efforts to summarize benchmarks in specific fields, like for text-to-image generation, but these surveys have had a narrow methodology scope. Therefore, a comprehensive survey of all benchmarks or even all benchmarks that could be relevant to Steel Thread is outside of the scope of this report. We chose some specific benchmarks to investigate in detail.

97 MATHEMATICS AND COMPUTING↗

Advancing the Frontiers of Deep Learning for Low-Dose 3D Cone-Beam CT Reconstruction

X-ray computed tomography (CT) is an important noninvasive medical imaging modality for studying the structural details of internal organs. Image reconstruction in CT is an inverse problem of recovering an object's internal structure from the absorption profile of X-ray beams (sinogram) measured using a detector. The classical variational approach for CT reconstruction minimizes an energy functional using an appropriate iterative algorithm. Motivated by the success of deep learning (DL), researchers have begun to leverage training data and enhanced computing capabilities in recent years to produce high-fidelity reconstructed images. Nonetheless, much of the academic research in DL algorithms for CT has focused primarily on the two-dimensional setting (with simplified forward operators and noise model) for proofs-of-concept, and a comprehensive benchmarking of various classical and data-driven CT reconstruction approaches has not beenundertaken. The key objective of our CT reconstruction grand challenge was to promote methodological advancements for both classical and DL-based approaches for clinical CT with a reasonably accurately simulated 3D CT forward operator and noise model. We have utilized the publicly available LIDC-IDRI dataset and simulated sinograms and FDK images corresponding to two dose levels (clinical- and low-dose, constituting two tracks of the challenge) starting from the normal-dose images as the ground truth. In this paper, we summarize the motivation, context, and results of our challenge, and highlight the future research directions in DL for clinical CT.

X-ray tomography↗

Model Development and Analysis of a High-Fidelity Neutron Transport Sensor: The Quadrupole Detector Concept for Measurement of the Neutron Flux Gradient

Accurate reconstruction of the neutron flux distribution within a reactor core is essential for safe and efficient reactor operation. Traditional power shape synthesis in Light Water Reactors relies on hundreds of in-core detectors. However, this approach becomes impractical for Advanced Reactors and Microreactors due to limited space and harsh environments. To address this challenge, we propose a data-driven methodology that combines high-fidelity modeling with real-time ex-core sensor measurements, enabling the reconstruction of core power distribution while minimizing the reliance on intrusive in-core instrumentation. This project began in FY24 and achieved two initial milestones: (1) the definition of a three-year development plan for a Digital Twin framework and (2) the development of high-fidelity neutronics models of the Purdue University Reactor One (PUR-1) using both MCNP6 and OpenMC. The PUR-1 reactor, a zero-power facility, was selected due to its suitability for neutronics-focused modeling and the availability of experimental data for validation. Both models were benchmarked using neutron flux measurements obtained from irradiated gold foils, which were strategically placed within the core during a dedicated campaign in July 2024. This report marks the continuation and completion of those foundational tasks. The OpenMC model has been refined (improved geometric accuracy, expanded cross-section libraries, and refined sampling) and validated using additional experimental data. An updated sensor design—based on quadrupole configuration—was designed to measure both ex-core flux and its spatial gradient. These measurements will serve as inputs to a neural network-based reconstruction algorithm. Finally, the methodology was demonstrated on a two-dimensional test case representative of the heterogeneous material composition of the PUR-1 reactor core. A neural network implementation of the Kirchhoff-Helmholtz integral equation was employed to solve the boundary value problem using peripheral sensor measurements. The preliminary results confirm the strong potential of the proposed approach for accurate and minimally invasive neutron flux reconstruction.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

An international benchmark for wind plant wakes from the American WAKE ExperimeNt (AWAKEN)

This article introduces the first benchmark study within the International Energy Agency Wind Task 57 framework, focusing on wind plant wakes. Leveraging data from the American WAKE ExperimeNt (AWAKEN), the benchmark aims to assess the accuracy of simulation tools in modeling wind plant wakes and their impact on the downstream flow under diverse inflow conditions. The AWAKEN field campaign, conducted in Oklahoma from 2022 to 2024, provides unprecedented observations of wind plant-atmosphere interactions, thus offering a large dataset to validate numerical models of different complexity. The benchmark will include three phases—code calibration, blind comparison, and iteration—allowing participants to refine their numerical models based on the feedback from the benchmark team. This article describes the benchmark case study selected from observations providing details on atmospheric conditions, wake evidence, and wind turbine operation. The benchmark’s structure and timeline, along with the expected publication of results, are discussed as well. This collaborative effort aims to enhance the accuracy of wind plant wake simulations, thus contributing to the improvement of wind energy production estimates.

17 WIND ENERGY↗

NASA POWER: Providing Analysis-Ready, Cloud-Optimized Data for AI /ML Training and Applications in Earth Science

As global demand for sustainable development grows, the integration of Earth Observation (EO) data into decision making frameworks has become a primary objective for the scientific community. The NASA Prediction of Worldwide Energy Resources (POWER) project serves as a bridge between NASA EO data and the specialized needs of the renewable energy, sustainable infrastructure and agroclimatology communities. In this poster presentation we will present an overview of POWER data products and services along with its use in diverse research to decision-making workflows. By providing over 40 years of high-resolution historical, hourly and daily solar and meteorological data, POWER transforms satellite observations and global model reanalysis into actionable, Analysis-Ready Dataset (ARD). Currently, the project delivers over 250 industry-friendly parameters to the users from different NASA datasets like CERES SYN1Deg, MERRA-2, and IMERG alongside downscaled CMIP6 climate model data, fulfilling over 16 million requests from 50,000 unique users monthly. To ensure data quality and traceability, these parameters are rigorously validated against the ground-based observations from the Baseline Surface Radiation Network (BSRN) and the Global Surface Summary of the Day (GSOD) – these results will be discussed in the presentation. A newly introduced web-based PaRameter Uncertainty ViEwer (PRUVE) tool will be presented that provides an online validation platform to the users that benchmarks satellite-based and assimilation data products against these surface measurements. To reduce technical barriers to data adoption, POWER data is accessible through RESTful APIs, ESRI ArcGIS Image Services, a web-based Data Access Viewer tool, allowing users to visualize, validate and apply the dataset. For efficient data delivery POWER data is cloud-optimized into Zarr datastore accessible through NASA managed Amazon S3 ensures high-performance allowing users to integrate EO directly into operational pipelines. These customized services will be presented. Use cases from application will be presented from the energy sector - such as for design of generation systems, performance monitoring of solar power plants, in infrastructure sector- optimizing building energy efficiency and thermal comfort, in agriculture – such as driving crop simulation and yield forecasting models to enable climate resilient farming. Furthermore, the shift toward machine learning (ML) in EO research that has positioned POWER as a key provider for training datasets which will be discussed. Use-cases will be presented to showcase how NASA data is enabling the development of predictive tools for climate variability and resource management. The poster will present POWER’s future plans including technology development to enhance data traceability and reproducibility and improving I/O performance to support the rapid integration of new EO products, ensuring that POWER remains a robust scalable backend for the evolving landscape of AI-driven Earth Science. Additionally, POWER is developing an AI Agent and an MCP-Server to enable industry AI-Agentic workflows.

Neha Khadka↗