Search NASASearch

SEARCH · Search NASA

Results for “initial data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Identifying preferential flow from soil moisture time series: Review of methodologies

Abstract Identifying and quantifying preferential flow (PF) through soil—the rapid movement of water through spatially distinct pathways in the subsurface—is vital to understanding how the hydrologic cycle responds to climate, land cover, and anthropogenic changes. In recent decades, methods have been developed that use measured soil moisture time series to identify PF. Because they allow for continuous monitoring and are relatively easy to implement, these methods have become an important tool for recognizing when, where, and under what conditions PF occurs. The methods seek to identify a pattern or quantification that indicates the occurrence of PF. Most commonly, the chosen signature is either (1) a nonsequential response to infiltrated water, in which soil moisture responses do not occur in order of shallowest to deepest, or (2) a velocity criterion, in which newly infiltrated water is detected at depth earlier than is possible by nonpreferential flow processes. Alternative signatures have also been developed that have certain advantages but are less commonly utilized. Choosing among these possible signatures requires attention to their pertinent characteristics, including susceptibility to errors, possible bias toward false negatives or false positives, reliance on subjective judgments, and possible requirements for additional types of data. We review 77 studies that have applied such methods to highlight important information for readers who want to identify PF from soil moisture data and to inform those who aim to develop new methods or improve existing ones. Core Ideas Soil moisture data can be used to identify the occurrence of preferential flow (PF) and its initiating conditions. Various data‐analysis methods to identify PF differ in susceptibility to error, bias, and subjectivity. These methods can utilize vast amounts of data from soil moisture monitoring networks to develop understanding of when, where, and under what conditions PF occurs. Newly developed methods may lead to better accuracy and reliability, and reduce the need for subjective judgments. Plain Language Summary Preferential flow through soil occurs when a large amount of water is suddenly available, as during an intense storm. This type of flow moves rapidly through the soil in distinct narrow pathways rather than moving evenly throughout the body of soil, with major consequences for groundwater resources, ecosystems, spreading of contaminants, and other vital concerns. Methods of detecting preferential flow have been developed that utilize measurements of soil water content made by sensors installed at various depths. This measurement technology has been widely implemented, many locations now having datasets years in length, and various methods have been developed for using these to identify preferential flow. The various methods are based on different features in the soil moisture records and vary in their advantages and shortcomings. In this review, we explain and evaluate these methods, highlighting important information for their implementation to identify preferential flow from soil moisture data and for efforts to develop new methods or improve existing ones.

Nimmo, John R

Fiber-Optic Sensing for Earthquake Hazards Research, Monitoring, and Early Warning

The use of fiber‐optic sensing systems in seismology has exploded in the past decade. Despite an ever‐growing library of ground‐breaking studies, questions remain about the potential of fiber‐optic sensing technologies as tools for advancing if not revolutionizing earthquake‐hazards‐related research, monitoring, and early warning systems. A working group convened to explore these topics; we comprehensively examined the application of fiber optics in various aspects of earthquake hazards, encompassing earthquake source processes, crustal imaging, data archiving, and technological challenges. There is great potential for fiber‐optic systems to advance earthquake monitoring and understanding, but to fully unlock their capabilities requires continued progress in key areas of research and development, including instrument testing and validation, increased dynamic range for applications focused on larger earthquakes, and continued improvement in subsurface and source imaging methods. A key current stumbling block results from the lack of clear data archiving requirements, and we propose an initial strategy that balances data volume requirements with preserving key data for a broad range of future studies. In addition, we demonstrate the potential for fiber‐optic sensing to impact monitoring efforts by documenting the data completeness in a number of long‐term experiments. Finally, we outline the features of a instrument testing facility that would enable progress toward reliable and standardized distributed acoustic sensing data. Overcoming these current obstacles would facilitate progress in fiber‐optic sensing and unlock its potential application to a broad range of earthquake hazard problems.

58 GEOSCIENCES

The Dark Energy Bedrock All-sky Supernova Program: Motivation, Design, Implementation, and Preliminary Data Release

Precise measurements of Type Ia supernovae (SNe Ia) at low redshifts (z) serve as one of the most viable keys to unlocking our understanding of cosmic expansion, isotropy, and growth of structure. The Dark Energy Bedrock All-Sky Supernovae (DEBASS) program will deliver a uniformly calibrated low-z dataset of more than 400 spectroscopically confirmed SNe Ia in the Southern Hemisphere. DEBASS utilizes the Dark Energy Camera to image supernovae in conjunction with the Wide-Field Spectrograph to gather comprehensive host-galaxy information. By using the same photometric instrument as both the Dark Energy Survey (DES) and the DECam Local Volume Exploration Survey, DEBASS not only benefits from a robust photometric pipeline and well-calibrated images across the Southern sky, but can replace the historic and external low-z samples that were used in the final DES supernova analysis. In this paper, along with a companion paper, we present an early data release of 77 DEBASS SNe within the DES footprint. We introduce the DEBASS program, discuss its scientific goals and the advantages it offers for supernova cosmology, and present our initial results demonstrating data quality. With this early data release, we find a robust median absolute standard deviation of Hubble diagram residuals of ∼0.10 mag and an initial measurement of the host-galaxy mass step of 0.06 ± 0.04 mag, both before performing bias corrections. This low scatter shows the promise of a low-z SN Ia program with a well-calibrated telescope and high signal-to-noise ratio across multiple bands.

Sherman, Nora F. [Boston U.] (ORCID:00000001539901

Outcomes of PAX sapiens-Supported Global Wildlife Data Sharing Conferences for Enhanced One Health Security (GWDSC)

Across two consecutive Global Wildlife Data Sharing Conferences supported by PAX sapiens—Year 1 (May 2024) at Pacific Northwest National Laboratory and Year 2 (2025) in Ciudad Real, Spain—the initiative converted wildlife data sharing from aspiration into operational reality, producing measurable impacts in platform development, data mobilization, standards harmonization, and international partnership formation. The conferences addressed a critical gap in global health security: while 75% of emerging infectious diseases affect both humans and animals and over 60% originate in wildlife, wildlife health surveillance has historically lagged behind human and agricultural sectors due to fragmented databases, inconsistent terminology, uneven capacity, and limited cross-border coordination. By convening practitioners, government agencies, international organizations, academic institutions, and NGOs, the GWDSC catalyzed trust-based relationships and practical workflows that enable earlier detection, better risk assessment, and more effective prevention of threats at the wildlife–domestic animal–human–environment interface.

54 ENVIRONMENTAL SCIENCES

Proposal from the NA61/SHINE Collaboration for update of European Strategy for Particle Physics

Building on the current program's success and driven by new physics challenges, the NA61/SHINE Collaboration proposes to continue measuring hadron production properties in reactions induced by hadron and ion beams after CERN Long Shutdown 3. These measurements are of significant interest to the heavy-ion, cosmic-ray, and neutrino physics communities and will focus on: - Investigating hadron production in the light-ion systems to explore the diagram of high-energy nuclear collisions, and to obtain new insight into the unexpected violation of isospin (flavor) symmetry recently observed by the experiment; - Measuring charm-anticharm correlations to gain unique insights into the production locality of charm and anticharm quark pairs; - Examining strangeness and multi-strangeness production to improve our understanding of the early Universe's evolution and neutron star formation; - Measuring cross sections relevant for cosmic-ray measurements, significantly boosting searches for new physics in our Galaxy; - Conducting hadron production measurements with proton, pion, and kaon beams for neutrino physics, enhancing the precision of hadron production data needed for initial neutrino flux predictions in neutrino oscillation experiments; - Measuring hadron production processes relevant for understanding the flux of atmospheric neutrinos, as well as neutrinos and muons from spallation sources. To achieve these objectives, a detector upgrade and a beam upgrade are required, with data-taking planned for the period 2029-2032 and beyond.

Adhikary, H. [Jan Kochanowski U.] (ORCID:000000025

Quantifying the Thermodynamic Impacts on the Atmospheric Boundary Layer due to the Sea Breeze in the Coastal Houston Region

The atmospheric boundary layer (ABL) is unique in coastal regions because of kinematic and thermodynamic influences from continental and marine environments. Sea-breeze (SB) circulations act to equilibrate the land–sea temperature gradient through advecting marine air onshore. The strength of the SB varies in terms of stability, temperature, and moisture advection and influences air quality and weather forecasts. The Tracking Aerosol Convection Interactions Experiment (TRACER) collected a wealth of data on coastal boundary layer evolution, including observations from uncrewed aerial systems (UASs). Vertical profiles of temperature, humidity, and winds were collected by the OU CopterSonde UAS from June to September in the coastal region of Houston. These profiles offer 5-m vertical resolution, on average, every 30 min through diurnal transitions, SB events, and nearby deep convection. During the campaign, CopterSonde observations were gathered through 17 SB events, six of which led to convection initiation. The UAS data can resolve the thermodynamic evolution and interactions between the SB and the preexisting convective boundary layer. Results show large variability across observed SBs and their impacts on temperature and moisture. The intensity of thermodynamic changes depends on the time of sea-breeze passage and influence from the Galveston Bay Breeze, a secondary marine circulation commonly observed in this region. In quantifying the spectrum of SB impacts, equivalent potential temperature θ e is used to contextualize its role in convection initiation and evolution. In conclusion, while all SBs tend to increase θ e from moisture advection, the rate and timing of the θ e rise can distinguish convective from nonconvective cases.

54 ENVIRONMENTAL SCIENCES

VA EDH Advanced Software Pipeline Framework Report: Enhancing Automation and Scalability

The VA Environmental Determinants of Health (EDH) Advanced Software Pipeline Framework is designed to enhance the efficiency, scalability, and security of geospatial data processing workflows. This framework integrates modern data orchestration and containerization technologies, including Prefect for workflow automation, Docker for containerization, and PostgreSQL/PostGIS for geospatial data storage and analysis. It ensures standardized, reproducible, and automated data processing, supporting VA objectives related to substance use risk assessment and recovery research. The pipeline addresses key scalability and performance challenges through horizontal and vertical scaling, high-performance computing (HPC) integration, parallel processing, task caching, and dynamic resource allocation. These optimizations improve throughput and reduce latency, allowing the system to efficiently manage large and complex datasets. Additionally, security and compliance measures—such as data encryption (SSL), Role-Based Access Control (RBAC), and adherence to GDPR and HIPAA standards—safeguard sensitive information throughout data transmission and storage. A key implementation of this framework includes the automation of shelter list geolocation workflows, ensuring that up-to-date data is readily available for VA decision-making. Lessons learned from this project include the transition from in-memory processing to incremental storage writes, improving resource management and reliability. Future enhancements aim to expand automation, integrate AI-driven anomaly detection, and incorporate high-performance computing resources. This framework provides a scalable, secure, and adaptable solution for managing geospatial datasets, reinforcing the VA’s ability to support clinical and strategic initiatives through data-driven decision-making.

97 MATHEMATICS AND COMPUTING

GLBRC Soil Yearlong Incubation 13C-SIP-Lipidomics

Data package for Lipids represent a dynamic, yet stable pool of microbially-derived soil carbon This data is published under a CC0 license. The authors encourage data reuse and request attribution by referencing the below citations for the data packages and associated manuscript. Please cite as: Rempfert KR, Bell SL, Kasanke CP, Kyle JE, Hofmockel KS. 2025. GLBRC Soil Yearlong Incubation 13C-SIP-Lipidomics. [Data Set] PNNL DataHub. doi: Rempfert KR, Bell SL, Kasanke CP, Kyle JE, Hofmockel KS. 2025. MSV000097435: GLBRC soil yearlong incubation 13C-SIP-Lipidomics [Data Set] MassIVE. doi:10.25345/C57659T3K Rempfert KR, Bell SL, Kasanke CP, Kyle JE, Hofmockel KS. 2025. Lipids represent a dynamic, yet stable pool of microbially-derived soil carbon. In Prep This data package consists of compound-specific 13C SIP-lipidomics data from a yearlong tracer incubation experiment designed to investigate microbial lipid persistence in switchgrass bioenergy crop soils. In order to explore how lipid structure may modulate the persistence of C in soil lipids, we leveraged soils from two sites (Michigan - sandy texture, Wisconsin - silty texture) operated by the U.S. Department of Energy-funded Great Lakes Bioenergy Research Center (GLBRC). These sites had comparable climates, identical management practices, but contrasting soil textures, allowing us to assess the variability of lipid accrual or degradation in soils as well as provide insight regarding the degree to which edaphic properties may regulate the retention of soil lipids. Untargeted lipidomics analyses were performed to identify 13C-labeled lipids in the soil microbiome after long-term incubation. Soils were supplemented with 100 micrograms glucose per gram dry soil (99 atom % 13C or natural abundance for paired control) and incubated; samples were collected two months and one year after glucose addition. Lipid extracts (MPLEx) were analyzed by LC-MS/MS and identified using LIQUID. Calculation of isotopic enrichment of lipids was performed by targeted approach using TarMet to quantify lipid isotopologues and IsoCorrectoR to correct for natural abundance isotopes. Contents: Data package contents reported here are the first version and contain downstream analysis files for the raw LC-MS mass spectrometry files (.mzXML) deposited at the MassIVE database repository under accession MSV000097435 (80 experimental runs; 5.85 GB) | MassIVE DOI: 10.25345/C57659T3K. Support files include the additional data download 'Read Me' file containing data descriptor information. Reported data download contents are structured for compliance with project data sharing guidelines, community standards initiatives, and sponsor stakeholder policies supporting FAIR data principles. Data processing software, analysis tools, and data workflows are listed below corresponding to the host repository long-term location. Available Data Downloads (0.3 GB): "GLBRC soil yearlong incubation 13C-SIP-Lipidomics_readme.txt" - 'Read Me' data package content file (txt) "GLBRC_DataPackage_analysis files" - Data processing files (Rmd) and saved intermediate data processing outputs (rds, csv, xlsx) "GLBRC_13C_lipidomics_dataset.xlsx" - processed data in tabular format (xlsx) Linked Software: LIQUID LC-MS Analysis Software | 10.5281/zenodo.6459462 Lipid Mini-On Software Tools | 10.5281/zenodo.1492803 pmartR Omics Statistical Software | 10.5281/zenodo.6108667 xcms (v4.3.3) TarMet (v1.1.1) IsoCorrectoR (1.24.0) Funding Acknowledgments: This research was supported by an Early Career Research Program award funded by the U.S. Department of Energy, Office of Science, Office of Biological and Environmental Research (OBER) Genomic Science program under FWP 68292, FWP 07880 and EMSL Exploratory Research Project 51095. A portion of this work was performed in the William R. Wiley Environmental Molecular Sciences Laboratory, a national scientific user facility sponsored by OBER and located at Pacific Northwest National Laboratory (PNNL). PNNL is a multi-program national laboratory operated by Battelle for the DOE under Contract DE-AC05-76RLO1830.

Rempfert, Kaitlin R [Pacific Northwest National La

Hyper Spectral Anomaly Detection

The HSA is a statistics based anomaly detection model. The model performs unsupervised anomaly detection, based on a datapoint's density and similarity within a dataset. Density and similarity data are encoded into an affinity matrix. The affinity matrix is evolved to summarize the data's structure on greater topographical scales within the data's function space. The set of evolved affinity matrices and an anomaly score vector are passed to a user defined penalized objective function. The penalized objective function of anomaly scores is then minimized. Data points where the absolute value of the z-scores of anomaly scores greater than a specified threshold are predicted as anomalies. A novel multi-filter feature has also been implemented. To reduce false positive rates, the multi-filter records the indexes of the HSA predictions. A new dataset and data loader are instantiated consisting of all the initial HSA predictions and non-anomalous data points in a 10% and 90% split respectively. The HSA is then run through this data set and a count of number of times a data point is predicted is kept. In this way the initial predictions may be compared with data spanning the entire dataset. After the multi-filter is complete, all datapoints will have an associated anomaly score, as well as a multi-filter prediction count to further filter the anomalous predictions.

Rogers, DempseyD [Idaho National Laboratory (INL),

Understanding coarsening of a post-corrosion microstructure in a molten salt by combining phase-field modeling and in situ tomography

Alloys corroding in molten salt have been observed to form bicontinuous, nanoporous microstructures via dealloying, which subsequently undergo coarsening due to facile transport in high-temperature conditions. In this work, we describe a methodology to elucidate the underlying transport mechanisms during coarsening of a bicontinuous microstructure via quantitative comparisons between phase-field simulations and four-dimensional in situ experiments, in this case X-ray nanotomography of the coarsening of a dealloyed 80 wt% Ni-20 wt% Cr microwire in molten KCl-MgCl 2 at 800°C. We conduct phase-field simulations initialized from experimental data to model coarsening via three different transport mechanisms: surface diffusion, solid bulk diffusion, and liquid bulk diffusion. These simulations reproduce key features of the experiment, such as the densification of the outer layer of the dealloyed wire and the reduction in radius over time. We quantitatively compare different microstructural characteristics between the simulations and experiment and extract temporal scaling factors that optimally match the time scales of the simulations to that of the experiment. This allows us to evaluate morphological similarity between the simulations and experiment and relate the experimental coarsening kinetics to fundamental material properties. We find that surface diffusion is most likely to be the dominant coarsening mechanism, and its kinetics imply a surface diffusivity of D S = 8.9 x 10 -20 m 3 /s, which is within the range of reported values for Ni-vacuum interfaces at 800°C. However, the difference between the experiment and the surface diffusion simulation increases substantially at late times, suggesting that other mechanisms, such as the dissolution of residual Cr, may be at play.

36 MATERIALS SCIENCE

nmRanalysis: An Open-Source Web Application for Semi-automated NMR Metabolite Profiling

Though data acquisition and initial signal pre-processing of nuclear magnetic resonance (NMR) spectra have achieved high degrees of automation, downstream processing - specifically the profiling of spectra - has bottlenecked the overall NMR analysis workflow. Several efforts have been made to mitigate this bottleneck, but these solutions often trade an increase in automation for limitations elsewhere. Here, in this technical note, we introduce nmRanalysis, a user-friendly web-application that integrates the strengths of existing profiling tools for a more automated profiling workflow. nmRa-nalysis additionally incorporates novel features, including a machine-learning-driven recommender system for me-tabolite identification, further increasing the utility of nmRanalysis over the individual tools that it incorporates.

Flores, Javier E. [Pacific Northwest National Labo

Probing the atmospheric boundary layer with integrated remote-sensing platforms during the American WAKE ExperimeNt (AWAKEN) campaign

The American WAKE ExperimeNt (AWAKEN) collaboration is an observational-based field campaign in northern Oklahoma intended to analyze the potential influence of onshore wind farms and their collective wakes on wind power production, turbine structural loads, and on the atmospheric boundary layer (ABL). Focusing on the ABL effects, the University of Oklahoma and the Lawrence Livermore National Laboratory collected continuous high-resolution kinematic and thermodynamic profile measurements during 2022 and Summer 2023. The deployment strategy for these campaigns is detailed first, followed by an initial comparison of data from two sites in the AWAKEN domain: a near-farm site to examine collective wake impacts on the ABL, and a far-field site remaining outside the wind farm-waked region. Here, we summarize the datasets available and demonstrate the benefits of these observations and multiple value-added products (VAPs) for investigation of ABL features observed during AWAKEN. We also highlight examples of preliminary analyses, including ABL height detection and nocturnal low-level jet examination, which are produced using novel VAPs based on optimal estimation to retrieve deeper Doppler lidar wind profiles than previously resolved, along with their uncertainty. By including the near-farm and far-field site in these analyses, we identified a pattern of stronger lower-atmospheric mixing at the near-farm site than the far-field site, motivating deeper investigation into the relationship between wind farms and general ABL characteristics. Future analysis will delve deeper into this relationship by examining other ABL characteristics, such as atmospheric stability and convection.

17 WIND ENERGY

Low-threshold response of a scintillating xenon bubble chamber to nuclear and electronic recoils

A device filled with pure xenon first demonstrated the ability to operate simultaneously as a bubble chamber and scintillation detector in 2017. Initial results from data taken at thermodynamic thresholds down to ∼4 keV showed sensitivity to ∼20 keV nuclear recoils with no observable bubble nucleation by 𝛾-ray interactions. Here, this paper presents results from further operation of the same device at thermodynamic thresholds as low as 0.50 keV, hardware limited. The bubble chamber has now been shown to have sensitivity to ∼1 keV nuclear recoils while remaining insensitive to bubble nucleation by 𝛾-rays. A data-driven calibration of the chamber’s nuclear recoil nucleation response, as a function of nuclear recoil energy and thermodynamic state, is presented. Stringent upper limits are established for the probability of bubble nucleation by 𝛾-ray-induced Auger cascades, with a limit of <1.1 ×10 −6 set at 0.50 keV, the lowest thermodynamic threshold explored.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND

PRIME: An evaluation framework for protein representation inference and generalization in viral mutation space

Background Protein language models (PLMs) have revolutionized protein fitness prediction, yet their application to rapidly evolving viral pathogens is often confounded by extreme sequence homology. This homology leads to “data leakage” in standard random validation splits, yielding inflated performance metrics that fail to translate into real-world biosurveillance utility. Results We present Protein Representation Inference for Mutation Evaluation (PRIME), a framework that integrates domain-specific fine-tuning with a rigorous position-stratified validation protocol to evaluate viral threats. Using a dataset of 347,432 SARS-CoV-2 receptor binding domain (RBD) sequences, we demonstrate that while random training data split yields deceptive R 2 values (> 0.90), they fail to generalize to novel mutational sites. By benchmarking models up to 650 M parameters, we show that domain-specific fine-tuning of the ESM-C 600 M model with correctly stratified data provides an initial demonstration of predictive signal for binding affinity and expression at unseen mutational sites of binding affinity and expression on unseen sites (R 2 ~0.23), a significant advancement over base foundation models which exhibit no predictive power (R 2 <0). PRIME’s embedding-based clustering identified 3.03% of bat coronavirus sequences as candidates for further experimental prioritization based on their functional similarity to human-infective strains in embedding space, offering a perspective complementary to traditional phylogenetic methods. Conclusion PRIME establishes a new benchmark for the application of PLMs in pathogen surveillance. Our findings demonstrate that state-of-the-art models and fine-tuning, when paired with stratified validation, provide biologically meaningful insights into pathogen evolution and zoonotic risk.

59 BASIC BIOLOGICAL SCIENCES

X-ray Absorption Spectroscopy and Neutron Scattering Data, Mineral Incubation Experiment, Walker Branch Watershed, TN (2020 - 2021)

Iron (Fe) (oxyhydr)oxides are well-recognized contributors to soil carbon (C) storage, but the effects of manganese (Mn) oxides on carbon storage and transformation are relatively unexplored. Here, the relative capacities of Fe and Mn oxides to bind and stabilize soil organic C were directly compared using an in-situ incubation experiment. Quartz sands coated with either poorly crystalline Mn(III/IV) oxides or Fe(III) oxides, or left uncoated, were buried in a temperate forest soil for up to one year. This data package contains processed data outputs of carbon near edge x-ray absorption fine structure spectroscopy (C NEXAFS), iron and manganese x-ray absorption near edge structure spectroscopy (XANES), and small angle neutron scattering (SANS) data collected for initial oxide-coated and uncoated quartz sands and for those buried in the temperate forest soil for 365 days.

Energy (eV)

Next-Level Energy Management in Manufacturing: Facility-Level Energy Digital Twin Framework Based on Machine Learning and Automated Data Collection

This research introduces an energy prediction framework at the facility level supported by automated data collection and machine learning models. It investigates whether reducing the prediction time scale allows for applying more complex machine learning techniques and if those techniques improve the prediction accuracy. The primary advantages of this framework lie in its automation of the energy prediction process and its provision of real-time energy data suitable for use in energy dashboards or digital twins. A sitewide dataset was created by combining 15 min energy and daily production data of five shops—assembly, battery, body (electric), body (gas), and paint—from a globally recognized electric vehicle manufacturer. Various machine learning models were evaluated on daily, weekly, and monthly datasets, including, in increasingly complex order: naïve, simple linear regression, net regularized generalized linear regression, principal component regression, k-nearest neighbor, random forest, and Bayesian regularized neural network. Compared to the current state-of-the-art energy consumption prediction for the industrial facility level, this research investigates more complex models and smaller time intervals for higher accuracy. The findings revealed that the more complex monthly models require a minimum of a year and a half of data to operate, while weekly models demand a year of data to achieve improved accuracy. Daily models can operate with only six months of data but exhibit poor performance due to reduced prediction accuracy of production. Key challenges identified include access to reliable, high-quality energy and production data and the initial demand for human labor.

digital twin

The AWAKEN wind farm benchmark, Part 2: Modeling results

Accurately modeling wind farm performance in complex atmospheric flows remains a challenge. This paper presents the modeling results of the American WAKE experimeNt (AWAKEN) wind farm benchmark, a collaborative effort involving 16 research groups from academia and industry within the International Energy Agency Wind Technology Collaboration Programme Task 57. The study evaluates a diverse suite of simulation tools, ranging from fast-running engineering wake models to high-fidelity large-eddy simulations, against a diurnal case study observed during the AWAKEN campaign. The benchmark utilized a three-phase structure to progressively assess model performance as observational data availability increased. Initial blind predictions showed that higher-fidelity models did not uniformly outperform simpler simulation tools. A distinct spatial bias was observed where models struggled to resolve the interplay between a low-level jet, wakes, and terrain-induced flow acceleration. In subsequent phases, leveraging additional measurements for model improvement led to a reduction in mean absolute error across the model ensemble; however, this effect was most pronounced in engineering wake models, where targeted calibration reduced error by up to 40~\%. Overall, the study demonstrates that inflow characterization remains a primary prerequisite for accuracy, particularly for models relying on coarse forcing datasets. While the limited ability to resolve local terrain-flow interactions under single-day conditions represent a recognized constraint, the overall findings on wake modeling and real-world validation still provide valuable guidance for model application and for mitigating this limitation.

Bodini, Nicola

Xanthos-Lake Dataset

The Xanthos-Lake v1.0 dataset provides the input data, trained machine-learning models, and simulation outputs needed to characterize lake water balance, snow and ice conditions, and mixing-layer temperature within the Xanthos global hydrological modeling framework. The dataset supports lake representation across a wide range of lake sizes and hydroclimatic conditions by combining xLSIM, a basin-specific machine-learning emulator of lake snow, ice, ice-cover fraction, and mixing-layer temperature, with the Xanthos-Lake water-balance model. The archive contains NetCDF datasets used to train and evaluate xLSIM, trained model weights, processed meteorological and lake-property inputs, and basin- and lake-category-specific simulation outputs. These materials are organized into four primary data groups, described below. Snowice_model_inputs: Contains the NetCDF input data used to train xLSIM. The xLSIM machine-learning framework uses three lake-based datasets. The meteorological forcing dataset provides monthly relative humidity, specific humidity, surface wind speed, maximum and minimum air temperature, downward longwave and shortwave radiation, snowfall, surface air pressure, and total precipitation. Lake surface area is included as an additional static predictor. The target-state dataset provides lake ice thickness, snow depth, snow cover, and lake mixing-layer temperature, while a companion lake-surface dataset provides the lake ice-cover fraction. Before training, ice thickness and snow depth are converted from meters to centimeters, mixing-layer temperature is converted from kelvin to degrees Celsius and constrained to nonnegative values, and ice-cover fraction is converted from a fraction to a percentage. The predictor variables are normalized using statistics calculated across the selected lakes and time steps. Snowice_model_outputs: Contains the NetCDF outputs generated by xLSIM. For each basin, xLSIM produces a file containing observed and predicted lake-state variables for the training, validation, and testing periods. The modeled variables include lake ice thickness, snow depth, snow cover, mixing-layer temperature, and lake ice-cover fraction. For basins without a sufficiently persistent snow-and-ice signal, the emulator predicts only mixing-layer temperature. The outputs also include training and validation loss histories, the selected model configuration, identifiers of the lakes used in training, and SHAP-based feature-importance information at the global, lake, and seasonal-regime levels. The trained machine-learning model weights are provided separately within the dataset archive. Together, these files support model evaluation and subsequent coupling with the Xanthos-Lake water-balance framework. XanthosLAKES: Contains the NetCDF input data used by the Xanthos-Lake framework. Monthly meteorological inputs include relative and specific humidity, downward shortwave and longwave radiation, mean, maximum, and minimum air temperature, wind speed, precipitation, snowfall, and surface air pressure. Static lake-property datasets provide lake identifiers, geographic locations, surface area, volume, mean depth, elevation, drainage area, fetch, outlet-routing information, and associated Xanthos grid-cell attributes. Separate bathymetric datasets provide the coefficients of the area–depth and volume–depth relationships for each aggregated lake unit. GLEV-based records provide observed lake surface area and evaporation data used to initialize lake states, define reference conditions, and calibrate and evaluate the model. Xanthos-Lake Outputs: Contains the basin- and lake-category-specific NetCDF outputs generated by Xanthos-Lake. Monthly variables include lake surface area, storage volume, outlet discharge, evaporation rate, evaporation volume, lake–groundwater exchange, lake inflow, ice thickness, snow depth, snow-cover fraction, ice-cover fraction, and mixing-layer temperature. The files also contain lake-specific calibration and validation statistics, including normalized root-mean-square error, mean absolute error, Nash–Sutcliffe efficiency, Kling–Gupta efficiency, and percent bias. Stored calibrated and derived parameters include the weir discharge coefficient, fractional freeboard, groundwater exchange coefficient, reference water level, corresponding reference surface area and storage volume, weir-width adjustment factor, and the fraction of routed inflow entering the lake. Basin identifiers, lake category, simulation period, calibration and validation periods, and parameter-schema information are retained as NetCDF metadata.

Abeshu, Guta [Pacific Northwest National Laborator