Search NASA⌕ Search

SEARCH · Search NASA

Results for “reference”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Automating the Analysis of Large Language Models Responses through Zero-Shot Question Answering

Recent advancements in Large Language Models (LLMs) have shown significant potential in various applications, yet their evaluation, particularly in zero-shot question answering scenarios, remains a challenging task. In this study, our objective was to explore precision metrics for Large Language Models (LLM) and design and implement a software pipeline to automatically evaluate LLMs' outputs under zero-shot question answering. Zero-shot question answering involves a model providing answers to questions about topics it hasn't seen during training. It leverages the principles of zero-shot learning by relying on semantic understanding and generalization from related knowledge. The data used was metadata from medical databases on congenital heart disease. We explored eleven LLM metrics and selected three for our evaluation: BLEU, BERTScore, and MoverScore. BLEU calculates a score based on the overlap of n-grams (contiguous sequences of n items, typically words) between the machine-generated translation and the reference translations. Higher BLEU scores indicate better correspondence between the machine-generated and human-generated translations. BERTScore is a metric used to evaluate the quality of machine-generated text by measuring the similarity of token embeddings produced by BERT (Bidirectional Encoder Representations from Transformers) between the generated text and reference text. MoverScore is a metric that quantifies the dissimilarity between the distributions of word embeddings from machine-generated text and reference text, emphasizing semantic similarity over exact token overlap. We also introduced HBKI, a composite metric summarizing these approaches. We tested five models —GPT-3, Llama-2, Gemini 1.5 Pro, Solar 10.7B, and Mixtral-8x7b. Our software pipeline, designed and implemented using Object-Oriented Programming principles, allows users to customize the selection and extraction of features for topics of interest in their own research. Our results show that MoverScore delivered the most precise evaluation of the LLM's outputs, while Mixtral-8x7b achieved the best overall performance in extracting metadata from the databases.

97 MATHEMATICS AND COMPUTING↗

Cybersecurity Considerations for Hydrogen Infrastructure in Airport Environments

This report explores key cybersecurity concerns and best practices within environments that serve as reference points for the development of hydrogen fueling infrastructure for aviation. This cybersecurity analysis leverages prior NREL studies: 1) hydrogen fueling station component validation to identify vulnerabilities and failure events documented in physical equipment, and 2) electric aircraft charging infrastructure analysis to explore primary cybersecurity vulnerabilities. It reviews the criticality of digitized technologies in sustaining hydrogen fuel production, storage, and fueling systems, noting cybersecurity concerns that are universal to power systems and industrial control systems in general. In considering cybersecurity vulnerabilities within a future landscape of hydrogen energy for aviation applications, a reference architecture was intended to reveal the points of connection between assets and the potential sensors that are vulnerable to manipulation in the event of compromised access or communication within a SCADA system. A generalized reference architecture can help stakeholders, engineers, or strategists understand connections, criticalities, and standard practices when it comes to designing and planning for new systems. There are several gaps to account for in assessing the future of hydrogen production, storage, and fueling for aviation. Engaging stakeholders, including aircraft manufacturers, electric utilities, site property owners, and local communities, will inform decision-making around site structure, operations, and resources for future hydrogen fueling infrastructure to understand operational needs and cybersecurity awareness. Cybersecurity mitigation strategy must consider physical attack vectors that emerge with the integration of hydrogen systems into existing airport security requirements. The cybersecurity risk assessment contained in this report is an entry point into potential future granular-level analyses to be conducted as part of hazard and risk assessments for safe aviation hydrogen infrastructure, determining how the scale of hydrogen fuel infrastructure for aviation impacts the volume of cyber attack vectors, and what, if any, are the vulnerabilities associated with different types of on-board hydrogen systems. In this nascent development phase, assessing how best to integrate cybersecurity practices into an evolving U.S. aviation landscape provides critical insights into building increased awareness and stakeholder engagement to support a cyber-resilient infrastructure.

08 HYDROGEN↗

Photovoltaic Mini-Module Soiling Stations: Cooperative Research and Development (Final Report)

Photovoltaic (PV) panels can become soiled due to a variety of environmental factors. This soiling reduces the amount of light reaching the cells and thus their energy output. The ratio of actual output of the soiled device that of a clean device is known as the soiling ratio. The daily change in soiling ratio is known as the soiling rate. To study soiling for different glass coatings. Two soiling stations were fabricated. The stations were designed to monitor the short-circuit current of PV cells encapsulated with glass featuring different coatings and measure the daily soiling ratio with a pair of reference cells, one of which is automatically brushed daily. This work meets the need to understand how the glass coatings perform in additional environments. The stations are designed to quantify both the soiling ratio (the ratio between actual/expected photovoltaic panel power output) as well as differences in anti-soiling performance of different glass coatings. The design of the dirty/clean reference cell system is described by Toth et al. The soiling ratio is calculated by calculating the average irradiance measured within one hour of solar noon by each of the reference cells and then taking the ratio of these daily near-noon averages. The soiling ratio time series for the station deployed in Georgia is shown in Figure 1. This figure shows that the soiling ratio at the site has not yet fallen below 0.99, indicating that peak daily soiling losses have been below 1%. The site was chosen because it is thought to be affected by pollen soiling. With continued monitoring over the course of at least a full year, we expect to be able to observe and quantify pollen soiling events.

14 SOLAR ENERGY↗

Industrial Electrification Technologies Booklet

Electrification refers to transitioning from fuel-powered systems to electrically-powered alternatives. In an industrial context, this primarily refers to the electrification of process heat, but there are also widespread opportunities to electrify space heating and cooling as well as onsite transportation (e.g., forklifts). Electrification does not refer to one particular technology, but a wide range of technologies that use different methods to transfer electrical energy into a usable form – usually heat.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Validating Greater Sage-Grouse Individual-based Model (IBM) Tool (Final Report)

The project focused on validating the previously developed Greater Sage-Grouse Individual-based Model (GrSG IBM; LaGory et al. 2012, 2021). The objective was to transform this predictive, spatially and temporally explicit model into a portable resource to assist siting/resource managers in proactively assessing the cumulative impacts of wind energy development on the greater sage-grouse. Utilizing a bottom-up, individual-based approach, the GrSG IBM accounts for landscape context and species behavior, aiming to reduce uncertainty in estimating development impacts and support ecologically mindful land-based wind energy development. The validation effort covered approximately 6,540 km 2 near the Seven Mile Hill Wind Project in Wyoming. The GrSG IBM tool, built on the NetLogo platform (Tisue and Wilensky 2004), was executed over a 50-year period, with the analysis focusing on years following a 10-year initialization phase. Key results demonstrated the tool’s biological soundness across five key biological metrics: non-chick age class distribution (older than 10 weeks), adult sex ratio, life expectancy, population size, and overall population growth. For instance, the tool estimated that 58.6% of the non-chick population was reproductively immature, while the reference ranges from 51.4% to 57.8% (Patterson 1952, Rogers 1964). Experts confirmed the tool’s estimate was within a reasonable range for the species. The tool estimated average life expectancy of 1.43 years, while the reference ranges from 0.9 years to 1.1 years (Ammann 1957, Hamerstrom 1949). Experts also supported the model’s life-expectancy estimate as ecologically sound for the species in the study area. In terms of population change, the model estimated an annual shift between a 0.6% decline and a 1.0% increase over 50 years. While the reference suggests 2.9% annual decline in range-wide populations (Cortes et al. 2023), that includes many at-risk populations in South Dakota and Washington, for example. Our study area—in the northeastern part of Carbon County and western-edge of Albany County, Wyoming—is one of the remaining greater sage-grouse habitats supporting some of the most stable populations. Experts confirmed that the range of the annual population change spanning from a 0.6% decline to a 1.0% increase estimated by the tool was reasonable for our study area for this reason and confirmed that aligned with population estimates from existing studies on the greater sage-grouse and wind energy development in the study area (LeBeau et al. 2017a, Smith et al. 2024). Furthermore, the project showed that temporally explicit biological metrics generated by the GrSG IBM tool can complement the USGS’ Prioritizing Restoration of Sagebrush Ecosystems Tool (PReSET; Duchardt et al. 2021) by incorporating habitat restoration strategies into seasonal habitat suitability models to visualize population responses over time.

17 WIND ENERGY↗

INL CRMO XRF Special Report No. 3: INL CRMO XRF Calibration Methods for the Analysis of Ceramics and Brick

This report details the analytical methods used by the Idaho National Laboratory Cultural Resource Management Office to characterize ceramic and brick artifacts using X-ray fluorescence spectrometry. More specifically, this report provides details on instrumentation, analysis times, and matrix-specific calibration protocols for ceramic materials. Calibration curves were developed to adjust raw instrumental output for 13 elements using a suite of reference standards made from historic brick borrowed from the Yale Peabody Museum. Calibration curves were validated by plotting predicted values against published reference values for the calibration set as well as a certified soil reference standard (SRM 2711a).

36 MATERIALS SCIENCE↗

Pressure Safety Training (Rev. 8)

This is the workbook for the Pressure Safety Training Course. It is intended as a reference manual and guide for all work with pressurized liquids or gases. This workbook contains basic references to make work with pressure safer. It is intended to supplement classroom instruction, rather than serve as a definitive text on pressure. Earlier versions of the manual were intended specifically for training at Lawrence Livermore National Laboratory. This revision is a generic version intended for training at all DOE facilities. The information in the Standards chapter is from the LLNL Health and Safety Manual. It is included here as a convenient reference and guide for developing similar standards at your own facility.

42 ENGINEERING↗

Develop and verify soil/structure interaction for pile/foundation interaction

Phase II of the Offshore Code Comparison Collaboration, Continued, with Correlation and unCertainty (OC6) project was used to verify the implementation of a new soil-structure interaction (SSI) model for use within offshore wind turbine modeling software. The REDWIN Macro-element model implemented and verified in this study enables a computationally efficient way to model the linear and nonlinear SSI problem, including hysteretic damping, of a monopile structure. The modeling approach was integrated into several modeling tools and a series of increasingly complex simulations was conducted using the IEA 10MW reference turbine mounted on a monopile support structure to verify the coupling between the tools and the REDWIN Macro-element SSI model. This campaign includes only numerical verification between various software and modeling approaches so no experimental measurements are available. The load cases (LC) considered include: LC1 – static response of the tower and substructure LC2 – frequency and mode-shape analysis of the tower and substructure LC3 – response of the tower and substructure due to wind-only loading LC4 – response of the tower and substructure due to wave-only loading LC5 – response of the tower and substructure due to wind and wave loading. Detailed properties of the modeled system are found in the following reference, “Bergua, Roger, Amy Robertson, Jason Jonkman, and Andy Platt. 2021. "Specification Document for OC6 Phase II: Verification of an Advanced Soil-Structure Interaction Model for Offshore Wind Turbines.” Golden, CO: National Renewable Energy Laboratory. NREL/TP-5000-79938. https://www.nlr.gov/docs/fy21osti/79938.pdf. Details on the results from the OC6 Phase II project can be found in the following reference, “Bergua R, Robertson A, Jonkman J, et al. OC6 Phase II: Integration and verification of a new soil–structure interaction model for offshore wind design.” Wind Energy. 2022;25(5):793-810. doi:10.1002/we.2698

17 WIND ENERGY↗

FOCAL Campaign II/III: Applying Active Hull Controls Using Tuned-mass Dampers/Hull Flexibility and Internal Loads

Campaign II and III of the Floating Offshore-wind Controls Advanced Laboratory Experimental Program (FOCAL) aimed to generate a dataset enabling the validation of the performance and loads of a scaled hull, with and without structural hull control. The floating platform was subjected to a variety of wave environments and controlled using tuned-mass dampers (TMDs) tuned to two of the systems natural frequencies (Platform Pitch and Tower-bending). The floating platform is fully instrumented to record a variety of parameters in real time such as platform dynamics, accelerations, and loads at different points in the structure. The test data considered was generated at the University of Maine's Harold Alfond Wind and Wave (W2) testing facility. This testing was focused only on validation of wave loading, and wind conditions were not considered. As such the platform does not support a working turbine, and instead supports a structure designed to have the same mass properties as the 1:70 IEA 15MW Reference turbine. The Load Cases (LC) considered in this testing campaign are as follows: LC 1.X - Platform Static Offset (TMD off); LC 2.X - Platform Free-decays (TMD off); LC 3.X - Wave Cases (Regular Wave, Irregular Wave, Pink Noise Wave) with and without TMDs active. Detailed properties on the model system are found in the following reference: Lenfest E., Floating Offshore-wind Controls Advanced Laboratory (FOCAL) Experimental Program - Campaigns 2 and 3: 1:70 Model-scale Testing of the IEA-Wind 15MW Reference Turbine and the VolturnUS-S Hull. UMaine ASCC Report Number 23-56-1183.

17 WIND ENERGY↗

Human RNome Project draft human RNome sequence of GM12878, B-cell line, obtained by mass-spectrometry sequencing, long-read sequencing and short-read sequencing.

Here we report the first draft of the human RNome sequence, a reference map of RNA chemical modifications in a human B-cell line. RNA carries a diverse repertoire of chemical modifications that regulate gene expression, cellular function, and responses to physiological and pathological cues. Yet, unlike the genome, no reference map of RNA modifications is available for any human cell. To generate this resource, the Human RNome Project Consortium analyzed a shared RNA preparation from the well-characterized GM12878 B-cell line using short-read sequencing, long-read direct RNA sequencing, and mass spectrometry, generating more than 7.1 billion sequencing reads spanning approximately 1.2 trillion nucleotides. The resulting maps of the human RNome reveal that RNA modifications are organized according to function, transcript architecture, and cellular identity. Modifications concentrate at functional centers of ribosomal and transfer RNAs, follow the canonical topology of N6-methyladenosine in coding transcripts, and form coordinated hotspots in immune regulatory genes. This first reference human RNome provides a foundation for understanding how RNA chemistry shapes cellular identity, human disease, and the development of RNA-based therapeutics.

59 BASIC BIOLOGICAL SCIENCES↗

Wind Turbine Sound Setbacks and Supply Curves: Ordinances and Extrapolated Trends, 110 Hub Height, 130 Rotor Diameter

This dataset provides a comprehensive set of wind turbine sound setbacks from every residential structure in the contiguous United States (CONUS). A sound setback is defined as the minimum required distance between a residential structure and a hypothetical turbine installation site to ensure that modeled sound levels received at the residence do not exceed local sound ordinances, which are commonly expressed in A-weighted decibels (dBA). Therefore, sound setbacks are a local spatial assessment combining multiple factors, including the sound pressure curve as a function of the observer location (distance and direction) relative to the turbine, local sound regulations, and the geographical distribution of residential structures. The dataset is organized into multiple scenario-based products, detailed as follows: 1. Existing and extrapolated sound setbacks. An existing scenario characterizes sound setbacks only in states or counties that have implemented sound regulations as of 2022. The extrapolated scenarios extend a constant sound threshold to counties that lack explicit sound regulations, with thresholds ranging from 35 to 60 dBA, in 5-dBA increments reflecting the variation observed in current sound ordinances. 2. Sound setbacks in directional and worst scenarios. The directional scenario accounts for the distance and orientation of residential structures relative to a hypothetical turbine location, utilizing the turbine's sound emissions in that specific direction. In contrast, the worst scenario takes loudest sound level at each distance step from the turbine, irrespective of directional considerations, which aligns with current industry practice. 3. Supply curves for Open and Reference Access scenarios. This dataset includes supply curves generated by the reV model, which integrates each of the above sound setbacks into both Open and Reference siting scenarios. In addition, two Open and Reference baselines scenarios were included which do not consider sound setbacks for comparative analysis. All sound setback data are stored in TIF files, with partial maps of the data provided in PNG format. The values in the sound setback raster range from 0 to 1, representing the fraction of developable land within a 90 meter by 90 meter pixel due to sound ordinances. A value of 0 indicates areas where wind energy development is prohibited, while a value of 1 signifies areas fully permissible. The wind turbine parameters used in the sound modeling are based on the land-based turbine from International Energy Agency (IEA), featuring a rated electrical power of 3.4 MW, a rotor diameter of 130 meters, and a hub height of 110 meters. The atmospheric conditions, including wind speed/direction, turbulence, air temperature, relative humidity, and air pressure, that drive the sound generation are obtained from the WIND Toolkit dataset.

Array↗

Microstructure, electrical resistivity, and tensile properties of neutron-irradiated Cu–Cr–Nb–Zr

High strength, high conductivity copper alloys that can resist creep at high temperatures are one of the primary candidates for efficient heat exchangers in fusion reactors. Cu–Cr–Nb–Zr (CCNZ) alloys, which were designed to improve the strength and creep life of ITER Cu–Cr–Zr (CCZ) reference alloys, have been found to have comparable electrical conductivity and tensile properties to CCZ alloys. The measured creep rupture times for these improved alloys is about ten times higher than the ITER reference alloys at 90–125 MPa at 500 °C. However, the effects of neutron irradiation on these alloys, and the ensuing material properties, have not been studied; thus, their utility in a fusion reactor environment is not well understood. This study characterizes the room temperature mechanical and electrical properties of a neutron-irradiated CCNZ alloy and compares them to a neutron-irradiated ITER reference heat sink CCZ alloy. Tensile specimens were neutron irradiated in the High Flux Isotope Reactor (HFIR) to 5 dpa between 250 °C and 325 °C. Post-irradiation characterization included electrical resistivity measurements, hardness, and tensile tests. Microstructural evaluation used scanning electron microscopy, energy dispersive x-ray spectroscopy, and atom probe tomography to characterize the irradiation-produced changes in the microstructure and investigate the mechanistic processes leading to post-irradiation properties. Transmutation calculations were validated with composition measurements from atom probe data and used to calculate contributions to the increased electrical resistivity measured after irradiation. Comparisons with CCZ alloys in the same irradiation heat found that the post-irradiated CCNZ and CCZ alloys had comparable electrical resistivity. Although CCNZ alloys suffered more irradiation hardening than CCZ, the overall tensile behavior deviated very little from non-irradiated values in the temperature range studied.

36 MATERIALS SCIENCE↗

Simulating Continuum-based Redshift Measurement in the Roman’s High Latitude Spectroscopic Survey

We investigate the capability of the Nancy Grace Roman Space Telescope’s (Roman) Wide-Field Instrument G150 slitless grism to detect red, quiescent galaxies based on the current reference survey. We simulate dispersed images for Roman reference High-Latitude Spectroscopic Survey (HLSS) and analyze two-dimensional spectroscopic data using the grism Redshift and Line Analysis (Grizli) software. This study focus on assessing Roman grism’s capability for continuum-level redshift measurement for a redshift range of 0.5 ≤ z ≤ 2.5. The redshift recovery is assessed by setting three requirements of: σ z = $\frac{|z–z_{true}|}{1+z}$ ≤ 0.01, signal-to-noise ratio≥ 5 and the presence of a single dominant peak in redshift likelihood function. We find that, for quiescent galaxies, the reference HLSS can reach a redshift recovery completeness of ≥50% for F158 magnitude brighter than 20.2 mag. We also explore how different survey parameters, such as exposure time and the number of exposures, influence the accuracy and completeness of redshift recovery, providing insights that could optimize future survey strategies and enhance the scientific yield of the Roman in cosmological research.

Astronomical simulations↗

DESI Emission-line Galaxies: Unveiling the Diversity of [O II ] Profiles and Its Links to Star Formation and Morphology

We study the [O II ] profiles of emission-line galaxies (ELGs) from the Early Data Release of the Dark Energy Spectroscopic Instrument (DESI). To this end, we decompose and classify the shape of [O II ] profiles with the first two eigenspectra derived from principal component analysis. Our results show that DESI ELGs have diverse line profiles, which can be categorized into three main types: (1) narrow lines with a median width of ∼50 km s −1 , (2) broad lines with a median width of ∼80 km s −1 , and (3) two redshift systems with a median velocity separation of ∼150 km s −1 , i.e., double-peak galaxies. To investigate the connections between the line profiles and galaxy properties, we utilize the information from the COSMOS data set and compare the properties of ELGs, including star formation rate (SFR) and galaxy morphology, with the average properties of reference star-forming galaxies with similar stellar mass, sizes, and redshifts. Our findings show that, on average, DESI ELGs have a higher SFR and more asymmetrical/disturbed morphology than the reference galaxies. Moreover, we uncover a relationship between the line profiles, the excess SFR, and the excess asymmetry parameter, showing that DESI ELGs with broader [O II ] line profiles have more disturbed morphology and higher SFR than the reference star-forming galaxies. Finally, we discuss possible physical mechanisms giving rise to the observed relationship and the implications of our findings on the galaxy clustering measurements, including the halo occupation distribution modeling of DESI ELGs and the observed excess velocity dispersion of the satellite ELGs.

79 ASTRONOMY AND ASTROPHYSICS↗

Marine and continental stratocumulus cloud microphysical properties obtained from routine ARM Cimel sunphotometer observations

This study investigates marine and continental stratocumulus (Sc) cloud properties obtained from an automated implementation of a multispectral photometer retrieval. Photometer methods simultaneously retrieve cloud optical depth (τ) and cloud droplet effective radius (r e ), with estimates for liquid water path (LWP) calculated on the availability of those quantities. These applied methods evaluate retrieved cloud properties for Sc identified during a recent 6 year period over the U.S. Department of Energy Atmospheric Radiation Measurement (ARM) program sites in Oklahoma, USA (SGP) and in the Azores, Portugal (ENA). Modest agreement in key quantity retrievals is found between the routine photometer products and multisensor collocated profiling references. Cumulative breakdowns contingent on cloud thickness indicate increases in all retrieved quantities in thicker clouds, with larger discrepancies in the relative performance between the retrievals collected in the presence of drizzle. Under continental cloud conditions, the clouds of a similar thickness and r e to those sampled under marine conditions report a factor of 1.5 larger τ and LWP. An r 2 ≅0.65 is found between photometer τ retrievals and shadowband radiometer measurements, with photometer retrievals reporting a high (relative) bias. The τ intercomparisons indicate that variability between retrievals is a factor of three larger than errors reported from individual retrieval input perturbation tests. Photometer r e retrievals suggest a low r 2 (< 0.1) having a standard deviation ≅ 3 µm when compared to ARM baseline multi-sensor radar/radiometer references (accounting for offsets in the cloud droplet number concentration assumptions of the latter). However, photometer LWP calculations remain relatively unbiased in non-drizzling conditions, with errors O (50 g m −2 ) and r 2 ≅0.5 to collocated radiometer and interferometer references. Additional sensitivity tests for island influences on marine Sc properties suggest that while island-influenced winds may promote larger cloud LWP or thickness, the influence could be within retrieval method uncertainty and/or collocated instrument variability.

54 ENVIRONMENTAL SCIENCES↗

Evaluating the potential of short-term instrument deployment to improve distributed wind resource assessment

Distributed wind projects, which are connected at the distribution level of an electricity system or in off-grid applications to serve specific or local energy needs, often rely solely on wind resource models to establish wind speed and energy generation expectations. Historically, anemometer loan programs have provided an affordable avenue for more accurate onsite wind resource assessment, and the lowering cost of lidar systems has shown similar advantages for more recent assessments. While a full 12 months of onsite wind measurement is the standard for correcting model-based long-term wind speed estimates for utility-scale wind farms, the time and capital investment involved in gathering onsite measurements must be reconciled with the energy needs and funding opportunities that drive expedient deployment of distributed wind projects. Much literature exists to quantify the performance of correcting long-term wind speed estimates with 1 or more years of observational data, but few studies explore the impacts of correcting with months-long observational periods. This study aims to answer the question of how short you can go in terms of the observational time period needed to make impactful improvements to model-based long-term wind speed estimates. Three algorithms, multivariable linear regression, adaptive regression splines, and regression trees, are evaluated for their skill at correcting long-term wind resource estimates from the European Centre for Medium-Range Weather Forecasts Reanalysis version 5 (ERA5) using months-long periods of observational data from 66 locations across the US. On average, correction with even 1 month of observations provides significant improvement over the baseline ERA5 wind speed estimates and produces median bias magnitudes and relative errors within 0.22 m s −1 and 4 percentage points of the median bias magnitudes and relative errors achieved using the standard 12 months of data for correction. However, in cases when the shortest observational periods (1 to 2 months) used for correction are not well correlated with the overlapping ERA5 reference, the resultant long-term wind speed errors are worse than those produced using ERA5 without correction. Summer months, which are characterized by weaker relative wind speeds and standard deviations for most of the evaluation sites, tend to produce the worst results for long-term correction using months-long observations. The three tested algorithms perform similarly for long-term wind speed bias; however, regression trees perform notably worse than multivariable linear regression and adaptive regression splines in terms of correlation when using 6 months or less of observational data for correction. Translating the analysis to wind energy, median relative errors in the capacity factor are on average within 10 % using 1 month of training. If the observation period used for correction is not well correlated with the reference data, however, misrepresentation of the observed capacity factor can be substantial. The risk associated with poor correlation between the observed and reference datasets decreases with increasing training period length. In the worst-correlation scenarios, the median capacity factor relative errors from using 1, 3, and 6 months are within 47 %, 26 %, and 16 %, respectively.

17 WIND ENERGY↗

IM3 Open Source Data Center Atlas

IM3 Open Source Data Center Atlas Description This dataset contains locations of existing data center facilities in the United States. Data center locations were derived from OpenStreetMap (OSM), a crowd-sourced database. Data points from OSM are processed in various ways to determine additional variables provided in the data including: facility area (square feet), associated US county, and US state. This dataset can be used to identify areas of concentrated data center development and inform government and private sector planning strategies for future buildout of data centers and the infrastructure necessary to support it. Usage Notes Validation of OSM-derived data center locations is an ongoing development under the IM3 project, and the database will be updated as new information becomes available. In some instances, both the data center area (e.g., campus) and individual data center buildings are included as overlapping areas in the database. Both values are retained. Data center points, buildings, and campus areas are provided as separate layers in the downloadable data package. Note that data items are not necessarily complete across layers. That is, a specific data center may only be present as a single point geometry in the "point" layer while other data centers are represented in both the campus and building layers. In some cases, data center campuses and/or buildings straddle a county boundary line. Mappings to both counties are retained in the database as separate rows. These data rows will have the same data center id information, but each will have different county information. Crowd-sourced data, by nature, relies on individuals and communities to provide information. As a result, some data may be missing where it has not yet been reported. As we collect information on additional data center locations and as OSM receives additional contributions, the database will be updated to capture additional data points not yet shown. Technical Information Data is available for download under the following formats: GeoPackage (GPKG) CSV Geospatial data is provided in the WGS84 (EPSG:4326) coordinate reference system. The GeoPackage download contains the following layers. See usage notes for more information. "point" "building" "campus" The "point" layer includes all data from OSM that had POINT geometry type (i.e., individual coordinates). The "building" layer includes all OSM data that did not have POINT geometry and where the building tag in the OSM export was neither equal to "no" or null. Data that did not meet the "point" or "building" qualification was assumed to be a facility campus and included in the "campus" layer. The dataset contains the following parameters. Variables provided by OSM are labeled with (OSM-provided). id - unique identification number (OSM-provided with prefix of "node/", "relation/" and similar attributes removed) state - name of US state state_abb - two letter US state abbreviation state_id - state ID number county - name of US county county_id - county ID number ref - reference numbers or codes (OSM-provided) operator - the name of the company, corporation, or person in charge facility (OSM-provided) name - name of facility (OSM-provided) sqft - surface area of facility polygon, measured in square feet. Only available for "building" and "campus" layers lat - latitude of data centroid point lon - longitude of data centroid point type – represented spatial information. One of "point", "building", or "campus". geometry – POLYGON geometry of area footprint (in "campus" and "building" layers) or POINT geometry of locations (in "point" layer). This parameter is not included in the csv download. Attribution Data center locations were derived from OpenStreetMap, which is made available at openstreetmap.org under the Open Database License (ODbL). US state and county boundary information was collected from the US Census Bureau for the year 2024, which is made publicly available at https://www.census.gov/geographies/mapping-files.html Acknowledgment IM3 is a multi-institutional effort led by Pacific Northwest National Laboratory and supported by the U.S. Department of Energy's Office of Science as part of research in MultiSector Dynamics, Earth and Environmental Systems Modeling Program. License The IM3 Open Source Data Center Atlas is made available under the Open Database License: http://opendatacommons.org/licenses/odbl/1.0/. Disclaimer This material was prepared as an account of work sponsored by an agency of the United States Government. Neither the United States Government nor the United States Department of Energy, nor the Contractor, nor any or their employees, nor any jurisdiction or organization that has cooperated in the development of these materials, makes any warranty, express or implied, or assumes any legal liability or responsibility for the accuracy, completeness, or usefulness or any information, apparatus, product, software, or process disclosed, or represents that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by trade name, trademark, manufacturer, or otherwise does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States Government or any agency thereof, or Battelle Memorial Institute. The views and opinions of authors expressed herein do not necessarily state or reflect those of the United States Government or any agency thereof. PACIFIC NORTHWEST NATIONAL LABORATORYoperated byBATTELLEfor theUNITED STATES DEPARTMENT OF ENERGYunder Contract DE-AC05-76RL01830

Mongird, Kendall [Pacific Northwest National Labor↗

Xanthos-Lake Dataset

The Xanthos-Lake v1.0 dataset provides the input data, trained machine-learning models, and simulation outputs needed to characterize lake water balance, snow and ice conditions, and mixing-layer temperature within the Xanthos global hydrological modeling framework. The dataset supports lake representation across a wide range of lake sizes and hydroclimatic conditions by combining xLSIM, a basin-specific machine-learning emulator of lake snow, ice, ice-cover fraction, and mixing-layer temperature, with the Xanthos-Lake water-balance model. The archive contains NetCDF datasets used to train and evaluate xLSIM, trained model weights, processed meteorological and lake-property inputs, and basin- and lake-category-specific simulation outputs. These materials are organized into four primary data groups, described below. Snowice_model_inputs: Contains the NetCDF input data used to train xLSIM. The xLSIM machine-learning framework uses three lake-based datasets. The meteorological forcing dataset provides monthly relative humidity, specific humidity, surface wind speed, maximum and minimum air temperature, downward longwave and shortwave radiation, snowfall, surface air pressure, and total precipitation. Lake surface area is included as an additional static predictor. The target-state dataset provides lake ice thickness, snow depth, snow cover, and lake mixing-layer temperature, while a companion lake-surface dataset provides the lake ice-cover fraction. Before training, ice thickness and snow depth are converted from meters to centimeters, mixing-layer temperature is converted from kelvin to degrees Celsius and constrained to nonnegative values, and ice-cover fraction is converted from a fraction to a percentage. The predictor variables are normalized using statistics calculated across the selected lakes and time steps. Snowice_model_outputs: Contains the NetCDF outputs generated by xLSIM. For each basin, xLSIM produces a file containing observed and predicted lake-state variables for the training, validation, and testing periods. The modeled variables include lake ice thickness, snow depth, snow cover, mixing-layer temperature, and lake ice-cover fraction. For basins without a sufficiently persistent snow-and-ice signal, the emulator predicts only mixing-layer temperature. The outputs also include training and validation loss histories, the selected model configuration, identifiers of the lakes used in training, and SHAP-based feature-importance information at the global, lake, and seasonal-regime levels. The trained machine-learning model weights are provided separately within the dataset archive. Together, these files support model evaluation and subsequent coupling with the Xanthos-Lake water-balance framework. XanthosLAKES: Contains the NetCDF input data used by the Xanthos-Lake framework. Monthly meteorological inputs include relative and specific humidity, downward shortwave and longwave radiation, mean, maximum, and minimum air temperature, wind speed, precipitation, snowfall, and surface air pressure. Static lake-property datasets provide lake identifiers, geographic locations, surface area, volume, mean depth, elevation, drainage area, fetch, outlet-routing information, and associated Xanthos grid-cell attributes. Separate bathymetric datasets provide the coefficients of the area–depth and volume–depth relationships for each aggregated lake unit. GLEV-based records provide observed lake surface area and evaporation data used to initialize lake states, define reference conditions, and calibrate and evaluate the model. Xanthos-Lake Outputs: Contains the basin- and lake-category-specific NetCDF outputs generated by Xanthos-Lake. Monthly variables include lake surface area, storage volume, outlet discharge, evaporation rate, evaporation volume, lake–groundwater exchange, lake inflow, ice thickness, snow depth, snow-cover fraction, ice-cover fraction, and mixing-layer temperature. The files also contain lake-specific calibration and validation statistics, including normalized root-mean-square error, mean absolute error, Nash–Sutcliffe efficiency, Kling–Gupta efficiency, and percent bias. Stored calibrated and derived parameters include the weir discharge coefficient, fractional freeboard, groundwater exchange coefficient, reference water level, corresponding reference surface area and storage volume, weir-width adjustment factor, and the fraction of routed inflow entering the lake. Basin identifiers, lake category, simulation period, calibration and validation periods, and parameter-schema information are retained as NetCDF metadata.

Abeshu, Guta [Pacific Northwest National Laborator↗