Search NASASearch

SEARCH · Search NASA

Results for “Dataset”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Meta-analysis of North American Arctic and boreal aboveground biomass datasets: assessing accuracy, dynamics, and similarities

The North American arctic and boreal regions (ABRs) are rapidly warming and experiencing intensifying disturbances. Accurately quantifying aboveground biomass (AGB) is critical for understanding the impacts of these changes on the carbon cycle and for designing climate change mitigation strategies. Several AGB maps have been developed for the North American ABRs, including recent contributions from National Aeronautics and Space Administration’s Arctic-Boreal Vulnerability Experiment (ABoVE) campaign. However, these maps differ widely in training data, methodology, and resulting AGB density estimates. Presently, a comprehensive comparative evaluation is lacking, making it difficult for users to select datasets suited to their research or management needs. Here, in this study, we conducted a comparative analysis of nine AGB density datasets across North American ABRs, specifically for Alaska and Canada. We (1) summarized AGB by ecoregion and Canadian provinces, (2) evaluated their accuracy against field-based measurements, (3) analyzed spatial and temporal similarities among datasets, and (4) assessed their ability to capture disturbance (fire and harvest) impacts on AGB. We found substantial variation in regional and local AGB estimates across datasets, with overall accuracy ranging from R 2 = 0.25–0.62 and Bias% from −47.8% to 69.9% when validated against field plots. Despite these differences, most datasets have comparatively consistent spatial patterns in AGB (r > 0.8 for most cases). In contrast, agreement on the temporal patterns of AGB change is generally low. We found datasets with spatial resolutions ⩽300 m are capable of capturing disturbance impacts on AGB dynamics, though sensitivity varies across products. Our findings and dataset summary provide guidance for selecting appropriate AGB datasets for different applications within our study area. Our analysis also highlights the need to decrease map bias and increase capability to detect temporal change to decrease uncertainty of AGB datasets potentially by using training data which is representative of major plant functional types within the mapped area.

ABoVE

Observational ozone datasets over the global oceans and polar regions (version 2024)

Studying tropospheric ozone over the remote areas of the planet, such as the open oceans and the polar regions, is crucial to understand the role of ozone as a global climate forcer and regulator of atmospheric oxidative capacity. A focus on the pristine oceanic and polar regions complements the available land-based datasets and provides insights into key photochemical and depositional loss processes that control the concentrations and spatiotemporal variability in ozone as well as the physicochemical mechanisms driving these patterns. However, an assessment of the role of ozone over the oceanic and polar regions has been hampered by a lack of comprehensive observational datasets. Here, we present the first comprehensive collection of ozone data over the oceans and the polar regions. The overall dataset consists of 77 ship cruises/buoy-based observations and 48 aircraft-based campaigns. The dataset, consisting of more than 630 000 independent ozone measurement data points covering the period from 1977 to 2022 and an altitude range from the surface to 5000 m (with a focus on the lowest 2000 m), allows systematic analyses of the spatiotemporal distribution and long-term trends over the 11 defined ocean/polar regions. The datasets from ships, buoys, and aircraft are complemented by ozonesonde data from 29 launch sites or field campaigns and by 21 non-polar and 17 polar ground-based station datasets. The datasets contain information on how long the observed air masses were isolated from land, as estimated by backward trajectories from the individual observation points. To extract observations representative of oceanic conditions, we recommend using a subset of the data with an isolation time of 72 h or longer, from the analysis with coincident radon observations. These filtered oceanic and polar data showed typically flat diurnal cycles at high latitudes, whereas daytime decreases in ozone (11 %–16 %) were observed at lower latitudes. The ship/buoy- and aircraft-based datasets presented here will supplement the land-based ones in the TOAR-II (Tropospheric Ozone Assessment Report Phase II) database to provide a fully global assessment of tropospheric ozone. The described dataset is available at https://doi.org/10.17596/0004044 (Kanaya et al., 2025).

Kanaya, Yugo [Japan Agency for Marine-Earth Scienc

Performance of wind assessment datasets in United States coastal areas

The atmospheric dynamics that occur near the intersection of land and water offer exciting and challenging opportunities for wind energy deployment in coastal locations. New models and tools are continually being developed in support of wind resource assessment, and three recent products are explored in this work for their performance in representing characteristics of the wind resource at coastal locations: the Global Wind Atlas 3 (GWA3), the 2023 National Offshore Wind dataset (NOW-23), and the wind climate simulations that are a component of the Wind Integration National Dataset (WIND) Toolkit Long-Term Ensemble Dataset (WTK-LED Climate). These relatively new products are freely available and user-friendly so that anyone – from a utility-scale developer to a resident or business owner – can evaluate the potential for wind energy generation at their location of interest. The validations in this work provide guidance on the accuracy of wind resource assessments for coastal customers interested in installing small or midsize wind turbines (≤ 1 MW in capacity) to support energy needs at the residential, business, or community scale, such as the island and remotely located participants of the U.S. Department of Energy's Energy Transitions Initiative Partnership Project. At 23 coastal locations across the United States, dataset performance varies according to different evaluation metrics. All three recent datasets tend to overestimate the observed coastal wind resource. GWA3 produces the smallest annual average wind speed relative errors, whereas WTK-LED Climate is in best agreement in terms of representing diurnal wind speed cycles. NOW-23 is the highest performing of the datasets for representing seasonal and interannual trends in the coastal wind resource. While GWA3 and WTK-LED Climate are relatively insensitive to the dataset output heights selected for wind resource assessment at small and midsize wind turbine hub heights (20–60 m), significant variation in the NOW-23 representation of wind shear across the wind profile in the lowest 100 m of the atmosphere leads to notable differences in wind speed estimates according to the dataset output heights selected for evaluation. GWA3 exhibits challenges in the representation of observed wind speed diurnal cycles at small and midsize turbine hub heights, likely due to the dataset's consistent treatment of hourly wind speed trends regardless of altitude.

17 WIND ENERGY

Dataset of U.S. School Bus Depots

A large body of public health literature describes how undesirable or dangerous facilities, such as truck depots and industrial plants, located in or near communities can lead to health harms. Research also describes the high levels of traffic-related air and noise pollution that is linked to health harms and may be disproportionately distributed near many schools. Therefore, a primary use case for this dataset is to analyze the location of school bus depots and to create an evidence base that would better enable the work of community members, advocates, and other stakeholders toward improving air quality and public health. Other possible uses for this school bus depot dataset include electricity grid planning and reliability, given recent momentum toward school bus electrification. This dataset was created using an object-based approach with remote sensing data. The primary source of aerial imagery was the National Agriculture Imagery Program (NAIP) dataset. NAIP imagery was analyzed to locate individual school buses based on their color and size, and then classified clusters of school buses as potential depots, which were then verified visually. The resulting dataset contains 11,309 depots across the 48 contiguous U.S. states and Washington, D.C. Fifty-one percent (5,730 depots) are at schools, defined as being 350 meters or less from the nearest school. The accuracy of the dataset was assessed by comparing it with independent reference datasets containing 506 depots from the records of two school transportation companies. We found good agreement, with an omission error rate of 15.2% (77 depots). This dataset represents one of the only remote sensing projects to conduct object detection using data at the sub-meter to 1-meter resolution for a continental-scale application.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Topsoil bulk geochemical compositions - An updated harmonized global dataset

Mineral weathering is a key biogeochemical process because of the capacity of minerals to stabilize organic matter. However, predicting soil weathering status across large spatial areas still isn’t possible due to a lack of global data and theoretical frameworks. To address this knowledge gap, multiple global datasets of bulk topsoil geochemical compositions have been harmonized using R. These datasets document topsoil bulk geochemical compositions across five continents (n = ~16,000 observations). Source data for these observations include the EuroGEOSurveys Geochemical Baseline Database (FOREGS), the US Geological Survey National Geochemical Database (NASGLP), the Geochemical Atlas of Australia (GAA), the US Geological Survey Alaska Geochemical Database (AGD84), the National Cooperative Soil Survey (NCSS), the European Geochemical Mapping of Agricultural Soil (GEMAS), Ecorespira-Amazon (ERA), the New Zealand Geochemical Baseline Survey (NZ_GBS), and the African Soil Information Service (AFSIS). Major elements observed include Aluminum (Al), Calcium (Ca), Iron (Fe), Potassium (K), Magnesium (Mg), Sodium (Na), Titanium (Ti), Manganese (Mn), Phosphorus (P), Carbon (C), and Sulfur (S). This data package includes the harmonized dataset itself, and the R scripts necessary to harmonize these datasets, in addition to metadata that describes all columns, files, and databases used in this project. Methods & Sampling Step 1 – Databases of geochemical data identified This study aimed to leverage existing measurements of topsoil geochemical data. Databases were first identified and deemed appropriate for inclusion if they were measuring soils and performed these measurements on the <2mm soil fraction. Databases such as NCSS and AGD84 needed more post processing to include in the database and this was done using the NCSS_datamerge_031626 R file and Alaska_USGSmerge_031626 R file, respectively. Step 2 – Database harmonization Once appropriate databases were identified, they were harmonized for ease of analysis using the R script Database_Harmonization_031826. This included removing columns from original datasets that would not be used in analysis (removed columns are noted in the code). Then, data cleaning procedures specific to each dataset were undertaken. This includes standardizing columns to include units and adding metadata columns regarding procedures for analyzing specific elements. Functions for standardizing measurements and units are outline in R files: calculate element_mg_kg_031626, calculate_oxide_wt_perc_031626, change_oxide_caps_031626, and conv_2_numeric_031626. This also included adding a unique identifier for each sample to identify it with its respective database (see CD_ID in data dictionary). Geographic information: Data reflect a compilation of datasets collected globally. Geographic areas covered by each of the datasets include: - EuroGEOSurveys Geochemical Baseline Database (FOREGS) - European continent - North American Soil Geochemical Landscapes (NASGLP) - continental United States and limited parts of Canada (see database key for more details) - National Geochemical Survey of Australia (GAA) - Australia - Alaska geochemical database (AGDB4) - Alaska - National Cooperative Soil Survey (NCSS) - Global measurements, but concentrated in the continental United States - Geochemical data for arable land and land under permanent grass cover in continental Europe (GEMAS) - continental Europe - Ecorespira-Amazon (ERA) - Geochemical data from the Amazon basin - Geochemical baseline data for New Zealand (NZGBS) - New Zealand - Geochemical data collected across continental Africa (AfSIS) - Measurements across Africa

EARTH SCIENCE > LAND SURFACE > SOILS

Dataset of Generative AI Workload Power Profiles

This dataset provides a collection of high-resolution (5/10 Hz or every 0.2/0.1 seconds) power consumption profiles for generative artificial intelligence (GenAI) workloads executed on NLR's High Performance Computing (HPC) platform Kestrel. The dataset also includes examples of representative whole-facility power profiles generated using a bottom-up, event-driven, data center energy model . This dataset is designed to support research in energy modeling, infrastructure planning, energy system integration, and sustainability analysis for AI-driven computing systems. The dataset captures time-resolved electrical power measurements across a diverse set of configurations, including variations in job type (inference vs. training), workload (LLM vs. image generation), datasets, and number of compute nodes. Power traces are provided in a standardized format and include both raw/instantaneous and aggregated files. Each profile is accompanied by metadata describing workload parameters, enabling reproducibility and cross-study comparison. The dataset is intended for use in applications such as data center infrastructure planning, energy modeling, demand response and grid impact studies, and development and validation of system-level simulation tools. By making these workload-specific power profiles publicly available, this dataset aims to address the current lack of open, empirical energy data for generative AI systems and to facilitate transparent, reproducible research on the energy and environmental impacts of large-scale AI deployment. If you use this dataset, please cite the associated publication: Vercellino et al., “Measurement of Generative AI Workload Power Profiles for Whole-Facility Data Center Infrastructure Planning,” arXiv:2604.07345 (2026).

97 MATHEMATICS AND COMPUTING

Citation network datasets for benchmarking spiking graph neural networks on experimental neuromorphic hardware

Spiking neural networks (SNNs) running on neuromorphic computers offer an energy-efficient alternative for AI tasks. Recently, spiking graph neural networks (S-GNNs) have been shown to produce encouraging results on benchmark citation network datasets such as Cora, CiteSeer, and PubMed for node classification tasks. These S-GNNs were run on SNN simulators only because they contain up to tens of thousands of neurons and up to millions of synapses, translating poorly to neuromorphic hardware. Therefore, in this paper, we create a suite of benchmark datasets from the CiteSeer dataset that can be accommodated on current neuromorphic hardware platforms. Our contribution consists of a collection of three datasets. First, we have an induced subgraph of CiteSeer, which we call MiniSeer, containing 2110 papers, 3604 binary features, and 6 topics. Second, MicroSeer is a very small dataset consisting of 84 papers, 1227 features, and 6 topics. Lastly, BiteSeer is a collection of 15 binary classification datasets. We present creation of these datasets along with accuracies, running times, and spike counts when simulated. We believe that our results in this paper will be used by the neuromorphic community to benchmark, test, and develop neuromorphic hardware and simulators.

Zhu, Kevin [George Mason University, Virginia]

Open Power System Datasets and Open Simulation Engines: A Survey Toward Machine Learning Applications

A major factor behind the success of machine learning (ML) models in multiple domains is the availability and accessibility of large, labeled, and well-organized datasets for training and benchmarking. In comparison, power grid datasets face three major challenges: (i) real-world data is often restricted by regulatory constraints, privacy reasons, or security concerns, making it difficult to obtain and work with; (ii) synthetic datasets, which are created to address these limitations, often have incomplete information and are released using specialized tools, making them inaccessible to the broader community; and, (iii) input-output datasets are difficult to generate through simulation for non-experts because open-source simulators are not known outside the power system community. This survey addresses these challenges by serving as an entry point to publicly available datasets and simulators for researchers venturing in this area. We review the current landscape of open-source power network data, machine models, consumer demand profiles, renewable generation data, and inverter models. We also examine open-source power system simulators, which are crucial for generating high-quality, high-fidelity power grid datasets. We aim to provide a foundation for overcoming data scarcity and advance towards a structured web of datasets and simulators to support the development of ML for power systems.

42 ENGINEERING

A Co-Registered In-Situ and Ex-Situ Dataset of Electrical, Acoustic, and CT Characteristics from Wire Arc Additive Manufacturing Process

Recent progress in sensing techniques and data analytics tools have significantly accelerated the development of Wire Arc Additive Manufacturing (WAAM) systems. This data centric approach emphasizes leveraging available data throughout the production process to optimize performance. Integration of extensive data analysis provides the opportunity to improve precision, reduce waste, and enhance the quality of produced parts. This method relies on AI/ML models and optimization techniques, which are developed using the data collected from various sources, including in-situ sensors, ex-situ imaging, and manufacturing process parameters. The quality and diversity of this data, along with the alignment between different data streams (achieved through spatiotemporal registration) are critical for the successful development of AI/ML and optimization models. In this work, we present a spatiotemporally registered dataset generated during the WAAM process of deposition of a rectangular block. The dataset includes the comprehensive description of deposition process, process parameters, in-situ collected welding characteristics, acoustic data, and X-Ray Computed Tomography analysis data for the build. Dataset A Co-Registered In-Situ and Ex-Situ Dataset of Electrical, Acoustic, and CT Characteristics from Wire Arc Additive Manufacturing Process has arisen under UT-Battelle, LLC’s Prime Contract No. DE-AC05-00OR22725 with the U.S. Department of Energy (DOE) to manage and operate the Oak Ridge National Laboratory. UT-Battelle, LLC will not assert any rights under United States law or under the Prime Contract it has in the dataset against any user of the dataset, including any copyrights or patent rights. UT-Battelle, LLC requests that attribution to the dataset is provided as academically appropriate.

42 ENGINEERING

WTK-LED: The WIND Toolkit Long-Term Ensemble Dataset

To satisfy a wide group of stakeholders across various wind energy disciplines, including but not limited to stakeholders in the distributed and utility scale wind industry, the new emerging airborne wind energy field, grid integration, power systems modeling, environmental modeling, and researchers in academia, and to close some of the gaps that current public datasets have, we aimed at developing an updated version of the meteorological WIND Toolkit, named WIND Toolkit Long-term Ensemble Dataset (WTK-LED), which is a meteorological dataset providing time series every 5 min and 2 km, including model uncertainty of wind speed at every modeling grid point so that users are provided with a range of possible wind speeds every 2 km. The data were produced using the Weather Research and Forecasting Model (WRF). The vertical grid used in WTK-LED includes many vertical layers in the atmospheric boundary layer to provide information of atmospheric quantities across the rotor layer of utility scale and distributed wind turbines. The WTK-LED includes: 1) Numerical simulations covering the continental United States, Alaska, and Hawaii, with high-resolution data being available for 3 years (2018-2020). 2) Climate simulations from Argonne National Laboratories covering the North American continent, including Alaska, Canada, and most of Mexico and the Caribbean Islands. These simulations complement the new WTK-LED to offer a 4-km dataset covering 20 years, from 2001-2020. 3) Specific long-term,high-resolution offshore simulations have been conducted separately for the US coasts, Hawaii, and the Great Lakes, leading to the 2023 National Offshore Wind data set. This report focuses on a description of the land-based WTK-LED for CONUS, Hawaii, and Alaska, for the 3-year 2-km/5-min dataset and the 20-year 4-km/hourly dataset, as well as the uncertainty quantification method. We also provide limited validation results. Based on our results to date, we suggest use cases and applications for each dataset of the WTK-LED.

17 WIND ENERGY

Pavement condition and climatic data in southeast Texas: A dataset for evaluating flood impacts on pavement performance

Effective pavement maintenance is essential for economic stability, optimal network performance, and roadway safety. Achieving this requires thorough evaluation of pavement conditions, including structural integrity, surface roughness, and distress characteristics. Pavement performance indicators play a critical role in influencing vehicle safety and ride quality. Recent advances have emphasized the use of data-driven modeling to anticipate pavement behavior, with the goal of optimizing resource allocation and refining Maintenance and Rehabilitation (M&R) strategies through accurate condition assessment. A foundational requirement for these modeling efforts is the availability of standardized, high-quality datasets that can support robust and reproducible infrastructure analysis. This data article presents a comprehensive dataset assembled to facilitate pavement performance prediction, with a geographic focus on Southeast Texas, particularly the flood-vulnerable area of Beaumont. The dataset encompasses pavement and traffic attributes, meteorological records, flood simulation outputs, ground deformation measurements, and topographic indices, enabling detailed examination of both load-associated and non-load-associated degradation mechanisms. Data preprocessing was performed using ArcGIS Pro, Microsoft Excel, and Python to ensure consistency and usability in data-driven modeling applications, including machine learning workflows. Key contributions of this dataset include its utility in analyzing the climatic and environmental factors affecting pavement conditions, identifying critical predictive features, and enabling in-depth correlation analysis across diverse variables. By filling existing gaps in input variable selection resources, this dataset supports the development of predictive tools for estimating future maintenance demand and enhancing the resilience of pavement networks in flood-impacted areas. The resource highlights the importance of standardized datasets for advancing pavement management practices and provides a robust foundation for ongoing infrastructure performance modeling.

42 ENGINEERING

Search for ultralight dark matter in the SuperMAG high-fidelity dataset

Ultralight dark matter, such as kinetically mixed dark-photon dark matter (DPDM) or axion-like-particle dark matter (axion DM), can source an oscillating magnetic-field signal at Earth’s surface. Previous work searched for this signal in a publicly available dataset of global magnetometer measurements maintained by the SuperMAG collaboration. This “low-fidelity” dataset reported measurements with a 1-min time resolution, allowing the search to set leading direct constraints on DPDM and axion DM with Compton frequencies f DM ≤ 1 / ( 1 min ) (corresponding to masses m DM ≤ 7 × 10 − 17 eV ). More recently, a dedicated experiment undertaken by the SNIPE Hunt collaboration has also searched for this same signal at higher frequencies f DM ≥ 0.5 Hz (or m DM ≥ 2 × 10 − 15 eV ). In this work, we search for this signal of ultralight DM in the SuperMAG “high-fidelity” dataset, which features a 1-sec time resolution, allowing us to probe the gap in parameter space between the low-fidelity dataset and the SNIPE Hunt experiment. The high-fidelity dataset exhibits lower geomagnetic noise than the low-fidelity dataset and features more data than the SNIPE Hunt experiment, making it a powerful probe of ultralight DM. Our search finds no robust DPDM or axion DM candidates. We set constraints on DPDM and axion DM parameter space for 10 − 3 Hz ≤ f DM ≤ 0.98 Hz (or 4 × 10 − 18 eV ≤ m DM ≤ 4 × 10 − 15 eV ). Our results are the leading direct constraints on both DPDM and axion DM in this mass range, and our DPDM constraint surpasses the leading astrophysical constraint in a narrow range around m A ′ ≈ 2 × 10 − 15 eV . Published by the American Physical Society 2024

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

Data Fusion for the Development of a Multimodal Freight Transload Facilities Dataset in the U.S.

To withstand the growing demand of commodity volume and its strain on the transportation infrastructure, it is necessary to identify the flow of commodities by route and mode. However, a national multimodal freight routing model does not exist for the U.S. The development of such model requires multiple building blocks, such as virtual representations of roadway, railway, and waterway networks, transload facilities (TFs), and access/egress links. Most of these blocks have a robust database in the U.S., except for the TFs. Here, this paper presents the fusion of dispersed and heterogeneous representations of multimodal TFs into a single, comprehensive, geospatial freight TF dataset. The TF dataset is derived from several sources, including the U.S. Army Corps of Engineers Master Docks Plus, the National Transportation Atlas Database, the Intermodal Association of North America, industry publications, and other public information. First, individual datasets were queried and reconciled. A geocoding/reverse geocoding process was applied to get the best street address and latitude/longitude location for each terminal. Then, duplicate terminals were identified by a fuzzy match algorithm based on terminal name and location, and removed. Validation was performed by visual inspection of random facilities. The main contributions of this work are: a publicly available version of the TF dataset, including facility location and multimodal transfer capability of 9,003 facilities, and an enterprise-version with the same facilities but including commodity handling capabilities. The main purpose of developing the TF dataset is to inform multimodal routing algorithms. The proposed TF dataset allows for credibly modeling the multimodal transfer of commodities within shipment routes.

Commodity Routing

U.S. Freight Transload Facilities Dataset

The U.S. Freight Transload Facilities Dataset provides location information (latitude, longitude, zip, city, county, state)for more than 9,000 facilities across 50 U.S. States where freight may be transferred between waterways, railways, and roadways. The dataset lists the known modes and available direction(s) for freight transfers at each facility as of 2024. The U.S. Freight Transload Facilities dataset was built by mining and fusing several public sources, such as the USACE Master Docks Plus, the USDOT National Transportation Atlas Database (NTAD), files from the Intermodal Association of North America (IANA), and the industry publication Bulk Transloader. The dataset constitutes a key piece of a multimodal freight transportation network and routing algorithm developed by USACE-ERDC. The dataset is shared as a .csv file. The dataset is published for research purposes and should not be considered exhaustive or authoritative.

Peterson, Steven [ORNL] (ORCID:0000000287672998)

Development of a 95-Year Solar Dataset for Resource Adequacy Studies

Long-term high-resolution solar data provides enhanced understanding of variability of solar generation and enhances our ability to develop strategies for a resilient and reliable electric grid under high deployment of solar energy. Therefore, it is important to develop long-term synthetic datasets that can provide multiple occurrences of various severe weather scenarios that are expected to test the limits of resource adequacy under scenarios contain various energy generation sources. Examples of such scenarios could be long periods of high temperatures when demand for electricity is high or periods where high winds could lead to a shut-down of transmission lines for long periods of time to ensure fire safety. NREL has developed the first version of such a dataset covering a 95-year period covering 2006-2100 at a 4km hourly resolution. This dataset contains all variables necessary to calculate solar generation. During development of this dataset, we focused on creating unbiased, high-resolution solar irradiance through statistical downscaling methods, using Regional Climate Model (RCM) simulations from the North American Coordinated Regional Climate Downscaling Experiment (NA-CORDEX) as input. The National Solar Radiation Database (NSRDB) containing over 25 years of observations was used to calibrate the statistical downscaling models. This presentation will outline the primary steps in developing this dataset, including (1) regridding RCM data to a common grid at 20-km resolution, (2) correcting RCM biases with NSRDB, (3) applying temporal and spatial downscaling methods to generate high-resolution (4-km, hourly) solar and ancillary data. Additionally, we will present an evaluation of the downscaled data against the NSRDB across various zones in the CONUS. Lastly, we will present a user guide for accessing the datasets.

14 SOLAR ENERGY

Evaluation of normalization strategies for mass spectrometry-based multi-omics datasets

Introduction Data normalization is crucial for multi-omics integration, reducing systematic errors and maximizing the likelihood of discovering true biological variation. Most studies assess normalization for a single omics type or use datasets from separate experiments. Few address time-course data, where normalization might bias temporal differentiation. In this study, we compared common normalization methods and a machine learning approach, Systematical Error Removal using Random Forest (SERRF), using multi-omics datasets generated from the same experiment—even from the same cell lysate. Objectives To develop a straightforward process to assess normalization effects and identify the most robust methods across multi-omics datasets. Methods We analyzed metabolomics, lipidomics, and proteomics datasets from primary human cardiomyocytes and motor neurons exposed to acetylcholine-active compounds over time. Normalization effectiveness was evaluated based on improvement in QC features consistency and observing the change in treatment and time-related variance. Results Probabilistic Quotient Normalization (PQN) and Locally Estimated Scatterplot Smoothing (LOESS) QC were identified as optimal for metabolomics and lipidomics, while PQN, Median, and LOESS normalization excelled for proteomics. These methods consistently enhanced QC feature consistency in metabolomics and lipidomics, and preserved time-related variance or treatment-related variance in proteomics, demonstrating their effectiveness and robustness. SERRF normalization, applied only to metabolomics in this study, outperformed other methods in some datasets but inadvertently masked treatment-related variance in others. Conclusion Our evaluation identified PQN and LoessQC as the top methods for metabolomics and lipidomics, and PQN, Median, and Loess normalization for proteomics, in multi-omics integration in a temporal study.

60 APPLIED LIFE SCIENCES

Dataset about Warming Effects on Carbon Cycling and Greenhouse Gas Fluxes in Permafrost Ecosystems

Field observations provide direct evidence of how does carbon cycling in permafrost ecosystems respond to climate change. This study provides a comprehensive dataset on the impact of warming on carbon cycling and greenhouse gas (GHG) fluxes in permafrost ecosystems. The dataset is extracted and integrated from 132 peer-reviewed studies with 1430 paired observations across eight major permafrost ecosystems, including Arctic and subarctic tundra and wetland, and alpine meadow, steppe, tundra and wetland. This dataset includes 17 variables from experiments conducted during the growing season, covering the plant and soil carbon pools, soil nitrogen pool, and GHG (i.e., CO 2 , CH 4 , and N 2 O) fluxes, among others. Background information on site climate conditions, vegetation and soil characteristics, and details of the warming experiments, including timing, methods, and warming magnitude, are also contained in the dataset. This dataset facilitates a comprehensive understanding of the impact of warming on carbon cycling and GHG fluxes in permafrost ecosystems, and provides supports for meta-analyses and literature reviews, remote sensing data validation, and land model development and parameterization.

Bao, Tao [Chinese Academy of Sciences (CAS), Beiji

Using multiple high-resolution datasets to benchmark the energy exascale earth system model (E3SM) for renewable resource assessment

The United States is accelerating its shift toward a renewable energy system. However, renewable resources, which harness energy from the Earth system, are susceptible to both present-day climate variability and future climate change. For example, variations in regional climate can alter renewable energy production patterns and site viability. The use of high-resolution climate model projections can therefore facilitate and may be critical to long-term planning of renewable energy investments. However, climate models must first be validated for renewable resource assessment. This research employs multiple high-spatiotemporal-resolution datasets to assess the capability of the Department of Energy’s (DOE) Energy Exascale Earth System Model version 2 North American Regionally Refined Model (E3SMv2-NARRM) for predicting multi-year climatological values of solar and wind energy capacity factors in the continental U.S., with a focus on regional and seasonal variability. Present-day E3SMv2-NARRM simulations are compared with reported utility-scale production data obtained from the Energy Information Administration (EIA). In addition, E3SMv2-NARRM data are evaluated against non-climate benchmark models from the National Renewable Energy Laboratory, including the Wind Integration National Dataset Toolkit and the National Solar Radiation Database (NSRDB), as well as three wind energy datasets from PLUSWIND. Our analysis indicates that solar capacity factors from E3SM closely match those from the NSRDB dataset. However, both datasets tend to overestimate values by 10% in comparison to EIA data. Furthermore, biases in wind capacity factors within E3SM are notably pronounced in the West Coast regions, where the seasonal cycle diverges from EIA data.

Energy forecasting, Capacity factor, Renewable ene