Search NASA⌕ Search

SEARCH · Search NASA

Results for “raw data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Public Reference Data for Megawatt-Scale Hydrogen Electrolysis - NLR Historical Wind

The U.S. Department of Energy and the National Laboratory of the Rockies (NLR) demonstrate hydrogen electrolysis from variable sources, hydrogen compression and storage, and hydrogen fuel cell power production using megawatt-scale equipment at NLR’s Flatirons Campus as part of the Advanced Research on Integrated Energy Systems (ARIES) initiative. This dataset represents part of that effort and is intended for academic, national laboratory, industrial, and other stakeholders to plan, design, and validate models of megawatt-scale hydrogen technologies and diverse energy infrastructure nationwide. These data provide a baseline for how existing hydrogen electrolysis technologies perform when coupled with various energy technologies. Future datasets will demonstrate how existing hydrogen fuel cell technologies can provide controllable, dispatchable, and variable power output for artificial intelligence (AI) data centers and other variable loads. This dataset entry describes hydrogen production by conducting a statistical analysis of historical wind data over a five-year period (2020-2025) from a single 1.5MW turbine manufactured by General Electric (GE) located at NLR’s Flatirons Campus, to generate an experimental test profile that was deployed on a 1.25-MW proton exchange membrane type MC250 electrolyzer system manufactured by Nel Hydrogen . [1] While the electrolyzer balance-of-plant supports up to 2.5 MW of electrolysis, NLR only has a single 1.25-MW electrolysis stack. The historical wind data provided several metrics, however, the analysis particularly focused on the measured power output by the wind turbine. The power output time series of data for each day was categorized by total energy generation and standard deviation, and the day that represented the highest combination of these two metrics was chosen – December 25th, 2022. This process was then repeated for a moving four-hour window within this day to identify the most statistically variable period. Finally, this four-hour period was scaled by 65% to match the 1.25 MW electrolyzer. The electrolysis system controls hydrogen production by varying DC current applied to the stack, from a maximum of 3000 A to a minimum safe operation of 300 A, or 10%. Because the current – voltage characteristic changes as the stack ages and efficiency degrades, the actual minimum safe operating power changes over time. The historical wind profiles were translated from power (kilowatts) to current (amperes) using a curve fit with calibration data and sent to the electrolyzer power supply at 1 Hz frequency. For more details on the statistical analysis process, see the presentation labeled “ Public Reference Data for Megawatt-Scale Hydrogen Electrolysis” provided with each data entry. These datasets report relevant hydrogen balance-of-plant and system data, all captured at 1 Hz, including hydrogen mass production measured with an Emerson Coriolis flow meter. Each .zip file represents a single wind turbine electrolysis experiment and is formatted as follows: {technology}_{scaling factor}-{electrolyzer ramp rate in amperes/second} For instance, “wind-GE1.5MW_0.65-400.zip” represents the hour-long experiment using historical data from the wind-GE1.5MW turbine, scaled to 65%, with the electrolyzer power supply set to a maximum ramp rate (gain and slew) of 400 A/s. Each .zip folder contains the following files: A .csv file containing raw data An .xlsx file explaining all the fields in the raw data. A .png plot showing the time series of hydrogen production, electrolysis power consumption, and wind power input. A PDF file detailing the historical wind data statistical analysis used to generate the wind profile. An experiment labeled “characterization_200.zip” demonstrates the MC250 electrolyzer steady-state response with 30-minute load steps for a total duration of 5 hours. Finally, a .csv file is provided with all simulated wind experiments combined into one dataset labeled "combined_historical_wind_experiments.csv". NLR also built an AI/machine-learning predictive model based on these datasets. The model ingests the electrolyzer current command in amperes, as well as various pressures and temperatures across the system, and predicts hydrogen output in kilograms per hour. The complete model can be found at https://huggingface.co/NatLabRockies/ptmelt-hydrogen-electrolysis [1] nelhydrogen.com/product/mc-series-electrolyser .

08 HYDROGEN↗

Public Reference Data for Megawatt-Scale Hydrogen Electrolysis - NLR Historical Solar PV

The U.S. Department of Energy and National Laboratory of the Rockies (NLR) demonstrate hydrogen electrolysis from variable sources, hydrogen compression and storage, and hydrogen fuel cell power production using megawatt-scale equipment at NLR’s Flatirons Campus as part of the Advanced Research on Integrated Energy Systems (ARIES) research platform. This dataset represents part of that effort and is intended for academic, national laboratory, industrial, and other stakeholders to plan, design, and validate models of megawatt-scale hydrogen technologies and diverse energy infrastructure nationwide. These data provide a baseline for how existing hydrogen electrolysis technologies perform when coupled with various energy technologies. Future datasets will demonstrate how existing hydrogen fuel cell technologies can provide controllable, dispatchable, and variable power output for artificial intelligence data centers and other variable loads. This dataset entry describes the behavior of a 1.25-MW proton exchange membrane MC250 electrolyzer system, manufactured by Nel Hydrogen , [1] when fed historical data generated by the 430-kW, fixed-axis solar photovoltaic (PV) array located at NLR’s Flatirons Campus. (While the electrolyzer balance of plant supports up to 2.5 MW of electrolysis, NLR only has a single 1.25-MW electrolysis stack.) Solar PV power output data for the 2020 calendar year were categorized on a daily basis by total energy generation and standard deviation. Each day was then ranked by these metrics, and the 25th, 50th, and 100th percentiles were selected. The 75th percentile day did not exhibit sufficient variability to make for a valuable experiment. A similar process was used for the related historical wind dataset . [2] The historical days in 2020 that represented these percentiles are Dec. 19, March 29, and May 4, respectively. The entire solar day’s power profile was then fed through the MC250 electrolyzer. Due to its length, the 100th percentile day experiment was split into two parts, and the final 3 hours of the solar day were not captured. These final 3 hours contained no spikes or dips of interest and simply represented a slow decay of input solar power. Also, a single timestamp (13:13:47 on Jan. 14, 2026) was lost in the hydrogen system supervisory control and data acquisition. Finally, during the 25th percentile experiment (solar day Dec. 19, 2020) data recording was lost from 11:00:13 to 11:14:45. The roughly 15 minutes of the solar profile were rerun at the end of the experiment and spliced into this time slot during post-processing. The electrolysis system controls hydrogen production by varying direct current applied to the stack, from a maximum of 3,000 A to a minimum safe operation of 300 A, or 10%. Because the current–voltage characteristic changes as the stack ages and efficiency degrades, the actual minimum safe operating power changes over time. The historical solar profiles were translated from power (kilowatts) to current (amperes) using a curve fit with calibration data and sent to the electrolyzer power supply at 1-Hz frequency. For more details on the statistical analysis process, see the slide deck “Public Reference Data for Megawatt-Scale Hydrogen Electrolysis: NLR Historical Solar PV Analysis and Profile Generation” accessible with this data entry. These datasets report relevant hydrogen balance-of-plant and system data, all captured at 1 Hz, including hydrogen mass production measured with an Emerson Coriolis flow meter. Each .zip file represents a single solar PV electrolysis experiment and is formatted as: {technology}_{percentile}_{scaling factor} For instance, “solarPV-430kW_25_2x.zip” reports the experiment using the 25th percentile solar data from the historical 2020 solar PV dataset, scaled to 200%. Scaling factors were applied to the generated solar PV power output files to more closely match the 1.25-MW capacity of the electrolyzer. Each .zip folder contains the following files: A .csv file containing raw data. An .xlsx file explaining all the fields in the raw data. A .png plot showing the time series of hydrogen production, electrolysis power consumption, and solar power input. A PDF file detailing the historical solar data statistical analysis used to generate the solar profile. An experiment labeled “characterization_200.zip” demonstrates the MC250 electrolyzer steady-state response with 30-minute load steps for a total duration of 5 hours. Finally, a .csv file is provided with all experiments combined into one dataset labeled "combined_solarPV_experiments.csv". [1] nelhydrogen.com/product/mc-series-electrolyser . [2] data.nlr.gov/submissions/316 .

08 HYDROGEN↗

WFIP3 / WHOI ASIT Lidar / Standardized Data

This dataset contains standardized raw data from the WHOI ASIT deployed for WFIP-3. A ZX 300M is currently installed at the tower and has been deployed there since September 2021.

17 WIND ENERGY↗

Radar / Processed Data

This dataset contains processed raw data from the UND 94 GHz W-band Cloud Doppler Radar.

17 WIND ENERGY↗

Public Reference Data for Megawatt-Scale Hydrogen Electrolysis - Simulated Wave

The U.S. Department of Energy and the National Laboratory of the Rockies (NLR) demonstrate hydrogen electrolysis, hydrogen compression and storage, and variable hydrogen fuel cell power production using megawatt-scale equipment at NLR’s Flatirons Campus as part of the Advanced Research on Integrated Energy Systems (ARIES) initiative. This dataset represents part of that effort and is intended for academic, national laboratory, industrial, and other stakeholders to plan, design, and validate models of megawatt-scale hydrogen technologies and diverse energy infrastructure nationwide. These data provide a baseline for how existing hydrogen electrolysis technologies perform when coupled with various energy technologies. Future datasets will demonstrate how existing hydrogen fuel cell technologies can provide controllable, dispatchable, and variable power output for artificial intelligence (AI) data centers and other variable loads. This dataset entry describes hydrogen production using a single, simulated wave energy conversion device. The electrolyzer is a 1.25-MW proton exchange membrane type MC250 system manufactured by Nel Hydrogen. While the unit supports up to 2.5 MW of electrolysis, NLR only has a single 1.25-MW electrolysis stack. For the wave energy, NLR used a wave energy converter model from PacWave. These devices can be equipped with accumulators and pressure relief values to smooth the power output by storing and releasing hydraulic energy. Using a peak power output of 10 MW, the model created two 25-minute profiles: one with and one without the accumulators and pressure relief valves. To down select the profile data from the native resolution of 20 Hz to 1 Hz, NLR took the mean of every 20 data points. NLR experimented with two simulated wave energy power plants: one that peaks at 10 MW, and one that peaks at 5 MW. These profiles were scaled for the physical 1.25 MW electrolyzer by multiplying the original profiles by one eighth and one quarter, respectively. The first profile matches the capacity rating of eight of the 1.25 MW electrolyzers, while the second matches four electrolyzers. Finally, NLR experimented with two settings for the electrolyzer power supply minimum and maximum current ramp rates (gain and slew): 200 and 400 amperes per second. The simulated profiles were translated from power (kilowatts) to current (amperes) using a curve fit with calibration data and sent to the electrolyzer power supply at 1-Hz frequency. These datasets report relevant hydrogen balance-of-plant and system data, all captured at 1 Hz, including hydrogen mass production measured with an Emerson Coriolis flow meter. Each .zip file represents a single wave electrolysis experiment and is formatted as follows: {technology}-{accumulator?}_{number of 1.25 MW electrolyzers connected}-{electrolyzer ramp rate in amperes/second} For instance, “wavePacWave-Noacc_4-400.zip” represents the 25 minute-long experiment using the PacWave’s wave energy converter model, equipped with no accumulator, connected to four 1.25-MW electrolyzers with their power supplies set to a maximum current ramp rate (gain and slew) of 400 A/s. Each .zip folder contains the following files: A .csv file containing raw data. An .xlsx file explaining all the fields in the raw data. A .png plot showing the time series of hydrogen production in kilograms per hour, electrolysis power consumption, and input wave power. An experiment, labeled “characterization_200.zip”, demonstrates the MC250 electrolyzer steady-state response with 30 minute load steps for a total duration of 5 hours. Finally, a .csv file is provided with all wave profiles combined into one dataset labeled "combined_wave_experiments.csv". NLR also built an AI/machine-learning predictive model based on these datasets. The model ingests the electrolyzer current command in amperes, as well as various pressures and temperatures across the system, and predicts hydrogen output in kilograms per hour. The complete model can be found at https://huggingface.co/NatLabRockies/ptmelt-hydrogen-electrolysis.

08 HYDROGEN↗

Public Reference Data for Megawatt-Scale Hydrogen Electrolysis – Simulated Wind

The U.S. Department of Energy and the National Laboratory of the Rockies (NLR) demonstrate hydrogen electrolysis, hydrogen compression and storage, and variable hydrogen fuel cell power production using megawatt-scale equipment at NLR’s Flatirons Campus as part of the Advanced Research on Integrated Energy Systems (ARIES) initiative. This dataset represents part of that effort and is intended for academic, national laboratory, industrial, and other stakeholders to plan, design, and validate models of megawatt-scale hydrogen technologies and diverse energy infrastructure nationwide. These data provide a baseline for how existing hydrogen electrolysis technologies perform when coupled with various energy technologies. Future datasets will demonstrate how existing hydrogen fuel cell technologies can provide controllable, dispatchable, and variable power output for artificial intelligence (AI) data centers and other variable loads. This dataset entry describes hydrogen production using a single, simulated wind turbine. The electrolyzer is a 1.25-MW proton exchange membrane type MC250 system manufactured by Nel Hydrogen . While the unit supports up to 2.5 MW of electrolysis, NLR only has a single 1.25-MW electrolysis stack. For the simulated wind energy profiles, NLR used OpenFAST to simulate a 3.4-MW International Energy Agency (IEA) reference wind turbine. The hour-long wind energy profiles varied over wind turbulence intensity (Class A or Class C) and average wind speed (5, 7, or 9 m/s). To match the power limits of the 1.25-MW electrolyzer and 3.4-MW IEA wind turbine most effectively and to maximize the efficiency of hydrogen production at a given average wind speed, the profiles were sometimes scaled by two times. This means that, in some cases, the experimental setup assumed two 1.25-MW electrolyzers were coupled with the wind turbine, representing a total maximum electrolysis load of 2.5 MW. Finally, NLR experimented with two settings for the electrolyzer power supply minimum and maximum current ramp rates (gain and slew): 200 and 400 amperes per second. The simulated profiles were translated from power (kilowatts) to current (amperes) using a curve fit with calibration data and sent to the electrolyzer power supply at 1-Hz frequency. These datasets report relevant hydrogen balance-of-plant and system data, all captured at 1 Hz, including hydrogen mass production measured with an Emerson Coriolis flow meter. Each .zip file represents a single wind turbine electrolysis experiment and is formatted as follows: {technology}-{average wind speed}-{turbulence class}_{number of 1.25 MW electrolyzers connected}-{electrolyzer ramp rate in amperes/second} For instance, “windIEA3.4-5ms-C_2-400.zip” represents the hour-long experiment using the IEA 3.4-MW turbine, subjected to an average wind speed of 5 m/s and Class C wind turbulence, and connected to two 1.25-MW electrolyzers with the power supply set to a maximum current ramp rate (gain and slew) of 400 A/s. Each .zip folder contains the following files: A .csv file containing raw data. An .xlsx file explaining all the fields in the raw data. A .png plot showing the time series of hydrogen production in kilograms per hour, electrolysis power consumption, and input wind turbine power. An experiment labeled “characterization_200.zip” demonstrates the MC250 electrolyzer steady-state response with 30 minute load steps for a total duration of 5 hours. Finally, a .csv file is provided with all simulated wind experiments combined into one dataset labeled "combined_wind_experiments.csv". NLR also built an AI/machine-learning predictive model based on these datasets. The model ingests the electrolyzer current command in amperes, as well as various pressures and temperatures across the system, and predicts hydrogen output in kilograms per hour. The complete model can be found at https://huggingface.co/NatLabRockies/ptmelt-hydrogen-electrolysis .

08 HYDROGEN↗

1000 Soils Pilot Dataset, version 8, May 2025

This record hosts data generated by the 1000 Soils Pilot. Data will be updated as more become available. Please see the most recent data upload for current data. A beta visualization tool is available for some data types at https://shinyproxy.emsl.pnnl.gov/app/1000soils. Please submit any suggestions or comments through the 'contact' tab. We are actively working to improve visualizations and value all feedback. Data completed include: Geochemistry, texture, respiration, and enzyme activities FTICR-MS organic matter chemistry Microbial biomass C and N TOC/TDN of water-extractable OM X-ray computed tomography (derived metrics available here, raw data available upon request) Metagenomes; a variety of data formats are available upon request Soil hydraulic properties Data in progress: LC-MS/MS in development, timeline TBD, inquire for status 1000S_processed_BGC_summary.csv contains all available biogeochemical data; microbial biomass C and N; and TOC/TDN of water-extractable OM; and 1000S_Tomography.xslx contains a summary of data generated via X-ray computed tomography. icr_v2_corems2.csv contains FTICR-MS data processed by CoreMS version 2. These data are merged by formula across instrument runs to enable cross-sample comparisons. Technical replicates are merged by retaining peaks present in 2 out of 3 replicates. 1000Soils_Metadata_Site_Mastersheet_v1.csv contains site information. Soil Hydraulics_corrected_02042025.xlsx contains soil hydraulics information. Readme File_v4.xlsx is the readme file. Please contact the MONet project (monet.emsl@pnnl.gov) or Emily Graham (emily.graham@pnnl.gov) with questions. The following file and all raw data are available upon request: icr_by_mass_for_single_sample_analysis_only.csv contains FTICR-MS data processed by CoreMS and is intended for usage in the calculation of biochemical transformations within samples only. These data are not acceptable for cross-sample comparison of masses because they are from multiple instrument runs. For more information, please see: https://www.emsl.pnnl.gov/monet and https://sc-data.emsl.pnnl.gov/monet Acknowledgment: Soil data were provided by the Molecular Observation Network (MONet) at the Environmental Molecular Sciences Laboratory (https://ror.org/04rc0xn13), a DOE Office of Science user facility sponsored by the Biological and Environmental Research program under Contract No. DE-AC05-76RL01830. The work (proposal: 10.46936/10.25585/60008970) conducted by the U.S. Department of Energy, Joint Genome Institute (https://ror.org/04xm1d337), a DOE Office of Science user facility, is supported by the Office of Science of the U.S. Department of Energy operated under Contract No. DE-AC02-05CH11231. The Molecular Observation Network (MONet) database is an open, FAIR, and publicly available compilation of the molecular and microstructural properties of soil. Data in the MONet open science database can be found at https://sc-data.emsl.pnnl.gov/.

biogeochemistry↗

High resolution characterization of soil dissolved organic matter with FTICR-MS (Fourier-transform ion cyclotron resonance mass spectrometry) from soil samples in control and warming plots in Blodgett Forest, CA (2014 and 2018)

The pathways of carbon transport and loss through and from soils—soil organic matter (SOM) depolymerization to dissolved organic carbon and mineralization to carbon dioxide (CO2)—are fundamentally driven by microbial activity, which is strongly regulated by environmental conditions. As part of Lawrence Berkeley National Laboratory (LBNL) Terrestrial Ecosystem Science (TES) Belowground Biogeochemistry Science Focus Area (SFA), we have established a novel whole-soil long-term warming experiment at the University of California (UC) Blodgett Forest Research Station (Sierra Nevada) in 2014, where we study the role of biogeochemical, microbial and geochemical process interactions in SOM decomposition and stabilization.This package contains Fourier transform ion cyclotron resonance mass spectrometry (21 Tesla FTICR-MS) data measured in negative and positive ionization mode from water and methanol soil extracts. Soil samples were collected in 2014/06/03 and 2018/06/04 from 3 replicated paired plots that had been subjected to experimental warming since June 2014 to simulate a predicted climate change scenario for northern California. The following files are included: (1) fticr_neg_h2oMeoh_data_raw.csv: raw data from combined water (H2O) and methanol (MeOH) extracts in negative ion mode, (2) fticr_neg_h2oMeoh_data_processed.csv: processed data from combined water (H2O) and methanol (MeOH) extracts in negative ion mode, (3) fticr_neg_metadata.csv: metadata for samples/measurements in negative ion mode, (4) fticr_pos_h2oMeoh_data_raw.csv: raw data from combined water (H2O) and methanol (MeOH) extracts in positive ion mode, (5) fticr_pos_h2oMeoh_data_processed.csv: processed data from combined water (H2O) and methanol (MeOH) extracts in positive ion mode, (6) fticr_pos_metadata.csv: metadata for samples/measurements in positive ion mode.

54 ENVIRONMENTAL SCIENCES↗

HarDWR - Harmonized Water Rights Records

A dataset within the Harmonized Database of Western U.S. Water Rights (HarDWR). For a detailed description of the database, please see the meta-record v2.0. Changelog v2.0 - Recalculated based on data sourced from WestDAAT - Changed using a Site ID column to identify unique records to using aa combination of Site ID and Allocation ID - Removed the Water Management Area (WMA) column from the harmonized records. The replacement is a separate file which stores the relationship between allocations and WMAs. This allows for allocations to contribute to water right amounts to multiple WMAs during the subsequent cumulative process. - Added a column describing a water rights legal status - Added "Unspecified" was a water source category - Added an acre-foot (AF) column - Added a column for the classification of the right's owner v1.02 - Added a .RData file to the dataset as a convenience for anyone exploring our code. This is an internal file, and the one referenced in analysis scripts as the data objects are already in R data objects. v1.01 - Updated the names of each file with an ID number less than 3 digits to include leading 0s v1.0 - Initial public release Description Here we present an updated database of Western U.S. water right records. This database provides consistent unique identifiers for each water right record, and a consistent categorization scheme that puts each water right record into one of seven broad use categories. These data were instrumental in conducting a study of the multi-sector dynamics of inter-sectoral water allocation changes though water markets (Grogan et al., *in review*). Specifically, the data were formatted for use as input to a process-based hydrologic model, Water Balance Model (WBM), with a water rights module (Grogan et al., *in review*). While this specific study motivated the development of the database presented here, water management in the U.S. West is a rich area of study (e.g., Anderson and Woosly, 2005; Tidwell, 2014; Null and Prudencio, 2016; Carney et al., 2021) so releasing this database publicly with documentation and usage notes will enable other researchers to do further work on water management in the U.S. West. We produced the water rights database presented here in four main steps: (1) data collection, (2) data quality control, (3) data harmonization, and (4) generation of cumulative water rights curves. Each of steps (1)-(3) had to be completed in order to produce (4), the final product that was used in the modeling exercise in Grogan et al. (*in review*). All data in each step is associated with a spatial unit called a Water Management Area (WMA), which is the unit of water right administration utilized by the state in which the right came from. Steps (2) and (3) required use to make assumptions and interpretation, and to remove records from the raw data collection. We describe each of these assumptions and interpretations below so that other researchers can choose to implement alternative assumptions an interpretation as fits their research aims. Motivation for Changing Data Sources The most significant change has been a switch from collecting the raw water rights directly from each state to using the water rights records presented in WestDAAT, a product of the Water Data Exchange (WaDE) Program under the Western States Water Council (WSWC). One of the main reasons for this is that each state of interest is a member of the WSWC, meaning that WaDE is partially funded by these states, as well as many universities. As WestDAAT is also a database with consistent categorization, it has allowed us to spend less time on data collection and quality control and more time on answering research questions. This has included records from water right sources we had previously not known about when creating v1.0 of this database. The only major downside to utilizing the WestDAAT records as our raw data is that further updates are tied to when WestDAAT is updated, as some states update their public water right records daily. However, as our focus is on cumulative water amounts at the regional scale, it is unlikely most records updates would have a significant effect on our results. The structure of WestDAAT led to several important changes to how HarWR is formatted. The most significant change is that WaDE has calculated a field known as `SiteUUID`, which is a unique identifier for the Point of Diversion (POD), or where the water is drawn from. This separate from `AllocationNativeID`, which is the identifier for the allocation of water, or the amount of water associated with the water right. It should be noted that it is possible for a single site to have multiple allocations associated with it and for an allocation to be able to be extracted from multiple sites. The site-allocation structure has allowed us to adapt a more consistent, and hopefully more realistic, approach in organizing the water right records than we had with HarDWR v1.0. This was incredibly helpful as the raw data from many states had multiple water uses within a single field within a single row of their raw data, and it was not always clear if the first water use was the most important, or simply first alphabetically. WestDAAT has already addressed this data quality issue. Furthermore, with v1.0, when there were multiple records with the same water right ID, we selected the largest volume or flow amount and disregarded the rest. As WestDAAT was already a common structure for disparate data formats, we were better able to identify sites with multiple allocations and, perhaps more importantly, allocations with multiple sites. This is particularly helpful when an allocation has sites which cross WMA boundaries, instead of just assigning the full water amount to a single WMA we are now able to divide the amount of water between the number of relevant WMAs. As it is now possible to identify allocations with water used in multiple WMAs, it is no longer practical to store this information within a single column. Instead the stAllocationToWMATab.csv file was created, which is an allocation by WMA matrix containing the percent Place of Use area overlap with each WMA. We then use this percentage to divide the allocation's flow amount between the given WMAs during the cumulation process to hopefully provide more realistic totals of water use in each area. However, not every state provides areas of water use, so like HarDWR v1.0, a hierarchical decision tree was used to assign each allocation to a WMA. First, if a WMA could be identified based on the allocation ID, then that WMA was used; typically, when available, this applied to the entire state and no further steps were needed. Second was the spatial analysis of Place of Use to WMAs. Third was a spatial analysis of the POD locations to WMAs, with the assumption that allocation's POD is within the WMA it should belong to; if an allocation still had multiple WMAs based on its POD locations, then the allocation's flow amount would be divided equally between all WMAs. The fourth, and final, process was to include water allocations which spatially fell outside of the state WMA boundaries. This could be due to several reasons, such as coordinate errors / imprecision in the POD location, imprecision in the WMA boundaries, or rights attached with features, such as a reservoir, which crosses state boundaries. To include these records, we decided for any POD which was within one kilometer of the state's edge would be assigned to the nearest WMA. Other Changes WestDAAT has Allowed In addition to a more nuanced and consistent method of assigning water right's data to WMAs, there are other benefits gained from using the WestDAAT dataset. Among those is a consistent categorization of a water right's legal status. In HarDWR v1.0, legal status was effectively ignored, which led to many valid concerns about the quality of the database related to the amounts of water the rights allowed to be claimed. The main issue was that rights with legal status' such as "application withdrawn", "non-active", or "cancelled" were included within HarDWR v1.0. These, and other water rights status' which were deemed to not be in use have been removed from this version of the database. Another major change has been the addition of the "unspecified water source category. This is water that can come from either surface water or groundwater, or the source of which is unknown. The addition of this source category brings the total number of categories to three. Due to reviewer feedback, we decided to add the acre-foot (AF) column so that the data may be more applicable to a wider audience. We added the ownerClassification column so that the data may be more applicable to a wider audience. File Descriptions The dataset is a series of various files organized by state sub-directories. In addition, each file begins with the state's name, in case the file is separate from its sub-directory for some reason. After the state name is the text which describes the contents of the file. Here is each file described in detail. Note that st is a placeholder for the state's name. stFullRecords_HarmonizedRights.csv: A file of the complete water records for each state. The column headers for each of this type of file are: state - The name of the state to which the allocations belong to. FIPS - The two digit numeric state ID code. siteID - The site location ID for POD locations. A site may have multiple allocations, which are the actual amount of water which can be drawn. In a simplified hypothetical, a farm stead may have an allocation for "irrigation" and an allocation for "domestic" water use, but the water is drawn from the same pumping equipment. It should be noted that many of the site ID appear to have been added by WaDE, and therefore may not be recognized by a given state's water rights database. allocationID - The allocation ID for the water right. For most states this is the water right ID, and what is recommended to use should a right be looked up on a given state's water rights database. The water amounts associated with these IDs tend to be finer scaled than those associated with siteID. It should be noted that some allocations may be extracted from multiple sites, particularly for larger Places of Use. ownerClassification - A classification of the types of owners for water rights. The most common is `Private` which incorporates a wide range of entities. Several classifications would be grouped into a government category, most of which are for the U.S. Federal Government. These allocations could be listed as "Federal", "United States of America", or as the names of any number of federal agencies. The last major grouping of entities is for "Native American"s. priorityDate - The date we use as the water right priority date for our modeling analysis. This is the legal priority date when it is available. However, for some rights, specifically from California and New Mexico, we used a pseudo priority date (e.g. well completion date or start of well drilling date) when a legal priority date was not available. The most questionable dates come from New Mexico, where the only date associated with certain water right records was the date the allocation was recorded in the database. As the allocation record creation tended to be within a few months of the filing of the application of the water right, from manually double checking the water rights, and our analysis focuses on aggregating water rights on the timescale of years, we determined it was acceptable to use such dates to include as many records as possible. primaryBeneficialUse - From the numerous state water use categories, WaDE categorized them into 21 categories WestDAAT. This column is the original WaDE category for the primary water use at the PoD site. allocationBeneficialUse - From the numerous state water use categories, WaDE categorized them into 21 categories for WestDAAT. This column is the original WaDE category

Economics↗

Fermilab PIP-II machine protection system digitized data noise elimination scheme and its FPGA implementation

In Fermilab's PIP-II machine protection system, beam loss signals from various detectors are digitized at 125 MS/s. Noise from both high-frequency sources and low-frequency 60 Hz AC power equipment can contaminate the data. To suppress noise across these ranges—especially 60 Hz and its harmonics, which overlap with beam loss signal frequencies—advanced digital processing beyond standard filtering is required. Several real-time functional blocks were simulated and tested on an FPGA: (1) a dual time-constant discharging integrator filter, (2) a de-ripple baseline extraction and storage block, and (3) a fast-recovery discharging integrator. The nonlinear IIR integrator filter removes high-frequency noise and feeds into the baseline extractor. Upon detecting abrupt beam loss, it switches to a longer time constant to prevent baseline distortion. The de-ripple block calculates a valid baseline by averaging over multiple 60 Hz periods, storing results in a 4096-word FPGA RAM. This baseline is subtracted from raw data before integration by the fast-recovery block, which resets quickly after use. All blocks achieved expected performance and were successfully implemented on a low-cost FPGA.

Wu, Jinyuan [Fermilab]↗

Generative large language models for predictive maintenance planning

Maintenance planning and the generation of necessary components for tasks can prove time-consuming and complex. Automating the creation of recurring or similar tasks by leveraging previous planning packages and data, while uncovering insights to automate planning package generation, presents an opportunity to conserve valuable time and resources. This work aims to harness the textual and probabilistic capabilities of large language models (LLMs) to automate the generation of planning packages. Utilizing diverse data sources ranging from raw data to handwritten text, both singular and collaborative LLMs are trained and tested. Results demonstrate their capability to generate essential planning package components, effectively replicating the statistical patterns in the data. This demonstrates the use of these tools inside a digital asset for automated planning. This work outlines a methodology for constructing datasets, a training suite, and evaluation methods for LLM-based textual and conversational planning tools utilized in an asset digital twin. Results indicate that the fine-tuned models generate estimated planning information within the statistical ranges observed in real maintenance data. The models achieve high accuracy (>90%) in document question-answering and instruction generation tasks. Furthermore, the conversational retrieval-augmented generation (RAG) assistant system achieves 100% document retrieval accuracy, while conversational information capture exceeds 98% across the majority of work-package assistant modules.

97 MATHEMATICS AND COMPUTING↗

Top-down proteomics

Proteoforms arising from posttranslational modifications, genetic polymorphisms, and RNA splice variants, play a pivotal role as the key drivers in biology. Thus, a comprehensive understanding of proteoforms is essential for unraveling the intricacies of biological systems and bridging the gap between genotype and phenotype. By analyzing whole proteins without digestion, top-down proteomics (TDP) provides a holistic view of the proteome and presents a next-generation approach for deciphering protein function, uncovering disease mechanisms, and advancing precision medicine. This Primer embarks on a journey into the world of TDP by encapsulating its historical context, underlying principles, recent advances, and an outlook on the future of TDP. The experimental section navigates instrumentation, sample preparation, intact protein separation, tandem mass spectrometry techniques, and data collection. Results decipher raw data, visualize intact protein spectra, unravel data analysis, and explain proteoform identification, characterization, and quantitation, as well as statistical analysis. Various applications of TDP spanning the human proteoform project, biomedical, biopharmaceutical, and clinical applications are described. These are complemented by discussions on measurement reproducibility, limitations, and a forward-looking perspective outlining uncharted waters where the field can advance, and potential exciting future applications of TDP.

Roberts, David S.↗

SALSA: a new versatile readout chip for MPGD detectors

The SALSA chip is a future readout ASIC foreseen for the MPGD detectors, developed in the framework of the EIC collider project, to equip the MPGD trackers of the EPIC experiment. It is designed to be versatile, to be adapted to other usages of MPGD detectors like TPC or photon detectors. It integrates a frontend block and an ADC for each of the 64 channels, associated to a configurable DSP processor meant to correct data and reduce the raw data flux to limit the output bandwidth. It will be compatible with the continuous readout foreseen for the EPIC DAQ, but will also work in a triggered environment. Several prototypes are already produced in order to qualify the different blocks of the chip, in particular the frontend, the ADC and the clock generation. The next 32-channel prototype is currently under development and is planed to be produced in 2025. In conclusion, the final prototype will be produced and tested from 2026 for a production of the SALSA chip at the horizon of 2027.

Data processing↗

DP-TwoLevel: two-stage gradient subspace learning for differentially private federated learning

Federated learning (FL) enables collaborative model training across distributed data sources without sharing raw data, but faces fundamental challenges in communication efficiency and privacy. Differentially private (DP) training mitigates information leakage but introduces noise that degrades model performance, especially in high-dimensional settings. We propose DP-TwoLevel, a hierarchical gradient projection method that improves utility under fixed DP constraints by exploiting low-dimensional structure in model updates. Our approach learns a two-level PCA-based representation of gradients and applies DP noise in a reduced-dimensional subspace, thereby lowering the effective noise magnitude while preserving dominant signal components. We evaluate the method across three datasets (MNIST, Fashion-MNIST, CIFAR-10) and three privacy regimes (ϵ∈0.5, 1.0, 2.0). Across nine experimental settings, DP-TwoLevel consistently outperforms DP-FedAvg, achieving an average accuracy improvement of 9.44%, with larger gains observed in lower ϵ(higher-noise) regimes (up to +22.31%). We further analyze scalability across models ranging from 100K to 1.49M parameters and identify a variance-based success criterion: performance remains strong when the projection preserves more than 75% of gradient variance, degrades in a marginal regime (65–75%), and fails below this threshold. Our results demonstrate that structure-aware dimensionality reduction can significantly improve the privacy–utility tradeoff in FL without modifying formal privacy guarantees. We also provide empirical evidence of scaling limitations for global projections and motivate per-layer extensions for larger models.

Kotevska, Olivera [ORNL] (ORCID:0000000316772243)↗

Automated point dendrometer, soil moisture and temperature, and meteorological variables datasets, Oct 2024 – Nov 2025, G.A. Pearson Natural Area, Flagstaff, AZ, USA

This data package includes parsed, cleaned, and calibrated data from 48 TOMST automated point dendrometers, 48 TOMST 15 cm soil moisture sensors, and 12 TOMST 30 cm soil moisture sensors. The point dendrometers were cleaned with the “dendRoAnalyst” package in RStudio. The soil sensors were cleaned and calibrated for volumetric water content (VWC) with the “myClim” package in RStudio using the soil texture of the site (sandy clay loam). Additionally, this data package also includes raw data from 2 METER weather stations. Dendrometers and soil sensors have both their sensor ID, as well as the ID for the specific tree they were instrumented on at the G.A. Pearson Natural Area (GPNA) site and their experimental group. The purpose of these data is to understand how ponderosa pine trees in restored (thinned and burned) vs. unrestored (no treatment) areas are responding to drought and seasonal precipitation. These data use radial growth and soil moisture data to answer the following question: how are active season length, growth on different time scales (weekly, monthly, seasonally, and annually), growth during dry periods and after precipitation events, and environmental and biological drivers of radial growth different between restored versus unrestored areas?

Air temperature↗

Vegetation Warming Experiment: Environmental Conditions, Utqiagvik (Barrow), Alaska, 2018

Environmental conditions measured in five warming chambers and paired ambient control plots located on the Barrow Environmental Observatory (BEO), Utqiagvik, Alaska from 16 June - 24 September, 2018. These data were recorded in support of the Zero Power Warming (ZPW) vegetation warming experiment, a series of single season vegetation warming treatments conducted over four years from 2017-2021 (no experiment in 2020). Air temperature and humidity, infrared surface (canopy) temperature, soil temperature, soil moisture, NDVI (normalized difference vegetation index), PRI (photochemical reflectance index), solar radiation and chamber venting were recorded in each chamber at 1 minute intervals. Ambient air temperature, humidity, solar radiation and uplooking PRI and NDVI were measured at a centrally located meteorology station. Vapor pressure deficit (VPD) was calculated and included in the final processed data products. Data has undergone full QA/QC and is presented as 1 minute data, and hourly and daily aggregate data products. This data package includes unprocessed raw data (*.dat files), processed data (*.csv) and metadata including a full description of sensors, calculations and processing (*.csv, *.pdf). See related NGEE-Arctic "Vegetation Warming Experiment" data packages for leaf-level gas exchange and other leaf trait data; chamber, plot and landscape phenocamera images; thaw depth, and GPS locations of chambers and ambient plots.The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic), was a 15-year research effort (2012-2027) to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research. The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska. Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

54 ENVIRONMENTAL SCIENCES↗

Vegetation Warming Experiment: Environmental conditions, Utqiagvik (Barrow), Alaska, 2019

Environmental conditions measured in five warming chambers and paired ambient control plots located on the Barrow Environmental Observatory (BEO), Utqiagvik, Alaska from 19 June – 25 September, 2019. These data were recorded in support of the Zero Power Warming (ZPW) vegetation warming experiment, a series of single season vegetation warming treatments conducted over four years from 2017–2021 (no experiment in 2020). Air temperature and humidity, infrared surface (canopy) temperature, soil temperature, soil moisture, NDVI (normalized difference vegetation index), PRI (photochemical reflectance index), solar radiation and chamber venting were recorded in each chamber at 1 minute intervals. Ambient air temperature, humidity, solar radiation and uplooking PRI and NDVI were measured at a centrally located meteorology station. Vapor pressure deficit (VPD) was calculated and included in the final processed data products. Data has undergone full QA/QC and is presented as 1 minute data, and hourly and daily aggregate data products. This data package includes unprocessed raw data (*.dat files), processed data (*.csv) and metadata including a full description of sensors, calculations and processing (*.csv, *.pdf). See related NGEE-Arctic "Vegetation Warming Experiment" data packages for leaf-level gas exchange and other leaf trait data; chamber, plot and landscape phenocamera images; thaw depth, and GPS locations of chambers and ambient plots.The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic), was a 15-year research effort (2012-2027) to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research. The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska. Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

54 ENVIRONMENTAL SCIENCES↗

Vegetation Warming Experiment: Environmental conditions, Utqiagvik (Barrow), Alaska, 2021

Environmental conditions measured in five warming chambers and paired ambient control plots located on the Barrow Environmental Observatory (BEO), Utqiagvik, Alaska from 19 June - 17 September, 2021. These data were recorded in support of the Zero Power Warming (ZPW) vegetation warming experiment, a series of single season vegetation warming treatments conducted over four years from 2017-2021 (no experiment in 2020). Air temperature and humidity, infrared surface (canopy) temperature, soil temperature, soil moisture, NDVI (normalized difference vegetation index), PRI (photochemical reflectance index), solar radiation and chamber venting were recorded in each chamber at 1 minute intervals. Ambient air temperature, humidity, solar radiation and uplooking PRI and NDVI were measured at a centrally located meteorology station. Vapor pressure deficit (VPD) was calculated and included in the final processed data products. Data has undergone full QA/QC and is presented as 1 minute data, and hourly and daily aggregate data products. This data package includes unprocessed raw data (*.dat files), processed data (*.csv) and metadata including a full description of sensors, calculations and processing (*.csv, *.pdf). See related NGEE-Arctic "Vegetation Warming Experiment" data packages for leaf-level gas exchange and other leaf trait data; chamber, plot and landscape phenocamera images; thaw depth, and GPS locations of chambers and ambient plots. The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic), was a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research. The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska. Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

54 ENVIRONMENTAL SCIENCES↗