Search NASA⌕ Search

SEARCH · Search NASA

Results for “labeled data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28

Bayes Error Rate Estimation Using Classifier Ensembles

The Bayes error rate gives a statistical lower bound on the error achievable for a given classification problem and the associated choice of features. By reliably estimating th is rate, one can assess the usefulness of the feature set that is being used for classification. Moreover, by comparing the accuracy achieved by a given classifier with the Bayes rate, one can quantify how effective that classifier is. Classical approaches for estimating or finding bounds for the Bayes error, in general, yield rather weak results for small sample sizes; unless the problem has some simple characteristics, such as Gaussian class-conditional likelihoods. This article shows how the outputs of a classifier ensemble can be used to provide reliable and easily obtainable estimates of the Bayes error with negligible extra computation. Three methods of varying sophistication are described. First, we present a framework that estimates the Bayes error when multiple classifiers, each providing an estimate of the a posteriori class probabilities, a recombined through averaging. Second, we bolster this approach by adding an information theoretic measure of output correlation to the estimate. Finally, we discuss a more general method that just looks at the class labels indicated by ensem ble members and provides error estimates based on the disagreements among classifiers. The methods are illustrated for artificial data, a difficult four-class problem involving underwater acoustic data, and two problems from the Problem benchmarks. For data sets with known Bayes error, the combiner-based methods introduced in this article outperform existing methods. The estimates obtained by the proposed methods also seem quite reliable for the real-life data sets for which the true Bayes rates are unknown.

Tumer, Kagan↗

Human Specimen Repository: Sample Collection for the Future [#133-000445]

In early 2014, an executive memo was released directing the improvement of the management of and access to scientific collections funded by government resources. In the spirit of this memo, NASA’s Human Research Program (HRP) funded a project to review existing residual samples in long-term storage for viability, identification, and scientific significance. These samples were to be collected in a centralized location and a detailed inventory was to be conducted to create the Human Sample Repository (HSR). HSR would work to review data from over 27,000 samples, from flight research projects dating back to 1998. The samples themselves would be reviewed, re-labeled, and organized into easy to retrieve, long-term storage. In addition to the older residual samples, HSR would expand to include pristine samples from all consenting crew who have flown on International Space Station (ISS) missions. All samples inventoried and managed by HSR are intended for the same purpose; to support future research projects which will utilize advanced technologies or techniques, by providing one-of-a-kind sample collections that can not be replicated. With this mission in mind, HSR will continue to collect residual samples from HRP funded projects and pristine samples from consenting crew, to further build this unique collection and increase the diversity of flight samples available to the scientific community.

Human Specimen Repository↗

An isotopic labeling investigation into the influence of the nitro group on LLM-105 thermal decomposition

Here, this work presents the first application of isotopically labeled LLM-105 (2,6-diamino-3,5-dinitropyrazine-1-oxide) to investigate thermal decomposition pathways. Specially synthesized LLM-105 isotopologues were utilized to isolate the influence of labeled 15 NO 2 nitro groups on the formation of lightgas products. Simultaneous differential scanning calorimetry, thermo-gravimetric, and mass spectrometry measurements were employed to track the evolution of product gases, enabling the direct comparison of isotopically shifted species with unlabeled LLM-105. Key findings show that C 2 N 2 production is mainly dependent on nitrogen sources from either the amine groups or the pyrazine ring (i.e., not the nitro groups). The formation of NO, N 2 , and N 2 O all involves the nitro groups to some extent. NO (nitric oxide) was found to be the predominant gas species directly formed from the nitro group of LLM-105. In contrast, mixed nitrogen isotopologues of N 2 and N 2 O (i.e., 14 N 15 N and 15 NNO) formed more readily in comparison to their pure counterparts (i.e., 15 N 2 and 15 N 2 O). This indicates the amine and/or pyrazine groups of LLM-105, in addition to the nitro group, are involved in the decomposition pathways forming N 2 and N 2 O. In addition, our investigation led to the discovery of two previously unreported decomposition products (CHO and HNCO), which were confirmed through hydrogen labelling utilizing deuterium isotopes. These results provide detailed speciation trends of gaseous products during LLM-105 decomposition, offering new insights into reaction pathways. Experimental data reported here will support the development of a detailed chemical kinetics model for LLM-105, essential for the safe handling of high explosives.

Chemistry - Chemical explosives↗

Predictive Modeling for Differential Diagnosis and Mortality Risk Assessment

The prevalence of electronic health record (EHR) systems has brought prodigious biomedical informatics opportunity. Automated machine learning methods can effectively utilize such data and have become common tools for healthcare predictive modeling. Researches in medical informatics have explored the potential of deep learning and classical models in emergent care scenarios. In particular, predicting differential diagnoses for admissions have proven useful in decreasing unnecessary lab tests and improving inpatient triage decision-making. Moreover, identification of high-risk patients for in-hospital mortality is vitally important to maximize allocation of medical resources.The Medical Information Mart for Intensive Care (MIMIC-III) database, containing de-identified critical care inpatient was used in our study. This data set captures hospital patient laboratory measurements, pharmacologic prescriptions, diagnostic data and procedure event recordings. When considering adult patients and discounting admissions with ICU length of stay less than 24 hours, there were 37,787 unique admissions and 30,414 total patients. We examined the top 25 most prevalent ICD-9 group-level disease specificities in MIMIC-III using a multi-label classification model. In-hospital mortality was modeled as binary classification with 4,155 (13%) adult patients that expired, of which 3,138 (75.5%) were in the ICU setting. The metrics AUC, F1 score, sensitivity and specificity values calculated for each disease label measured prediction performance.The usage of ICD-9 group codes reduced feature dimension from 14,567 to 942 and greatly improved distribution of patient diagnostic categories. Disease temporal patterns were captured by considering the most frequently sampled 6 vital signs and 13 laboratory values. Missing data were imputed at each time-stamp. Time-series raw hourly average values were converted into 5 summary features (mean, standard deviation, number of observations, min & max values). Patient demographic variables such as age, gender, marital status and ethnicity were also factored into the modeling. Choi et al showed that contextual embedding of medical data, diagnostic and procedural codes alone can predict future diagnoses with sensitivity as high as 0.79. We utilized an embedding technique called word2vec which allowed sparse representations of medical history to be transformed into dense word vectors. The mappings captured contextual information by treating each admission as a sentence and learning the most likely neighboring words in a sliding window fashion. Binary and multi-label classification was achieved via collapse models, which do not consider temporal information, as well as recurrent neural networks with regularization, Softmax output layer activation together with categorical cross-entropy as the loss function.

US Army collaboration↗

Laboratory time series moisture manipulative experiment from sediment across the contiguous US: time series aerobic respiration and geochemistry (v2)

This dataset supports a broader study examining the effects of wetting and drying on hyporheic zone respiration across the contiguous United States (CONUS). The dataset provides data generated from a laboratory moisture manipulation experiment. The contents include time series aerobic respiration and moisture; dissolved oxygen; sediment geochemistry data; and field metadata (including qualitative information on instream and river corridor characteristics). Samples were collected as part of the WHONDRS CONUS-Scale Model-Sample Study (CM). This study was designed following ICON (integrated, coordinated, open, and networked) principles to facilitate a model-experiment (ModEx) iteration approach, leveraging crowdsourced sampling across the CONUS. The data package associated with the CM study is available at https://data.ess-dive.lbl.gov/view/doi:10.15485/1923689. CM sampling began in April 2022 and ended in October 2023. This study uses subsamples from a subset of CM samples collected between June 2022 and June 2023. The original field samples were labeled as CM_###. Subsequent subsamples for this study were labeled as EC_###. The labels from the field samples and the EC subsamples can be mapped directly based on the digits following the prefix and underscore (i.e., EC_001 is a subsample from CM_001). See the critical details section below for more details on sample naming. This data package was originally published in August 2024. It was updated in February 2026 (v2; new and modified files). See the change history section in the readme for more details. For details on how to navigate this data package, see this infographic from the River Corridor SFA https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. This dataset is comprised of one folder of raw Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) data and one main data folder containing (1) file-level metadata; (2) data dictionary; (3) field metadata; (4) readme; (5) field protocol; and a (6) a subfolder with sediment sample data from the incubation experiment. The sample data subfolder contains (1) dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC); (2) total nitrogen (TN); (3) adenosine triphosphate (ATP); (4) percent carbon and nitrogen; (5) effect size; (6) iron (II); (7) gravimetric moisture; (8) respiration rates and raw dissolved oxygen values; (9) specific conductance; (10) pH; (11) temperature; (12) a summary containing median values of each data type for each treatment (wet and dry); (13) methods codes; (14) FTICR-MS methods; and (15) a subfolder of 9.4 Tesla FTICR-MS data. This folder contains three subfolders, one containing the sediment .xml data files, one containing the sediment CoreMS output files, the other containing instructions and scripts for processing the files in CoreMS (https://github.com/EMSL-Computing/CoreMS). All files are .csv, .pdf, .R, .ref, or .xml.

54 ENVIRONMENTAL SCIENCES↗

Using Classical Reliability Models and Single Event Upset (SEU) Data to Determine Optimum Implementation Schemes for Triple Modular Redundancy (TMR) in SRAM-Based Field Programmable Gate Array (FPGA) Devices

Space applications are complex systems that require intricate trade analyses for optimum implementations. We focus on a subset of the trade process, using classical reliability theory and SEU data, to illustrate appropriate TMR scheme selection.

Field Programmable Gate Array (FPGA)↗

Open Power System Datasets and Open Simulation Engines: A Survey Toward Machine Learning Applications

A major factor behind the success of machine learning (ML) models in multiple domains is the availability and accessibility of large, labeled, and well-organized datasets for training and benchmarking. In comparison, power grid datasets face three major challenges: (i) real-world data is often restricted by regulatory constraints, privacy reasons, or security concerns, making it difficult to obtain and work with; (ii) synthetic datasets, which are created to address these limitations, often have incomplete information and are released using specialized tools, making them inaccessible to the broader community; and, (iii) input-output datasets are difficult to generate through simulation for non-experts because open-source simulators are not known outside the power system community. This survey addresses these challenges by serving as an entry point to publicly available datasets and simulators for researchers venturing in this area. We review the current landscape of open-source power network data, machine models, consumer demand profiles, renewable generation data, and inverter models. We also examine open-source power system simulators, which are crucial for generating high-quality, high-fidelity power grid datasets. We aim to provide a foundation for overcoming data scarcity and advance towards a structured web of datasets and simulators to support the development of ML for power systems.

42 ENGINEERING↗

Development of a universal water signature for the LANDSAT-3 Multispectral Scanner, part 1

A generalized four channel hyperplane to discriminate water from nonwater was developed using LANDSAT-3 multispectral scaner (MSS) scenes and matching same/next day color infrared aerial photography. The MSS scenes varied in sun elevation angle from 40 to 58 deg. The 28 matching air photo frames contained over 1400 water bodies larger than one surface acre. A preliminary water discriminant, was used to screen the data and eliminate from further consideration all pixels distant from water in MSS spectral space. A linear discriminant was iteratively fitted to the labelled pixels. This discriminant correctly classified 98.7% of the water pixels and 98.6% of the nonwater pixels. The discriminant detected 91.3% of the 414 water bodies over 10 acres in surface area, and misclassified as water 36 groups of contiguous nonwater pixels.

Schlosser, E. H.↗

Some approaches to optimal cluster labeling of aerospace imagery

Some approaches are presented to the problem of labeling clusters using information from a given set of labeled and unlabeled aerospace imagery patterns. The assignment of class labels to the clusters is formulated as the determination of the best assignment over all possible ones with respect to some criterion. Cluster labeling is also viewed as the probability of correct labeling with a maximization of likelihood function. Results of the application of these techniques in the processing of remotely sensed multispectral scanner imagery data are presented.

Chittineni, C. B.↗

The spectroscopic, chemical, and photophysical properties of Martian soils and their analogs (MERC, phase 2)

A series of variably proportioned iron/calcium smectite clays and iron loaded smectite clays containing iron up to the level found in the Martian soil were prepared from a typical montomorillonite clay using the Banin method. Evidence was obtained which supports the premise that these materials provide a unique and appropriate model soil system for the Martian surface in that they are consistent with the constraints imposed by the Viking surface elemental analysis, the reflectance data obtained by various spacecraft instruments and ground based telescopes, and the chemical reactivity measured by one of the Viking biology experiments, the Labeled Release (LR) experiment.

Banin, Amos↗

Estimating new production in the equatorial Pacific Ocean at 150 deg W

A major goal of the WEC88 cruise of the R/V Wecoma to the equatorial Pacific (made in February-March 1988) was to establish rates of new production along a meridional section at 150 deg W and to compare these measured rates with the relatively high values for the equatorial Pacific that had been reported previously using indirect methods and models. Production values were obtained from the traditional approach using N-15 labeled nitrate uptake, and by using C-14 fixation values multiplied by f (proportion of new production) from various sources: from N-15 data, from a C-14 fixation-versus-f relationship, or from a nitrate-versus-f relationship. The ratios of directly measured nitrate and carbon uptake and the ratios of nitrate to nitrate plus ammonium uptake, i.e., values of f, agree well; values of f calculated from carbon uptake or from nitrate concentration are overestimates for the equatorial upwelling region. Carbon-to-nitrogen uptake ratios measured with C-14 and N-15, respectively, approximate the Redfield molar ratio, 6.6 C:N. The overall mean value of f (0.17) helps confirm the view that the low primary production in the enriched eastern equatorial Pacific is due to failure of the nitrate-uptake system.

Dugdale, Richard C.↗

Streamlining Payload Integration

Payload integration onto space transport vehicles and the International Space Station (ISS) is a complex process. Yet, cargo transport is the sole reason for any space mission, be it for ferrying humans, science, or hardware. As the largest such effort in history, the ISS offers a wide variety of payload experience. However, for any payload to reach the Space Station under the current process, Payload Developers face a list of daunting tasks that go well beyond just designing the payload to the constraints of the transport vehicle and its stowage topology. Payload customers are required to prove their payload s functionality, structural integrity, and safe integration - including under less than nominal situations. They must also plan for or provide training, procedures, hardware labeling, ground support, and communications. In addition, they must deal with negotiating shared consumables, integrating software, obtaining video, and coordinating the return of data and hardware. All the while, they must meet export laws, launch schedules, budget limits, and the consensus of more than 12 panel and board reviews. Despite the cost and infrastructure overhead, payload proposals have increased. Just in the span from FY08 to FY09, the NASA Payload Space Station Support Office budget rose from $78M to $96M in attempt to manage the growing manifest, but the potential number of payloads still exceeds available Payload Integration Management manpower. The growth has also increased management difficulties due to the fact that payloads are more frequently added to a flight schedule late in the flow. The current standard ISS template for payload integration from concept to payload turn-over is 36 months, or 18 months if the payload already has a preliminary design. Customers are increasingly requiring a turn-around of 3 to 6-months to meet market needs. The following paper suggests options for streamlining the current payload integration process in order to meet customer schedule needs and reduce costs for both the integration support teams and the developers, without reducing quality or compromising safety. Issues for the key integration areas of planning, training, verification, and safety are presented in a Root-Cause Analysis study, with plausible solutions provided that involve technology and tools already available to the ISS community. Although based upon the ISS process, the payload integration techniques outlined herein also offer an integration template for any space transport endeavor.

Lufkin, Susan N.↗

Demonstration of Aerosol Property Profiling by Multiwavelength Lidar Under Varying Relative Humidity Conditions

During the months of July-August 2007 NASA conducted a research campaign called the Tropical Composition, Clouds and Climate Coupling (TC4) experiment. Vertical profiles of ozone were measured daily using an instrument known as an ozonesonde, which is attached to a weather balloon and launch to altitudes in excess of 30 km. These ozone profiles were measured over coastal Las Tablas, Panama (7.8N, 80W) and several times per week at Alajuela, Costa Rica (ION, 84W). Meteorological systems in the form of waves, detected most prominently in 100-300 in thick ozone layer in the tropical tropopause layer, occurred in 50% (Las Tablas) and 40% (Alajuela) of the soundings. These layers, associated with vertical displacements and classified as gravity waves ("GW," possibly Kelvin waves), occur with similar stricture and frequency over the Paramaribo (5.8N, 55W) and San Cristobal (0.925, 90W) sites of the Southern Hemisphere Additional Ozonesondes (SHADOZ) network. The gravity wave labeled layers in individual soundings correspond to cloud outflow as indicated by the tracers measured from the NASA DC-8 and other aircraft data, confirming convective initiation of equatorial waves. Layers representing quasi-horizontal displacements, referred to as Rossby waves, are robust features in soundings from 23 July to 5 August. The features associated with Rossby waves correspond to extra-tropical influence, possibly stratospheric, and sometimes to pollution transport. Comparison of Las Tablas and Alajuela ozone budgets with 1999-2007 Paramaribo and San Cristobal soundings shows that TC4 is typical of climatology for the equatorial Americas. Overall during TC4, convection and associated meteorological waves appear to dominate ozone transport in the tropical tropopause layer.

Veselovskii, I.↗

Convective and Wave Signatures in Ozone Profiles Over the Equatorial Americas: Views from TC4 (2007) and SHADOZ

During the months of July-August 2007 NASA conducted a research campaign called the Tropical Composition, Clouds and Climate Coupling (TC4) experiment. Vertical profiles of ozone were measured daily using an instrument known as an ozonesonde, which is attached to a weather balloon and launch to altitudes in excess of 30 km. These ozone profiles were measured over coastal Las Tablas, Panama (7.8N, 80W) and several times per week at Alajuela, Costa Rica (ION, 84W). Meteorological systems in the form of waves, detected most prominently in 100- 300 in thick ozone layer in the tropical tropopause layer, occurred in 50% (Las Tablas) and 40% (Alajuela) of the soundings. These layers, associated with vertical displacements and classified as gravity waves ("GW," possibly Kelvin waves), occur with similar stricture and frequency over the Paramaribo (5.8N, 55W) and San Cristobal (0.925, 90W) sites of the Southern Hemisphere Additional Ozonesondes (SHADOZ) network. The gravity wave labeled layers in individual soundings correspond to cloud outflow as indicated by the tracers measured from the NASA DC-8 and other aircraft data, confirming convective initiation of equatorial waves. Layers representing quasi-horizontal displacements, referred to as Rossby waves, are robust features in soundings from 23 July to 5 August. The features associated with Rossby waves correspond to extra-tropical influence, possibly stratospheric, and sometimes to pollution transport. Comparison of Las Tablas and Alajuela ozone budgets with 1999-2007 Paramaribo and San Cristobal soundings shows that TC4 is typical of climatology for the equatorial Americas. Overall during TC4, convection and associated meteorological waves appear to dominate ozone transport in the tropical tropopause layer.

Thompson, Anne M.↗

Feature Acquisition with Imbalanced Training Data

This work considers cost-sensitive feature acquisition that attempts to classify a candidate datapoint from incomplete information. In this task, an agent acquires features of the datapoint using one or more costly diagnostic tests, and eventually ascribes a classification label. A cost function describes both the penalties for feature acquisition, as well as misclassification errors. A common solution is a Cost Sensitive Decision Tree (CSDT), a branching sequence of tests with features acquired at interior decision points and class assignment at the leaves. CSDT's can incorporate a wide range of diagnostic tests and can reflect arbitrary cost structures. They are particularly useful for online applications due to their low computational overhead. In this innovation, CSDT's are applied to cost-sensitive feature acquisition where the goal is to recognize very rare or unique phenomena in real time. Example applications from this domain include four areas. In stream processing, one seeks unique events in a real time data stream that is too large to store. In fault protection, a system must adapt quickly to react to anticipated errors by triggering repair activities or follow- up diagnostics. With real-time sensor networks, one seeks to classify unique, new events as they occur. With observational sciences, a new generation of instrumentation seeks unique events through online analysis of large observational datasets. This work presents a solution based on transfer learning principles that permits principled CSDT learning while exploiting any prior knowledge of the designer to correct both between-class and withinclass imbalance. Training examples are adaptively reweighted based on a decomposition of the data attributes. The result is a new, nonparametric representation that matches the anticipated attribute distribution for the target events.

Thompson, David R.↗

Selection of Next Priority IMPACT Medical Conditions Based on Available Terrestrial and Spaceflight Data

BACKGROUND: As the era of exploration class missions begins, identification of medical conditions that may occur and require management becomes essential for the modeling of medical risk. To this end, NASA has developed IMPACT (Informing Mission Planning via Analysis of Complex Tradespaces), a suite of tools to assist in assessment of medical risk analysis. It has incorporated an initial list of the 120 conditions of highest concern, labeled the IMPACT condition list 1.0 (ICL 1.0). This abstract describes a method for prioritizing the 92 conditions included on the Proposed Future Conditions List (PFCL) for inclusion in future iterations of the ICL. OVERVIEW: To construct the Prioritized Proposed Future Conditions List (P-PFCL), each condition on the PFCL was scored as “low,” “medium,” or “high” on each of four variables: incidence, likelihood of significant task impairment, diagnostic and treatment complexity, and treatment futility. Qualitative assessment using clinical judgement was utilized to score complexity, futility, and likelihood of impairment. Incidence was assessed quantitatively using spaceflight data and/or analog populations where available then assigned a score using established cutoffs. Logarithmic numerical values were assigned to each category label. A Prioritization Score was generated for each condition by taking the product of incidence and likelihood of task impairment (risk) divided by the product of complexity and futility (difficulty of care), with higher values corresponding to higher priority for future inclusion in the ICL. DISCUSSION: The described methods allow for the generation of a ranked P-PFCL to act as a decision support tool for selection of the next generation of modeled medical conditions. Some of the conditions ranked highly on the P-PFCL include EVA-related upper and lower extremity sprain/strain, iron deficiency, delirium, and hypertension, among others. While this effort does not attempt to quantify the absolute risk associated with each condition, it does attempt to semi-quantitatively estimate the risk of each condition relative to the other possible conditions. This tool in concert with subject matter expert opinion could optimize the future use of limited resources thereby producing a more accurate medical risk model, which will be essential to the upcoming exploration class missions.

Michael Pohlen↗

Datasets for Custom-trained Machine-learning Interatomic Potentials: Nitric Acid Aqueous Solution

This dataset was generated using an iterative active learning strategy with the ArcaNN software package (https://github.com/arcann-chem/arcann_training) to train machine-learning interatomic potentials (MLIPs) for aqueous nitric acid. Each active-learning cycle consisted of three stages: (1) training, (2) exploration, and (3) labeling. The initial training set comprised approximately 800 randomly selected configurations from a previous study by Lewis et al. (https://doi.org/10.1021/jp205510q), which investigated nitric acid solutions at 2, 3, 4, and 5 mol/L. For all configurations, single-point calculations of atomic forces and total energies were performed at the quantum density functional theory BLYP-D2 and PBE-D3 levels of theory using the CP2K Quickstep module. Valence electrons were treated explicitly, while core electrons on all atoms were represented by norm-conserving Goedecker–Teter–Hutter (GTH) pseudopotentials. Long-range dispersion interactions were accounted for using Grimme dispersion corrections. Wave functions were expanded in a mixed Gaussian-and-plane-wave scheme using TZV2P-MOLOPT basis sets for all elements and an 800 Ry auxiliary plane-wave cutoff for the electron density. Self-consistent field convergence was accelerated using orbital transformation and Direct Inversion in the Iterative Subspace, with a convergence threshold of 10^{-6}. All single-point calculations were carried out in periodic orthorhombic cells whose dimensions match those of the molecular configurations sampled from earlier trajectories. The CELL_REF keyword in CP2K was used to define a fixed reference cell, ensuring consistency in the reference data used for MLIP training, particularly when cell fluctuations are present in NpT simulations. The resulting high-fidelity energies and forces constitute the ground-truth labels used to train the MLIPs contained in this dataset.

Dinpajooh, Mohammadhasan [Pacific Northwest Nation↗

Public Reference Data for Megawatt-Scale Hydrogen Electrolysis - Simulated Wave

The U.S. Department of Energy and the National Laboratory of the Rockies (NLR) demonstrate hydrogen electrolysis, hydrogen compression and storage, and variable hydrogen fuel cell power production using megawatt-scale equipment at NLR’s Flatirons Campus as part of the Advanced Research on Integrated Energy Systems (ARIES) initiative. This dataset represents part of that effort and is intended for academic, national laboratory, industrial, and other stakeholders to plan, design, and validate models of megawatt-scale hydrogen technologies and diverse energy infrastructure nationwide. These data provide a baseline for how existing hydrogen electrolysis technologies perform when coupled with various energy technologies. Future datasets will demonstrate how existing hydrogen fuel cell technologies can provide controllable, dispatchable, and variable power output for artificial intelligence (AI) data centers and other variable loads. This dataset entry describes hydrogen production using a single, simulated wave energy conversion device. The electrolyzer is a 1.25-MW proton exchange membrane type MC250 system manufactured by Nel Hydrogen. While the unit supports up to 2.5 MW of electrolysis, NLR only has a single 1.25-MW electrolysis stack. For the wave energy, NLR used a wave energy converter model from PacWave. These devices can be equipped with accumulators and pressure relief values to smooth the power output by storing and releasing hydraulic energy. Using a peak power output of 10 MW, the model created two 25-minute profiles: one with and one without the accumulators and pressure relief valves. To down select the profile data from the native resolution of 20 Hz to 1 Hz, NLR took the mean of every 20 data points. NLR experimented with two simulated wave energy power plants: one that peaks at 10 MW, and one that peaks at 5 MW. These profiles were scaled for the physical 1.25 MW electrolyzer by multiplying the original profiles by one eighth and one quarter, respectively. The first profile matches the capacity rating of eight of the 1.25 MW electrolyzers, while the second matches four electrolyzers. Finally, NLR experimented with two settings for the electrolyzer power supply minimum and maximum current ramp rates (gain and slew): 200 and 400 amperes per second. The simulated profiles were translated from power (kilowatts) to current (amperes) using a curve fit with calibration data and sent to the electrolyzer power supply at 1-Hz frequency. These datasets report relevant hydrogen balance-of-plant and system data, all captured at 1 Hz, including hydrogen mass production measured with an Emerson Coriolis flow meter. Each .zip file represents a single wave electrolysis experiment and is formatted as follows: {technology}-{accumulator?}_{number of 1.25 MW electrolyzers connected}-{electrolyzer ramp rate in amperes/second} For instance, “wavePacWave-Noacc_4-400.zip” represents the 25 minute-long experiment using the PacWave’s wave energy converter model, equipped with no accumulator, connected to four 1.25-MW electrolyzers with their power supplies set to a maximum current ramp rate (gain and slew) of 400 A/s. Each .zip folder contains the following files: A .csv file containing raw data. An .xlsx file explaining all the fields in the raw data. A .png plot showing the time series of hydrogen production in kilograms per hour, electrolysis power consumption, and input wave power. An experiment, labeled “characterization_200.zip”, demonstrates the MC250 electrolyzer steady-state response with 30 minute load steps for a total duration of 5 hours. Finally, a .csv file is provided with all wave profiles combined into one dataset labeled "combined_wave_experiments.csv". NLR also built an AI/machine-learning predictive model based on these datasets. The model ingests the electrolyzer current command in amperes, as well as various pressures and temperatures across the system, and predicts hydrogen output in kilograms per hour. The complete model can be found at https://huggingface.co/NatLabRockies/ptmelt-hydrogen-electrolysis.

08 HYDROGEN↗