Search NASASearch

SEARCH · Search NASA

Results for “data statistics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Evaluating the Impacts of Autonomous Electric Vehicles Adoption on Vehicle Miles Traveled and CO2 Emissions

Autonomous electric vehicles (AEVs) can potentially revolutionize the transportation landscape, offering a safer, contact-free, easily accessible, and more eco-friendly mode of travel. Prior to the market uptake of AEVs, it is critical to understand the consumer segments that are most likely to adopt these vehicles. Beyond market adoption, it is also important to quantify the impact of AEVs on broader transportation systems and the environment, such as impacts on the annual vehicle miles traveled (VMT) and greenhouse gas (GHG) emissions. In this pilot study, using survey data, a statistical model correlating AEV adoption intention and socioeconomic and built environment attributes was estimated, and a sensitivity analysis was conducted to understand the importance of factors impacting AEV adoption. We found that the market segments range from early adopters who are wealthy, technologically savvy, and relatively young to non-adopters who are more cautious to new technologies. This is followed by a synthetic population microsimulation of market penetration for the San Francisco Bay Area. With five household vehicle replacement scenarios, we assessed the annual VMT and tailpipe carbon dioxide (CO2) emissions change associated with vehicle replacement. It is found that adopting AEVs can potentially reduce more than 5 megatons of CO2 yearly, which is approximately 30% of the total CO2 emitted by internal combustion engine (ICE) cars in the region.

33 ADVANCED PROPULSION SYSTEMS

A bi-level spatiotemporal clustering approach and its application to drought extraction

We present a novel flexible bi-level spatiotemporal clustering algorithm to extract events based on their intensity and spatiotemporal structures. Our algorithm consists of using (i) a novel space-time k-means clustering to obtain spatiotemporally coherent intensity clusters, and (ii) a density-based spatial clustering of applications with noise (DBSCAN) to spatiotemporally section the intensity clusters into individual events. We discuss the development of the algorithm, the selection, tuning and meaning of the parameters within each step, as well as its validation. Finally, we apply the algorithm to a spatiotemporal drought index, standardized vapor pressure deficit drought index (SVDI), over the continental United States (US) from 1980–2021 and show that it captures historical drought events over the continental United States and their spatiotemporal extents.

17 WIND ENERGY

ResStock Measure Documentation: Residential Two-Stage Geothermal Heat Pump (4.0 COP, 20.5 EER) With Envelope Improvements and Advanced Air Sealing

The goal of this work is to develop energy efficiency, demand flexibility, and other retrofit end-use load shapes (electricity, gas, propane, or fuel oil) that cover a majority of the high-impact, market-ready (or nearly market-ready) measures. "Measures" refers to retrofits that can be applied to buildings during modeling. An "end-use savings shape" is the difference in energy consumption between a baseline building and a building with an energy efficiency, demand flexibility, or other retrofit measure applied. It results in a time-series profile that is broken down by end use and fuel (electricity or on-site gas, propane, or fuel oil use) at each time step. ResStock is a highly granular, physics-based, bottom-up model that uses multiple data sources, statistical sampling methods, and advanced building energy simulations to estimate the annual subhourly energy consumption of the residential building stock across the United States. The baseline model intends to represent the U.S. residential building stock as it existed in 2018. Technical documentation for the inputs and assumptions in the baseline building stock model is available in Reyna et al. (2025). Calibration and validation of the baseline model results are available in the final technical report of the End-Use Load Profiles project (Wilson, et al. 2022). This document focuses on a single end-use savings shape measure: Residential Two-Stage Geothermal Heat Pump (GHP) (4.0 COP, 20.5 EER) With Envelope Improvements. This measure combines a two-stage GHP with envelope improvements as a single package. As this package is a combination of two other measures, this document focused on documenting the results associated with this combination of technologies, with individual measure documents for two-stage GHPs and envelope improvements providing the information on the details of these measures. When the two technologies are combined, envelope improvements can modestly reduce energy consumption by a further 10%-15%, but also reduce the required size of the ground heat exchanger and heat pump by approximately 33% on average across all sites. The cost of installing envelope improvements in these homes is likely to be more than paid for by the reduction in equipment and drilling costs in these buildings for the majority of the stock.

15 GEOTHERMAL ENERGY

ResStock Measure Documentation: Residential Single-Stage Geothermal Heat Pump (3.8 COP, 18.6 EER)

The goal of this work is to develop energy efficiency, demand flexibility, and other retrofit end-use load shapes (electricity, gas, propane, or fuel oil) that cover a majority of the high-impact, market-ready (or nearly market-ready) measures. "Measures" refers to retrofits that can be applied to buildings during modeling. An "end-use savings shape" is the difference in energy consumption between a baseline building and a building with an energy efficiency, demand flexibility, or other retrofit measure applied. It results in a time-series profile that is broken down by end use and fuel (electricity or on-site gas, propane, or fuel oil use) at each time step. ResStock is a highly granular, physics-based, bottom-up model that uses multiple data sources, statistical sampling methods, and advanced building energy simulations to estimate the annual subhourly energy consumption of the residential building stock across the United States. The baseline model intends to represent the U.S. residential building stock as it existed in 2018. Technical documentation for the inputs and assumptions in the baseline building stock model is available in Reyna et al. (2025). Calibration and validation of the baseline model results are available in the final technical report of the End-Use Load Profiles project (Wilson et al. 2022). This documentation focuses on a single end-use savings shape measure: Residential Single-Stage Geothermal Heat Pump (GHP).?Single-stage GHPs are able to reduce energy consumption by 31% for the entire stock. Additional results provided below detail how savings changes for sections of the housing stock with different base heating fuel and in different climate zones, as well as the savings potential by state for both heating and cooling. Utility bills and electric panel impacts are also shown and discussed.

15 GEOTHERMAL ENERGY

ResStock Measure Documentation: Residential Variable-Speed Geothermal Heat Pump (4.4 COP, 30.9 EER)

The goal of this work is to develop energy efficiency, demand flexibility, and other retrofit end-use load shapes (electricity, gas, propane, or fuel oil) that cover a majority of the high-impact, market-ready (or nearly market-ready) measures. "Measures" refers to retrofits that can be applied to buildings during modeling. An "end-use savings shape" is the difference in energy consumption between a baseline building and a building with an energy efficiency, demand flexibility, or other retrofit measure applied. It results in a time-series profile that is broken down by end use and fuel (electricity or on-site gas, propane, or fuel oil use) at each time step. ResStock (TM) is a highly granular, physics-based, bottom-up model that uses multiple data sources, statistical sampling methods, and advanced building energy simulations to estimate the annual subhourly energy consumption of the residential building stock across the United States. The baseline model intends to represent the U.S. residential building stock as it existed in 2018. Technical documentation for the inputs and assumptions in the baseline building stock model is available in Reyna et al. (2025). Calibration and validation of the baseline model results are available in the final technical report of the End-Use Load Profiles project (Wilson et al. 2022). This documentation focuses on a single end-use savings shape measure: Residential Variable-Speed Geothermal Heat Pump (GHP). This document provides the relevant new modeling information for variable-speed systems not previously covered in either the single-stage or two-stage documents. Variable-speed GHPs represent the most efficient option available for this technology: They provide the most savings, with up to 46% for the applicable portion of the housing stock, compared to 31% for less efficient single-stage GHPs. Additional results shown here detail how the savings change for sections of the housing stock with different base heating fuels and in different climate zones, and they show the savings potential by state for both heating and cooling. Utility bills and electric panel impacts are also shown and discussed.

15 GEOTHERMAL ENERGY

ResStock Measure Documentation: Residential Two-Stage Geothermal Heat Pump (4.0 COP, 20.5 EER)

The goal of this work is to develop energy efficiency, demand flexibility, and other retrofit end-use load shapes (electricity, gas, propane, or fuel oil) that cover a majority of the high-impact, market-ready (or nearly market-ready) measures. "Measures" refers to retrofits that can be applied to buildings during modeling. An "end-use savings shape" is the difference in energy consumption between a baseline building and a building with an energy efficiency, demand flexibility, or other retrofit measure applied. It results in a time-series profile that is broken down by end use and fuel (electricity or on-site gas, propane, or fuel oil use) at each time step. ResStock is a highly granular, physics-based, bottom-up model that uses multiple data sources, statistical sampling methods, and advanced building energy simulations to estimate the annual subhourly energy consumption of the residential building stock across the United States. The baseline model intends to represent the U.S. residential building stock as it existed in 2018. Technical documentation for the inputs and assumptions in the baseline building stock model is available in Reyna et al. (2025). Calibration and validation of the baseline model results are available in the final technical report of the End-Use Load Profiles project (Wilson et al. 2022). This document focuses on a single end-use savings shape measure: Residential Two-Stage Geothermal Heat Pump (4.0 COP, 20.5 EER). This document builds on details established in the single-stage document (Maguire et al. 2025) to detail differences in the approach to modeling this higher efficiency, but more commonly deployed, type of geothermal heat pump. Specific EnergyPlus objects and product specific curves used are highlighted along with showing the results of this measure compared to the baseline and single-speed geothermal heat pumps. Two-speed geothermal heat pumps are able to save even more energy and on utility bills than single-speed products, albeit at the expense of a higher first cost.

15 GEOTHERMAL ENERGY

System Engineers and Decisions: It?s All about Knowledge

In order to guarantee that a system meets adequate levels of reliability and availability, system performances are continuously monitored and analyzed thanks to the technological advancements driving the Industry 4.0 revolution. An Industry 4.0 approach is typically based on advanced statistical, big data mining, machine learning, and internet-of-things methods designed to detect anomalies in the behavior of system, detect the most likely failure modes, and provide indications to system engineers on when maintenance activities should be performed before system performance are deemed unacceptable (which can be generated by diagnostic and prognostic methods). However, these analyses, which are designed to automatize and increase the efficacy of the system maintenance program, require large amount of data which can come in various forms: numeric, textual, images, sounds etc. Such data constitutes the historic knowledge benchmark to track system performances and support system engineer decisions. Here we claim that data is not sufficient to support this kind of analyses when applied to systems characterized by complex architectures and behaviors. Robust system engineer decisions require the ability to understand the system operational context that lies behind the observed data elements. In this respect, system models are in fact necessary to “put data in context” and capture relationships between data elements. Industry 4.0 methods require in fact contextual knowledge as a basis upon which hypotheses can be generated and assumptions tested. In our view, for complex systems, model-based system engineering (MBSE) models can afford this contextual knowledge, as they are typically used to describe systems architecture and dynamic behaviors. System knowledge is here intended as the blending of collected data and system architecture which takes the form of a “knowledge graph”. A knowledge graph is a database which consists of a large set of nodes (in our case an entity can be either a data or an MBSE element) which are linked to each other. The types of nodes and links follow a pre-defined topology, sometimes also refers as an ontology, that is designed to fit the actual decisions that needs to be performed. We show here how a knowledge graph can be defined to support system engineer maintenance decisions and how the same graph can be built based on system MBSE models and pre-processed data from numeric (through anomaly detections and diagnostic methods) and textual elements (through technical language processing TLP).

97 - MATHEMATICS AND COMPUTING

Selection Algorithm Improvement for MicroBooNE

Data selection is an extremely important part of data analysis for any experiment. Finding a physics result is often the result of sifting through a massive amount of data, keeping data that we believe to be signal and throwing out data we do not. This process is called data selection. Creating a selection algorithm is an intensive process that must balance keeping enough data to have statistics and maximizing the signal purity of that data. In this study, we used three different reconstruction tools, Pandora, WireCell, and LANTERN, for the MicroBooNE experiment in conjunction to improve the selection algorithm for analysis. For the case of this study, we look into the charged current N proton 0 pions (CCNp0$\pi$) interaction channel. This is the dominant channel for the Short Baseline Neutrino (SBN) program and is expected to be a large contributor to the Deep Underground Neutrino Experiment (DUNE). We first investigated each of the three tools to find out more about their strengths and weaknesses as reconstructions. We then put together a direct comparison of the three methods to find which method or combination of methods would return the best result for us. While the study is ongoing, we have learned a lot about data selection for the experiment and the differences between the reconstruction tools.

Dillon, Brayden [Michigan State U.]

BASIN-3D Data Integration for Selected ARM Data Field Campaign Report

The purpose of this data services request was to demonstrate integration of the Atmospheric Radiation Measurement (ARM) User Facility’s “met” datastreams with time series data from other earth science data sources using the BASIN-3D data synthesis software tool. BASIN-3D is an open-source Python library that enables researchers to integrate data across configured public and private data sources. It provides a common query language for researchers to request measurement locations and time series data based on specified locations, variables, time period, statistics, aggregation, and data quality. BASIN-3D acquires the data that match the query from each configured data source and translates the results into harmonized vocabularies, thus reducing researchers' data-wrangling effort. In addition, because the queries are executed on demand, researchers can easily regenerate their synthesized data sets as new data and/or data updates become available, eliminating one-off data products. BASIN-3D can output data using a variety of different data structures for end-user applications including Python pandas data frames and hdf5 output formats.

54 ENVIRONMENTAL SCIENCES

Data Selection Improvement For MicroBooNE

Data selection is an extremely important part of data analysis for any experiment. Finding a physics result is often the result of sifting through a massive amount of data, keeping data that we believe to be signal, and throwing out data we do not. This process is called data selection. Creating a selection algorithm is an intensive process that must balance keeping enough data to have statistics and maximizing the signal purity of that data. We also need to choose the right reconstruction method, a tool to take raw data from the detector and convert it into physics results. In this study, we used three different reconstruction tools, Pandora, WireCell, and LANTERN, for the MicroBooNE experiment in conjunction to improve the selection algorithm for analysis. For the case of this study, we look into the charged current N proton 0 pions (CCNp0$\pi$) interaction channel. This is the dominant channel for the Short Baseline Neutrino (SBN) program and is expected to be a large contributor to the Deep Underground Neutrino Experiment (DUNE). We first investigated each of the three tools to find out more about their strengths and weaknesses as reconstructions, and compared them to the truth information directly from the MicroBooNE simulation pipeline. We then put together a direct comparison of the three methods to find which method or combination of methods would return the best result for us. While the study is ongoing, we have learned a lot about data selection for the experiment and the differences between the reconstruction tools.

Dillon, Brayden [Fermilab]

A Parameter-masked Mock Data Challenge for Beyond-two-point Galaxy Clustering Statistics

The past few years have seen the emergence of a wide array of novel techniques for analyzing high-precision data from upcoming galaxy surveys, which aim to extend the statistical analysis of galaxy clustering data beyond the linear regime and the canonical two-point (2pt) statistics. We test and benchmark some of these new techniques in a community data challenge named “Beyond-2pt,” initiated during the Aspen 2022 Summer Program “Large-Scale Structure Cosmology beyond 2-Point Statistics,” whose first round of results we present here. The challenge data set consists of high-precision mock galaxy catalogs for clustering in real space, in redshift space, and on a light cone. Participants in the challenge have developed end-to-end pipelines to analyze mock catalogs and extract unknown (“masked”) cosmological parameters of the underlying ΛCDM models with their methods. The methods represented are density-split clustering, nearest neighbor statistics, BACCO power spectrum emulator, void statistics, LEFTfield field-level inference using effective field theory (EFT), and joint power spectrum and bispectrum analyses using both EFT and simulation-based inference. In this work, we review the results of the challenge, focusing on problems solved, lessons learned, and future research needed to perfect the emerging beyond-2pt approaches. The unbiased parameter recovery demonstrated in this challenge by multiple statistics and the associated modeling and inference frameworks supports the credibility of cosmology constraints from these methods. The challenge data set is publicly available, and we welcome future submissions from methods that are not yet represented.

Krause, Elisabeth [Univ. of Arizona, Tucson, AZ (U

The Statistical Spread of Transmission Outages on a Fast Protection Time Scale Based on Utility Data

When there is a fault, the protection system automatically removes one or more transmission lines on a fast time scale of less than one minute. The outaged lines form a pattern in the transmission network. We extract these patterns from utility outage data, determine some key statistics of these patterns, and then show how to generate new patterns consistent with these statistics. The generated patterns provide a new and easily feasible way to model the overall effect of the protection system at the scale of a large transmission system. This new data-driven generative modeling of protection is expected to contribute to simulations of disturbances in large grids so that they can better quantify the risk of blackouts. Analysis of the pattern sizes suggests an index that describes how much outages spread in the transmission network at the fast timescale.

Transmission

Statistical Analysis of Fermilab’s Safety Data

In the past ten months, since the Safety and Security pauses of May 31$^{st}$ and June 29$^{th}$ of 2023, there has been an increase in the number of Total Recordable Counts to the point that there are already more incidents in ten months that at any given year in the past decade. The purpose of this study is to try to determine if this increase could be explained as a statistical fluctuation. We will show that it is very unlikely (essentially 4$\sigma$ unlikely) that the increase is a fluctuation, and therefore there must be a cause. The Safety and Security pauses of the past May and June, and many events thereafter, have been impactful events at Fermilab, so we should feel an obligation to assess their benefit, or detriment, to the safety of our people.

99 GENERAL AND MISCELLANEOUS

A road map to cosmological parameter analysis with third-order shear statistics: III. Efficient estimation of third-order shear correlation functions and an application to the KiDS-1000 data

Context. Third-order lensing statistics contain a wealth of cosmological information that is not captured by second-order statistics. However, the computational effort it takes to estimate such statistics in forthcoming stage IV surveys is prohibitively expensive. Aims. We derive and validate an efficient estimation procedure for the three-point correlation function (3PCF) of polar fields such as weak lensing shear. We then use our approach to measure the shear 3PCF and the third-order aperture mass statistics on the KiDS-1000 survey. Methods We constructed an efficient estimator for third-order shear statistics that builds on the multipole decomposition of the 3PCF. We then validated our estimator on mock ellipticity catalogs obtained from N -body simulations. Finally, we applied our estimator to the KiDS-1000 data and presented a measurement of the third-order aperture statistics in a tomographic setup. Results. Our estimator provides a speedup of a factor of ∼100–1000 compared to the state-of-the-art estimation procedures. It is also able to provide accurate measurements for squeezed and folded triangle configurations without additional computational effort. We report a significant detection of tomographic third-order aperture mass statistics in the KiDS-1000 data (S/N = 6.69). Conclusions. Our estimator will make it computationally feasible to measure third-order shear statistics in forthcoming stage IV surveys. Furthermore, it can be used to construct empirical covariance matrices for such statistics.

Astronomy & Astrophysics

Data from: 'Abiotic influences on continuous conifer forest structure across a subalpine watershed'

This package archives the core data used for analysis and inference in 'Abiotic influences on continuous conifer forest structure across a subalpine watershed' (Worsham et al., 2025). All data were collected in the East River, Washington Gulch, Slate River, and Coal Creek watersheds of Colorado. In the paper, we quantified the relative influence of climate, topographic, edaphic, and geologic factors on conifer stand structure and composition, and their functional relationships, at the watershed scale. We used waveform LiDAR data to derive spatially continuous stand structure metrics. We fused these with a species-level classification map to estimate tree species abundance. We applied generalized additive and generalized boosted models to evaluate the covariability of structural and compositional metrics with abiotic variables. The package contains the essential products required for reproducing our analysis and the tables and figures reported in the publication. The products comprise four classes: (1) geospatial data, (2) tabular data used for inferential analysis, (3) tabular data describing analytical results and performance statistics, and (4) a data user guide. (1) includes discretized waveform LiDAR data, locations and attributes of individual tree crowns, sampling locations and domain boundaries, a canopy height model, and raster files of estimated forest structural and compositional metrics at 100 m grid scale. (2) includes all response and explanatory variable values applied in inferential models. Response variables include conifer forest stand density, basal area, 95th percentile height, quadratic mean diameter, and others. Explanatory variables include climatic water deficit, actual evapotranspiration, elevation, heat load, soil available water content, and others. (3) includes results of training and testing several individual tree detection (ITD) algorithms, as well as inferential modeling results. (4) is a PDF user guide for this data package, including detailed descriptions and data dictionaries for all files. The data package root contains 17 assets: 8 compressed tape archive (.tar.gz) files, 5 comma-separated values (.csv) files, 3 Geographic Tagged Image File Format (GeoTIFF) (.tif) files, and 1 Portable Document Format (.pdf) file. The compressed .tar.gz archives contain ESRI shapefiles (.shp) .tif, compressed LASer (.laz), and .csv files. The archives must first be decompressed using the widely distributed command-line software utility TAR. All other files, including constituent files within the .tar.gz archives, can be opened in the open-source R statistical computing environment. Alternatively, .csv files may also be read in any simple text editor software or Microsoft Excel. Geospatial files including .shp and .tif files can also be opened in GIS software, such as QGIS (open-source) or ESRI ArcGIS (proprietary). The .pdf Data User Guide can be read with Adobe Acrobat Reader or other compatible readers.

2018 NEON and 2025 CHESS Campaigns

Public Reference Data for Megawatt-Scale Hydrogen Electrolysis - NLR Historical Solar PV

The U.S. Department of Energy and National Laboratory of the Rockies (NLR) demonstrate hydrogen electrolysis from variable sources, hydrogen compression and storage, and hydrogen fuel cell power production using megawatt-scale equipment at NLR’s Flatirons Campus as part of the Advanced Research on Integrated Energy Systems (ARIES) research platform. This dataset represents part of that effort and is intended for academic, national laboratory, industrial, and other stakeholders to plan, design, and validate models of megawatt-scale hydrogen technologies and diverse energy infrastructure nationwide. These data provide a baseline for how existing hydrogen electrolysis technologies perform when coupled with various energy technologies. Future datasets will demonstrate how existing hydrogen fuel cell technologies can provide controllable, dispatchable, and variable power output for artificial intelligence data centers and other variable loads. This dataset entry describes the behavior of a 1.25-MW proton exchange membrane MC250 electrolyzer system, manufactured by Nel Hydrogen , [1] when fed historical data generated by the 430-kW, fixed-axis solar photovoltaic (PV) array located at NLR’s Flatirons Campus. (While the electrolyzer balance of plant supports up to 2.5 MW of electrolysis, NLR only has a single 1.25-MW electrolysis stack.) Solar PV power output data for the 2020 calendar year were categorized on a daily basis by total energy generation and standard deviation. Each day was then ranked by these metrics, and the 25th, 50th, and 100th percentiles were selected. The 75th percentile day did not exhibit sufficient variability to make for a valuable experiment. A similar process was used for the related historical wind dataset . [2] The historical days in 2020 that represented these percentiles are Dec. 19, March 29, and May 4, respectively. The entire solar day’s power profile was then fed through the MC250 electrolyzer. Due to its length, the 100th percentile day experiment was split into two parts, and the final 3 hours of the solar day were not captured. These final 3 hours contained no spikes or dips of interest and simply represented a slow decay of input solar power. Also, a single timestamp (13:13:47 on Jan. 14, 2026) was lost in the hydrogen system supervisory control and data acquisition. Finally, during the 25th percentile experiment (solar day Dec. 19, 2020) data recording was lost from 11:00:13 to 11:14:45. The roughly 15 minutes of the solar profile were rerun at the end of the experiment and spliced into this time slot during post-processing. The electrolysis system controls hydrogen production by varying direct current applied to the stack, from a maximum of 3,000 A to a minimum safe operation of 300 A, or 10%. Because the current–voltage characteristic changes as the stack ages and efficiency degrades, the actual minimum safe operating power changes over time. The historical solar profiles were translated from power (kilowatts) to current (amperes) using a curve fit with calibration data and sent to the electrolyzer power supply at 1-Hz frequency. For more details on the statistical analysis process, see the slide deck “Public Reference Data for Megawatt-Scale Hydrogen Electrolysis: NLR Historical Solar PV Analysis and Profile Generation” accessible with this data entry. These datasets report relevant hydrogen balance-of-plant and system data, all captured at 1 Hz, including hydrogen mass production measured with an Emerson Coriolis flow meter. Each .zip file represents a single solar PV electrolysis experiment and is formatted as: {technology}_{percentile}_{scaling factor} For instance, “solarPV-430kW_25_2x.zip” reports the experiment using the 25th percentile solar data from the historical 2020 solar PV dataset, scaled to 200%. Scaling factors were applied to the generated solar PV power output files to more closely match the 1.25-MW capacity of the electrolyzer. Each .zip folder contains the following files: A .csv file containing raw data. An .xlsx file explaining all the fields in the raw data. A .png plot showing the time series of hydrogen production, electrolysis power consumption, and solar power input. A PDF file detailing the historical solar data statistical analysis used to generate the solar profile. An experiment labeled “characterization_200.zip” demonstrates the MC250 electrolyzer steady-state response with 30-minute load steps for a total duration of 5 hours. Finally, a .csv file is provided with all experiments combined into one dataset labeled "combined_solarPV_experiments.csv". [1] nelhydrogen.com/product/mc-series-electrolyser . [2] data.nlr.gov/submissions/316 .

08 HYDROGEN

Algorithms for Non-Negative Matrix Factorization on Noisy Data With Negative Values

Non-negative matrix factorization (NMF) is a dimensionality reduction technique that has shown promise for analyzing noisy data, especially astronomical data. For these datasets, the observed data may contain negative values due to noise even when the true underlying physical signal is strictly positive. Prior NMF work has not treated negative data in a statistically consistent manner, which becomes problematic for low signal-to-noise data with many negative values. In this paper we present two algorithms, Shift-NMF and Nearly-NMF, that can handle both the noisiness of the input data and also any introduced negativity. Both of these algorithms use the negative data space without clipping or masking and recover non-negative signals without any introduced positive offset that occurs when clipping or masking negative data. We demonstrate this numerically on both simple and more realistic examples, and prove that both algorithms have monotonically decreasing update rules.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

Statistical Analysis of Imaging Laser Scan Data of an Exhaust Tunnel at the SRS

• The SRS H-Canyon Building is a critical facility under the responsibility of DOE-EM. • It includes an Air Exhaust Tunnel (HCAEX) that allows for ventilation of the process airflow. • Inspections are performed remotely because of hazards, e.g. radioactivity, debris, high airflow, and nitric acid vapors.

Wells, William Willie [Savannah River National Lab