Data processing and quality verification for improved photovoltaic performance and reliability analytics
Not Available
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Not Available
These data encompass 5-minute air monitoring and meteorological observations collected in the greater Houston, Texas metropolitan region, at nine (9) Continuous Ambient Monitoring Stations (CAMS) operated by the Texas Commission on Environmental Quality (TCEQ) between November 1, 2021 and October 31, 2022. The CAMS sites (CAMS 1, 8, 35, 45, 148, 403, 405, 410, and 1052) were chosen because their instrumentation includes measurements of PM2.5. These sites also provide continuous multi-parameter air-quality and meteorological measurements. Particulate matter (PM2.5, PM10) was sampled along with several trace gases, including ozone (O3), nitrogen oxides (NO, NO2, NOx), sulfur dioxide (SO2), and carbon monoxide (CO). The data set also contains standard surface meteorological parameters (temperature, humidity, pressure, wind speed, and wind direction). Several sites also include AutoGC-based measurements of volatile organic compounds (VOCs). Air monitoring instruments deployed at the selected sites comprise the following systems: BAM-1020 or TEOM (PM2.5), Thermo Scientific TEI 49i (O3), TEI 42i (NOx), and AutoGCs (VOCs). This data set is similar to the data included within the houairq5mX1.00 datastream, except for a few additional quality control steps. A systematic data cleaning and verification process was performed on the data set to ensure its quality and preparation for analysis. Removal of non-numeric status flags (e.g., [LIM], [QAS], [SPZ], [CAL], [PMA], [AQI], [SPN], [MAL]) was accomplished by employing rule-based string parsing to extract valid numerical values. Missing entries were set to -9999; however, invalid or anomalous values (e.g., 99999) were retained as originally reported by the TCEQ to preserve data provenance. The time sequence was verified for completeness, removal of duplicates, and uniformity at 5-minute intervals. Column labeling was standardized, and corresponding values were assessed for physical plausibility. All timestamps in the data set were reported in Coordinated Universal Time (UTC) as provided by the TCEQ. Further, the latitude and longitude coordinates were added for each CAMS site. A subset of the data (June 1–September 30, 2022) has been used in the following publication: Subba et al. 2025. “Implications of sea breeze circulations on boundary layer aerosols in the southern coastal Texas region.” EGUsphere 2025: 1–49, https://doi.org/10.5194/egusphere-2025-2659.
This data package contains raw output from two Campbell Scientific CR1000 data loggers each connected to three UMS T4e tensiometer sensors in Manaus, Brazil. Tensiometers at each data logger location were installed to depths of 10 cm, 20 cm and 50 cm below the ground surface for the purpose of measuring soil matric potential at each depth. “Station 3” tensiometers were installed at approximately 2°36'31.68"S, 60°12'34.70"W. “Station 4” tensiometers were installed at approximately 2°36'32.46"S, 60°12'34.92"W. The files “tensio3.edit.202109.csv” and “tensio4.edit.202109.csv” contain raw data (with invalid data removed). A text editor will be required to view the raw files in .csv (comma-separated values) format. Data processing was limited to: 1) Editing column headers for clarity; 2) Applying a correction that converts voltage read by the data loggers to pressure units (kPa); 3) removing all data greater than 50 kPa and less than -115 kPa; 4) Adding columns entitled “Caution.10cm”, “Caution.20cm” and “Caution.50cm” that contains values of either “NA” or “Caution”, the latter of which is applied when data exceeds the manufacturer reported data validity range of less than -85 kPa to suggest use of these data with caution; 5) Adding columns entitled “Invalid.10cm”, “Invalid.20cm”, and “Invalid.50cm” that contains values of either “NA” or “Invalid”, the latter of which is applied when data accuracy is deemed compromised due to the introduction of air bubbles into the instrument. Data marked “Invalid” may still indicate general trends of wet/high matric potential and dry/low matric potential, but values are inaccurate and dry/low matric potential values are likely not low enough. Wet/high matric potential values are likely to be more accurate than dry/low matric potential values; 6) Data format changed from .dat to .csv. This research was performed as a part of the NGEE Tropics project, which aims to advance model predictions of tropical forest carbon cycle responses to a changing climate over the 21st Century.
Qubit performance is often reported in terms of a variety of single-value metrics, each providing a facet of the underlying noise mechanism limiting performance. However, the value of these metrics may drift over long timescales, and reporting a single number for qubit performance fails to account for the low-frequency noise processes that give rise to this drift. Here, in this work, we demonstrate how we can use the distribution of these values to validate or invalidate candidate noise models. We focus on the case of randomized benchmarking (RB), where typically a single error rate is reported but this error rate can drift over time when multiple passes of RB are performed. We show that using a statistical test as simple as the Kolmogorov-Smirnov statistic on the distribution of RB error rates can be used to rule out noise models, assuming the experiment is performed over a long enough time interval to capture relevant low frequency noise. With confidence in a noise model, we show how care must be exercised when performing error attribution using the distribution of drifting RB error rate.
The validation of PU-MET-FAST-016 contributed to the centralized LANL benchmark repository currently under development. The revisions made the MCNP models statistically, significantly more similar to the benchmark models. The overall impact on the USL is negligible. The revisions to the PU-MET-FAST-016 models provide value to the Los Alamos Benchmark Suite without invalidating past Whisper results.
Accurate estimates of earthquake ground shaking rely on uncertain ground-motion models derived from limited instrumental recordings of historical earthquakes. A critical issue is that there is currently no method to empirically validate the resultant ground-motion estimates of these models at the timescale of rare, large earthquakes; this lack of validation causes great uncertainty in ground-motion estimates. Here, we address this issue and validate ground-motion estimates for southern California utilizing the unexceeded ground motions recorded by 20 precariously balanced rocks. We used cosmogenic 10Be exposure dating to model the age of the precariously balanced rocks, which ranged from ca. 1 ka to ca. 50 ka, and calculated their probability of toppling at different ground-motion levels. With this rock data, we then validated the earthquake ground motions estimated by the Uniform California Earthquake Rupture Forecast, Version 3 (UCERF3) seismic-source characterization and the Next Generation Attenuation (NGA)-West2 ground-motion models. We found that no ground-motion model estimated levels of earthquake ground shaking consistent with the observed continued existence of all 20 precariously balanced rocks. The ground-motion model I14 estimated ground-motion levels that were inconsistent with the most rocks; therefore, I14 was invalidated and removed. At a 2475 year mean return period, the removal of this invalid ground-motion model resulted in a 2−7% reduction in the mean and a 10−36% reduction in the 5th−95th fractile uncertainty of the ground-motion estimates. Our findings demonstrate the value of empirical data from precariously balanced rocks as a validation tool for removing invalid ground-motion models and, in turn, reducing the uncertainty in earthquake ground-motion estimates.
Sequences 8 and 9: Downwind Sonics (F,P) and Downwind Sonics Parked (P) This test sequence used an upwind, rigid turbine with a 0° cone angle. The wind speed ranged from 5 m/s to 25 m/s. Yaw angles of 0° to 60° were achieved. The blade tip pitch was 3°. The rotor rotated at 72 RPM during Sequence 8, but it was parked during Sequence 9. Blade pressure measurements were collected. The five-hole probes were removed and the plugs were installed. Plastic tape 0.03-mm-thick was used to smooth the interface between the plugs and the blade. The teeter dampers were replaced with rigid links, and these two channels were flagged as not applicable by setting the measured values in the data file to –99999.99 Nm. The teeter link load cell was pre-tensioned to 40,000 N. During post-processing, the probe channels were set to read -99999.99. Sonic anemometers were mounted on a strut downwind of the turbine. The strut was mounted to the T-frame, which was rotated to align the anemometers aft of the 9% and 49% radius locations at hub height. Because of this configuration, the tunnel balance data are considered invalid. Sequence 9 was designed to compare the downwind sonic anemometer readings with the upwind sonic anemometers without interference from the turbine. The rotor was parked with the instrumented blade at 0° azimuth. All pressure measurements obtained in Sequence 9 are invalid because sufficient time for temperature stabilization did not occur, thus all associated data values were flagged as not applicable by setting the measured values in the data file to 0.0000 Pa. This test is further described in Appendix G.
Sequences 8 and 9: Downwind Sonics (F,P) and Downwind Sonics Parked (P) This test sequence used an upwind, rigid turbine with a 0° cone angle. The wind speed ranged from 5 m/s to 25 m/s. Yaw angles of 0° to 60° were achieved. The blade tip pitch was 3°. The rotor rotated at 72 RPM during Sequence 8, but it was parked during Sequence 9. Blade pressure measurements were collected. The five-hole probes were removed and the plugs were installed. Plastic tape 0.03-mm-thick was used to smooth the interface between the plugs and the blade. The teeter dampers were replaced with rigid links, and these two channels were flagged as not applicable by setting the measured values in the data file to –99999.99 Nm. The teeter link load cell was pre-tensioned to 40,000 N. During post-processing, the probe channels were set to read -99999.99. Sonic anemometers were mounted on a strut downwind of the turbine. The strut was mounted to the T-frame, which was rotated to align the anemometers aft of the 9% and 49% radius locations at hub height. Because of this configuration, the tunnel balance data are considered invalid. Sequence 9 was designed to compare the downwind sonic anemometer readings with the upwind sonic anemometers without interference from the turbine. The rotor was parked with the instrumented blade at 0° azimuth. All pressure measurements obtained in Sequence 9 are invalid because sufficient time for temperature stabilization did not occur, thus all associated data values were flagged as not applicable by setting the measured values in the data file to 0.0000 Pa. This test is further described in Appendix G.
We investigate the basis for the Paradox where the Urbach-slope disorder parameter disagrees with the Raman width in amorphous diamond-like carbon (DLC) materials (lower Urbach-Eo yet greater Raman width). We examined the bandgap and Urbach-slope measurement issues. While significant errors are identified, ultimately these cannot resolve the Paradox. This resolution involves large changes in shape of DLC's absorption band resulting from large energy shifts. To solidify the understanding of band energy shifts, we examined their properties using a-Si:H, well known for its agreement between Raman and Urbach-slope disorder parameters. Using Tauc-Lorentz (T-L) dispersion functions, the Urbach method works for a-Si:H because small changes in shape occur, as measured by (i) peak position, (ii) peak width, (iii) peak amplitude, and (iv) and gap energies. Moreover, for a-Si:H, a 5th condition (v) there is no other feature below the principal band to interfere with its fall. All 5 of these conditions fail for DLCs. Because T-L dispersions accommodate asymmetry, we traced the Paradox to a steepening of the band shape as it pushes toward lower energy. This steepening is unrelated to disorder, but to a boundary value problem where energy cannot be negative. We conclude that Urbach analysis is invalid for DLC materials, but the T-L width is a good measure instead. We also examined another disorder measure, where π bonds should be considered as disorder relative to the σ bonds in an idealized “amorphous diamond” sp 3 network: superconvergence can be used to quantify these π transitions. Finally, further molecular orbital calculations are needed to properly interpret the changes in T-L widths.
Severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) and its variants are a continuous threat to human life. An urgent need remains for simple and fast tests that reliably detect active infections with SARS-CoV-2 and its variants in the early stage of infection. Here we introduce a simple and rapid activity-based diagnostic (ABDx) test that identifies SARS-CoV-2 infections by measuring the activity of a viral enzyme, Papain-Like protease (PLpro). The test system consists of a peptide that fluoresces when cleaved by SARS PLpro that is active in crude, unprocessed lysates from human tongue scrapes and saliva. Test results are obtained in 30 minutes or less using widely available fluorescence plate readers, or a battery-operated portable instrument for on-site testing. Proof-of-concept was obtained in a study on clinical specimens collected from patients with COVID-19 like symptoms who tested positive (n = 10) or negative (n = 10) with LIAT RT-PCR using nasal mid turbinate swabs. When saliva from these patients was tested with in-house endpoint RT-PCR, 17 were positive and only 5 specimens were negative, of which 2 became positive when tested 5 days later. PLpro activity correlated in 17 of these cases (3 out of 3 negatives and 14 out of 16 positives, with one invalid specimen). Despite the small number of samples, the agreement was significant (p value = 0.01). Two false negatives were detected, one from a sample with a late Ct value of 35 in diagnostic RT-PCR, indicating that an active infection was no longer present. The PLpro assay is easily scalable and expected to detect all viable SARS-CoV-2 variants, making it attractive as a screening and surveillance tool. Additionally, we show feasibility of the platform as a new homogeneous phenotypic assay for rapid screening of SARS-CoV-2 antiviral drugs and neutralizing antibodies.
Reconstructing complex, high-dimensional global fields from limited data points is a challenge across various scientific and industrial domains. This is particularly important for recovering spatio-temporal fields using sensor data from, for example, laboratory-based scientific experiments, weather forecasting, or drone surveys. Given the prohibitive costs of specialized sensors and the inaccessibility of
certain regions of the domain, achieving full field coverage is typically not feasible. Therefore, the development of machine learning algorithms trained to reconstruct fields given a limited dataset is of critical importance. In this study, we introduce a general
approach that employs moving sensors to enhance data exploitation during the training of an attention based neural network, thereby improving field reconstruction. The training of sensor locations is accomplished using an end-to-end workflow, ensuring
differentiability in the interpolation of field values associated to the sensors, and is simple to implement using differentiable programming. Additionally, we have incorporated a correction mechanism to prevent sensors from entering invalid regions within the domain. We evaluated our method using two distinct datasets; the results show that our approach enhances learning, as evidenced by improved test scores.
The Active Thermochemical Tables approach produces the enthalpy of formation of gas phase boron atom: Δ f H° 298 (B (g) ) = 570.48±0.61 kJ/mol and Δ f H° 0 (B (g) ) = 565.38±0.61 kJ/mol. This is about 5 kJ/mol higher and nearly an order of magnitude more accurate than the CODATA value. Here, while the ATcT value is in excellent agreement with the revisions proposed by Bauschlicher, Martin, and Taylor [J. Phys. Chem. A 103 (1999) 7715] and by Karton and Martin [J. Phys. Chem. A 111 (2007) 5936], it invalidates several earlier theoretical revisions that are too high by up to 5 kJ/mol.
Representing energy-limited resources in power system probabilistic resource adequacy assessment introduces new considerations that invalidate classical modeling assumptions. In particular, such resources have multi-period operating objectives and constraints that in real systems are addressed via a sequence of rolling intertemporal optimizations. Ideally, adequacy models would develop dispatch decisions by solving a similar sequence of problems, but this approach has historically been too computationally intensive for practical use in Monte Carlo simulations, with studies making use of simplifying approximations instead. These simplifications have the potential to distort the assessed value of energy-limited resources on the system.This work describes three classes of storage dispatch assumptions in current use and discusses their theoretical differences. It then provides an empirical analysis of their differences on test systems with different levels of storage, assessing the potential for a study's modeling assumptions to influence the perceived contribution of energy-limited resources.
Poly(vinylidene fluoride) (PVDF) and its random copolymers exhibit the most distinctive ferroelectric properties; however, their spontaneous polarization (60–105 mC m -2 ) is still inferior to those (>200 mC m -2 ) of the ceramic counterparts. In this work, we report an unprecedented spontaneous polarization ( P s = 140 mC m -2 ) for a highly poled biaxially oriented PVDF (BOPVDF) film, which contains a pure β crystalline phase. Given the crystallinity of ~0.52, the P s for the β phase ( P s,β ) is calculated to be 279 mC m -2 , if a simple two-phase model of semicrystalline polymers is assumed. This high P s,β is invalid, because the theoretical limit of P s,β is 185 mC m -2 , as calculated by density functional theory. To explain such a high P s for the poled BOPVDF, a third component in the amorphous phase must participate in the ferroelectric switching to contribute to the P s . Namely, an oriented amorphous fraction (OAF) links the lamellar crystal and the mobile amorphous fraction. From the hysteresis loop study, the OAF content was determined to be ~0.28, more than 50% of the amorphous phase. Because of the high polarizability of the OAFs, the dielectric constant of the poled BOPVDF reached nearly twice the value of conventional PVDF. Overall, the fundamental knowledge obtained from this study will provide a solid foundation for the future development of PVDF-based high performance electroactive polymers for wearable electronics and soft robotic applications.
Understanding the statistics of fluctuation driven flows in the boundary layer of magnetically confined plasmas is desired to accurately model the lifetime of the vacuum vessel components. Mirror Langmuir probes (MLPs) are a novel diagnostic that uniquely allow us to sample the plasma parameters on a time scale shorter than the characteristic time scale of their fluctuations. Sudden large-amplitude fluctuations in the plasma degrade the precision and accuracy of the plasma parameters reported by MLPs for cases in which the probe bias range is of insufficient amplitude. While some data samples can readily be classified as valid and invalid, we find that such a classification may be ambiguous for up to 40% of data sampled for the plasma parameters and bias voltages considered in this study. In this contribution, we employ an autoencoder (AE) to learn a low-dimensional representation of valid data samples. By definition, the coordinates in this space are the features that mostly characterize valid data. Ambiguous data samples are classified in this space using standard classifiers for vectorial data. In this way, we avoid defining complicated threshold rules to identify outliers, which require strong assumptions and introduce biases in the analysis. By removing the outliers that are identified in the latent low-dimensional space of the AE, we find that the average conductive and convective radial heat fluxes are between approximately 5% and 15% lower as when removing outliers identified by threshold values. For contributions to the radial heat flux due to triple correlations, the difference is up to 40%.
The H0LiCOW collaboration inferred via strong gravitational lensing time delays a Hubble constant value of H0 = 73.3−1.8+1.7 km s−1 Mpc−1, describing deflector mass density profiles by either a power-law or stars (constant mass-to-light ratio) plus standard dark matter halos. The mass-sheet transform (MST) that leaves the lensing observables unchanged is considered the dominant source of residual uncertainty in H0. We quantify any potential effect of the MST with a flexible family of mass models, which directly encodes it, and they are hence maximally degenerate with H0. Our calculation is based on a new hierarchical Bayesian approach in which the MST is only constrained by stellar kinematics. The approach is validated on mock lenses, which are generated from hydrodynamic simulations. We first applied the inference to the TDCOSMO sample of seven lenses, six of which are from H0LiCOW, and measured H0 = 74.5−6.1+5.6 km s−1 Mpc−1. Secondly, in order to further constrain the deflector mass density profiles, we added imaging and spectroscopy for a set of 33 strong gravitational lenses from the Sloan Lens ACS (SLACS) sample. For nine of the 33 SLAC lenses, we used resolved kinematics to constrain the stellar anisotropy. From the joint hierarchical analysis of the TDCOSMO+SLACS sample, we measured H0 = 67.4−3.2+4.1 km s−1 Mpc−1. This measurement assumes that the TDCOSMO and SLACS galaxies are drawn from the same parent population. The blind H0LiCOW, TDCOSMO-only and TDCOSMO+SLACS analyses are in mutual statistical agreement. The TDCOSMO+SLACS analysis prefers marginally shallower mass profiles than H0LiCOW or TDCOSMO-only. Without relying on the form of the mass density profile used by H0LiCOW, we achieve a ∼5% measurement of H0. While our new hierarchical analysis does not statistically invalidate the mass profile assumptions by H0LiCOW – and thus the H0 measurement relying on them – it demonstrates the importance of understanding the mass density profile of elliptical galaxies. The uncertainties on H0 derived in this paper can be reduced by physical or observational priors on the form of the mass profile, or by additional data.Key words: gravitational lensing: strong / galaxies: general / galaxies: kinematics and dynamics / distance scale / cosmological parameters / cosmology: observations⋆ The full analysis is available at https://github.com/TDCOSMO/hierarchy_analysis_2020_public.