Search NASA⌕ Search

SEARCH · Search NASA

Results for “data quality”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Methods for Incorporating Model Uncertainty into Exoplanet Atmospheric Analysis

A key goal of exoplanet spectroscopy is to measure atmospheric properties, such as abundances of chemical species, in order to connect them to our understanding of atmospheric physics and planet formation. In this new era of high-quality JWST data, it is paramount that these measurement methods are robust. When comparing atmospheric models to observations, multiple candidate models may produce reasonable fits to the data. Typically, conclusions are reached by selecting the best-performing model according to some metric. This ignores model uncertainty in favor of specific model assumptions, potentially leading to measured atmospheric properties that are overconfident and/or incorrect. In this paper, we compare three ensemble methods for addressing model uncertainty by combining posterior distributions from multiple analyses: Bayesian model averaging, a variant of Bayesian model averaging using leave-one-out predictive densities, and stacking of predictive distributions. We demonstrate these methods by fitting the Hubble Space Telescope (HST) + Spitzer transmission spectrum of the hot Jupiter HD 209458b using models with different cloud and haze prescriptions. All of our ensemble methods lead to uncertainties on retrieved parameters that are larger but more realistic and consistent with physical and chemical expectations. Since they have not typically accounted for model uncertainty, uncertainties of retrieved parameters from HST spectra have likely been underreported. We recommend stacking as the most robust model combination method. Our methods can be used to combine results from independent retrieval codes and from different models within one code. They are also widely applicable to other exoplanet analysis processes, such as combining results from different data reductions.

79 ASTRONOMY AND ASTROPHYSICS↗

Data from: A high-quality genome assembly of the tetraploid Teucrium chamaedrys unveils a recent whole genome duplication and a large biosynthetic gene cluster for diterpenoid metabolism

Teucrium is well known for making clerodane-type diterpenoids that are produced from the backbone kolavanyl diphosphate. In order to begin to elucidate some of the complex biosynthetic pathways of these medicinal compounds, we identified and functionally characterized several kolavanyl diphosphate synthases from T. chamaedrys . Along the way, we discovered the genome of this species to be one of the largest genomes published from the Lamiaceae family, to which it belongs. This tetraploid, 3 Gbp genome is especially rich in diterpene synthase genes, with 74 putative sequences identified.

biosynthetic gene cluster (BGC)↗

Monthly Quality-filtered Aggregation of NOAA Climate Data Record (CDR) of AVHRR Leaf Area Index (LAI) and Fraction of Absorbed Photosynthetically Active Radiation (FAPAR), Version 5

This dataset contains gridded monthly Leaf Area Index (LAI) derived from the daily NOAA Climate Data Record (CDR) of AVHRR Leaf Area Index (LAI) and Fraction of Absorbed Photosynthetically Active Radiation (FAPAR), Version 5. This data record spans from 1981 to 2018 using data from eight NOAA polar orbiting satellites: NOAA-7, -9, -11, -14, -16, -17, -18 and -19. The data are projected on a 0.05 degree x 0.05 degree global grid, as in the original CDR. The original CDR is one of the Land Surface CDR Version 5 products produced by the NASA Goddard Space Flight Center (GSFC) and the University of Maryland (UMD), which is accompanied by algorithm documentation, data flow diagram and source code for the NOAA CDR Program. This dataset is in the netCDF-4 file format following ACDD and CF Conventions. This dataset has applied quality assurance information to only include "OK" data from the original CDR in the monthly aggregation.

Vermote, Eric [NASA Goddard Space Flight Center (G↗

Monthly Quality-filtered Aggregation of NOAA Climate Data Record (CDR) of AVHRR (Version 5) and VIIRS (Version 1) Leaf Area Index (LAI) and Fraction of Absorbed Photosynthetically Active Radiation (FAPAR)

This dataset contains gridded monthly Leaf Area Index (LAI) derived from the daily NOAA Climate Data Record (CDR) of AVHRR (Version 5) and VIIRS (Version 1) Leaf Area Index (LAI) and Fraction of Absorbed Photosynthetically Active Radiation (FAPAR). This data record spans from 1981 to 2024 using data from NOAA polar orbiting satellites: NOAA-7, -9, -11, -14, -16, -17, -18, -19 and S-NPP. The data are projected on a 0.05 degree x 0.05 degree global grid, as in the original CDR. The original CDR is one of the Land Surface CDR products produced by the NASA Goddard Space Flight Center (GSFC) and the University of Maryland (UMD), which is accompanied by algorithm documentation, data flow diagram and source code for the NOAA CDR Program. This dataset is in the netCDF-4 file format following ACDD and CF Conventions. This dataset has applied quality assurance information to only include "OK" data from the original CDR in the monthly aggregation.

Vermote, Eric [NASA Goddard Space Flight Center (G↗

Photonuclear Reaction Types Map to ENDF Reaction Type Numbers MT

In the ENDF format, photonuclear cross secCons are stored in File = 3 under the NSUB = 0 sub-library. The type of reacCon data and the products resulCng from that reacCon are idenCfied by an integer number from 1 to 999, called MT number. However, the ENDF manual is wriPen with a strong focus on neutron induced reacCons and the meaning of MT numbers in the context of photonuclear reacCons can be at Cmes confusing and the assessment of evaluated photonuclear data may require summing various MT numbers to obtain the desired cross secCons to compare with experimental data. An important step in assessing the quality of evaluated data for photonuclear cross secCons is therefore to clarify the mapping between MT numbers and the nomenclature used to describe photonuclear reacCons in the literature. The intent of this document is to provide such a map.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

SITCOMTN-162: Testing the implementation of Metadetection and Cell-Based Coadds on Abell 360 LSSTComCam data

The purpose of this technote is to test the technical quality of LSSTComCam commissioning data, specifically the Rubin_SV_38_7 field, by utilizing cell-based coadds and Metadetection by measuring the tangential and cross weak lensing shear profiles of the massive cluster Abell 360 (called A360 throughout the technote). The process entails generating the cell-based coadds for Metadetection to run on, identifying and removing cluster member galaxies, applying quality cuts, calibrating the shear measurements, and validation. Cell-based coadds and Metadetection are both currently in the process of being implemented within the LSST Science Pipelines at the time of this technote. There is substantial technical value in attempting a difficult measurement prior to full implementation. Measuring the tangential shear around A360 will showcase the current abilities of these algorithms, as well as highlight where work is still needed. As seen from the resulting shear profile of A360, the cell-based coadds and Metadetection are able to work in tandem to produce a shear catalog and resulting reduced shear profile. This technote is one part of a series studying A360 in order to both stress test the commissioning camera and demonstrate the technical capabilities of the Vera Rubin Observatory. We study the quality of the PSF modeling and impact it can have on cluster WL in [Combet et al., 2025], implementation of cell-based coadds and subsequent use for Metadetect [Sheldon et al., 2023] in this technote, photometric calibration in (in prep), source selection and photometric redshifts in [Adari et al., 2025], use of Anacal [Li et al., 2024] to produce a cluster shear profile in [Li et al., 2025], and background subtraction in this field and Fornax in [Zhou et al., 2025].

79 ASTRONOMY AND ASTROPHYSICS↗

Roadmap for transforming heterogeneous catalysis with artificial intelligence

Artificial intelligence (AI) is poised to transform heterogeneous catalysis, opening avenues for catalytic materials discovery. By uncovering intricate patterns in high-dimensional data, AI has been reshaping our pursuit of sustainable catalytic processes across the energy, environmental and chemical sectors. This promise, however, hinges on overcoming fundamental barriers, including limitations in data availability and quality, challenges in the generalizability and interpretability of data-augmented decisions, and the persistent gap between in silico predictions and experiments. Furthermore, we outline a forward-looking roadmap for deeply integrating AI into heterogeneous catalysis with an AI-ready data ecosystem, multimodal foundation models, and ultimately autonomous laboratories to accelerate the development of next-generation catalytic technologies via AI-empowered human–machine collaboration.

Computational methods↗

Quality Assurance Program Plan for SFR Metallic Fuel Data Qualification

This document contains an evaluation of the applicability of the current Quality Assurance Standards from the American Society of Mechanical Engineers Standard NQA-1 (NQA-1) criteria and identifies and describes the quality assurance process(es) by which attributes of historical, analytical, and other data associated with sodium-cooled fast reactor [SFR] metallic fuel will be evaluated. This process is being instituted to facilitate validation of data to the extent that such data may be used to support future licensing efforts associated with advanced reactor designs. The initial data to be evaluated under this program were generated during the US Integral Fast Reactor program between 1984-1994, where the data include, but are not limited to, research and development data and associated documents, test plans and associated protocols, operations and test data, technical reports, and information associated with past United States Nuclear Regulatory Commission reviews of SFR designs. It is recognized that managing the data generated by large research and development projects presents a significant challenge for retaining data integrity and availability. American Society of Mechanical Engineers Standard NQA-1 (NQA-1) 2008/2009a provides appropriate requirements for this plan.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Effect of likelihood misspecification in Gaussian process-driven autonomous experimentation

In recent years, several groups have designed Autonomous Experiment (AE) models with the aim of using them as an alternative method for neutron scattering scanning. In an AE, Gaussian processes (GPs) are most frequently used due to their interpretability, their non-parametric nature, their universal approximation, and their closed-form predictive distribution. GPs have two key components, namely, the model for the likelihood of a neutron count knowing the underlying dynamic structure factor and the acquisition function. In this paper, we investigate the impact, on the quality of an AE, of the likelihood and acquisition function choices, in energy scans and (Q, ω) ones, with respect to the signal-over-noise ratio. While we hypothesized that the quality of GP predictions would decrease when the normal to Poisson likelihood approximation breaks down at low count rates, we found that the use of the correct Poisson likelihood does not improve the quality of the data collected, as well as yields very poor results in (Q, ω) scans at low count rates. In fact, the best results are obtained with a combination of normal likelihood, including the observation noise, and the change in variance acquisition function. In addition, we find that the performance, or quality of the predictive distribution, is a misleading measure of efficiency, that is, of the quality of the data collected.

Perryman, David Elliott [Inst. Laue-Langevin (ILL)↗

Size-resolved Eddy-Covariance Particle Flux Measurement during the TRACER Campaign (Final Report)

The main goal of the TRacking Aerosol Convection interactions ExpeRiment (TRACER) campaign was to study aerosol–cloud interactions during deep convection over the Houston area. This project deployed a suite of instrumentation with the aim to (1) quantify turbulent vertical particle fluxes during at DOE-ARM sites, including TRACER, (2) assess hygroscopic growth factors and hygroscopicity parameters of the material driving modal aerosol growth during new particle formation and growth events, (3) derive turbulent aerosol mass fluxes using co-located Doppler LIDAR measurements, and (4) create quality-controlled PI data products to support future research utilizing data collected during the TRACER campaign. This report summarized the main findings from the deployments at two DOE-ARM sites. Briefly, we found that new particle formation may occur aloft, in a residual layer, near the top of the boundary layer. Small grown particles appear later due to downward mixing with daytime turbulence. The species that are responsible for aerosol modal growth had hygroscopicity parameters varying between 0.05 and 0.34. These values systematically depended on the wind sector, suggesting that the chemical composition of the precursors differed. This work demonstrated that lidar retrievals of the elastic backscatter and Doppler velocity can be used to obtain surface number emissions of particles with a diameter greater than 0.53 µm. During TRACER, emission particle number fluxes peaked near ∼ 100 cm−2 s−1. Multiple quality-controlled PI data products that will support future TRACER related science were generated and made publically available.

54 ENVIRONMENTAL SCIENCES↗

Surrogate-driven design optimization with uncertainty constraints in Monte Carlo simulations

In multi-objective design tasks, the computational cost increases rapidly when high-fidelity simulations are used to evaluate objective functions. Surrogate models help mitigate this cost by approximating the simulation output, simplifying the design process. However, under high uncertainty, surrogate models trained on noisy data can produce inaccurate predictions, as their performance depends heavily on the quality of training data. This study investigates the impact of data uncertainty on two multi-objective design problems modelled using Monte Carlo transport simulations: a neutron moderator and an ion-to-neutron converter. For each, a grid search was performed using five different tally uncertainty levels to generate training data for neural network surrogate models. These models were then optimized using NSGA-III. The recovered Pareto-fronts were analyzed across uncertainty levels: in the moderator problem, normalized hypervolume dropped from 0.886 at 1.0% uncertainty to 0.748 at 10% uncertainty, while in the converter problem it remained near 0.50 for all cases. Average simulation times were also compared to evaluate the trade-off between accuracy and computational cost. Results show that the influence of simulation uncertainty is strongly problem-dependent. In the neutron moderator case, higher uncertainties led to exaggerated objective sensitivities and distorted Pareto-fronts, reducing normalized hypervolume. In contrast, the ion-to-neutron converter task was less affected—low-fidelity simulations produced results similar to those from high-fidelity data. These findings suggest that a fixed-fidelity approach is not optimal. Surrogate models can recover the Pareto-front under noisy conditions, and multi-fidelity studies help identify suitable uncertainty levels for each problem to balance efficiency and accuracy.

07 ISOTOPE AND RADIATION SOURCES↗

Alaska Observed Hydropower Generation

This dataset contains compiled observed hydropower generation for hydropower plants in Alaska. Data have been compiled from data provided to the Energy Information Administration by asset owners, data contained in annual reports produced by the Institute of Social and Economic Research at the University of Alaska Anchorage (Alaska Electric Power Statistics and Alaska Energy Statistics) and data provided to the Federal Energy Regulatory Commission by asset owners. This dataset provides available generation data from all sources in monthly and annual files, with quality flags, and generation data identifying the highest quality source in monthly and annual files.

hydropower datasets↗

Alaska Observed Hydropower Generation

This dataset contains compiled observed hydropower generation for hydropower plants in Alaska. Data have been compiled from data provided to the Energy Information Administration by asset owners, data contained in annual reports produced by the Institute of Social and Economic Research at the University of Alaska Anchorage (Alaska Electric Power Statistics and Alaska Energy Statistics) and data provided to the Federal Energy Regulatory Commission by asset owners. This dataset provides available generation data from all sources in monthly and annual files, with quality flags, and generation data identifying the highest quality source in monthly and annual files.

Broman, Daniel [Pacific Northwest National Laborat↗

The high level trigger and express data production at STAR

To meet the demands of the Beam Energy Scan phase-II (BES-II) program, the STAR experiment at the Relativistic Heavy Ion Collider (RHIC) developed a dual real-time framework consisting of a High Level Trigger (HLT) and an Express Data Production system (xProduction). The HLT operates online within the Data Acquisition (DAQ) chain on a dedicated multi-core CPU cluster with the option to offload compute-intensive kernels to Xeon Phi coprocessors. It uses parallelized algorithms, such as the Cellular Automaton (CA) Track Finder, to perform rapid tracking, vertexing, and event filtering. This allows it to select events of interest in real time and provide immediate feedback on detector and beam conditions. In contrast, the xProduction workflow runs concurrently and independently of the DAQ loop. It applies near offline-quality calibration and reconstruction within hours of data collection. The xProduction input is the express data stream, whose content can be enriched by HLT trigger/priority selections under DAQ/HLT resource constraints, and it uses the STAR calibration/conditions framework, incorporating online calibration/QA information when available. This enables early preliminary physics analysis, including the reconstruction of rare signals, such as hyperons and hypernuclei. It also provides collaboration-wide access to analysis-ready datasets. Together, the HLT and xProduction systems form a complementary architecture: the HLT performs online event selection while the xProduction chain delivers high-quality results within a short amount of time. This integrated framework has enabled the prompt reconstruction of the $^5_Λ$ He hypernucleus with high statistical significance and the efficient processing of hundreds of millions of heavy-ion collision events. In conclusion, its demonstrated scalability and robustness establish a model for future high-luminosity experiments requiring both online event filtering and rapid access to analysis-quality data.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Identification and Photometric Classification of Extragalactic Transients in the Vera C. Rubin Observatory’s Data Preview 1

The Vera C. Rubin Observatory will soon survey the southern sky, delivering a depth and sky coverage that is unprecedented in time-domain astronomy. As part of commissioning, Data Preview 1 (DP1) has been released. It comprises a Legacy Survey of Space and Time (LSST) Commissioning Camera observing campaign between 2024 November and December with multiband imaging of seven fields, covering roughly 0.4 deg 2 each, providing a first glimpse into the data products that will become available once the LSST begins. In this work, we search three fields for extragalactic transients. We identify eight new likely supernovae (SNe), and three known ones from a sample of 369,644 difference image analysis objects. Photometric classification using Superphot+ assigns subclasses with >95% confidence to only one SN Ia and one SN II in this sample. Our findings are in agreement with SN detection rate predictions of 15 ± 4 SNe from simulations using simsurvey. The SN detection rate in the data is possibly affected by the lack of suitable templates. Nevertheless, this work demonstrates the quality of the data products delivered in DP1 and indicates that the Rubin Observatory’s LSST is well placed to fulfill its discovery potential in time-domain astronomy.

Freeburn, James [University of North Carolina, Cha↗

Downloadable Dynamometer Database (D3): Public Test Data on Advanced-Technology Vehicles

Access to high-quality, independent vehicle test data is critical to advancing energy-efficient transportation research. The Downloadable Dynamometer Database (D3) is a public repository of dynamometer test data on advanced-technology vehicles, generated at the Advanced Mobility Technology Laboratory (AMTL) at Argonne National Laboratory and hosted by the Transportation and Power Systems Division. The database has been made available to support researchers, students, and professionals engaged in energy-efficient vehicle research, development, and education. A wide range of vehicle categories has been tested (i.e., alternative fuel vehicles, conventional gasoline and diesel vehicles, all-electric vehicles, hybrid electric vehicles, and plug-in hybrid electric vehicles), as well as various drive cycles and test conditions documented in the accompanying D3 user presentation. Stakeholders can select a vehicle type, identify a vehicle of interest, and download the associated test data for use in their own analyses. Data downloaded from D3 must be accompanied by the required attribution: "This data is from the Downloadable Dynamometer Database and was generated at the Advanced Mobility Technology Laboratory (AMTL) at Argonne National Laboratory." These data are critical to vehicle modeling, validation, technology assessment, and educational use.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗