Search NASA⌕ Search

SEARCH · Search NASA

Results for “Statistical calibration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Rapid measurement of soluble xylo-oligomers using near-infrared spectroscopy (NIRS) and multivariate statistics: calibration model development and practical approaches to model optimization

Rapid monitoring of biomass conversion processes using techniques such as near-infrared (NIR) spectroscopy can be substantially quicker and less labor-, resource-, and energy-intensive than conventional measurement techniques such as gas or liquid chromatography (GC or LC) due to the lack of solvents and preparation methods, as well as removing the need to transfer samples to an external lab for analytical evaluation. The purpose of this study was to determine the feasibility of rapid monitoring of a biomass conversion process using NIR spectroscopy combined with multivariate statistical modeling, and to examine the impact of (1) subsetting the samples in the original dataset by process location and (2) reducing the spectral range used in the calibration model on model performance. We develop multivariate calibration models for the concentrations of soluble xylo-oligosaccharides (XOS), monomeric xylose, and total solids at multiple points in a biomass conversion process which produces and then purifies XOS compounds from sugar cane bagasse. A single model using samples from multiple locations in the process stream showed acceptable performance as measured by standard statistical measures. However, compared to the single model, we show that separate models built by segregating the calibration samples according to process location show improved performance. We also show that combining an understanding of the sample spectra with simple multivariate analysis tools can result in a calibration model with a substantially smaller spectral range that provides essentially equal performance to the full-range model. We demonstrate that real-time monitoring of soluble xylo-oligosaccharides (XOS), monomeric xylose, and total solids concentration at multiple points in a process stream using NIR spectroscopy coupled with multivariate statistics is feasible. Segregation of sample populations by process location improves model performance. Models using a reduced spectral range containing the most relevant spectral signatures show very similar performance to the full-range model, reinforcing the importance of performing robust exploratory data analysis before beginning multivariate modeling.

09 BIOMASS FUELS↗

Stochastic modeling and statistical calibration with model error and scarce data

This paper introduces a procedure to assess the predictive accuracy of stochastic models subject to model error and sparse data. Model error is introduced as uncertainty on the coefficients of appropriate polynomial chaos expansions (PCE). The error associated with finite sample size allows us to conceive of these coefficients as statistics of the data that we describe as random variables whose influence on output quantities of interest is evaluated through the extended polynomial chaos expansion (EPCE). A Bayesian data assimilation scheme is introduced to update these expansions by considering the resulting nested chaos expansion as a hierarchical probabilistic model. Stochastic models of quantities of interest (QoI) are thus constructed and efficiently evaluated. Here, the Metropolis–Hastings Markov chain Monte Carlo procedure is used to sample the posterior. Two illustrative analytical and numerical problems are used to demonstrate the proposed approach.

Bayesian inference↗

Statistical Calibration and Validation of a Homogeneous Ventilated Wall-Interference Correction Method for the National Transonic Facility

Wind tunnel experiments will continue to be a primary source of validation data for many types of mathematical and computational models in the aerospace industry. The increased emphasis on accuracy of data acquired from these facilities requires understanding of the uncertainty of not only the measurement data but also any correction applied to the data. One of the largest and most critical corrections made to these data is due to wall interference. In an effort to understand the accuracy and suitability of these corrections, a statistical validation process for wall interference correction methods has been developed. This process is based on the use of independent cases which, after correction, are expected to produce the same result. Comparison of these independent cases with respect to the uncertainty in the correction process establishes a domain of applicability based on the capability of the method to provide reasonable corrections with respect to customer accuracy requirements. The statistical validation method was applied to the version of the Transonic Wall Interference Correction System (TWICS) recently implemented in the National Transonic Facility at NASA Langley Research Center. The TWICS code generates corrections for solid and slotted wall interference in the model pitch plane based on boundary pressure measurements. Before validation could be performed on this method, it was necessary to calibrate the ventilated wall boundary condition parameters. Discrimination comparisons are used to determine the most representative of three linear boundary condition models which have historically been used to represent longitudinally slotted test section walls. Of the three linear boundary condition models implemented for ventilated walls, the general slotted wall model was the most representative of the data. The TWICS code using the calibrated general slotted wall model was found to be valid to within the process uncertainty for test section Mach numbers less than or equal to 0.60. The scatter among the mean corrected results of the bodies of revolution validation cases was within one count of drag on a typical transport aircraft configuration for Mach numbers at or below 0.80 and two counts of drag for Mach numbers at or below 0.90.

Walker, Eric Lee↗

Statistical Calibration and Validation of a Homogeneous Ventilated Wall-Interference Correction Method for the National Transonic Facility

Wind tunnel experiments will continue to be a primary source of validation data for many types of mathematical and computational models in the aerospace industry. The increased emphasis on accuracy of data acquired from these facilities requires understanding of the uncertainty of not only the measurement data but also any correction applied to the data. One of the largest and most critical corrections made to these data is due to wall interference. In an effort to understand the accuracy and suitability of these corrections, a statistical validation process for wall interference correction methods has been developed. This process is based on the use of independent cases which, after correction, are expected to produce the same result. Comparison of these independent cases with respect to the uncertainty in the correction process establishes a domain of applicability based on the capability of the method to provide reasonable corrections with respect to customer accuracy requirements. The statistical validation method was applied to the version of the Transonic Wall Interference Correction System (TWICS) recently implemented in the National Transonic Facility at NASA Langley Research Center. The TWICS code generates corrections for solid and slotted wall interference in the model pitch plane based on boundary pressure measurements. Before validation could be performed on this method, it was necessary to calibrate the ventilated wall boundary condition parameters. Discrimination comparisons are used to determine the most representative of three linear boundary condition models which have historically been used to represent longitudinally slotted test section walls. Of the three linear boundary condition models implemented for ventilated walls, the general slotted wall model was the most representative of the data. The TWICS code using the calibrated general slotted wall model was found to be valid to within the process uncertainty for test section Mach numbers less than or equal to 0.60. The scatter among the mean corrected results of the bodies of revolution validation cases was within one count of drag on a typical transport aircraft configuration for Mach numbers at or below 0.80 and two counts of drag for Mach numbers at or below 0.90.

Walker, Eric L.↗

A Method for Producing Hierarchical and Statistically Calibrated Predictions of Nuclear Material Properties from Existing Models

Computer vision-based analysis of micrographs of nuclear materials is an emerging technique for property prediction, synthetic route identification, and other material analysis tasks. These analysis tasks play a pivotal role in many material characterization applications such as signature development for treaty verification, process optimization, etc. The backbone in many of the recent computer vision-based techniques is a deep learning model, which takes a fixed-size set of pixels and provides a class prediction for that set of pixels. For example, previous work developed a deep convolutional neural network (CNN) to predict the synthetic route from a 256 px x 256 px patch taken from a larger image of uranium ore concentrates. In this work, we present several methods for first calibrating these models in a manner that they can provide accurate probabilities of their predictions’ veracity, and several methods of combining these probabilities. Overall, the combination of these two steps into a pipeline allows for full-image and even full-sample (where a sample has many images) predictions with associated confidence values. Finally, we show that one can also use the patch predictions and confidence to produce a visualization to map predicted constituents through the image. Results and examples for predicting and mapping uranium ore concentrates’ synthetic process from imagery will be presented.

artificial intelligence↗

Statistical Partial Calibration Of Polarimetric SAR Imagery

Mathematical technique for partial calibration of quadpolarization synthetic-aperture-radar (SAR) image data makes possible to remove partially those contaminating cross-polarization effects (crosstalk) arising in antenna(s) and/or other transmitting and receiving channels. Quadpolarization data processed with help of technique better approximates backscattering properties of target area. If corner reflectors or other known discrete radar targets placed in target area, all of data channels calibrated both relatively to each other and absolutely.

Klein, Jeffrey D.↗

National Climate Database (NCDB)

The National Climate Database (NCDB) is a high resolution, bias-corrected climate dataset consisting of the three most widely used variables of solar radiation- global horizontal (GHI), direct normal (DNI), and diffuse horizontal irradiance (DHI)- as well as other meteorological data. The goal of the NCDB is to provide unbiased high temporal and spatial resolution climate data needed for renewable energy modeling. The NCDB is modeled using a statistical downscaling approach with Regional Climate Model (RCM)-based climate projections obtained from the North American Coordinated Regional Climate Downscaling Experiment (NA-CORDEX; linked below). Daily climate projections simulated by the Canadian Regional Climate Model 4 (CanRCM4) forced by the second-generation Canadian Earth System Model (CanESM2) for two Representative Concentration Pathways (RCP4.5 or moderate emissions scenario and RCP8.5 or highest baseline emission scenario) are selected as inputs to the statistical downscaling models. The National Solar Radiation Database (NSRDB) is used to build and calibrate statistical models.

Array↗

Development of an Unbiased Future Solar Dataset for Solar Resource Adequacy Research Over CONUS

A high-resolution, long-term solar dataset is essential for capturing the variability of solar energy resources and informing strategies to ensure grid reliability and resilience in systems with high levels of solar energy integration. This study focuses on generating unbiased, high-resolution projections of solar irradiance through a statistical downscaling framework, using Earth system model (ESM) simulations obtained from the North American Coordinated Regional Climate Downscaling Experiment (NA-CORDEX). The National Solar Radiation Database (NSRDB) is used to calibrate statistical downscaling models. The newly developed dataset provides solar irradiance, surface air temperature, and surface wind speed at 4-km and hourly resolutions across the contiguous United States (CONUS), based on two future scenarios (RCP4.5 and RCP8.5). This study outlines key steps in developing the high-resolution future solar dataset, including (1) regridding ESM data to a common 20-km resolution grid, (2) correcting ESM biases using the NSRDB, and (3) applying temporal and spatial downscaling methods to generate high-resolution (4-km, hourly) solar projections. Preliminary results indicate that downscaled projections (4-km) captured reasonable spatial patterns when compared to observations across CONUS for four variables. On average across all pixels, 4-km daily-total GHI and DNI projections showed normalized bias (nBias) less than 1% and 6% for GHI and DNI against NSRDB, respectively (nBias less than 1% and 5% for daily-average surface air temperature and surface wind speed). In terms of long-term trend for GHI and DNI, there was no strong increasing or decreasing trend (when compared to surface air temperature), but it showed a very weak decreasing trend.

14 SOLAR ENERGY↗

Constraining gravity with a new precision 𝐸 𝐺 estimator using Planck + SDSS BOSS data

The 𝐸 𝐺 statistic is a discriminating probe of gravity developed to test the prediction of general relativity (GR) for the relation between gravitational potential and clustering on the largest scales in the observable Universe. We present a novel high-precision estimator for the 𝐸 𝐺 statistic using CMB lensing and galaxy clustering correlations that carefully matches the effective redshifts across the different measurement components to minimize corrections. A suite of detailed tests is performed to characterize the estimator’s accuracy, its sensitivity to assumptions and analysis choices, and the non-Gaussianity of the estimator’s uncertainty is characterized. After finalization of the estimator, it is applied to Planck CMB lensing and SDSS CMASS and LOWZ galaxy data. We report the first harmonic space measurement of 𝐸 𝐺 using the LOWZ sample and CMB lensing and also updated constraints using the final CMASS sample and the latest Planck CMB lensing map. We find $\hat{𝐸}$$^{Planck+CMASS}_{𝐺}$ = 0.3⁢6$^{+0.06}_{−0.05}$⁢(68.27%) and $\hat{𝐸}$$^{Planck+LOWZ}_{𝐺}$ = 0.4⁢0$^{+0.11}_{−0.09}$⁢(68.27%), with additional subdominant systematic error budget estimates of 2% and 3%, respectively. Using Ω m,0 constraints from Planck and SDSS BAO observations, Λ⁢CDM-GR predicts 𝐸$^{GR}_ {𝐺}$⁡(𝑧 =0.555) = 0.401 ± 0.005 and 𝐸$^{GR}_{𝐺}$⁡(𝑧 =0.316) = 0.452 ± 0.005 at the effective redshifts of the CMASS and LOWZ based measurements. We report the measurement to be in good statistical agreement with the Λ⁢CDM-GR prediction and report that the measurement is also consistent with the more general GR prediction of scale independence for 𝐸 𝐺 . Furthermore, this work provides a carefully constructed and calibrated statistic with which 𝐸 𝐺 measurements can be confidently and accurately obtained with upcoming survey data.

79 ASTRONOMY AND ASTROPHYSICS↗

CAHS: Context-Aware Homology Search

Protein homology search is foundational to bioinformatics: it supports annotation transfer, structure/function inference, and evolutionary analysis over rapidly expanding sequence repositories (e.g., UniProtKB). Profile hidden Markov models (pHMMs), as implemented in HMMER, remain the most widely trusted approach because they provide statistically calibrated E-values; however, their gap behavior is fixed once a profile is trained, despite biological evidence that insertion/deletion tolerance varies across flexible loops and intrinsically disordered regions. We present CAHS (Context-Aware Homology Search), a lightweight query-time adapter for pHMM search that incorporates learned and biologically motivated signals without changing HMMER's downstream search pipeline or its calibrated E-value reporting. Given a query sequence, CAHS computes per-residue representations from a protein language model and a disorder predictor, maps these to profile coordinates, and modulates only match-state transition rows (gap-open and gap-extension probabilities) while preserving Plan7 constraints. We comprehensively evaluate CAHS across six structurally diverse protein families and multi-domain architectures against a 570k-sequence target corpus. CAHS expands detection capability, retrieving thousands of additional remote homologs at relaxed thresholds by maintaining alignment quality through flexible regions. For multi-domain proteins, context-aware modulation resolves 94% of fragmented alignments. Crucially, CAHS preserves hit-set invariance at stringent operating points (E<10-10), demonstrating increased statistical confidence without inflating false positives. Furthermore, sharper statistical distinction between homologs and background noise during early filter stages yields up to a 3.87× acceleration in end-to-end wall-clock time on high-performance computing clusters. Overall, CAHS illustrates a practical AI-for-science design pattern: augmenting a trusted probabilistic model with query-specific learned signals to improve interpretable, reproducible inference in data-rich biology.

Bhattaram, Swethasree [Georgia Institute of Techno↗

Vertical Scales of Turbulence at the Mount Wilson Observatory

The vertical scales of turbulence at the Mount Wilson Observatory are inferred from data from the University of California at Berkeley Infrared Spatial Interferometer (ISI), by modeling path length fluctuations observed in the interferometric paths to celestial objects and those in instrumental ground-based paths. The correlations between the stellar and ground-based path length fluctuations and the temporal statistics of those fluctuations are modeled on various timescales to constrain the vertical scales. A Kolmogorov-Taylor turbulence model with a finite outer scale was used to simulate ISI data. The simulation also included the white instrumental noise of the interferometer, aperture-filtering effects, and the data analysis algorithms. The simulations suggest that the path delay fluctuations observed in the 1992-1993 ISI data are largely consistent with being generated by refractivity fluctuations at two characteristic vertical scales: one extending to a height of 45 m above the ground, with a wind speed of about 1 m/ s, and another at a much higher altitude, with a wind speed of about 10 m/ s. The height of the lower layer is of the order of the dimensions of trees and other structures near the interferometer, which suggests that these objects, including elements of the interferometer, may play a role in generating the lower layer of turbulence. The modeling indicates that the high- attitude component contributes primarily to short-period (less than 10 s) fluctuations, while the lower component dominates the long-period (up to a few minutes) fluctuations. The lower component turbulent height, along with outer scales of the order of 10 m, suggest that the baseline dependence of long-term interferometric, atmospheric fluctuations should weaken for baselines greater than a few tens of meters. Simulations further show that there is the potential for improving the seeing or astrometric accuracy by about 30%-50% on average, if the path length fluctuations in the lower component are directly calibrated. Statistical and systematic effects induce an error of about 15 m in the estimate of the lower component turbulent altitude.

Treuhaft, Robert N.↗

Assessment of model parameters in MFiX particle-in-cell approach

The limitations in numerical treatment of solids-phase in conventional methods like Discrete Element Model and Two-Fluid Model have facilitated the development of alternative techniques such as Particle-In-Cell (PIC). However, a number of parameters are involved in PIC due to its empiricism. In this work, global sensitivity analysis of PIC model parameters is performed under three distinct operating regimes common in chemical engineering applications, viz. settling bed, bubbling fluidized bed and circulating fluidized bed. Simulations were performed using the PIC method in Multiphase Flow with Interphase eXchanges (MFiX) developed by National Energy Technology Laboratory (NETL). A non-intrusive uncertainty quantification (UQ) based approach is applied using Nodeworks to first construct an adequate surrogate model and then identify the most influential parameters in each case. This knowledge will aid in developing an effective design of experiments and determine optimal parameters through techniques such as deterministic or statistical calibration.

01 COAL, LIGNITE, AND PEAT↗

Bayesian Parameter Estimation of the k-ω Shear Stress Transport Model for Accurate Simulations of Impinging-Jet Heat Transfer

Given the lack of fusion-relevant component test facilities, current estimates of the thermo-fluid performance of plasma-facing components are based for the most part on numerical simulations. A major source of uncertainty in these simulations is the semiempirical turbulence (closure) models for the Reynolds stresses appearing in the governing Reynolds-averaged Navier-Stokes equations, which involve a set of constants that depend upon the flow. The objective of this study is to evaluate Bayesian parameter estimation of turbulence closure constants in ANSYS Fluent to model heat transfer in impinging jets. The Bayesian statistical calibration produces a probability distribution for these constants from experimental data; the maximum a posteriori estimates are then taken to be the calibrated constants, or parameters. The turbulence model constants are calibrated using an experimental study of a submerged jet of air impinging on a flat heated surface at Reynolds numbers Re = O(10 4 ) and impingement distance in jet diameters H/d = 2. Numerical predictions using the calibrated model parameters are then compared with those generated using the default constants. Predictions obtained with model parameters calibrated on datasets of two different sizes are compared to evaluate the effect of the number of calibration samples. Lastly, the extrapolative ability of the calibrated model is examined by predictions at a Re beyond the calibration values.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Q-BEEP: Quantum Bayesian Error Mitigation Employing Poisson Modeling over the Hamming Spectrum

Quantum computing technology has grown rapidly in recent years, with new technologies being explored, error rates being reduced, and quantum processor’s qubit capacity growing. However, near-term quantum algorithms are still unable to be induced without compounding consequential levels of noise, leading to non-trivial erroneous results. Quantum Error Correction (in-situ error mitigation) and Quantum Error Mitigation (post-induction error mitigation) are promising fields of research within the quantum algorithm scene, aiming to alleviate quantum errors, increasing the overall fidelity and hence the overall quality of circuit induction. Earlier this year, a pioneering work, namely HAMMER, published in ASPLOS-22 demonstrated the existence of a latent structure regarding post-circuit induction errors when mapping to the Hamming spectrum. However, they intuitively assumed that errors occur in local clusters, and that at higher average Hamming distances this structure falls away. In this work, we show that such a correlation structure is not only local but extends certain non-local clustering patterns which can be precisely described by a Poisson distribution model taking the input circuit, the device run time status (i.e., calibration statistics) and qubit topology into consideration. Using this quantum error characterizing model, we developed an iterative algorithm over the generated Bayesian network state-graph for post-induction error mitigation. Thanks to more precise modeling of the error distribution latent structure and the new iterative method, our Q-Beep approach provides state of the art performance and can boost circuit execution fidelity by up to 234.6% on Bernstein-Vazirani circuits and on average 71.0% on QAOA solution quality, using 16 practical IBMQ quantum processors. For other benchmarks such as those in QASMBench, the fidelity improvement is up to 17.8%. Q-Beep is a light-weight post-processing technique that can be performed offline and remotely, making it a useful tool for quantum vendors to integrate and provide more reliable circuit induction results.

Stein, Samuel A.↗

Portable, heterogeneous ensemble workflows at scale using libEnsemble

libEnsemble is a Python-based toolkit for running dynamic ensembles, developed as part of the DOE Exascale Computing Project. The toolkit utilizes a unique generator–simulator–allocator paradigm, where generators produce input for simulators, simulators evaluate those inputs, and allocators decide whether and when a simulator or generator should be called. The generator steers the ensemble based on simulation results. Generators may, for example, apply methods for numerical optimization, machine learning, or statistical calibration. libEnsemble communicates between a manager and workers. Flexibility is provided through multiple manager–worker communication substrates each of which has different benefits. These include Python’s multiprocessing, mpi4py, and TCP. Multisite ensembles are supported using Balsam or Globus Compute. We overview the unique characteristics of libEnsemble as well as current and potential interoperability with other packages in the workflow ecosystem. We highlight libEnsemble’s dynamic resource features: libEnsemble can detect system resources, such as available nodes, cores, and GPUs, and assign these in a portable way. These features allow users to specify the number of processors and GPUs required for each simulation; and resources will be automatically assigned on a wide range of systems, including Frontier, Aurora, and Perlmutter. Such ensembles can include multiple simulation types, some using GPUs and others using only CPUs, sharing nodes for maximum efficiency. We also describe the benefits of libEnsemble’s generator–simulator coupling, which easily exposes to the user the ability to cancel, and portably kill, running simulations based on models that are updated with intermediate simulation output. We demonstrate libEnsemble’s capabilities, scalability, and scientific impact via a Gaussian process surrogate training problem for the longitudinal density profile at the exit of a plasma accelerator stage. In conclusion, the study uses gpCAM for the surrogate model and employs either Wake-T or WarpX simulations, highlighting efficient use of resources that can easily extend to exascale.

Dynamic ensembles↗

Total Measurement Uncertainty in Neutron Coincidence Multiplicity Analysis

Neutron multiplicity counting is the most commonly used nondestructive assay technique for determining the plutonium mass within containers of scrap PuO 2 or mixed oxide (MOX). In multiplicity analysis, the 240 Pu eff mass, leakage multiplication, and alpha ratio (the ratio of [α, n]-to-spontaneous fission neutron production) are the three primary unknown sample properties. They must be determined simultaneously. To solve for these three unknowns in a multiplicity assay, three measured values are needed: the singles, doubles, and triples neutron count rates. While the analysis is limited to solving for three unknowns, there are many additional factors that impact the observed count rates and contribute to the measurement uncertainty. In this study we investigate the various uncertainty contributors for the multiplicity analysis through a combination of traditional uncertainty propagation techniques supplemented by Monte Carlo simulations to address the dependences not explicitly expressed by the point source model. Uncertainties arising from counting statistics, calibration parameters, calibration method, nuclear data, and various material characteristics (isotopic abundances, chemical form, density, and impurities) are considered. A Total Measurement Uncertainty (TMU) estimate is then developed from these uncertainty contributors. This study is confined to multiplicity analysis of items commonly encountered in international safeguards applications. That is, the study focused on Pu oxides and MOX materials for the masses ranging up to 4000 grams total Pu. Multiplicity measurements were simulated using MCNP V6 based on the Plutonium Scrap Multiplicity Counter (PSMC), Epithermal Multiplicity Counter (ENMC), Pyrochemical Multiplicity Counter, and Large Epithermal Multiplicity Counter (LEMC) for this study; however, this report focuses on the parameterization of the uncertainties for the PSMC. The performance differences between the PSMC and the other multiplicity counting systems are relatively small, primarily manifesting in the impact on measurement precision so that the evaluation developed for the PSMC can be applied to the other multiplicity counting systems. Finally an analysis tool, the Multiplicity TMU Estimator, was developed from this study to serve as an aid for evaluation of the total measurement uncertainty of multiplicity assay results obtained from the commonly used INCC acquisition and analysis software.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Uncertainty Analysis of Inertial Model Attitude Sensor Calibration and Application with a Recommended New Calibration Method

Statistical tools, previously developed for nonlinear least-squares estimation of multivariate sensor calibration parameters and the associated calibration uncertainty analysis, have been applied to single- and multiple-axis inertial model attitude sensors used in wind tunnel testing to measure angle of attack and roll angle. The analysis provides confidence and prediction intervals of calibrated sensor measurement uncertainty as functions of applied input pitch and roll angles. A comparative performance study of various experimental designs for inertial sensor calibration is presented along with corroborating experimental data. The importance of replicated calibrations over extended time periods has been emphasized; replication provides independent estimates of calibration precision and bias uncertainties, statistical tests for calibration or modeling bias uncertainty, and statistical tests for sensor parameter drift over time. A set of recommendations for a new standardized model attitude sensor calibration method and usage procedures is included. The statistical information provided by these procedures is necessary for the uncertainty analysis of aerospace test results now required by users of industrial wind tunnel test facilities.

Tripp, John S.↗