Search NASA⌕ Search

SEARCH · Search NASA

Results for “Variance Inflation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Differential credibility assessment for statistical downscaling

Climate science is increasingly using (i) ensembles of climate projections from multiple models derived using different assumptions and/or scenarios and (ii) process-oriented diagnostics of model fidelity. Efforts to assign differential credibility to projections and/or models are also rapidly advancing. A framework to quantify and depict the credibility of statistically downscaled model output is presented and demonstrated. Here, the approach employs transfer functions in the form of robust and resilient generalized linear models applied to downscale daily minimum and maximum temperature anomalies at 10 locations using predictors drawn from ERA-Interim reanalysis and two global climate models (GCM; GFDL-ESM2M and MPI-ESM-LR). The downscaled time series are used to derive several impact relevant CLIMDEX temperature indices that are assigned credibility based on (1) the reproduction of relevant large-scale predictors by the GCMs (i.e. fraction of regression beta-weights derived from predictors that are well-reproduced) and (2) the degree of variance in the observations reproduced in the downscaled series following application of a new variance inflation technique. Credibility of the downscaled predictands varies across locations, between the two GCM and is generally higher for minimum temperature than maximum temperature. The differential credibility assessment framework demonstrated here is easy to use and flexible. It can be applied as is to inform decision makers regarding projection confidence, and/or extended to include other components of the transfer functions, and/or used to weight members of a statistically downscaled ensemble.

54 ENVIRONMENTAL SCIENCES↗

Development of Short-Term Forecasting Models Using Plant Asset Data and Feature Selection

Nuclear power plants collect and store large volumes of heterogeneous data from various components and systems. With recent advances in machine learning (ML) techniques, these data can be leveraged to develop diagnostic and short-term forecasting models to better predict future equipment condition. Maintenance operations can then be planned in advance whenever degraded performance is predicted, thus resulting in fewer unplanned outages and the optimization of maintenance activities. This enables lower maintenance costs and improves the overall economics of nuclear power. This paper focuses on developing a short-term forecasting process that leverages a feature selection process to distill large volumes of heterogeneous data and predict specific equipment parameters. A variety of feature selection methods, including Shapley Additive Explanations (SHAP) and variance inflation factor (VIF), were used to select the optimal features as inputs for three ML methods: long short-term memory (LSTM) networks, support vector regression (SVR), and random forest (RF). Each combination of model and input features was used to predict a pump bearing temperature both 1 and 24 hours in advance, based on actual plant system data. The optimal inputs for the LSTM and SVR were selected using the SHAP values, while the optimal input for the RF consisted solely of the response variable itself. Each model produced similar 1-hour-ahead predictions, with root mean square errors (RMSEs) of roughly 0.006. For the 24-hour-ahead predictions, differences could be seen between LSTM, SVR, and RF, as reflected by model performances of 0.036 +- 0.014, 0.0026 +- 0, and 0.063 +- 0.004 RMSE, respectively. As big data and continuous online monitoring become more widely available, the proposed feature selection process can be used for many applications beyond the prediction of process parameters within nuclear infrastructure.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Analytics-at-scale of Sensor Data for Digital Monitoring in Nuclear Plants (3 rd Annual Report)

Nuclear power plants collect and store large volumes of heterogeneous data from various components and systems. With recent advances in machine learning (ML) techniques, these data can be leveraged to develop diagnostic and short-term forecasting models to better predict future equipment condition. Maintenance operations can then be planned in advance whenever degraded performance is predicted, thus resulting in fewer unplanned outages and the optimization of maintenance activities. This enables lower maintenance costs and improves the overall economics of nuclear power. This report primarily focuses on developing a short-term forecasting process that leverages a feature selection process to distill large volumes of heterogeneous data and predict specific equipment parameters. A variety of feature selection methods, including Shapley Additive Explanations (SHAP) and variance inflation factor (VIF), were used to select the optimal features as inputs for three ML methods: long short-term memory (LSTM) networks, support vector regression (SVR), and random forest (RF). Each combination of model and input features was used to predict a pump bearing temperature both 1 and 24 hours in advance, based on actual plant system data. The optimal inputs for the LSTM and SVR were selected using the SHAP values, while the optimal input for the RF consisted solely of the response variable itself. Each model produced similar 1-hour-ahead predictions, with root mean square errors (RMSEs) of roughly 0.006. For the 24-hour-ahead predictions, differences could be seen between LSTM, SVR, and RF, as reflected by model performances of 0.036 ± 0.014, 0.0026 ± 0, and 0.063 ± 0.004 RMSE, respectively. As big data and continuous online monitoring become more widely available, the proposed feature selection process can be used for many applications beyond the prediction of process parameters within nuclear infrastructure. This report summarizes the Fiscal Year 2021 research progress encompassing the (1) data cleaning and feature selection necessary for ML applications; (2) development of short-term forecasting models to predict future plant process parameters for both single and multiple time steps ahead; and (3) validation of the feature selection methods and short-term forecasting models given new data from different systems.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Mountain Basin Controls on the Snow-to-Streamflow Signal: An AIC-Weighted Multiple Linear Regression Framework

A regression-based analysis quantifies how basin characteristics modulate the snow-to-streamflow signal. First, we use the ERA5-Land reanalysis gridded product (European Centre for Medium Range Weather Forecasts reanalysis 5 -Land component) for 4,655 hydrologic unit code - 10 (HUC10) mountain basins across the western United States (US) for water years 1987–2024. Linear regressions are performed for peak snow water equivalent (SWE) and annual streamflow for each mountain basin. Models use ordinary least squares in Python’s statsmodels package. After which, an Akaike Information Criterion (AIC)–weighted ensemble multiple linear regression (MLR) framework with 47 watershed traits is used to predict the linear regression coefficient of determination (r-squared) defining the ability of peak SWE to predict annual streamflow across all mountain basin. Predictor sets are constrained to avoid multicollinearity by excluding models with variance inflation factors (VIF) greater than 5. Mountain basin traits included in the MLR include seasonal climate, topography, vegetation type and structure, and bedrock geology. Accepted models are considered if their AIC is within 2.0 of the model with the minimum AIC, or best model. To compare predictor influence across acceptable models, we computed standardized regression coefficients. To evaluate structural redundancy among models, we constructed binary inclusion vectors for each acceptable model, denoting whether a predictor was present (1) or absent (0). Core predictor variables are defined as occurring in at least 67% of the acceptable models. For this regional analysis, only one model was found acceptable, with higher snow-to-streamflow translation (higher r-squared) occurring in colder mountain basins with higher relative winter precipitation, more snow accumulation and a lower fraction of annual precipitation that falls in the spring and summer. The second component of the data package uses previously published, high-resolution output from an integrated hydrological model of the East River watershed using the U.S. Geological Survey Groundwater and Surface water Flow model (GSFLOW, doi:10.15485/1998576). East River MLR expands upon the approach described above to explore the response of five streamflow metrics—annual streamflow, runoff efficiency, 7-day minimum flow, low-flow duration, and non-perennial stream fraction to snow system indicators including peak SWE, snow-covered area, snow disappearance date, and the fraction of basin area characterized by low-to-no snow, as well as seasonal precipitation and temperature, and annual hydrologic variables representing soil moisture, evapotranspiration (ET), the partitioning of incoming precipitation to evapotranspiration (ET/P), groundwater storage, and groundwater inflow to streams. MLR was done on all water years (P0: 1987-2024) and for each period as determined in the split analysis using pooled regression techniques (P1: 1987-2011 and P2: 2012-2024) to evaluate shifting predictor variable emphasis on streamflow generation. Results indicate that since 2012, peak SWE has lost statistical strength in its prediction of annual streamflow and runoff efficiency, and the indirect influence of spring temperature has emerged as critically important. Low-flow metrics remain largely influenced by soil moisture, vegetation water use and groundwater inflows with summer precipitation becoming a direct influence on minimum summer flow. Together, these data and Python-based analysis tools provide a framework for identifying the key watershed characteristics that control how streamflow responds to snow from year to year. The package also helps quantify uncertainty in statistical models and assess how snow–streamflow relationships vary across regions and over time. This dataset contains comma-separated values files (.csv), text files (.txt), python code files (.py), figure files (.png), and shapefiles (.cpg, .dbf, .prj, .sbn, .sbx, .shp, .xml). Further details on file contents and MLR execution can be found in the readme file and the FLMD files. Work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

54 ENVIRONMENTAL SCIENCES↗

TBASS: A Robust Adaptation of Bayesian Adaptive Spline Surfaces

The R package TBASS is an extension of the BASS package created by Francom and Sansó (2019). The package is used to fit a Bayesian multivariate adaptive spline to a dataset that either follows a Student’s t-distribution or has outliers. Much of the framework for TBASS is adapted from the concepts of Bayesian Multivariate Adaptive Regression Splines (BMARS), specifically the work done by Denison, Mallick, and Smith (1998). The spline function is fit using a Reversible-Jump Markov Chain Monte Carlo algorithm,. By including this more robust generalization, a dataset with outliers can be accurately fit using the BMARS model, without the possibility of overfitting or variance inflation.

97 MATHEMATICS AND COMPUTING↗

Anomalies of cosmic anisotropy from holographic universality of great-circle variance

We examine all-sky cosmic microwave background temperature maps on large angular scales to compare their consistency with two scenarios: the standard inflationary quantum picture, and a distribution constrained to have a universal variance of primordial curvature perturbations on great circles. The latter symmetry is not a property of standard quantum inflation, but may be a symmetry of holographic models with causal quantum coherence on null surfaces. Since the variation of great-circle variance is dominated by the largest angular scale modes, in the latter case the amplitude and direction of the unobserved intrinsic dipole (that is, the ℓ = 1 harmonics) can be estimated from measured ℓ = 2, 3 harmonics by minimizing the variance of great-circle variances including only ℓ = 1, 2, 3 modes. It is found that including the estimated intrinsic dipole leads to a nearly-null angular correlation function over a wide range of angles, in agreement with a null anti-hemispherical symmetry independently motivated by holographic causal arguments, but highly anomalous in standard cosmology. Simulations are used here to show that simultaneously imposing the constraints of universal great-circle variance and the vanishing of the angular correlation function over a wide range of angles tends to require patterns that are unusual in the standard picture, such as anomalously high sectorality of the ℓ = 3 components, and a close alignment of principal axes of ℓ = 2 and ℓ = 3 components, that have been previously noted on the actual sky. The precision of these results appears to be primarily limited by errors introduced by models of Galactic foregrounds.

79 ASTRONOMY AND ASTROPHYSICS↗

Planck 2018 results

We report on the implications for cosmic inflation of the 2018 release of the Planck cosmic microwave background (CMB) anisotropy measurements. The results are fully consistent with those reported using the data from the two previous Planck cosmological releases, but have smaller uncertainties thanks to improvements in the characterization of polarization at low and high multipoles. Planck temperature, polarization, and lensing data determine the spectral index of scalar perturbations to be n s = 0.9649 ± 0.0042 at 68% CL. We find no evidence for a scale dependence of n s , either as a running or as a running of the running. The Universe is found to be consistent with spatial flatness with a precision of 0.4% at 95% CL by combining Planck with a compilation of baryon acoustic oscillation data. The Planck 95% CL upper limit on the tensor-to-scalar ratio, r0.002 < 0.10, is further tightened by combining with the BICEP2/Keck Array BK15 data to obtain r 0.002 < 0.056. In the framework of standard single-field inflationary models with Einstein gravity, these results imply that: (a) the predictions of slow-roll models with a concave potential, V"(Φ) < 0, are increasingly favoured by the data; and (b) based on two different methods for reconstructing the inflaton potential, we find no evidence for dynamics beyond slow roll. Three different methods for the non-parametric reconstruction of the primordial power spectrum consistently confirm a pure power law in the range of comoving scales 0.005 Mpc -1 ≲ k ≲ 0.2 Mpc -1 . A complementary analysis also finds no evidence for theoretically motivated parameterized features in the Planck power spectra. For the case of oscillatory features that are logarithmic or linear in k, this result is further strengthened by a new combined analysis including the Planck bispectrum data. The new Planck polarization data provide a stringent test of the adiabaticity of the initial conditions for the cosmological fluctuations. In correlated, mixed adiabatic and isocurvature models, the non-adiabatic contribution to the observed CMB temperature variance is constrained to 1.3%, 1.7%, and 1.7% at 95% CL for cold dark matter, neutrino density, and neutrino velocity, respectively. Planck power spectra plus lensing set constraints on the amplitude of compensated cold dark matter-baryon isocurvature perturbations that are consistent with current complementary measurements. The polarization data also provide improved constraints on inflationary models that predict a small statistically anisotropic quadupolar modulation of the primordial fluctuations. However, the polarization data do not support physical models for a scale-dependent dipolar modulation. All these findings support the key predictions of the standard single-field inflationary models, which will be further tested by future cosmological observations.

79 ASTRONOMY AND ASTROPHYSICS↗

A demonstration of improved constraints on primordial gravitational waves with delensing

We present a constraint on the tensor-to-scalar ratio, $r$, derived from measurements of cosmic microwave background (CMB) polarization $B$-modes with "delensing,'' whereby the uncertainty on $r$ contributed by the sample variance of the gravitational lensing $B$-modes is reduced by cross-correlating against a lensing $B$-mode template. This template is constructed by combining an estimate of the polarized CMB with a tracer of the projected large-scale structure. The large-scale-structure tracer used is a map of the cosmic infrared background derived from Planck satellite data, while the polarized CMB map comes from a combination of South Pole Telescope, BICEP/Keck, and Planck data. We expand the BICEP/Keck likelihood analysis framework to accept a lensing template and apply it to the BICEP/Keck data set collected through 2014 using the same parametric foreground modelling as in the previous analysis. From simulations, we find that the uncertainty on $r$ is reduced by $\sim10\%$, from $\sigma(r)$= 0.024 to 0.022, which can be compared with a $\sim26\%$ reduction obtained when using a perfect lensing template. Applying the technique to the real data, the constraint on $r$ is improved from $r_{0.05} < 0.090$ to $r_{0.05} < 0.082$ (95% C.L.). Furthermore, this is the first demonstration of improvement in an $r$ constraint through delensing.

79 ASTRONOMY AND ASTROPHYSICS↗