A likelihood ratio test for shrinkage covariance estimators
Explore the source record for details and available documents.
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
The likelihood ratio is a crucial quantity for statistical inference in science that enables hypothesis testing, construction of confidence intervals, reweighting of distributions, and more. Many modern scientific applications, however, make use of data- or simulation-driven models for which computing the likelihood ratio can be very difficult or even impossible. By applying the so-called “likelihood ratio trick,” approximations of the likelihood ratio may be computed using clever parametrizations of neural network-based classifiers. A number of different neural network setups can be defined to satisfy this procedure, each with varying performance in approximating the likelihood ratio when using finite training data. We present a series of empirical studies detailing the performance of several common loss functionals and parametrizations of the classifier output in approximating the likelihood ratio of two univariate and multivariate Gaussian distributions as well as simulated high-energy particle physics datasets.
We propose an intuitive, machine-learning approach to multiparameter inference, dubbed the InferoStatic Networks (ISN) method, to model the score and likelihood ratio estimators in cases when the probability density can be sampled but not computed directly. The ISN uses a backend neural network that models a scalar function called the inferostatic potential \varphi φ . In addition, we introduce new strategies, respectively called Kernel Score Estimation (KSE) and Kernel Likelihood Ratio Estimation (KLRE), to learn the score and the likelihood ratio functions from simulated data. We illustrate the new techniques with some toy examples and compare to existing approaches in the literature. We mention en passant some new loss functions that optimally incorporate latent information from simulations into the training procedure.
Studies in condition monitoring literature often aim to detect rolling element bearing faults because they have one of the biggest shares among defects in turbo machinery. Accordingly, several prognosis and diagnosis methods have been devised to identify fault signatures from vibration signals. A recently proposed method to capture the rolling element bearing degradation provides the groundwork for new indicator families utilizing the generalized likelihood ratio test. This novel approach exploits the cyclostationarity and the impulsiveness of vibration signals independently in order to estimate the most suitable indicators for a given fault. However, the method has yet to be tested on complex experimental vibration signals such as those of a wind turbine gearbox. In this study, the approach is applied to the National Renewable Energy Laboratory Wind Turbine Gearbox Condition Monitoring Round Robin Study data set for bearing fault detection purposes. The data set is measured on an experimental test rig of a wind turbine gearbox; hence the complexity of the vibration signals is similar to a real case. The outcome demonstrates that the proposed method is capable of distinguishing between healthy and damaged vibration signals measured on a complex wind turbine gearbox.
We propose an intuitive, machine-learning approach to multiparameter inference, dubbed the InferoStatic Networks (ISN) method, to model the score and likelihood ratio estimators in cases when the probability density can be sampled but not computed directly. The ISN uses a backend neural network that models a scalar function called the inferostatic potential \varphi φ . In addition, we introduce new strategies, respectively called Kernel Score Estimation (KSE) and Kernel Likelihood Ratio Estimation (KLRE), to learn the score and the likelihood ratio functions from simulated data. We illustrate the new techniques with some toy examples and compare to existing approaches in the literature. We mention en passant some new loss functions that optimally incorporate latent information from simulations into the training procedure.
Most of the fundamental, emergent, and phenomenological parameters of particle and nuclear physics are determined through parametric template fits. Simulations are used to populate histograms which are then matched to data. This approach is inherently lossy, since histograms are binned and low-dimensional. Deep learning has enabled unbinned and high-dimensional parameter estimation through neural likelihood(-ratio) estimation. We compare two approaches for neural simulation-based inference (NSBI): one based on discriminative learning (classification) and one based on generative modeling. These two approaches are directly evaluated on the same datasets, with a similar level of hyperparameter optimization in both cases. In addition to a Gaussian dataset, we study NSBI using a Higgs boson dataset from the FAIR Universe Challenge. We find that both the direct likelihood and likelihood ratio estimation are able to effectively extract parameters with reasonable uncertainties. For the numerical examples and within the set of hyperparameters studied, we found that the likelihood ratio method is more accurate and/or precise. Both methods have a significant spread from the network training and would require ensembling or other mitigation strategies in practice.
A growing number of applications in particle physics and beyond use neural networks as unbinned likelihood ratio estimators applied to real or simulated data. Precision requirements on the inference tasks demand a high-level of stability from these networks, which are affected by the stochastic nature of training. We show how physics concepts can be used to stabilize network training through a physics-inspired optimizer. In particular, the energy conserving descent (ECD) optimization framework uses classical Hamiltonian dynamics on the space of network parameters to reduce the dependence on the initial conditions while also stabilizing the result near the minimum of the loss function. We develop a version of this optimizer known as , which has few free hyperparameters with limited ranges guided by physical reasoning. We apply to representative likelihood-ratio estimation tasks in particle physics and find on average that it out-performs the widely used Adam optimizer. We expect that ECD will be a useful tool for wide array of data-limited problems, where it is computationally expensive to exhaustively optimize hyperparameters and mitigate fluctuations with ensembling.
Discriminating between quark- and gluon-initiated jets has long been a central focus of jet substructure, leading to the introduction of numerous observables and calculations to high perturbative accuracy. At the same time, there have been many attempts to fully exploit the jet radiation pattern using tools from statistics and machine learning. We propose a new approach that combines a deep analytic understanding of jet substructure with the optimality promised by machine learning and statistics. After specifying an approximation to the full emission phase space, we show how to construct the optimal observable for a given classification task. This procedure is demonstrated for the case of quark and gluons jets, where we show how to systematically capture sub-eikonal corrections in the splitting functions, and prove that linear combinations of weighted multiplicity is the optimal observable. In addition to providing a new and powerful framework for systematically improving jet substructure observables, we demonstrate the performance of several quark versus gluon jet tagging observables in parton-level Monte Carlo simulations, and find that they perform at or near the level of a deep neural network classifier. Combined with the rapid recent progress in the development of higher order parton showers, we believe that our approach provides a basis for systematically exploiting subleading effects in jet substructure analyses at the Large Hadron Collider (LHC) and beyond.
In system prognostics and health management, multivariable degradation models have been widely developed to predict the life of complex systems using degradation data of multiple Performance Characteristics (PCs). Recent studies have detected a Long-Term Memory (LTM) effect among the degradation process of various PCs, implying a strong coupling phenomenon between the future degradation behavior and historical degradation trajectory. Although the LTM has been widely integrated into single-PC-based degradation modeling, it has not been considered in multi-PC-based scenarios. To capture LTM among multiple PCs, this article proposes a novel LTM-integrated Multivariate Degradation Model (MDM) for system life prediction based on multivariate fractional Brownian motion, which simultaneously incorporates the cross-correlation among different PCs. To estimate parameters of the LTM-integrated MDM, a maximum likelihood method is developed. Here, two likelihood-ratio hypothesis tests are developed to test the existence of the overall and individual LTM effect among multiple PCs. Both simulation studies and physical experiments on the performance degradation of solar energy conversion and storage devices are conducted to validate the proposed model. Results reveal that the proposed LTM-integrated MDM significantly outperforms existing MDMs in life prediction, while the lifetime uncertainty is heavily underestimated by those traditional approaches that neglect the LTM.
We present an automated and probabilistic method to make prediscovery detections of near-Earth asteroids (NEAs) in archival survey images, with the goal of reducing orbital uncertainty immediately after discovery. We refit the Minor Planet Center's astrometry and propagate the full six-parameter covariance to survey epochs to define search regions. We build low-threshold source catalogs for viable images and evaluate every detected source in a search region as a candidate prediscovery. We eliminate false positives by refitting a new orbit to each candidate and probabilistically linking detections across images using a likelihood ratio. Applied to the Zwicky Transient Facility's (ZTF) imaging, we identify approximately 3000 recently discovered NEAs with prediscovery potential, including a doubling of the observational arc for about 500. We use archival ZTF imaging to make prediscovery detections of the potentially hazardous asteroid 2021 DG1, extending its arc by 2.5 yr and reducing future apparition sky plane uncertainty from many degrees to arcseconds. We also recover 2025 FU24 nearly 7 yr before its first known observation, when its sky plane uncertainty covers hundreds of square degrees across thousands of ZTF images. The method is survey agnostic and scalable, enabling rapid orbit refinement for new discoveries from Rubin, NEO Surveyor, and NEOMIR.
Here, we perform a novel multi-messenger analysis for the identification and parameter estimation of the Standing Accretion Shock Instability (SASI) in a core collapse supernova with neutrino and gravitational wave (GW) signals. In the neutrino channel, this method performs a likelihood ratio test for the presence of SASI in the frequency domain. For gravitational wave signals we process an event with a modified constrained likelihood method. Using simulated supernova signals, the properties of the Hyper-Kamiokande neutrino detector, and O3 LIGO Interferometric data, we produce the two-dimensional probability density distribution (PDF) of the SASI activity indicator and calculate the probability of detection P D as well as the false identification probability P FI . We discuss the probability to establish the presence of the SASI as a function of the source distance in each observational channel, as well as jointly. Compared to a single-messenger approach, the joint analysis results in P D (at P FI = 0.1) of SASI activities that is larger by up to ≈ 40% for a distance to the supernova of 5 kpc. We also discuss how accurately the frequency and duration of the SASI activity can be estimated in each channel separately. Our methodology is suitable for implementation in a realistic data analysis and a multi-messenger setting.
Determining the form of the Higgs potential is one of the most exciting challenges of modern particle physics. Higgs pair production directly probes the Higgs self-coupling and should be observed in the near future at the High-Luminosity LHC. We explore how to improve the sensitivity to physics beyond the Standard Model through per-event kinematics for di-Higgs events. In particular, we employ machine learning through simulation-based inference to estimate per-event likelihood ratios and gauge potential sensitivity gains from including this kinematic information. In terms of the Standard Model Effective Field Theory, we find that adding a limited number of observables can help to remove degeneracies in Wilson coefficient likelihoods and significantly improve the experimental sensitivity.
A measurement of off-shell Higgs boson production in the $H^*\to ZZ\to 4\ell$ decay channel is presented. The measurement uses 140 fb −1 of proton–proton collisions at $\sqrt{s} = 13$ TeV collected by the ATLAS detector at the Large Hadron Collider and supersedes the previous result in this decay channel using the same dataset. The data analysis is performed using a neural simulation-based inference method, which builds per-event likelihood ratios using neural networks. The observed (expected) off-shell Higgs boson production signal strength in the $ZZ\to 4\ell$ decay channel at 68% CL is $0.87^{+0.75}_{-0.54}$ ($1.00^{+1.04}_{-0.95}$ ). The evidence for off-shell Higgs boson production using the $ZZ\to 4\ell$ decay channel has an observed (expected) significance of 2.5σ (1.3σ). The expected result represents a significant improvement relative to that of the previous analysis of the same dataset, which obtained an expected significance of 0.5σ. When combined with the most recent ATLAS measurement in the $ZZ\to 2\ell 2\nu$ decay channel, the evidence for off-shell Higgs boson production has an observed (expected) significance of 3.7σ (2.4σ). The off-shell measurements are combined with the measurement of on-shell Higgs boson production to obtain constraints on the Higgs boson total width. The observed (expected) value of the Higgs boson width at 68% CL is $4.3^{+2.7}_{-1.9}$ ($4.1^{+3.5}_{-3.4}$ ) MeV.
Knowledge of how to define and estimate the dependence among multivariate hydrological series is essential for detecting abrupt changes in the dependence. Here, in this paper, a new method (BMCTC) is proposed to detect all possible abrupt change points in the dependence among multivariate hydrological series. The total correlation estimated by the matrix-based Renyi's alpha-order entropy functional is firstly introduced to define and measure the dependence strength among multivariate hydrological series. Then, the moving cut total correlation (MCTC) sequence is built by the moving window technique, which is used to measure changes in the dependence strength among multivariate hydrological series. Finally, the Bernaola-Galvan algorithm is used to detect all change points of the MCTC sequence. Simulations are performed to compare the effectiveness of BMCTC with Pearson correlation (BMCPC) and Spearman correlation (BMCSC), Cramer-von Mises (CvM) and copula-based likelihood-ratio (CLR). The results show that all change points are detected by BMCTC regardless of the samples size, but wrong change points or no change points are detected by other methods in most cases. BMCTC is applied to detect change points in the dependence among annual runoff, precipitation and sediment discharge series in the Xiliugou and the Kuyehe River, China. It is found that the dependence among runoff, precipitation and sediment discharge changed abruptly in 1980 and 1996 in the Kuyehe River and in 1999 in the Xiliugou River. These changes are mainly caused by human activities such as construction of water conservancy projects and coal mining.
Modern cosmological surveys probe the Universe deep into the nonlinear regime, where massive neutrinos suppress cosmic structure. Traditional cosmological analyses, which use the 2-point correlation function to extract information, are no longer optimal in the nonlinear regime, and there is thus much interest in extracting beyond-2-point information to improve constraints on neutrino mass. Quantifying and interpreting the beyond-2-point information is thus a pressing task. We study the field-level information in weak lensing convergence maps using convolution neural networks. We find that the network performance increases as higher source redshifts and smaller scales are considered — investigating up to a source redshift of 2.5 and ℓ max ≃ 10 4 — verifying that massive neutrinos leave a distinct effect on weak lensing. However, the performance of the network significantly drops after scaling out the 2-point information from the maps, implying that most of the field-level information can be found in the 2-point correlation function alone. We quantify these findings in terms of the likelihood ratio and also use Integrated Gradient saliency maps to interpret which parts of the map the network is learning the most from. We find that, in the absence of noise, the network extracts a similar amount of information from the most overdense and underdense regions. However, upon adding noise, the information in underdense regions is distorted as noise disproportionately washes out void-like structures.
IceCube has observed a diffuse astrophysical neutrino flux over the energy region from a few TeV to a few PeV. At PeV energies, the spectral shape is not yet well measured due to the low statistics of the data. This analysis probes the gap between 1 and 10 PeV by using high-energy downgoing muon neutrinos. Here, to reject the large atmospheric muon background, two complementary techniques are combined. The first technique selects events with high stochasticity to reject atmospheric muon bundles whose stochastic energy losses are smoothed due to high muon multiplicity. The second technique vetoes atmospheric muons with the IceTop surface array. Using 9 yrs of data, we found two neutrino candidate events in the signal region, consistent with expectation from background, each with relatively high signal probabilities. A joint maximum likelihood estimation is performed using this sample and an independent 9.5-yr sample of tracks to measure the neutrino spectrum. A likelihood ratio test is done to compare the single power-law (SPL) vs SPL+cutoff hypothesis; the SPL+cutoff model is not significantly better than the SPL. High-energy astrophysical objects from four source catalogs are also checked around the direction of the two events. No significant coincidence was found.
We explore the possibility of evaluating flow harmonics by employing the maximum likelihood estimator (MLE). For a given finite multiplicity, the MLE simultaneously furnishes estimations for all the parameters of the underlying distribution function while efficiently suppressing the variance of measures. Also, the method provides a means to assess a specific class of mixed harmonics, which is not straightforwardly feasible by the approaches primarily based on particle correlations. The results are analyzed using the Wald, likelihood ratio, and score tests of hypotheses. Besides, the resultant flow harmonics obtained using MLE are compared with those derived using particle correlations and event plane methods. Here, the dependencies of extracted flow harmonics on the multiplicity of individual events and the total number of events are analyzed. It is shown that the proposed approach works efficiently to deal with the deficiency in detector acceptability. Moreover, we elaborate on a fictitious scenario where the event plane is not a well-defined quantity in the distribution function. For the latter case, the MLE is shown to largely perform better than the two-particle correlation estimator. In this regard, one concludes that the MLE furnishes a meaningful alternative to the existing approaches for flow analysis.
Here we introduce a systematic and quantitative methodology for establishing the presence of neutrino oscillatory signals due to the hadron-quark phase transition (PT) in failing core-collapse supernovae from the observed neutrino event rate in water- or ice-based neutrino detectors. The methodology uses a likelihood ratio in the frequency domain as a test-statistic; it is employed for quantitative analysis of neutrino signals without assuming the frequency, amplitude, starting time, and duration of the PT-induced oscillations present in the neutrino events and thus it is suitable for analyzing neutrino signals from a wide variety of numerical simulations. We test the validity of this method by using a core-collapse simulation of a 17 solar-mass star by Zha et al. [Astrophys. J. 911, 74 (2021) ]. Based on this model, we further report the presence of a PT-induced oscillations quantitatively for a core-collapse supernovae out to a distance of ~10 kpc, ~5 kpc for IceCube and to a distance of ~10 kpc, ~5 kpc, and ~1 kpc for a 0.4 Mt mass water Cherenkov detector. This methodology will aid the investigation of a future galactic supernova and the study of hadron-quark phase in the core of core-collapse supernovae.