Search NASASearch

SEARCH · Search NASA

Results for “Statistical Methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Knowledge graph-aided Bayesian active learning for top- K genetic interaction discovery

In silico methods for predicting the effects of multi-gene perturbations hold great promise for advancing functional genomics, computational drug discovery, and disease modeling. However, the development of these predictive algorithms for mammalian systems has been hampered by limited datasets and high experimental costs. In this study, we present a Bayesian active learning framework designed to discover pairwise host gene knockdowns that effectively inhibit viral proliferation in an in vitro HIV-1 infection model. Our method leverages a biological knowledge graph as side information and employs a computationally efficient batch diversification approach. We evaluated this framework using a dataset of viral load measurements obtained from multi-day dual-gene depletion experiments, encompassing all possible pairwise knockdowns of over 350 host genes associated with HIV infection. We demonstrate that our framework rapidly identifies the most effective gene knockdown pairs for reducing viral load. Furthermore, we show that incorporating side information enhances performance during the early stages of active learning (low data regime), while our batch diversification strategy significantly boosts performance in later stages (high data regime). This framework is general and can be adapted to explore gene interactions in other contexts, such as synthetic lethality prediction and mapping epistatic effects across quantitative trait loci.

Computational biology and bioinformatics

Boosting H I -Galaxy Cross-Clustering Signal through Higher-Order Cross-Correlations

After reionization, neutral hydrogen (${\rm H\, \small {I}}$) traces the large-scale structure (LSS) of the Universe, enabling ${\rm H\, \small {I}}$ intensity mapping (IM) to capture the LSS in 3D and constrain key cosmological parameters. We present a new framework utilizing higher-order cross-correlations to study ${\rm H\, \small {I}}$ clustering around galaxies, tested using real-space data from the IllustrisTNG300 simulation. This approach computes the joint distributions of k-nearest neighbor (kNN) optical galaxies and the ${\rm H\, \small {I}}$ brightness temperature field smoothed at relevant scales (the kNN-field framework), providing sensitivity to all higher-order cross-correlations, unlike two-point statistics. To simulate ${\rm H\, \small {I}}$ data from actual surveys, we add random thermal noise and apply a simple foreground cleaning model, filtering out Fourier modes of the brightness temperature field with k ∥ < k min,∥ . Under current levels of thermal noise and foreground cleaning, typical of a Canadian Hydrogen Intensity Mapping Experiment (CHIME)-like survey, the ${\rm H\, \small {I}}$-galaxy cross-correlation signal in our simulations, using the kNN-field framework, is detectable at >30σ across r = [3, 12] h –1 Mpc. In contrast, the detectability of the standard two-point correlation function (2PCF) over the same scales depends strongly on the foreground filter: a sharp k ∥ filter can spuriously boost detection to 8σ due to position-space ringing, whereas a less sharp filter yields no detection. Nonetheless, we conclude that kNN-field cross-correlations are robustly detectable across a broad range of foreground filtering and thermal noise conditions, suggesting their potential for enhanced constraining power over 2PCFs.

79 ASTRONOMY AND ASTROPHYSICS

Monte Carlo method for constructing confidence intervals with unconstrained and constrained nuisance parameters in the NOvA experiment

Measuring observables to constrain models using maximum-likelihood estimation is fundamental to many physics experiments. Wilks' theorem provides a simple way to construct confidence intervals on model parameters, but it only applies under certain conditions. These conditions, such as nested hypotheses and unbounded parameters, are often violated in neutrino oscillation measurements and other experimental scenarios. Monte Carlo methods can address these issues, albeit at increased computational cost. In the presence of nuisance parameters, however, the best way to implement a Monte Carlo method is ambiguous. Furthermore, this paper documents the method selected by the NOvA experiment, the profile construction. It presents the toy studies that informed the choice of method, details of its implementation, and tests performed to validate it. It also includes some practical considerations which may be of use to others choosing to use the profile construction.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

Redshift evolution and covariances for joint lensing and clustering studies with DESI Y1

ABSTRACT Galaxy–galaxy lensing (GGL) and clustering measurements from the Dark Energy Spectroscopic Instrument Year 1 (DESI Y1) data set promise to yield unprecedented combined-probe tests of cosmology and the galaxy–halo connection. In such analyses, it is essential to identify and characterize all relevant statistical and systematic errors. We forecast the covariances of DESI Y1 GGL + clustering measurements and the systematic bias due to redshift evolution in the lens samples. Focusing on the projected clustering and GGL correlations, we compute a Gaussian analytical covariance, using a suite of N-body and lognormal simulations to characterize the effect of the survey footprint. Using the DESI one percent survey data, we measure the evolution of galaxy bias parameters for the DESI luminous red galaxy (LRG) and bright galaxy survey (BGS) samples. We find mild evolution in the LRGs in $0.4 < z < 0.8$, subdominant to the expected statistical errors. For BGS, we find less evolution for brighter absolute magnitude cuts, at the cost of reduced sample size. We find that for a redshift bin width $\Delta z = 0.1$, evolution effects on DESI Y1 GGL is negligible across all scales, all fiducial selection cuts, all fiducial redshift bins. Galaxy clustering is more sensitive to evolution due to the bias squared scaling. Nevertheless the redshift evolution effect is insignificant for clustering above the 1-halo scale of $0.1h^{-1}$ Mpc. For studies that wish to reliably access smaller scales, additional treatment of redshift evolution is likely needed. This study serves as a reference for GGL and clustering studies using the DESI Y1 sample.

79 ASTRONOMY AND ASTROPHYSICS

Accuracy versus precision in boosted top tagging with the ATLAS detector

The identification of top quark decays where the top quark has a large momentum transverse to the beam axis, known as top tagging , is a crucial component in many measurements of Standard Model processes and searches for beyond the Standard Model physics at the Large Hadron Collider. Machine learning techniques have improved the performance of top tagging algorithms, but the size of the systematic uncertainties for all proposed algorithms has not been systematically studied. This paper presents the performance of several machine learning based top tagging algorithms on a dataset constructed from simulated proton-proton collision events measured with the ATLAS detector at $\sqrt{s}$ = 13 TeV. The systematic uncertainties associated with these algorithms are estimated through an approximate procedure that is not meant to be used in a physics analysis, but is appropriate for the level of precision required for this study. The most performant algorithms are found to have the largest uncertainties, motivating the development of methods to reduce these uncertainties without compromising performance. To enable such efforts in the wider scientific community, the datasets used in this paper are made publicly available.

47 OTHER INSTRUMENTATION

Concurrent Inter-Model Spread of Boreal Winter Westerly Jet Meridional Positions Between the Northern and Southern Hemispheres in CMIP6 Models

Here, this study investigates the inter-model spread of climatological extratropical westerly jets in boreal winter, using the historical simulation of 52 Coupled Model Intercomparison Project phase 6 (CMIP6) models from 1851 to 2014. The results show that there is a substantial spread in the latitude of the upper-tropospheric westerly jet across models, characterised by large inter-model standard deviations to both the poleward and equatorward sides of the jet axis, although the multi-model ensemble mean (MME) performs well in simulating meridional position of westerly jets. Furthermore, we detect the consistency of inter-model jet position spread between the Northern and Southern Hemispheres, based on the inter-model empirical orthogonal function (EOF) decomposition and correlation of regional-averaged zonal winds. Specifically, the models that simulate the westerly jets poleward/equatorward relative to the MME position in one hemisphere also tend to simulate the jets poleward/equatorward in the other hemisphere. Accordingly, we define a global jet spread index to depict the concurrence of jet shift in the two hemispheres. The results of inter-model regression analyses based on this index indicate that the models positioning the jets poleward than the MME tend to simulate a wider Hadley Cell, a poleward-shifted Ferrel Cell in the Southern Hemisphere, enhanced precipitation in the subtropics and suppressed precipitation in the tropics, and warmer sea surface temperatures in the subtropics and mid-latitudes. The present results suggest that improving the simulation of jet positions in climate models requires a comprehensive consideration of thermal states in the tropics and subtropics/mid latitudes.

54 ENVIRONMENTAL SCIENCES

First-Principles Statistical Mechanics Study of Magnetic Fluctuations and Order–Disorder in the Spinel LiNi 0.5 Mn 1.5 O 4 Cathode

While significant magnetic interactions exist in lithium transition metal oxides, commonly used as Li-ion cathodes, the interplay between magnetic couplings, disorder, and redox processes remains poorly understood. In this work, we focus on the high-voltage spinel LiNi 0.5 Mn 1.5 O 4 (LNMO) cathode as a model system on which to apply a computational framework that uses first principles-based statistical mechanics methods to predict the finite temperature magnetic properties of materials and provide insights into the complex interplay between magnetic and chemical degrees of freedom. Density functional theory calculations on multiple distinct Ni–Mn orderings within the LNMO system, including the ordered ground-state structure (space group P4332), reveal a preference for a ferrimagnetic arrangement of the Ni and Mn sublattices due to strong antiferromagnetic superexchange interactions between neighboring Mn 4+ and Ni 2+ ions and ferromagnetic Mn–Mn and Ni–Ni couplings, as revealed by magnetic cluster expansions. These results are consistent with qualitative predictions using the Goodenough-Kanamori-Anderson rules. Simulations of the finite temperature magnetic properties of LNMO are conducted using Metropolis Monte Carlo. We find that a “semiclassical” Monte Carlo sampling method based on the Heisenberg Hamiltonian accurately predicts experimental magnetic transition temperatures observed in magnetometry measurements. This study highlights the importance of a robust computational toolkit that accurately captures the complex chemomagnetic interactions and predicts finite temperature magnetic behavior to help analyze experimental magnetic and magnetic resonance spectroscopy data acquired ex situ and operando.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

RNA-Puzzles Round V: blind predictions of 23 RNA structures

RNA-Puzzles is a collective endeavor dedicated to the advancement and improvement of RNA three-dimensional structure prediction. With agreement from structural biologists, RNA structures are predicted by modeling groups before publication of the experimental structures. We report a large-scale set of predictions by 18 groups for 23 RNA-Puzzles: 4 RNA elements, 2 Aptamers, 4 Viral elements, 5 Ribozymes and 8 Riboswitches. We describe automatic assessment protocols for comparisons between prediction and experiment. Our analyses reveal some critical steps to be overcome to achieve good accuracy in modeling RNA structures: identification of helix-forming pairs and of non-Watson–Crick modules, correct coaxial stacking between helices and avoidance of entanglements. Three of the top four modeling groups in this round also ranked among the top four in the CASP15 contest.

59 BASIC BIOLOGICAL SCIENCES

Are light curve classification metrics good proxies for SN Ia cosmological constraining power?

Context. When selecting a light curve classifier for use as part of a photometric supernova Ia (SN Ia) cosmological analysis, it is common to make decisions based on metrics of classification performance, such as the contamination within the photometrically classified SN Ia sample, rather than a measure of cosmological constraining power. If the former is an appropriate proxy for the latter, this practice would eliminate the computational expense of a full cosmology forecast in the analysis pipeline design process. Aims. This study tests the assumption that light curve classification metrics are an appropriate proxy for cosmology metrics. Methods. We emulated photometric SN Ia cosmology light curve samples with controlled contamination rates of individual contaminant classes and evaluated each of them under a set of classification metrics. We then derived cosmological parameter constraints from all samples under two common analysis approaches and quantified the impact of contamination by each contaminant class on the resulting cosmological parameter estimates. Results. We observe that cosmology metrics are sensitive to both the contamination rate and the class of the contaminating population, whereas the classification metrics are shown to be insensitive to the latter. Conclusions. Based on these findings, we discourage any exclusive reliance on light curve classification-based metrics for analysis design decisions, which (counterintuitively) include but are not limited to the classifier choice. Instead, we recommend optimising science analysis pipeline design choices using a metric of the information gained about the physical parameters of interest.

79 ASTRONOMY AND ASTROPHYSICS

CIRCLEZ : Reliable photometric redshifts for active galactic nuclei computed solely using photometry from Legacy Survey Imaging for DESI

Photometric redshifts for galaxies hosting an accreting supermassive black hole in their center, known as active galactic nuclei (AGNs), are notoriously challenging. At present, they are most optimally computed via spectral energy distribution (SED) fittings, assuming that deep photometry for many wavelengths is available. However, for AGNs detected from all-sky surveys, the photometry is limited and provided by a range of instruments and studies. This makes the task of homogenizing the data challenging, presenting a dramatic drawback for the millions of AGNs that wide surveys such as SRG/eROSITA are poised to detect. This work aims to compute reliable photometric redshifts for X-ray-detected AGNs using only one dataset that covers a large area: the tenth data release of the Imaging Legacy Survey (LS10) for DESI. LS10 provides deep grizW1-W4 forced photometry within various apertures over the footprint of the eROSITA-DE survey, which avoids issues related to the cross-calibration of surveys. We present the results from CIRCLEZ, a machine-learning algorithm based on a fully connected neural network. CIRCLEZ is built on a training sample of 14 000 X-ray-detected AGNs and utilizes multi-aperture photometry, mapping the light distribution of the sources. The accuracy (σNMAD) and the fraction of outliers (η) reached in a test sample of 2913 AGNs are equal to 0.067 and 11.6%, respectively. The results are comparable to (or even better than) what was previously obtained for the same field, but with much less effort in this instance. We further tested the stability of the results by computing the photometric redshifts for the sources detected in CSC2 and Chandra-COSMOS Legacy, reaching a comparable accuracy as in eFEDS when limiting the magnitude of the counterparts to the depth of LS10. The method can be applied to fainter samples of AGNs using deeper optical data from future surveys (for example, LSST, Euclid), granting LS10-like information on the light distribution beyond the morphological type. Along with this paper, we have released an updated version of the photometric redshifts (including errors and probability distribution functions) for eROSITA/eFEDS.

79 ASTRONOMY AND ASTROPHYSICS

Selection function of clusters in Dark Energy Survey year 3 data from cross-matching with South Pole Telescope detections

Context. Galaxy clusters selected based on overdensities of galaxies in photometric surveys provide the largest cluster samples. However, modeling the selection function of such samples is complicated by noncluster members projected along the line of sight (projection effects) and the potential detection of unvirialized objects (contamination). Aims. We empirically constrained the magnitude of these effects by cross-matching galaxy clusters selected in the Dark Energy Survey data with the redMaPPer algorithm with significant detections in three South Pole Telescope surveys (SZ, pol-ECS, pol-500d). Methods. For matched clusters, we augmented the redMaPPer catalog with the SPT detection significance. For unmatched objects we used the SPT detection threshold as an upper limit on the SZe signature. Using a Bayesian population model applied to the collected multiwavelength data, we explored various physically motivated models to describe the relationship between observed richness and halo mass. Results. Our analysis reveals a clear preference for models with an additional skewed scatter component associated with projection effects over a purely log-normal scatter model. We rule out significant contamination by unvirialized objects at the high-richness end of the sample. While dedicated simulations offer a well-fitting calibration of projection effects, our findings suggest the presence of redshift-dependent trends that these simulations may not have captured. Our findings highlight that modeling the selection function of optically detected clusters remains a complicated challenge that requires a combination of simulation and data-driven approaches.

79 ASTRONOMY AND ASTROPHYSICS

Counterpart identification and classification for eRASS1 and characterisation of the active galactic nuclei content

Context. Accurately accounting for the Active Galactic Nucleus (AGN) phase in galaxy evolution requires a large, clean AGN sample. This is now possible with SRG/eROSITA, which completed its first all-sky X-ray survey (eRASS1) on June 12, 2020. The public Data Release 1 (DR1, Jan 31, 2024) includes 930,203 sources from the western Galactic hemisphere. Aims. The data enable the selection of a large AGN sample and the discovery of rare sources. However, scientific return depends on accurate characterisation of the X-ray emitters, requiring high-quality multi-wavelength data. This paper presents the identification and classification of optical and infrared counterparts to eRASS1 sources. Methods. Counterparts to eRASS1 X-ray point sources were identified using Gaia DR3, CatWISE2020, and Legacy Survey DR10 (LS10) with the Bayesian NWAY algorithm and trained priors. Sources were classified as Galactic or extragalactic via a machine-learning model combining optical/IR and X-ray properties, trained on a reference sample. For extragalactic LS10 sources, photometric redshifts were computed using CIRCLEZ. Results. Within the LS10 footprint, all 656,614 eROSITA/DR1 sources have at least one possible optical counterpart; ∼570 000 are extragalactic and likely AGN. Half are new detections compared to AllWISE, Gaia, and Quaia AGN catalogues. Gaia and CatWISE2020 counterparts are less reliable, due to the survey’s shallowness and the limited amount of features available to assess the probability of being an X-ray emitter. In the Galactic plane, where the overdensity of stellar sources also increases the chance of associations, using conservative reliability cuts, we identified approximately 18 000 Gaia and 55 000 CatWISE2020 extragalactic sources. Conclusions. We have released three high-quality counterpart catalogues – plus the training and validation sets – as a benchmark for the field. These datasets have many applications, but in particular, they empower researchers to build AGN samples tailored for completeness and purity, accelerating the hunt for the Universe’s most energetic engines.

X-rays: general

Model independent approach for calculating galaxy rotation curves for low S/N MaNGA galaxies

Internal kinematics of galaxies, traced through the stellar rotation curve or two dimensional velocity map, carry important information on galactic structure and dark matter. With upcoming surveys, the velocity map may play a key role in the development of kinematic lensing as an astrophysical probe. Here, we improve techniques for extracting velocity information from integral field spectroscopy at low signal-to-noise (S/N), without a template, and demonstrate substantial advantages over the standard Penalized PiXel-Fitting method (pPXF) approach. Robust rotation curves can be derived down to S/N ≈ 2 using our method.

79 ASTRONOMY AND ASTROPHYSICS

Photon classification with Gradient Boosted Trees at CLAS12

Dihadron semi-inclusive deep inelastic scattering (SIDIS) of 10.6 GeV longitudinally polarized electrons off the proton has been measured using the CLAS12 detector at Jefferson Lab. Two separate channels, π + π 0 and π - π 0 , were analyzed, requiring the reconstruction of diphoton pairs. Here, in this analysis, we addressed the problem of false neutral particles being reconstructed by CLAS12's event builder, polluting the otherwise physical combinatorial background underneath the π 0 peak. A photon classifier using a Gradient Boosted Trees (GBTs) architecture was trained with Monte Carlo simulations to reduce the amount of background π 0 's. We show that the nearest-neighbor features learned by the model lead to a substantial increase in signal vs. background discrimination compared to previous CLAS12 π^0 analyses. The machine learning approach recovers several times more dihadron statistics for the dataset.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND

Batch VUV4 characterization for the SBC-LAr10 scintillating bubble chamber

The Scintillating Bubble Chamber (SBC) collaboration purchased 32 Hamamatsu VUV4 silicon photomultipliers (SiPMs) for use in SBC-LAr10, a bubble chamber containing 10 kg of liquid argon. A dark-count characterization technique, which avoids the use of a single-photon source, was used at two temperatures to measure the VUV4 SiPMs breakdown voltage (V BD ), the SiPM gain (g SiPM ), the rate of change of g SiPM with respect to voltage (m), the dark count rate (DCR), and the probability of a correlated avalanche (P CA ) as well as the temperature coefficients of these parameters. A Peltier-based chilled vacuum chamber was developed at Queen's University to cool down the Quads to 233.15 ± 0.2 K and 255.15 ± 0.2 K with average stability of ±20 mK. An analysis framework was developed to estimate V BD to tens of mV precision and DCR close to Poissonian error. The temperature dependence of V BD was found to be 56 ± 2 mV K -1 , and m on average across all Quads was found to be (459 ± 3(stat.)±23(sys.))× 10 3 e- PE -1 V -1 . The average DCR temperature coefficient was estimated to be 0.099 ± 0.008 K -1 corresponding to a reduction factor of 7 for every 20 K drop in temperature. The average temperature dependence of P CA was estimated to be 4000 ± 1000 ppm K -1 . P CA estimated from the average across all SiPMs is a better estimator than the P CA calculated from individual SiPMs, for all of the other parameters, the opposite is true. All the estimated parameters were measured to the precision required for SBC-LAr10, and the Quads will be used in conditions to optimize the signal-to-noise ratio.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND

Beam-beam backgrounds for the Cool Copper Collider

In this paper, we present a comprehensive characterization of beam-beam backgrounds for the Cool Copper Collider (C 3 ), a proposed linear e + e - collider designed for precision Higgs studies at center-of-mass energies of 250 and 550 GeV. Using a simulation pipeline based on the Key4hep framework, we evaluate incoherent pair production and hadron photoproduction backgrounds through the SiD detector for baseline, power-efficiency, and high-luminosity C 3 operating scenarios. The occupancy induced by the beam-beam background is evaluated for each scenario, validating the compatibility of the existing SiD detector design with operations at C 3 without substantial modifications. Furthermore, at the same time, the modular simulation framework and analysis methodology presented in this paper offer a versatile toolkit for background studies in future collider proposals, contributing to a common platform for different machine designs.

Analysis and statistical methods

Validating sequential Monte Carlo for gravitational-wave inference

Nested sampling (NS) is the preferred stochastic sampling algorithm for gravitational-wave inference for compact binary coalescences. It can handle the complex nature of the gravitational-wave likelihood surface and provides an estimate of the Bayesian model evidence. However, there is another class of algorithms that meets the same requirements, but has not been used for gravitational-wave analyses: sequential Monte Carlo (SMC), an extension of importance sampling that maps samples from an initial density to a target density via a series of intermediate densities. In this work, we validate a type of SMC algorithm, called persistent sampling (PS), for gravitational-wave inference. We consider a range of different scenarios including binary black holes and binary neutron stars and real and simulated data and show that PS produces results that are consistent with NS whilst being, on average, 2 times more efficient and 2.74 times faster. This demonstrates that PS is a viable alternative to NS that should be considered for future gravitational-wave analyses.

black hole mergers

Unraveling emission line galaxy conformity at z ∼ 1 with DESI early data

Emission line galaxies (ELGs) are now the preeminent tracers of large-scale structure at z > 0.8 due to their high density and strong emission lines, which enable accurate redshift measurements. However, relatively little is known about ELG evolution and the ELG–halo connection, exposing us to potential modelling systematics in cosmology inference using these sources. In this paper, we use a variety of observations and simulated galaxy models to propose a physical picture of ELGs and improve ELG–halo connection modelling in a halo occupation distribution framework. We investigate Dark Energy Spectroscopic Instrument (DESI)-selected ELGs in COSMOS data, and infer that ELGs are rapidly star-forming galaxies with a large fraction exhibiting disturbed morphology, implying that many of them are likely to be merger-driven starbursts. We further postulate that the tidal interactions from mergers lead to correlated star formation in central–satellite ELG pairs, a phenomenon dubbed ‘conformity’. We argue for the need to include conformity in the ELG–halo connection using galaxy models such as IllustrisTNG, and by combining observations such as the DESI ELG autocorrelation, ELG cross-correlation with luminous red galaxies, and ELG–cluster cross-correlation. We also explore the origin of conformity using the UniverseMachine model and elucidate the difference between conformity and the well-known galaxy assembly bias effect.

79 ASTRONOMY AND ASTROPHYSICS