Search NASA⌕ Search

SEARCH · Search NASA

Results for “Astrostatistics techniques”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Disentangling the Black Hole Mass Spectrum with Photometric Microlensing Surveys

Abstract From the formation mechanisms of stars and compact objects to nuclear physics, modern astronomy frequently leverages surveys to understand populations of objects to answer fundamental questions. The population of dark and isolated compact objects in the Galaxy contains critical information related to many of these topics, but is only practically accessible via gravitational microlensing. However, photometric microlensing observables are degenerate for different types of lenses, and one can seldom classify an event as involving either a compact object or stellar lens on its own. To address this difficulty, we apply a Bayesian framework that treats lens type probabilistically and jointly with a lens population model. This method allows lens population characteristics to be inferred despite intrinsic uncertainty in the lens class of any single event. We investigate this method’s effectiveness on a simulated ground-based photometric survey in the context of characterizing a hypothetical population of primordial black holes (PBHs) with an average mass of 30 M ⊙ . On simulated data, our method outperforms current black hole (BH) lens identification pipelines and characterizes different subpopulations of lenses while jointly constraining the PBH contribution to dark matter to ≈25%. Key to robust inference, our method can marginalize over population model uncertainty. We find the lower mass cutoff for stellar origin BHs, a key observable in understanding the BH mass gap, particularly difficult to infer in our simulations. This work lays the foundation for cutting-edge PBH abundance constraints to be extracted from current photometric microlensing surveys.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A New Constraint on the Nuclear Equation of State from Statistical Distributions of Compact Remnants of Supernovae

Abstract Understanding how matter behaves at the highest densities and temperatures is a major open problem in both nuclear physics and relativistic astrophysics. Our understanding of such behavior is often encapsulated in the so-called high-temperature nuclear equation of state (EOS), which influences compact binary mergers, core-collapse supernovae, and other phenomena. Our focus is on the type (either black hole or neutron star) and mass of the remnant of the core collapse of a massive star. For each six candidates of equations of state, we use a very large suite of spherically symmetric supernova models to generate a sample of synthetic populations of such remnants. We then compare these synthetic populations to the observed remnant population. Our study provides a novel constraint on the high-temperature nuclear EOS and describes which EOS candidates are more or less favored by an information-theoretic metric.

79 ASTRONOMY AND ASTROPHYSICS↗

How Many Elements Matter?

Some studies of stars' multielement abundance distributions suggest at least 5–7 significant dimensions, but others show that many elemental abundances can be predicted to high accuracy from [Fe/H] and [Mg/Fe] (or [Fe/H] and age) alone. We show that both propositions can be, and are, simultaneously true. We adopt a machine-learning technique known as normalizing flow to reconstruct the probability distribution of Milky Way disk stars in the space of 15 elemental abundances measured by APOGEE. Conditioning on T eff and $\mathrm{log}\,g$ minimizes the differential systematics. After further conditioning on [Fe/H] and [Mg/Fe], the residual scatter for most abundances is σ [X/H] ≲ 0.02 dex, consistent with APOGEE's reported statistical uncertainties of ~0.01–0.015 dex and intrinsic scatter of 0.01–0.02 dex. Despite the small scatter, residual abundances display clear correlations between elements, which we show are too large to be explained by measurement uncertainties or by the finite sampling noise. We must condition on at least seven elements to reduce the correlations to a level consistent with the observational uncertainties. Our results demonstrate that cross-element correlations are a much more sensitive probe of a hidden structure than dispersion, and they can be measured precisely in a large sample even if the star-by-star measurement noise is comparable to the intrinsic scatter. We conclude that many elements have an independent story to tell, even for the mundane disk stars and elements produced by the core-collapse and Type Ia supernovae. The only way to learn these lessons is to measure the abundances directly, and not merely infer them.

79 ASTRONOMY AND ASTROPHYSICS↗

STag. II. Classification of Serendipitous Supernovae Observed by Galaxy Redshift Surveys

With the number of supernovae observed expected to drastically increase thanks to large-scale surveys like the Dark Energy Spectroscopic Instrument (DESI), it is necessary that the tools we use to classify these objects keep up with this increase. We previously created Supernova Tagging and Classification (STag) to address this problem by employing machine learning techniques alongside logistic regression in order to assign “tags” to spectra based on spectral features. STag II is a continuation of this work, which now makes use of model supernova spectra combined with real DESI spectra in order to train STag to better deal with realistic data. Furthermore, we also make use of the rlap score as a trustworthiness cut, making for a more robust and accurate supernova classifier than before.

Astrostatistics techniques↗

Inferring dark matter substructure with astrometric lensing beyond the power spectrum

Abstract Astrometry—the precise measurement of positions and motions of celestial objects—has emerged as a promising avenue for characterizing the dark matter population in our Galaxy. By leveraging recent advances in simulation-based inference and neural network architectures, we introduce a novel method to search for global dark matter-induced gravitational lensing signatures in astrometric datasets. Our method based on neural likelihood-ratio estimation shows significantly enhanced sensitivity to a cold dark matter population and more favorable scaling with measurement noise compared to existing approaches based on two-point correlation statistics. We demonstrate the real-world viability of our method by showing it to be robust to non-trivial modeled as well as unmodeled noise features expected in astrometric measurements. This establishes machine learning as a powerful tool for characterizing dark matter using astrometric data.

convolutional neural networks (1938)↗

Functional Data Analysis for Extracting the Intrinsic Dimensionality of Spectra: Application to Chemical Homogeneity in the Open Cluster M67

High-resolution spectroscopic surveys of the Milky Way have entered the Big Data regime and have opened avenues for solving outstanding questions in Galactic archeology. However, exploiting their full potential is limited by complex systematics, whose characterization has not received much attention in modern spectroscopic analyses. In this work, we present a novel method to disentangle the component of spectral data space intrinsic to the stars from that due to systematics. Using functional principal component analysis on a sample of 18,933 giant spectra from APOGEE, we find that the intrinsic structure above the level of observational uncertainties requires ≈10 functional principal components (FPCs). Our FPCs can reduce the dimensionality of spectra, remove systematics, and impute masked wavelengths, thereby enabling accurate studies of stellar populations. To demonstrate the applicability of our FPCs, we use them to infer stellar parameters and abundances of 28 giants in the open cluster M67. We employ Sequential Neural Likelihood, a simulation-based Bayesian inference method that learns likelihood functions using neural density estimators, to incorporate non-Gaussian effects in spectral likelihoods. By hierarchically combining the inferred abundances, we limit the spread of the following elements in M67: Fe ≲ 0.02 dex; C ≲ 0.03 dex; O, Mg, Si, Ni ≲ 0.04 dex; Ca ≲ 0.05 dex; N, Al ≲ 0.07 dex (at 68% confidence). Our constraints suggest a lack of self-pollution by core-collapse supernovae in M67, which has promising implications for the future of chemical tagging to understand the star formation history and dynamical evolution of the Milky Way.

79 ASTRONOMY AND ASTROPHYSICS↗

Modeling Redshift-space Clustering with Abundance Matching

Abstract We explore the degrees of freedom required to jointly fit projected and redshift-space clustering of galaxies selected in three bins of stellar mass from the Sloan Digital Sky Survey Main Galaxy Sample (SDSS MGS) using a subhalo abundance matching (SHAM) model. We employ emulators for relevant clustering statistics in order to facilitate our analysis, leading to large speed gains with minimal loss of accuracy. We are able to simultaneously fit the projected and redshift-space clustering of the two most massive galaxy samples that we consider with just two free parameters: scatter in stellar mass at fixed SHAM proxy, and the dependence of the SHAM proxy on dark matter halo concentration. We find some evidence for models that include velocity bias, but including orphan galaxies improves our fits to the lower-mass samples significantly. We also model the clustering signals of specific star formation rate (sSFR) selected samples using conditional abundance matching (CAM). We obtain acceptable fits to projected and redshift-space clustering as a function of sSFR and stellar mass using two CAM variants, although the fits are worse than for stellar-mass-selected samples alone. By incorporating nonunity correlations between the CAM proxy and sSFR, we are able to resolve previously identified discrepancies between CAM predictions and SDSS observations of the environmental dependence of quenching for isolated central galaxies.

79 ASTRONOMY AND ASTROPHYSICS↗

From Images to Dark Matter: End-to-end Inference of Substructure from Hundreds of Strong Gravitational Lenses

Abstract Constraining the distribution of small-scale structure in our universe allows us to probe alternatives to the cold dark matter paradigm. Strong gravitational lensing offers a unique window into small dark matter halos (<10 10 M ⊙ ) because these halos impart a gravitational lensing signal even if they do not host luminous galaxies. We create large data sets of strong lensing images with realistic low-mass halos, Hubble Space Telescope (HST) observational effects, and galaxy light from HST’s COSMOS field. Using a simulation-based inference pipeline, we train a neural posterior estimator of the subhalo mass function (SHMF) and place constraints on populations of lenses generated using a separate set of galaxy sources. We find that by combining our network with a hierarchical inference framework, we can both reliably infer the SHMF across a variety of configurations and scale efficiently to populations with hundreds of lenses. By conducting precise inference on large and complex simulated data sets, our method lays a foundation for extracting dark matter constraints from the next generation of wide-field optical imaging surveys.

79 ASTRONOMY AND ASTROPHYSICS↗

How the Galaxy–Halo Connection Depends on Large-scale Environment

We investigate the connection between galaxies, dark matter halos, and their large-scale environments at z = 0 with Illustris TNG300 hydrodynamic simulation data. We predict stellar masses from subhalo properties to test two types of machine learning (ML) models: explainable boosting machines (EBMs) with simple galaxy environment features and E(3)-invariant graph neural networks (GNNs). The best-performing EBM models leverage spherically averaged overdensity features on 3 Mpc scales. Interpretations via SHapley Additive exPlanations also suggest that in the context of the TNG300 galaxy–halo connection, simple spherical overdensity on ∼3 Mpc scales is more important than cosmic web distance features measured using the DisPerSE algorithm. Meanwhile, a GNN with connectivity defined by a fixed linking length, L, outperforms the EBM models by a significant margin. As we increase the linking length scale, GNNs learn important environmental contributions up to the largest scales we probe (L = 10 Mpc). We conclude that 3 Mpc distance scales are most critical for describing the TNG galaxy–halo connection using the spherical overdensity parameterization, but that information on larger scales, which is not captured by simple environmental parameters or cosmic web features, can further augment these models. Our study highlights the benefits of using interpretable ML algorithms to explain models of astrophysical phenomena, and the power of using GNNs to flexibly learn complex relationships directly from data while imposing constraints from physical symmetries.

79 ASTRONOMY AND ASTROPHYSICS↗

Applications of Machine Learning to Predicting Core-collapse Supernova Explosion Outcomes

Most existing criteria derived from progenitor properties of core-collapse supernovae are not very accurate in predicting explosion outcomes. We present a novel look at identifying the explosion outcome of core-collapse supernovae using a machine-learning approach. Informed by a sample of 100 2D axisymmetric supernova simulations evolved with F ornax , we train and evaluate a random forest classifier as an explosion predictor. Furthermore, we examine physics-based feature sets including the compactness parameter, the Ertl condition, and a newly developed set that characterizes the silicon/oxygen interface. With over 1500 supernovae progenitors from 9-27 M ⊙ , we additionally train an autoencoder to extract physics-agnostic features directly from the progenitor density profiles. We find that the density profiles alone contain meaningful information regarding their explodability. Both the silicon/oxygen and autoencoder features predict the explosion outcome with ≈90% accuracy. In anticipation of much larger multidimensional simulation sets, we identify future directions in which machine-learning applications will be useful beyond the explosion outcome prediction.

79 ASTRONOMY AND ASTROPHYSICS↗

How to Obtain the Redshift Distribution from Probabilistic Redshift Estimates

Abstract A reliable estimate of the redshift distribution n ( z ) is crucial for using weak gravitational lensing and large-scale structures of galaxy catalogs to study cosmology. Spectroscopic redshifts for the dim and numerous galaxies of next-generation weak-lensing surveys are expected to be unavailable, making photometric redshift (photo- z ) probability density functions (PDFs) the next best alternative for comprehensively encapsulating the nontrivial systematics affecting photo- z point estimation. The established stacked estimator of n ( z ) avoids reducing photo- z PDFs to point estimates but yields a systematically biased estimate of n ( z ) that worsens with a decreasing signal-to-noise ratio, the very regime where photo- z PDFs are most necessary. We introduce Cosmological Hierarchical Inference with Probabilistic Photometric Redshifts ( CHIPPR ), a statistically rigorous probabilistic graphical model of redshift-dependent photometry that correctly propagates the redshift uncertainty information beyond the best-fit estimator of n ( z ) produced by traditional procedures and is provably the only self-consistent way to recover n ( z ) from photo- z PDFs. We present the chippr prototype code, noting that the mathematically justifiable approach incurs computational cost. The CHIPPR approach is applicable to any one-point statistic of any random variable, provided the prior probability density used to produce the posteriors is explicitly known; if the prior is implicit, as may be the case for popular photo- z techniques, then the resulting posterior PDFs cannot be used for scientific inference. We therefore recommend that the photo- z community focus on developing methodologies that enable the recovery of photo- z likelihoods with support over all redshifts, either directly or via a known prior probability density.

79 ASTRONOMY AND ASTROPHYSICS↗

Red Noise–based False Alarm Thresholds for Astrophysical Periodograms via Whittle’s Approximation to the Likelihood

Astronomers who search for periodic signals using Lomb–Scargle periodograms rely on false alarm level (FAL) estimates to identify statistically significant peaks. Although FALs are often calculated from white noise models, many astronomical time series suffer from red noise. Prewhitening is a statistical technique in which a continuum model is subtracted from the log power spectrum estimate, after which the observer can proceed with a white-noise treatment. Here we present a prewhitening-based method of calculating frequency-dependent FALs. We fit power laws and autoregressive models of order 1 to each Lomb–Scargle periodogram by minimizing the Whittle approximation to the negative log-likelihood (NLL), then calculate FALs based on the best-fit model power spectrum. Our technique is a novel extension of the Whittle NLL to datasets with uneven time sampling. We demonstrate FAL calculations using observations of α Cen B, GJ 581, HD 192310, synthetic data from the radial velocity (RV) fitting challenge, and Kepler observations of a differential rotator. The Kepler data analysis shows that only true rotation signals are detected by red noise FALs, while white noise FALs suggest all spurious peaks in the low-frequency range are significant. A high-frequency sinusoid injected into α Cen B logR$'$ HK observations exceeds the 1% red noise FAL despite having only 8.9% of the power of the dominant rotation signal. In a periodogram of HD 192310 RVs, peaks associated with differential rotation and planets are detected against the 5% red noise FAL without iterative model fitting or subtraction. The software for calculating red noise–based FALs is available on GitHub.

Astrostatistics (1882)↗

Results of the Photometric LSST Astronomical Time-series Classification Challenge (PLAsTiCC)

Abstract Next-generation surveys like the Legacy Survey of Space and Time (LSST) on the Vera C. Rubin Observatory (Rubin) will generate orders of magnitude more discoveries of transients and variable stars than previous surveys. To prepare for this data deluge, we developed the Photometric LSST Astronomical Time-series Classification Challenge (PLAsTiCC), a competition that aimed to catalyze the development of robust classifiers under LSST-like conditions of a nonrepresentative training set for a large photometric test set of imbalanced classes. Over 1000 teams participated in PLAsTiCC, which was hosted in the Kaggle data science competition platform between 2018 September 28 and 2018 December 17, ultimately identifying three winners in 2019 February. Participants produced classifiers employing a diverse set of machine-learning techniques including hybrid combinations and ensemble averages of a range of approaches, among them boosted decision trees, neural networks, and multilayer perceptrons. The strong performance of the top three classifiers on Type Ia supernovae and kilonovae represent a major improvement over the current state of the art within astronomy. This paper summarizes the most promising methods and evaluates their results in detail, highlighting future directions both for classifier development and simulation needs for a next-generation PLAsTiCC data set.

79 ASTRONOMY AND ASTROPHYSICS↗

SDSS-IV MaNGA: Unveiling Galaxy Interaction by Merger Stages with Machine Learning

We use machine-learning techniques to classify galaxy merger stages, which can unveil physical processes that drive the star formation and active galactic nucleus (AGN) activities during galaxy interaction. The sample contains 4690 galaxies from the integral field spectroscopy survey SDSS-IV MaNGA and can be separated into 1060 merging galaxies and 3630 nonmerging or unclassified galaxies. For the merger sample, there are 468, 125, 293, and 174 galaxies (1) in the incoming pair phase, (2) in the first pericentric passage phase, (3) approaching or just passing the apocenter, and (4) in the final coalescence phase or post-mergers. With the information of projected separation, line-of-sight velocity difference, Sloan Digital Sky Survey (SDSS) gri images, and MaNGA Hα velocity map, we are able to classify the mergers and their stages with good precision, which is the most important score to identify interacting galaxies. For the two-phase classification (binary; nonmerger and merger), the performance can be high (precision > 0.90) with LGBMClassifier . We find that sample size can be increased by rotation, so the five-phase classification (nonmerger, and merger stages 1, 2, 3, and 4) can also be good (precision > 0.85). The most important features come from SDSS gri images. The contribution from the MaNGA Hα velocity map, projected separation, and line-of-sight velocity difference can further improve the performance by 0%–20%. In other words, the image and the velocity information are sufficient to capture important features of galaxy interactions, and our results can apply to all the MaNGA data, as well as future all-sky surveys.

79 ASTRONOMY AND ASTROPHYSICS↗