Search NASA⌕ Search

SEARCH · Search NASA

Results for “astrostatistics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

The Statistical Consulting Center for Astronomy (SCCA)

The process by which raw astronomical data acquisition is transformed into scientifically meaningful results and interpretation typically involves many statistical steps. Traditional astronomy limits itself to a narrow range of old and familiar statistical methods: means and standard deviations; least-squares methods like chi(sup 2) minimization; and simple nonparametric procedures such as the Kolmogorov-Smirnov tests. These tools are often inadequate for the complex problems and datasets under investigations, and recent years have witnessed an increased usage of maximum-likelihood, survival analysis, multivariate analysis, wavelet and advanced time-series methods. The Statistical Consulting Center for Astronomy (SCCA) assisted astronomers with the use of sophisticated tools, and to match these tools with specific problems. The SCCA operated with two professors of statistics and a professor of astronomy working together. Questions were received by e-mail, and were discussed in detail with the questioner. Summaries of those questions and answers leading to new approaches were posted on the Web (www.state.psu.edu/ mga/SCCA). In addition to serving individual astronomers, the SCCA established a Web site for general use that provides hypertext links to selected on-line public-domain statistical software and services. The StatCodes site (www.astro.psu.edu/statcodes) provides over 200 links in the areas of: Bayesian statistics; censored and truncated data; correlation and regression, density estimation and smoothing, general statistics packages and information; image analysis; interactive Web tools; multivariate analysis; multivariate clustering and classification; nonparametric analysis; software written by astronomers; spatial statistics; statistical distributions; time series analysis; and visualization tools. StatCodes has received a remarkable high and constant hit rate of 250 hits/week (over 10,000/year) since its inception in mid-1997. It is of interest to scientists both within and outside of astronomy. The most popular sections are multivariate techniques, image analysis, and time series analysis. Hundreds of copies of the ASURV, SLOPES and CENS-TAU codes developed by SCCA scientists were also downloaded from the StatCodes site. In addition to formal SCCA duties, SCCA scientists continued a variety of related activities in astrostatistics, including refereeing of statistically oriented papers submitted to the Astrophysical Journal, talks in meetings including Feigelson's talk to science journalists entitled "The reemergence of astrostatistics" at the American Association for the Advancement of Science meeting, and published papers of astrostatistical content.

Akritas, Michael↗

VEXT: A Virtual Observatory Exploration Toolkit

This final report consists of two main parts. The first is taken from a paper by the PiCA (Pittsburgh Computational Astrostatistics) Group which describes our ongoing work in fast computation of n-point correlation functions. We present here a new algorithm for the fast computation of N-point correlation functions in large astronomical data sets. The algorithm is based on kd-trees which are decorated with cached sufficient statistics thus allowing for orders of magnitude speed-ups over the naive non-tree-based implementation of correlation functions. We further discuss the use of controlled approximations within the computation which allows for further acceleration. In summary, our algorithm now makes it possible to compute exact, all-pairs, measurements of the two, three and four-point correlation functions for cosmological data sets like the Sloan Digital Sky Survey and the next generation of Cosmic Microwave Background experiments. The second part summarizes the progress made by the PiCA Group in this area through the AISR grant.

Schneider, Jeff↗

Functional Data Analysis for Extracting the Intrinsic Dimensionality of Spectra: Application to Chemical Homogeneity in the Open Cluster M67

High-resolution spectroscopic surveys of the Milky Way have entered the Big Data regime and have opened avenues for solving outstanding questions in Galactic archeology. However, exploiting their full potential is limited by complex systematics, whose characterization has not received much attention in modern spectroscopic analyses. In this work, we present a novel method to disentangle the component of spectral data space intrinsic to the stars from that due to systematics. Using functional principal component analysis on a sample of 18,933 giant spectra from APOGEE, we find that the intrinsic structure above the level of observational uncertainties requires ≈10 functional principal components (FPCs). Our FPCs can reduce the dimensionality of spectra, remove systematics, and impute masked wavelengths, thereby enabling accurate studies of stellar populations. To demonstrate the applicability of our FPCs, we use them to infer stellar parameters and abundances of 28 giants in the open cluster M67. We employ Sequential Neural Likelihood, a simulation-based Bayesian inference method that learns likelihood functions using neural density estimators, to incorporate non-Gaussian effects in spectral likelihoods. By hierarchically combining the inferred abundances, we limit the spread of the following elements in M67: Fe ≲ 0.02 dex; C ≲ 0.03 dex; O, Mg, Si, Ni ≲ 0.04 dex; Ca ≲ 0.05 dex; N, Al ≲ 0.07 dex (at 68% confidence). Our constraints suggest a lack of self-pollution by core-collapse supernovae in M67, which has promising implications for the future of chemical tagging to understand the star formation history and dynamical evolution of the Milky Way.

79 ASTRONOMY AND ASTROPHYSICS↗

How Many Elements Matter?

Some studies of stars' multielement abundance distributions suggest at least 5–7 significant dimensions, but others show that many elemental abundances can be predicted to high accuracy from [Fe/H] and [Mg/Fe] (or [Fe/H] and age) alone. We show that both propositions can be, and are, simultaneously true. We adopt a machine-learning technique known as normalizing flow to reconstruct the probability distribution of Milky Way disk stars in the space of 15 elemental abundances measured by APOGEE. Conditioning on T eff and $\mathrm{log}\,g$ minimizes the differential systematics. After further conditioning on [Fe/H] and [Mg/Fe], the residual scatter for most abundances is σ [X/H] ≲ 0.02 dex, consistent with APOGEE's reported statistical uncertainties of ~0.01–0.015 dex and intrinsic scatter of 0.01–0.02 dex. Despite the small scatter, residual abundances display clear correlations between elements, which we show are too large to be explained by measurement uncertainties or by the finite sampling noise. We must condition on at least seven elements to reduce the correlations to a level consistent with the observational uncertainties. Our results demonstrate that cross-element correlations are a much more sensitive probe of a hidden structure than dispersion, and they can be measured precisely in a large sample even if the star-by-star measurement noise is comparable to the intrinsic scatter. We conclude that many elements have an independent story to tell, even for the mundane disk stars and elements produced by the core-collapse and Type Ia supernovae. The only way to learn these lessons is to measure the abundances directly, and not merely infer them.

79 ASTRONOMY AND ASTROPHYSICS↗

Disentangling the Black Hole Mass Spectrum with Photometric Microlensing Surveys

Abstract From the formation mechanisms of stars and compact objects to nuclear physics, modern astronomy frequently leverages surveys to understand populations of objects to answer fundamental questions. The population of dark and isolated compact objects in the Galaxy contains critical information related to many of these topics, but is only practically accessible via gravitational microlensing. However, photometric microlensing observables are degenerate for different types of lenses, and one can seldom classify an event as involving either a compact object or stellar lens on its own. To address this difficulty, we apply a Bayesian framework that treats lens type probabilistically and jointly with a lens population model. This method allows lens population characteristics to be inferred despite intrinsic uncertainty in the lens class of any single event. We investigate this method’s effectiveness on a simulated ground-based photometric survey in the context of characterizing a hypothetical population of primordial black holes (PBHs) with an average mass of 30 M ⊙ . On simulated data, our method outperforms current black hole (BH) lens identification pipelines and characterizes different subpopulations of lenses while jointly constraining the PBH contribution to dark matter to ≈25%. Key to robust inference, our method can marginalize over population model uncertainty. We find the lower mass cutoff for stellar origin BHs, a key observable in understanding the BH mass gap, particularly difficult to infer in our simulations. This work lays the foundation for cutting-edge PBH abundance constraints to be extracted from current photometric microlensing surveys.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A New Constraint on the Nuclear Equation of State from Statistical Distributions of Compact Remnants of Supernovae

Abstract Understanding how matter behaves at the highest densities and temperatures is a major open problem in both nuclear physics and relativistic astrophysics. Our understanding of such behavior is often encapsulated in the so-called high-temperature nuclear equation of state (EOS), which influences compact binary mergers, core-collapse supernovae, and other phenomena. Our focus is on the type (either black hole or neutron star) and mass of the remnant of the core collapse of a massive star. For each six candidates of equations of state, we use a very large suite of spherically symmetric supernova models to generate a sample of synthetic populations of such remnants. We then compare these synthetic populations to the observed remnant population. Our study provides a novel constraint on the high-temperature nuclear EOS and describes which EOS candidates are more or less favored by an information-theoretic metric.

79 ASTRONOMY AND ASTROPHYSICS↗

Probing the consistency of cosmological contours for supernova cosmology

As the scale of cosmological surveys increases, so does the complexity in the analyses. This complexity can often make it difficult to derive the underlying principles, necessitating statistically rigorous testing to ensure the results of an analysis are consistent and reasonable. This is particularly important in multi-probe cosmological analyses like those used in the Dark Energy Survey (DES) and the upcoming Legacy Survey of Space and Time, where accurate uncertainties are vital. In this paper, we present a statistically rigorous method to test the consistency of contours produced in these analyses and apply this method to the Pippin cosmological pipeline used for type Ia supernova cosmology with the DES. We make use of the Neyman construction, a frequentist methodology that leverages extensive simulations to calculate confidence intervals, to perform this consistency check. A true Neyman construction is too computationally expensive for supernova cosmology, so we develop a method for approximating a Neyman construction with far fewer simulations. We find that for a simulated dataset, the 68% contour reported by the Pippin pipeline and the 68% confidence region produced by our approximate Neyman construction differ by less than a percent near the input cosmology; however, they show more significant differences far from the input cosmology, with a maximal difference of 0.05 in $Ω$ M and 0.07 in w. In conclusion, this divergence is most impactful for analyses of cosmological tensions, but its impact is mitigated when combining supernovae with other cross-cutting cosmological probes, such as the cosmic microwave background.

79 ASTRONOMY AND ASTROPHYSICS↗

A Morphological Model to Separate Resolved–Unresolved Sources in the DESI Legacy Surveys: Application in the LS4 Alert Stream

Separating resolved and unresolved sources in large imaging surveys is a fundamental step to enable downstream science, such as searching for extragalactic transients in wide-field time-domain surveys. Here we present our method to effectively separate point sources from the resolved, extended sources in the Dark Energy Spectroscopic Instrument (DESI) Legacy Surveys (LS). We develop a supervised machine learning model based on the Gradient Boosting algorithm XGBoost. The features input to the model are purely morphological and are derived from the tabulated LS data products. We train the model using ∼2 × 10 5 LS sources in the COSMOS field with HST morphological labels and evaluate the model performance on LS sources with spectroscopic classification from the DESI Data Release 1 (∼2 × 10 7 objects) and the Sloan Digital Sky Survey Data Release 17 (∼3 × 10 6 objects), as well as on ∼2 × 10 8 Gaia stars. A significant fraction of LS sources are not observed in every LS filter, and we therefore build a “Hybrid” model as a linear combination of two XGBoost models, each containing features combining aperture flux measurements from the “blue” (gr) and “red” (iz) filters. The Hybrid model shows a reasonable balance between sensitivity and robustness, and achieves higher accuracy and flexibility compared to the LS morphological typing. With the Hybrid model, we provide classification scores for ∼3 × 10 9 LS sources, making this the largest ever machine learning catalog separating resolved and unresolved sources. The catalog has been incorporated into the real-time pipeline of the La Silla Schmidt Southern Survey (LS4), enabling the identification of extragalactic transients within the LS4 alert stream.

astrostatistics↗

STag. II. Classification of Serendipitous Supernovae Observed by Galaxy Redshift Surveys

With the number of supernovae observed expected to drastically increase thanks to large-scale surveys like the Dark Energy Spectroscopic Instrument (DESI), it is necessary that the tools we use to classify these objects keep up with this increase. We previously created Supernova Tagging and Classification (STag) to address this problem by employing machine learning techniques alongside logistic regression in order to assign “tags” to spectra based on spectral features. STag II is a continuation of this work, which now makes use of model supernova spectra combined with real DESI spectra in order to train STag to better deal with realistic data. Furthermore, we also make use of the rlap score as a trustworthiness cut, making for a more robust and accurate supernova classifier than before.

Astrostatistics techniques↗

Inferring dark matter substructure with astrometric lensing beyond the power spectrum

Abstract Astrometry—the precise measurement of positions and motions of celestial objects—has emerged as a promising avenue for characterizing the dark matter population in our Galaxy. By leveraging recent advances in simulation-based inference and neural network architectures, we introduce a novel method to search for global dark matter-induced gravitational lensing signatures in astrometric datasets. Our method based on neural likelihood-ratio estimation shows significantly enhanced sensitivity to a cold dark matter population and more favorable scaling with measurement noise compared to existing approaches based on two-point correlation statistics. We demonstrate the real-world viability of our method by showing it to be robust to non-trivial modeled as well as unmodeled noise features expected in astrometric measurements. This establishes machine learning as a powerful tool for characterizing dark matter using astrometric data.

convolutional neural networks (1938)↗

SDSS-IV MaStar: Data-driven Parameter Derivation for the MaStar Stellar Library

The Mapping Nearby Galaxies at Apache Point Observatory (MaNGA) Stellar Library (MaStar) is a large collection of high-quality empirical stellar spectra designed to cover all spectral types and ideal for use in the stellar population analysis of galaxies observed in the MaNGA survey. The library contains 59,266 spectra of 24,130 unique stars with spectral resolution R ~ 1800 and covering a wavelength range of 3622–10,354 Å. In this work, we derive five physical parameters for each spectrum in the library: effective temperature (T eff ), surface gravity ($\mathrm{log} g$), metallicity ([Fe/H]), microturbulent velocity ($\mathrm{log}({v}_{\mathrm{micro}})$), and alpha-element abundance ([α/Fe]). These parameters are derived with a flexible data-driven algorithm that uses a neural network model. We train a neural network using the subset of 1675 MaStar targets that have also been observed in the Apache Point Observatory Galactic Evolution Experiment (APOGEE), adopting the independently-derived APOGEE Stellar Parameter and Chemical Abundance Pipeline parameters for this reference set. For the regions of parameter space not well represented by the APOGEE training set (7000 ≤ T ≤ 30,000 K), we supplement with theoretical model spectra. We present our derived parameters along with an analysis of the uncertainties and comparisons to other analyses from the literature.

79 ASTRONOMY AND ASTROPHYSICS↗

Red Noise–based False Alarm Thresholds for Astrophysical Periodograms via Whittle’s Approximation to the Likelihood

Astronomers who search for periodic signals using Lomb–Scargle periodograms rely on false alarm level (FAL) estimates to identify statistically significant peaks. Although FALs are often calculated from white noise models, many astronomical time series suffer from red noise. Prewhitening is a statistical technique in which a continuum model is subtracted from the log power spectrum estimate, after which the observer can proceed with a white-noise treatment. Here we present a prewhitening-based method of calculating frequency-dependent FALs. We fit power laws and autoregressive models of order 1 to each Lomb–Scargle periodogram by minimizing the Whittle approximation to the negative log-likelihood (NLL), then calculate FALs based on the best-fit model power spectrum. Our technique is a novel extension of the Whittle NLL to datasets with uneven time sampling. We demonstrate FAL calculations using observations of α Cen B, GJ 581, HD 192310, synthetic data from the radial velocity (RV) fitting challenge, and Kepler observations of a differential rotator. The Kepler data analysis shows that only true rotation signals are detected by red noise FALs, while white noise FALs suggest all spurious peaks in the low-frequency range are significant. A high-frequency sinusoid injected into α Cen B logR$'$ HK observations exceeds the 1% red noise FAL despite having only 8.9% of the power of the dominant rotation signal. In a periodogram of HD 192310 RVs, peaks associated with differential rotation and planets are detected against the 5% red noise FAL without iterative model fitting or subtraction. The software for calculating red noise–based FALs is available on GitHub.

Astrostatistics (1882)↗

How to Obtain the Redshift Distribution from Probabilistic Redshift Estimates

Abstract A reliable estimate of the redshift distribution n ( z ) is crucial for using weak gravitational lensing and large-scale structures of galaxy catalogs to study cosmology. Spectroscopic redshifts for the dim and numerous galaxies of next-generation weak-lensing surveys are expected to be unavailable, making photometric redshift (photo- z ) probability density functions (PDFs) the next best alternative for comprehensively encapsulating the nontrivial systematics affecting photo- z point estimation. The established stacked estimator of n ( z ) avoids reducing photo- z PDFs to point estimates but yields a systematically biased estimate of n ( z ) that worsens with a decreasing signal-to-noise ratio, the very regime where photo- z PDFs are most necessary. We introduce Cosmological Hierarchical Inference with Probabilistic Photometric Redshifts ( CHIPPR ), a statistically rigorous probabilistic graphical model of redshift-dependent photometry that correctly propagates the redshift uncertainty information beyond the best-fit estimator of n ( z ) produced by traditional procedures and is provably the only self-consistent way to recover n ( z ) from photo- z PDFs. We present the chippr prototype code, noting that the mathematically justifiable approach incurs computational cost. The CHIPPR approach is applicable to any one-point statistic of any random variable, provided the prior probability density used to produce the posteriors is explicitly known; if the prior is implicit, as may be the case for popular photo- z techniques, then the resulting posterior PDFs cannot be used for scientific inference. We therefore recommend that the photo- z community focus on developing methodologies that enable the recovery of photo- z likelihoods with support over all redshifts, either directly or via a known prior probability density.

79 ASTRONOMY AND ASTROPHYSICS↗

Measuring Chemical Likeness of Stars with Relevant Scaled Component Analysis

Identification of chemically similar stars using elemental abundances is core to many pursuits within Galactic archeology. However, measuring the chemical likeness of stars using abundances directly is limited by systematic imprints of imperfect synthetic spectra in abundance derivation. We present a novel data-driven model that is capable of identifying chemically similar stars from spectra alone. We call this relevant scaled component analysis (RSCA). RSCA finds a mapping from stellar spectra to a representation that optimizes recovery of known open clusters. By design, RSCA amplifies factors of chemical abundance variation and minimizes those of nonchemical parameters, such as instrument systematics. The resultant representation of stellar spectra can therefore be used for precise measurements of chemical similarity between stars. We validate RSCA using 185 cluster stars in 22 open clusters in the Apache Point Observatory Galactic Evolution Experiment survey. We quantify our performance in measuring chemical similarity using a reference set of 151,145 field stars. We find that our representation identifies known stellar siblings more effectively than stellar-abundance measurements. Using RSCA, 1.8% of pairs of field stars are as similar as birth siblings, compared to 2.3% when using stellar-abundance labels. We find that almost all of the information within spectra leveraged by RSCA fits into a two-dimensional basis, which we link to [Fe/H] and α-element abundances. We conclude that chemical tagging of stars to their birth clusters remains prohibitive. However, using the spectra has noticeable gain, and our approach is poised to benefit from larger data sets and improved algorithm designs.

79 ASTRONOMY AND ASTROPHYSICS↗

The DESI PRObabilistic Value-added Bright Galaxy Survey (PROVABGS) Mock Challenge

The PRObabilistic Value-added Bright Galaxy Survey (PROVABGS) catalog will provide measurements of galaxy properties, such as stellar mass ($M*$), star formation rate (SFR), stellar metallicity (Z), and stellar age ($t_{age}$), for >10 million galaxies of the Dark Energy Spectroscopic Instrument (DESI) Bright Galaxy Survey. Full posterior distributions of the galaxy properties will be inferred using state-of-the-art Bayesian spectral energy distribution (SED) modeling of DESI spectroscopy and Legacy Surveys photometry. In this work, we present the SED model, the neural emulator for the model, and the Bayesian inference framework of PROVABGS. Furthermore, we apply the PROVABGS SED modeling on realistic synthetic DESI spectra and photometry, constructed using the L-Galaxies semi-analytic model. We compare the inferred galaxy properties to the true values of the simulation using a hierarchical Bayesian framework to quantify accuracy and precision. Overall, we accurately infer the true $M*$, SFR, Z, and $t_{age}$ of the simulated galaxies. However, the priors on galaxy properties induced by the SED model have a significant impact on the posteriors, which we characterize in detail. This work also demonstrates that a joint analysis of spectra and photometry significantly improves the constraints on galaxy properties over photometry alone and is necessary to mitigate the impact of the priors. With the methodology presented and validated in this work, PROVABGS will maximize information extracted from DESI observations and extend current galaxy studies to new regimes and unlock cutting-edge probabilistic analyses. https://github.com/changhoonhahn/provabgs/

79 ASTRONOMY AND ASTROPHYSICS↗

SDSS-IV MaNGA: Unveiling Galaxy Interaction by Merger Stages with Machine Learning

We use machine-learning techniques to classify galaxy merger stages, which can unveil physical processes that drive the star formation and active galactic nucleus (AGN) activities during galaxy interaction. The sample contains 4690 galaxies from the integral field spectroscopy survey SDSS-IV MaNGA and can be separated into 1060 merging galaxies and 3630 nonmerging or unclassified galaxies. For the merger sample, there are 468, 125, 293, and 174 galaxies (1) in the incoming pair phase, (2) in the first pericentric passage phase, (3) approaching or just passing the apocenter, and (4) in the final coalescence phase or post-mergers. With the information of projected separation, line-of-sight velocity difference, Sloan Digital Sky Survey (SDSS) gri images, and MaNGA Hα velocity map, we are able to classify the mergers and their stages with good precision, which is the most important score to identify interacting galaxies. For the two-phase classification (binary; nonmerger and merger), the performance can be high (precision > 0.90) with LGBMClassifier . We find that sample size can be increased by rotation, so the five-phase classification (nonmerger, and merger stages 1, 2, 3, and 4) can also be good (precision > 0.85). The most important features come from SDSS gri images. The contribution from the MaNGA Hα velocity map, projected separation, and line-of-sight velocity difference can further improve the performance by 0%–20%. In other words, the image and the velocity information are sufficient to capture important features of galaxy interactions, and our results can apply to all the MaNGA data, as well as future all-sky surveys.

79 ASTRONOMY AND ASTROPHYSICS↗

Modeling Redshift-space Clustering with Abundance Matching

Abstract We explore the degrees of freedom required to jointly fit projected and redshift-space clustering of galaxies selected in three bins of stellar mass from the Sloan Digital Sky Survey Main Galaxy Sample (SDSS MGS) using a subhalo abundance matching (SHAM) model. We employ emulators for relevant clustering statistics in order to facilitate our analysis, leading to large speed gains with minimal loss of accuracy. We are able to simultaneously fit the projected and redshift-space clustering of the two most massive galaxy samples that we consider with just two free parameters: scatter in stellar mass at fixed SHAM proxy, and the dependence of the SHAM proxy on dark matter halo concentration. We find some evidence for models that include velocity bias, but including orphan galaxies improves our fits to the lower-mass samples significantly. We also model the clustering signals of specific star formation rate (sSFR) selected samples using conditional abundance matching (CAM). We obtain acceptable fits to projected and redshift-space clustering as a function of sSFR and stellar mass using two CAM variants, although the fits are worse than for stellar-mass-selected samples alone. By incorporating nonunity correlations between the CAM proxy and sSFR, we are able to resolve previously identified discrepancies between CAM predictions and SDSS observations of the environmental dependence of quenching for isolated central galaxies.

79 ASTRONOMY AND ASTROPHYSICS↗

From Images to Dark Matter: End-to-end Inference of Substructure from Hundreds of Strong Gravitational Lenses

Abstract Constraining the distribution of small-scale structure in our universe allows us to probe alternatives to the cold dark matter paradigm. Strong gravitational lensing offers a unique window into small dark matter halos (<10 10 M ⊙ ) because these halos impart a gravitational lensing signal even if they do not host luminous galaxies. We create large data sets of strong lensing images with realistic low-mass halos, Hubble Space Telescope (HST) observational effects, and galaxy light from HST’s COSMOS field. Using a simulation-based inference pipeline, we train a neural posterior estimator of the subhalo mass function (SHMF) and place constraints on populations of lenses generated using a separate set of galaxy sources. We find that by combining our network with a hierarchical inference framework, we can both reliably infer the SHMF across a variety of configurations and scale efficiently to populations with hundreds of lenses. By conducting precise inference on large and complex simulated data sets, our method lays a foundation for extracting dark matter constraints from the next generation of wide-field optical imaging surveys.

79 ASTRONOMY AND ASTROPHYSICS↗