Search NASA⌕ Search

SEARCH · Search NASA

Results for “Statistical Learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

767 records · Page 43

Evidence for tWZ production in proton-proton collisions at s = 13 TeV in multilepton final states

The first evidence for the standard model production of a top quark in association with a W boson and a Z boson is reported. The measurement is performed in multilepton final states, where the Z boson is reconstructed via its decays to electron or muon pairs. At least one W boson, associated or from top quark decay, decays leptonically, too. The analysed data were recorded by the CMS experiment at the CERN LHC in 2016–2018 in proton-proton collisions at s = 13 TeV, and correspond to an integrated luminosity of 138 fb −1 . The measured cross section is 354 ± 54 ( stat ) ± 95 ( syst ) fb, and corresponds to a statistical significance of 3.4 standard deviations.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

CIRCLEZ : Reliable photometric redshifts for active galactic nuclei computed solely using photometry from Legacy Survey Imaging for DESI

Photometric redshifts for galaxies hosting an accreting supermassive black hole in their center, known as active galactic nuclei (AGNs), are notoriously challenging. At present, they are most optimally computed via spectral energy distribution (SED) fittings, assuming that deep photometry for many wavelengths is available. However, for AGNs detected from all-sky surveys, the photometry is limited and provided by a range of instruments and studies. This makes the task of homogenizing the data challenging, presenting a dramatic drawback for the millions of AGNs that wide surveys such as SRG/eROSITA are poised to detect. This work aims to compute reliable photometric redshifts for X-ray-detected AGNs using only one dataset that covers a large area: the tenth data release of the Imaging Legacy Survey (LS10) for DESI. LS10 provides deep grizW1-W4 forced photometry within various apertures over the footprint of the eROSITA-DE survey, which avoids issues related to the cross-calibration of surveys. We present the results from CIRCLEZ, a machine-learning algorithm based on a fully connected neural network. CIRCLEZ is built on a training sample of 14 000 X-ray-detected AGNs and utilizes multi-aperture photometry, mapping the light distribution of the sources. The accuracy (σNMAD) and the fraction of outliers (η) reached in a test sample of 2913 AGNs are equal to 0.067 and 11.6%, respectively. The results are comparable to (or even better than) what was previously obtained for the same field, but with much less effort in this instance. We further tested the stability of the results by computing the photometric redshifts for the sources detected in CSC2 and Chandra-COSMOS Legacy, reaching a comparable accuracy as in eFEDS when limiting the magnitude of the counterparts to the depth of LS10. The method can be applied to fainter samples of AGNs using deeper optical data from future surveys (for example, LSST, Euclid), granting LS10-like information on the light distribution beyond the morphological type. Along with this paper, we have released an updated version of the photometric redshifts (including errors and probability distribution functions) for eROSITA/eFEDS.

79 ASTRONOMY AND ASTROPHYSICS↗

Mitigating imaging systematics for DESI 2024 emission Line Galaxies and beyond

Emission Line Galaxies (ELGs) are one of the main tracers that the Dark Energy Spectroscopic Instrument (DESI) uses to probe the universe. However, they are afflicted by strong spurious correlations between target density and observing conditions known as imaging systematics. In this paper, we present the imaging systematics mitigation applied to the DESI Data Release 1 (DR1) large-scale structure catalogs used in the DESI 2024 cosmological analyses. We also explore extensions of the fiducial treatment. This includes a combined approach, through forward image simulations (Obiwan) in conjunction with neural network-based regression, to obtain an angular selection function that mitigates the imaging systematics observed in the DESI DR1 ELGs target density. We further derive a line of sight selection function from the forward model that removes the strong redshift dependence between imaging systematics and low redshift ELGs. Combining both angular and redshift-dependent systematics, we construct a three-dimensional selection function and assess the impact of all selection functions on clustering statistics. We quantify differences between these extended treatments and the fiducial treatment in terms of the measured 2-point statistics. We find that the results are generally consistent with the fiducial treatment and conclude that the differences are far less than the imaging systematics uncertainty included in DESI 2024 full-shape measurements. We extend our investigation to the ELGs at 0.6 < z < 0.8, i.e., beyond the redshift range (0.8 < z < 1.6) adopted for the DESI clustering catalog, and demonstrate that determining the full three-dimensional selection function is necessary in this redshift range. Our tests showed that all changes are consistent with statistical noise for BAO analyses indicating they are robust to even severe imaging systematics. Specific tests for the full-shape analysis will be presented in a companion paper.

79 ASTRONOMY AND ASTROPHYSICS↗

Navigating the Noise: Bringing Clarity to ML Parameterization Design With O $\boldsymbol{\mathcal{O}}$(100) Ensembles

Abstract Machine‐learning (ML) parameterizations of subgrid processes (here of turbulence, convection, and radiation) may one day replace conventional parameterizations by emulating high‐resolution physics without the cost of explicit simulation. However, uncertainty about the relationship between offline and online performance (i.e., when integrated with a large‐scale general circulation model) hinders their development. Much of this uncertainty stems from limited sampling of the noisy, emergent effects of upstream ML design decisions on downstream online hybrid simulation. Our work rectifies the sampling issue via the construction of a semi‐automated, end‐to‐end pipeline for size ensembles of hybrid simulations, revealing important nuances in how systematic reductions in offline error manifest in changes to online error and online stability. For example, removing dropout and switching from a Mean Squared Error to a Mean Absolute Error loss both reduce offline error, but they have opposite effects on online error and online stability. Other design decisions, like incorporating memory, converting moisture input from specific humidity to relative humidity, using batch normalization, and training on multiple climates do not come with any such compromises. Finally, we show that ensemble sizes of may be necessary to reliably detect causally relevant differences online. By enabling rapid online experimentation at scale, we can empirically settle debates regarding subgrid ML parameterization design that would have otherwise remained unresolved in the noise.

Lin, Jerry [Department of Earth System Sciences Un↗

Hourly PM 2.5 Estimates across California from 2018 to 2023

This study presents a new data set of hourly PM 2.5 concentrations across California from 2018 to 2023 at a three-kilometer resolution. This data set was developed by assimilating observations from PurpleAir and the U.S. EPA Air Quality System monitors into wildfire smoke forecasts from the High-Resolution Rapid Refresh Smoke (HRRR-Smoke) model using the Gridpoint Statistical Interpolation (GSI) three-dimensional variational data assimilation framework. Archived forecasts of modeled wildfire smoke PM 2.5 from HRRR-Smoke create the background field for assimilation, which is then corrected using surface observations of total PM 2.5 . The resulting reanalysis from GSI provides an estimate of total PM 2.5 that minimizes error from both the observational and the model data. Validation results indicate strong performance, with monthly R 2 values ranging from 0.73 to 0.91 across the six-year data set, comparable to other PM 2.5 data sets. Case studies are presented for three major fire events, the 2018 Camp Fire, 2019 Kincade Fire, and 2020 Lightning Complex Fires to demonstrate the data set’s fidelity in resolving plume dynamics and local exposure patterns. Root-mean-squared error averaged over each month scales with average PM 2.5 concentrations, resulting in a low error under typical conditions but higher absolute errors during extreme smoke events. This is the first long-term, hourly PM 2.5 data set of its kind for California and enables the generation of subdaily exposure metrics, such as peak hourly concentrations, exceedance durations, and time-of-day exposure peaks. The novelty and strong validation of this data set make it a compelling resource for future studies on the impact and significance of subdaily PM 2.5 exposure.

PM2.5↗

New graph-neural-network flavor tagger for Belle II and measurement of sin 2⁢𝜙 1 in 𝐵 0 → 𝐽/𝜓⁢𝐾$^0_ S$ decays

We present GFlaT, a new algorithm that uses a graph-neural-network to determine the flavor of neutral 𝐵 mesons produced in ϒ⁡(4⁢𝑆) decays. It improves previous algorithms by using the information from all charged final-state particles and the relations between them. We evaluate its performance using 𝐵 decays to flavor-specific hadronic final states reconstructed in a 362 fb −1 sample of electron-positron collisions collected at the ϒ⁡(4⁢𝑆) resonance with the Belle II detector at the SuperKEKB collider. We achieve an effective tagging efficiency of (37.40 ± 0.43 ± 0.36%), where the first uncertainty is statistical and the second systematic, which is 18% better than the previous Belle II algorithm. Demonstrating the algorithm, we use 𝐵 0 →𝐽/𝜓⁢𝐾$^0_ S$ decays to measure the mixing-induced and direct 𝐶⁢𝑃 violation parameters, 𝑆 = (0.724 ± 0.035 ± 0.009) and 𝐶 = (−0.035 ± 0.026 ± 0.029).

CP violation↗

ARCH: Large-scale knowledge graph via aggregated narrative codified health records analysis

Objective: Electronic health record (EHR) systems contain a wealth of clinical data stored as both codified data and free-text narrative notes (NLP). The complexity of EHR presents challenges in feature representation, information extraction, and uncertainty quantification. Here, to address these challenges, we proposed an efficient Aggregated naRrative Codified Health (ARCH) records analysis to generate a large-scale knowledge graph (KG) for a comprehensive set of EHR codified and narrative features. Methods: Using data from 12.5 million Veterans Affairs patients, ARCH first derives embedding vectors and generates similarities along with associated p-values to measure the strength of relatedness between clinical features with statistical certainty quantification. Next, ARCH performs a sparse embedding regression to remove indirect linkage between features to build a sparse KG. Finally, ARCH was validated on various clinical tasks, including detecting known relationships between entity pairs, predicting drug side effects, disease phenotyping, as well as sub-typing Alzheimer’s disease patients. Results: ARCH produces high-quality clinical embeddings and KG for over 60,000 codified and narrative EHR concepts. The KG and embeddings are visualized in the R-shiny powered web-API.3 ARCH achieved high accuracy in detecting EHR concept relationships, with AUCs of 0.926 (codified) and 0.861 (NLP) for similar EHR concepts, and 0.810 (codified) and 0.843 (NLP) for related pairs. It detected drug side effects with a 0.723 AUC, which improved to 0.826 after fine-tuning. Using both codified and NLP features, the detection power increased significantly. Compared to other methods, ARCH has superior accuracy and enhances weakly supervised phenotyping algorithms’ performance. Notably, it successfully categorized Alzheimer’s patients into two subgroups with varying mortality rates. Conclusion: The proposed ARCH algorithm generates large-scale high-quality semantic representations and knowledge graph for both codified and NLP EHR features, useful for a wide range of predictive modeling tasks.

Electronic health records↗

Accuracy versus precision in boosted top tagging with the ATLAS detector

The identification of top quark decays where the top quark has a large momentum transverse to the beam axis, known as top tagging , is a crucial component in many measurements of Standard Model processes and searches for beyond the Standard Model physics at the Large Hadron Collider. Machine learning techniques have improved the performance of top tagging algorithms, but the size of the systematic uncertainties for all proposed algorithms has not been systematically studied. This paper presents the performance of several machine learning based top tagging algorithms on a dataset constructed from simulated proton-proton collision events measured with the ATLAS detector at $\sqrt{s}$ = 13 TeV. The systematic uncertainties associated with these algorithms are estimated through an approximate procedure that is not meant to be used in a physics analysis, but is appropriate for the level of precision required for this study. The most performant algorithms are found to have the largest uncertainties, motivating the development of methods to reduce these uncertainties without compromising performance. To enable such efforts in the wider scientific community, the datasets used in this paper are made publicly available.

47 OTHER INSTRUMENTATION↗

Measurement of neutron production in atmospheric neutrino interactions at Super-Kamiokande

We present measurements of total neutron production from atmospheric neutrino interactions in water, analyzed as a function of electron-equivalent visible energy over a range of 30 MeV to 10 GeV. These results are based on 4,270 days of data collected by Super-Kamiokande, including 564 days with 0.011 wt% gadolinium added to enhance neutron detection. Neutron signal selection is based on a neural network trained on simulation, with its performance validated using an Am/Be neutron point source. The measurements are compared to predictions from neutrino event generators combined with various hadron-nucleus interaction models, which include an intranuclear cascade model and a nuclear deexcitation model. We observe significant variations in the predictions depending on the choice of hadron-nucleus interaction model. We discuss key factors that contribute to describing our data, such as in-medium effects in the intranuclear cascade and the accuracy of statistical evaporation modeling.

Cherenkov detectors↗

Counterpart identification and classification for eRASS1 and characterisation of the active galactic nuclei content

Context. Accurately accounting for the Active Galactic Nucleus (AGN) phase in galaxy evolution requires a large, clean AGN sample. This is now possible with SRG/eROSITA, which completed its first all-sky X-ray survey (eRASS1) on June 12, 2020. The public Data Release 1 (DR1, Jan 31, 2024) includes 930,203 sources from the western Galactic hemisphere. Aims. The data enable the selection of a large AGN sample and the discovery of rare sources. However, scientific return depends on accurate characterisation of the X-ray emitters, requiring high-quality multi-wavelength data. This paper presents the identification and classification of optical and infrared counterparts to eRASS1 sources. Methods. Counterparts to eRASS1 X-ray point sources were identified using Gaia DR3, CatWISE2020, and Legacy Survey DR10 (LS10) with the Bayesian NWAY algorithm and trained priors. Sources were classified as Galactic or extragalactic via a machine-learning model combining optical/IR and X-ray properties, trained on a reference sample. For extragalactic LS10 sources, photometric redshifts were computed using CIRCLEZ. Results. Within the LS10 footprint, all 656,614 eROSITA/DR1 sources have at least one possible optical counterpart; ∼570 000 are extragalactic and likely AGN. Half are new detections compared to AllWISE, Gaia, and Quaia AGN catalogues. Gaia and CatWISE2020 counterparts are less reliable, due to the survey’s shallowness and the limited amount of features available to assess the probability of being an X-ray emitter. In the Galactic plane, where the overdensity of stellar sources also increases the chance of associations, using conservative reliability cuts, we identified approximately 18 000 Gaia and 55 000 CatWISE2020 extragalactic sources. Conclusions. We have released three high-quality counterpart catalogues – plus the training and validation sets – as a benchmark for the field. These datasets have many applications, but in particular, they empower researchers to build AGN samples tailored for completeness and purity, accelerating the hunt for the Universe’s most energetic engines.

X-rays: general↗

Recommendations on the Use of Commercial-Off-The-Shelf (COTS) Electrical, Electronic, and Electromechanical (EEE) Parts for NASA Missions - Phase II

This assessment had two Phases. Phase I captured NASA Centers’ current practices for commercial-off-the-shelf (COTS) Electrical, Electronic, and Electromechanical (EEE) parts 1 used in spaceflight systems and ground support equipment (available at https://ntrs.nasa.gov/citations/20205011579) [ref. 1]. The Phase II report provides guidance for selecting and using COTS parts in NASA missions. The approaches proposed in this report differ from current agency practices. This top-level executive summary touches on these new approaches for using COTS parts but does not provide the detailed information that is critical in understanding the rationale behind these new approaches. Readers will need to read the entire report to gain full understanding and effectively use the recommendations herein. NASA’s historical approach to selecting and applying parts has been to define certain parts, primarily specific classes of military specification (MIL-SPEC) parts, as “standard”, leaving all others, including COTS parts, as nonstandard. Standard parts typically are used without further testing (“use-as-is”). Nonstandard parts are subjected to initial screening and subsequent lot acceptance testing of representative samples from each procured lot per MIL-SPEC or similar requirements. Decades later, top-tier commercial part manufacturers have evolved significant manufacturing, statistical control, and technological improvements that can now provide parts as reliable or more reliable than MIL-SPEC parts, when used within their datasheet limits. Concurrently, the space science and exploration community’s needs demand technological advances unavailable with MIL-SPEC parts. This ongoing change necessitates using COTS parts for space missions. Properly selected COTS parts in appropriate applications can offer performance and supply availability advantages compared to MIL-SPEC parts. Their utility and demonstrated reliability result from large volumes and automated production and testing processes. However, careful review and a thorough understanding of their specifications (i.e., datasheet limitations) is needed, and verifying that manufacturer specifications and reliability meet space hardware application needs are necessary. This report recommends MIL-SPEC screening and non-radiation-related lot acceptance testing be reduced or eliminated in cases where evidence of sufficient quality and reliability exists for COTS parts. The extent of NASA's insight into COTS manufacturers and the amount and nature of the needed evidence will differ by mission and will likely be driven by a mission's resources and associated risk posture. To facilitate this goal, two new terminologies have been defined and described: “Industry Leading Parts Manufacturer (ILPM)” and “Established COTS parts.” An ILPM is a COTS manufacturer that produces high quality and reliable parts. Some parts produced by ILPMs, defined as Established COTS parts, do not need any additional MIL-SPEC or NASA screening and lot acceptance testing to be used in space applications. This report provides guidance for selecting, procuring, and applying COTS parts and for performing part-, board-, and system-level COTS parts verification. The recommendation to select Established COTS parts from ILPMs will assure those COTS parts will have comparable quality to corresponding MIL-SPEC parts. Selecting, applying, and verifying Established COTS parts from ILPMs requires a holistic team approach, engaging parts engineers, circuit designers, quality, reliability, and systems engineers, procurement specialists, radiation specialists, avionics leads, and program/project managers. A mission-specific approach tailored to a project’s Mission, Environment, Applications and Lifetime (MEAL) [ref. 2] requirements should be developed and approved by program/project managers. Any associated risks should be clearly identified, quantified, mitigated, and/or accepted. Different approaches are recommended according to program/project Risk Classes A, B, C, and D [ref. 3] and human-rated missions [ref. 4]: 1. Recommend Classes A and B and human-rated missions consider a “MIL-SPEC parts- based design” approach. ”MIL-SPEC parts-based design” approach is one in which most parts are MIL-SPEC parts and Established COTS parts from ILPMs are used only when an equivalent MIL-SPEC part does not meet functional or size, weight, and power (SWaP) or performance requirements, or is not available. 2. Recommend Classes D and Sub-D missions consider a “System of COTS” approach. “System of COTS” approach is one which most parts are Established COTS parts from ILPMs. 3. Recommend Class C missions determine which approach is the best for their projects; that is, use either a “MIL-SPEC parts-based design” approach, “System of COTS” approach, or a combined approach utilizing elements of both. This report intends to provide guidance in using COTS parts for NASA missions with risk classifications of A through D and human-rated missions; but it does not address the costs of using COTS parts. Costs of using COTS parts in different NASA mission classes can vary significantly even if the same parts are used in different risk postures, due to differing verification levels needed. The guidance does not distinguish between critical or non-critical systems, and a given project will need to apply the appropriate guidance based on their risk posture. The intended audience of this report are NASA personnel and commercial practitioners who support NASA’s spaceflight missions, including spaceflight program or project managers, parts engineers, parts manufacturers, radiation engineers, avionics engineers, system engineers, circuit design engineers, reliability engineers, safety and mission assurance (SMA) personnel, and parts procurement specialists. The NEPP Program will perform a pathfinder study to explore implementing the guidance in this NESC report. An ILPM verification process is not the same as conventional vendor qualification processes performed according to military standards and specifications. This NESC report intends to provide guidance in utilizing available parts data from ILPM manufacturers for parts assurance assessments needed for NASA missions. The report also captured the current practices from DoD and Federal Aviation Administration (FAA) in Section 10. Note each DoD and FAA report was provided by the corresponding agencies regarding their practices, which are independent from the NESC recommendations in the report.

Commercial-Off-The-Shelf↗