Search NASASearch

SEARCH · Search NASA

Results for “Statistical Algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Uncertainty Quantification Enabled by Automatic Differentiation for Hydrodynamic Simulation of Shock‐to‐Detonation Transition in High Explosives

Quantifying the effects of uncertainty in a reactive burn model on the run-to-detonation time in high explosives (HEs) provides a robust methodology for assessing the probability of an HE failing the IHE qualification standard. Moreover, uncertainty quantification helps evaluate whether the model calibration accurately represents data outside the calibration set. This study uses a specialized hydrodynamic simulation code for modeling detonation to determine the run-to-detonation time of the HE PBX 9502 for various impact velocities. To quickly approximate uncertainties in the model, a surrogate was constructed using a Taylor series expansion centered at the mean of the input parameters. To obtain the sensitivities required for constructing the Taylor series, HYP-percomplex Automatic Differentiation (HYPAD) was implemented. HYPAD is a methodology for infusing existing codes with automatic differentiation capabilities by augmenting variables with one or more imaginary units to compute step-size independent partial derivatives. These derivatives are accurate to machine precision with respect to the implemented numerical algorithm, meaning their accuracy reflects that of the underlying method (e.g., integration or discretization schemes). Using reduced order modeling techniques, the mean and standard deviation of the run-to-detonation time of a shock within PBX 9502 were computed for a number of initial impact velocities. A weighted least squares regression was then performed to obtain a best fit curve and prediction interval for the computed statistics. Historical data points from explosively driven wedge tests were utilized to validate the prediction interval, ensuring its reliability in predicting future outcomes. With this prediction interval and a known safety constraint curve, the most probable point of failure and the probability of failure for the HE PBX 9502 were determined.

97 MATHEMATICS AND COMPUTING

Infrared-enhanced Photometric Redshifts for the Dark Energy Survey Y6 Gold catalogue

The Dark Energy Survey (DES) provides optical data across 5000 square degrees of the southern sky, enabling a broad range of extragalactic and cosmological studies. Combining DES data with infrared surveys offers the opportunity to improve its photometric redshift (photo-z) estimates. We aim to investigate improvements in photometric redshift estimation achieved by combining DES optical data with infrared measurements from the VISTA Hemisphere Survey (VHS) and the Wide-field Infrared Survey Explorer (WISE), and release an updated version of the catalogue. We performed a positional sky cross-match between the DES Y6 Gold catalogue matched to a spectroscopic dataset, the 2013 AllWISE Data Release, and VHS Data Release 5, in order to test these improvements using the Directional Neighbourhood Fitting (DNF) algorithm (Y6 Gold catalogue reference estimator). We additionally matched it to the unWISE catalogue to verify the performance against this deeper dataset. Adding infrared data reduces all the metrics (scatter, bias and outlier fraction) in photo-z estimates, particularly at higher redshifts in comparison with only using optical data from DES. The obtained results are globally better for the DES+WISE sample, with improvements that are statistically significant. On the other hand, the addition of the VHS bands to available depth is only marginal. The combined use of DES and WISE W1 and W2 data improves the photometric redshift metrics analysed here. The addition of VHS data at the DES and VHS depths explored here does not provide any further improvement at z less than 1.5, indicating that, under these constraints, WISE data may already capture the key infrared features and depth needed for accurate photo-z estimation. In addition, low signal-to-noise (less than 10) infrared data does not contribute to any improvement beyond the DES optical dataset.

Puebla, M. M. [Madrid U.; Madrid, CIEMAT]

Deep Koopman operators for causal discovery

Causal discovery aims to identify cause-effect mechanisms for better scientific understanding, explainable decision-making, and more accurate modeling. Standard statistical frameworks, such as Granger causality, lack the ability to quantify causal relationships in nonlinear dynamics due to the presence of complex feedback mechanisms, timescale mixing, and nonstationarity. Thus, applying these methods to study causal dynamics in real-world systems, such as the Earth, is a major challenge. Addressing this shortcoming, we leverage deep learning and a Koopman operator-theoretic formalism to present a class of causal discovery algorithms. Kausal uses deep Koopman operator methods to approximate nonlinear dynamics in a linearized vector space in which traditional causal inference methods such as Granger causality can be more easily applied. Our idealized experiments demonstrate Kausal’s superior ability in discovering and characterizing causal signals compared to existing deep learning and non-deep learning state-of-the-art approaches. Finally, the successful identification of major El Niño and La Niña events in observations showcases Kausal’s skill to handle real-world applications.

54 ENVIRONMENTAL SCIENCES

Power and Limitations of Linear Programming Decoder for Quantum LDPC Codes

Decoding quantum error-correcting codes is a key challenge in enabling fault-tolerant quantum computation. In the classical setting, linear programming (LP) decoders offer provable performance guarantees and can leverage fast practical optimization algorithms. Although LP decoders have been proposed for quantum codes, their performance and limitations remain relatively underexplored. In this work, we uncover a key limitation of LP decoding for quantum low-density parity-check (LDPC) codes: certain constant-weight error patterns lead to ambiguous fractional solutions that cannot be resolved through independent rounding. To address this issue, we incorporate a post-processing technique known as ordered statistics decoding (OSD), which significantly enhances LP decoding performance in practice. Our results show that LP decoding, when augmented with OSD, can outperform belief propagation with the same post-processing for intermediate code sizes of up to hundreds of qubits. These findings suggest that LP-based decoders, equipped with effective post-processing, offer a promising approach for decoding near-term quantum LDPC codes.

Gu, Shouzhen [Yale U.]

Power and Limitations of Linear Programming Decoder for Quantum LDPC Codes

Decoding quantum error-correcting codes is a key challenge in enabling fault-tolerant quantum computation. In the classical setting, linear programming (LP) decoders offer provable performance guarantees and can leverage fast practical optimization algorithms. Although LP decoders have been proposed for quantum codes, their performance and limitations remain relatively underexplored. In this work, we uncover a key limitation of LP decoding for quantum low-density parity-check (LDPC) codes: certain constant-weight error patterns lead to ambiguous fractional solutions that cannot be resolved through independent rounding. To address this issue, we incorporate a post-processing technique known as ordered statistics decoding (OSD), which significantly enhances LP decoding performance in practice. Our results show that LP decoding, when augmented with OSD, can outperform belief propagation with the same post-processing for intermediate code sizes of up to hundreds of qubits. These findings suggest that LP-based decoders, equipped with effective post-processing, offer a promising approach for decoding near-term quantum LDPC codes.

Gu, Shouzhen [Yale U.]

Power and Limitations of Linear Programming Decoder for Quantum LDPC Codes

Decoding quantum error-correcting codes is a key challenge in enabling fault-tolerant quantum computation. In the classical setting, linear programming (LP) decoders offer provable performance guarantees and can leverage fast practical optimization algorithms. Although LP decoders have been proposed for quantum codes, their performance and limitations remain relatively underexplored. In this work, we uncover a key limitation of LP decoding for quantum low-density parity-check (LDPC) codes: certain constant-weight error patterns lead to ambiguous fractional solutions that cannot be resolved through independent rounding. To address this issue, we incorporate a post-processing technique known as ordered statistics decoding (OSD), which significantly enhances LP decoding performance in practice. Our results show that LP decoding, when augmented with OSD, can outperform belief propagation with the same post-processing for intermediate code sizes of up to hundreds of qubits. These findings suggest that LP-based decoders, equipped with effective post-processing, offer a promising approach for decoding near-term quantum LDPC codes.

Gu, Shouzhen [Yale U.]

Universal energy-speed-accuracy trade-offs in driven nonequilibrium systems

The connection between measure theoretic optimal transport and dissipative nonequilibrium dynamics provides a language for quantifying nonequilibrium control costs, leading to a collection of thermodynamic speed limits, which rely on the assumption that the target probability distribution is perfectly realized. This is almost never the case in experiments or numerical simulations, so here we address the situation in which the external controller is imperfect. We obtain a lower bound for the dissipated work in generic nonequilibrium control problems that (1) is asymptotically tight and (2) matches the thermodynamic speed limit in the case of optimal driving. Along with analytically solvable examples, we refine this imperfect driving notion to systems in which the controlled degrees of freedom are slow relative to the nonequilibrium relaxation rate, and identify independent energy contributions from fast and slow degrees of freedom. Furthermore, we develop a strategy for optimizing minimally dissipative protocols based on optimal transport flow matching, a generative machine learning technique. Furthermore, this latter approach ensures the scalability of both the theoretical and computational framework we put forth. Crucially, we demonstrate that we can compute the terms in our bound numerically using efficient algorithms from the computational optimal transport literature and that the protocols we learn saturate the bound.

59 BASIC BIOLOGICAL SCIENCES

A flexible class of priors for orthonormal matrices with basis function-specific structure

Statistical modeling of high-dimensional matrix-valued data motivates the use of a low-rank representation that simultaneously summarizes key characteristics of the data and enables dimension reduction. Low-rank representations commonly factor the original data into the product of orthonormal basis functions and weights, where each basis function represents an independent feature of the data. However, the basis functions in these factorizations are typically computed using algorithmic methods that cannot quantify uncertainty or account for basis function correlation structure a priori. While there exist Bayesian methods that allow for a common correlation structure across basis functions, empirical examples motivate the need for basis function-specific dependence structure. We propose a prior distribution for orthonormal matrices that can explicitly model basis function-specific structure. The prior is used within a general probabilistic model for singular value decomposition to conduct posterior inference on the basis functions while accounting for measurement error and fixed effects. We discuss how the prior specification can be used for various scenarios and demonstrate favorable model properties through synthetic data examples. Finally, we apply our method to two-meter air temperature data from the Pacific Northwest, enhancing our understanding of the Earth system’s internal variability.

97 MATHEMATICS AND COMPUTING

Associated Particle Imaging of Neutron Inelastic Scatter: 3-D Reconstruction, Capabilities, and Challenges

Associated particle imaging (API) offers unique advantages for 3-D imaging of neutron inelastic scatter, including single-view tomographic imaging and data acquisition when access is limited to only one side of the interrogated object. However, widespread adoption of neutron inelastic scatter imaging has been impeded by several inherent challenges, most prominently spatial resolution, self-attenuation, and statistical noise. Here, in this work, the capabilities and challenges of neutron inelastic scatter imaging are investigated. Instead of focusing on a single imaging application, we identify imaging principles that hold for various neutron inelastic scatter imaging techniques. The primary challenges for 3-D imaging are characterized. The inherent spatial resolution in the time-of-flight (TOF) dimension is derived based on the known system timing resolution and scan geometry. Three reconstruction algorithms are described and demonstrated, including the introduction of modern iterative reconstruction incorporating a physics-based system model. Simulation is leveraged to demonstrate imaging capability with varying coincidence count levels and system timing resolution. An example of measured data with both back-scatter and forward-scatter detector positioning is presented. System design characteristics and their effects on image quality are identified. The imaging framework presented in this article has the potential to facilitate growth of 3-D neutron inelastic scatter API by identifying applications that are a good match for the technique and by targeting system development resources toward the requirements of a specific imaging task.

Associated particle imaging (API)

Differentially Private Adaptive Noise Injection (DP-ANI) v1.0

Location data is collected from users continuously to understand their mobility patterns. Releasing the user trajectories may compromise user privacy. Therefore, the general practice is to release aggregated location datasets. However, private information may still be inferred from an aggregated version of location trajectories. Differential privacy (DP) protects the query output against inference attacks regardless of background knowledge. This software implements a differential privacy-based privacy model that protects the user's origins and destinations from being inferred from aggregated mobility datasets. This is achieved by injecting Planar Laplace noise to the user origin and destination GPS points. The noisy GPS points are then transformed into a link representation using a link-matching algorithm. Finally, the link trajectories form an aggregated mobility network. The injected noise level is selected using the Sparse Vector Mechanism. This DP selection mechanism considers the link density of the location and the functional category of the localized links. Compared to the different baseline models, including a k-anonymity method, our differential privacy-based aggregation model offers query responses that are close to the raw data in terms of aggregate statistics at both the network and trajectory-levels with maximum 9% deviation from the baseline in terms of network length.

Peisert, Sean [Lawrence Berkeley National Laborato

Search for 2p2h Interactions in the NOνA Near Detector

The physics of 2p2h interactions and their contribution to the NO$\nu$A near detector data are not fully understood. This study attempts to shed some light on these interactions and the accuracy of the models used to simulate them through a search for a specific 2p2h interaction in the NO$\nu$A near detector. By performing an event selection algorithm based on particle identifier algorithms run over reconstructed data, a signal region is created to minimize the background while maximizing the number of 2p2h events where a muon neutrino interacts with a neutron and a proton coupled by a meson exchange current and produces two protons and one muon. In the signal region, separation is found between the signal events and the background in plots of the angles between the protons and the muon. Although a full statistical analysis is not completed in this study, comparing the angle plots for simulation and data shows that the model used to simulate the events reasonably approximates reality and that the near detector data likely includes signal events. Signal events are also identified in event displays, further indicating that there is some contribution of the signal to the overall NO$\nu$A near detector data.

Gable, Kyle

Evaluating cosmological biases using photometric redshifts for Type Ia Supernova cosmology with the Dark Energy Survey Supernova Program

Cosmological analyses with Type Ia Supernovae (SNe Ia) have traditionally been reliant on spectroscopy for both classifying the type of supernova and obtaining reliable redshifts to measure the distance–redshift relation. While obtaining a host-galaxy spectroscopic redshift for most SNe is feasible for small-area transient surveys, it will be too resource intensive for upcoming large-area surveys such as the Vera Rubin Observatory Legacy Survey of Space and Time, which will observe on the order of millions of SNe. Here, we use data from the Dark Energy Survey (DES) to address this problem with photometric redshifts (photo-z) inferred directly from the SN light curve in combination with Gaussian and full p(z) priors from host-galaxy photo-z estimates. Using the DES 5-yr photometrically classified SN sample, we consider several photo-z algorithms as host-galaxy photo-z priors, including the Self-Organizing Map redshifts (SOMPZ), Bayesian Photometric Redshifts (BPZ), and Directional-Neighbourhood Fitting (DNF) redshift estimates employed in the DES 3 × 2 point analyses. With detailed catalogue-level simulations of the DES 5-yr sample, we find that the simulated w can be recovered within ±0.02 when using SN+SOMPZ or DNF prior photo-z, smaller than the average statistical uncertainty for these samples of 0.03. With data, we obtain biases in w consistent with simulations within ~1σ for three of the five photo-z variants. We further evaluate how photo-z systematics interplay with photometric classification and find classification introduces a subdominant systematic component. This work lays the foundation for next-generation fully photometric SNe Ia cosmological analyses.

(cosmology:) dark energy

Snow Distribution Patterns Revisited: A Physics-Based and Machine Learning Hybrid Approach to Snow Distribution Mapping in the Sub-Arctic

Snowpack distribution in Arctic and alpine landscapes often occurs in repeating, year-to-year patterns due to local topographic, weather, and vegetation characteristics. Previous studies have suggested that with years of observational data, these snow distribution patterns can be statistically integrated into a snow process modeling workflow. Recent advances in snow hydrology and machine learning (ML) have increased our ability to predict snowpack distribution using in-situ observations, remote sensing data sets, and simple landscape characteristics that can be easily obtained for most environments. Here, we propose a hybrid approach to couple a ML snow distribution pattern (MLSDP) map with a physics-based, snow process model. We trained a random forest ML algorithm on tens of thousands of snow survey observations from a subarctic study area on the Seward Peninsula, Alaska, collected during peak snow water equivalent (SWE). We validated hybrid model outputs using in-situ snow depth and SWE observations, as well as a light detection and ranging data set and a distributed temperature profiling sensor data set. When the hybrid results were compared with the physics-based method, the hybrid method more accurately depicted the spatial patterns of the snowpack, areas of drifting snow, and years when no in-situ observations were used in the random forest ML training data set. The hybrid method also showed improvements in root mean squared error at 61% of locations where time-series estimations of snow depth were observed. These results can be applied to any physics-based model to improve the snow distribution patterning to reflect observed conditions in high latitude and high elevation cold region environments.

54 ENVIRONMENTAL SCIENCES

Plan Position Indicator Hydrometeor Field Statistics (PPIHYD) Evaluation Data Product Version 1.0

The PPIHYD evaluation data product provides distinct hydrometeor field statistics calculated from U.S. Department of Energy Atmospheric Radiation Measurement (ARM) user facility scanning radar plan position indicator (PPI) scans. These statistics include the equivalent reflectivity factor and Doppler spectral width percentiles, min/max values, and first four moments (mean, standard deviation, skewness, and kurtosis) of distinct hydrometeor features (clustered hydrometeor fields). Statistics also include morphological properties, water content and precipitation rate parameterization-based estimates, and thermodynamic properties interpolated using the Interpolated Sonde value-added product (INTERPSONDE VAP). The data set is organized in tabular form and is accompanied by mask arrays with corresponding indices. This straightforward file structure simplifies scanning radar data processing and renders this data set useful for process understanding and model evaluation studies. This report describes the data set and its processing algorithm and provides some examples.

54 ENVIRONMENTAL SCIENCES

Functional protein mining with conformal guarantees

Molecular structure prediction and homology detection offer promising paths to discovering protein function and evolutionary relationships. However, current approaches lack statistical reliability assurances, limiting their practical utility for selecting proteins for further experimental and in-silico characterization. To address this challenge, we introduce a statistically principled approach to protein search leveraging principles from conformal prediction, offering a framework that ensures statistical guarantees with user-specified risk and provides calibrated probabilities (rather than raw ML scores) for any protein search model. Our method (1) lets users select many biologically-relevant loss metrics (i.e. false discovery rate) and assigns reliable functional probabilities for annotating genes of unknown function; (2) achieves state-of-the-art performance in enzyme classification without training new models; and (3) robustly and rapidly pre-filters proteins for computationally intensive structural alignment algorithms. Our framework enhances the reliability of protein homology detection and enables the discovery of uncharacterized proteins with likely desirable functional properties.

59 BASIC BIOLOGICAL SCIENCES

Designing an Optimal Sensor Network via Minimizing Information Loss

Optimal experimental design is a classic topic in statistics, with many well-studied problems, applications, and solutions. The design problem we study is the placement of sensors to monitor spatiotemporal processes, explicitly accounting for the temporal dimension in our modeling and optimization. We observe that recent advancements in computational sciences often yield large datasets based on physics-based simulations, which are rarely leveraged in experimental design. We introduce a novel model-based sensor placement criterion, along with a highly-efficient optimization algorithm, which integrates physics-based simulations and Bayesian experimental design principles to identify sensor networks that “minimize information loss” from simulated data. Our technique relies on sparse variational inference and (separable) Gauss-Markov priors, and thus may adapt many techniques from Bayesian experimental design. We validate our method through a case study monitoring air temperature in Phoenix, Arizona, using state-of-the-art physics-based simulations. Our results show our framework to be superior to random or quasi-random sampling, particularly with a limited number of sensors. We conclude by discussing practical considerations and implications of our framework, including more complex modeling tools and real-world deployments.

54 ENVIRONMENTAL SCIENCES

Testing convolutional neural network based deep learning systems: a statistical metamorphic approach

Machine learning technology spans many areas and today plays a significant role in addressing a wide range of problems in critical domains,i.e., healthcare, autonomous driving, finance, manufacturing, cybersecurity,etc. Metamorphic testing (MT) is considered a simple but very powerful approach in testing such computationally complex systems for which either an oracle is not available or is available but difficult to apply. Conventional metamorphic testing techniques have certain limitations in verifying deep learning-based models (i.e., convolutional neural networks (CNNs)) that have a stochastic nature (because of randomly initializing the network weights) in their training. In this article, we attempt to address this problem by using a statistical metamorphic testing (SMT) technique that does not require software testers to worry about fixing the random seeds (to get deterministic results) to verify the metamorphic relations (MRs). We propose seven MRs combined with different statistical methods to statistically verify whether the program under test adheres to the relation(s) specified in the MR(s). We further use mutation testing techniques to show the usefulness of the proposed approach in the healthcare space and test two CNN-based deep learning models (used for pneumonia detection among patients). The empirical results show that our proposed approach uncovers 85.71% of the implementation faults in the classifiers under test (CUT). Furthermore, we also propose an MRs minimization algorithm for the CUT, thus saving computational costs and organizational testing resources.

Computer Science

Generation of random geological models using multi-randomization for machine learning

Generating high-fidelity geological models is essential for advancing machine learning (ML) methods in automated seismic interpretation. For instance, seismic images paired with corresponding fault labels are foundational for ML-based fault detection from seismic migration sections. While several open-access datasets of random geological models exist, open-source tools specifically designed to produce large volumes of such models for ML applications remain scarce. To address this gap, we present RGM (Random Geological Model), an open-source software package for efficiently generating 2D and 3D synthetic geological models tailored for ML workflows. RGM supports the creation of diverse model components, including medium property distributions (P-/S-wave velocities and density), seismic reflectivity images (i.e., synthetic migration sections), relative geological time, and discrete fault attributes such as probability, dip, strike, rake, and displacement. It also accommodates the creation of complex geological features such as salt bodies and unconformities. The model generation algorithm employs a multi-randomization strategy, yielding an effectively infinite-dimensional model space that encompasses a wide range of geological scenarios and associated seismic features. Furthermore, RGM incorporates a method to generate synthetic elastic migration images using analytical elastic reflection coefficients combined with frequency-dependent scaling. This functionality enables the creation of training datasets for ML models that leverage elastic seismic images. RGM is implemented in modern object-oriented Fortran, allowing users to flexibly control statistical parameters governing model variability. We demonstrate the capability, performance, and geological realism of the package through comprehensive 2D and 3D examples.

58 GEOSCIENCES