Search NASA⌕ Search

SEARCH · Search NASA

Results for “Statistical methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

A Frequency Domain Methodology for Quantitative Evaluation of Diffuse Wavefield With Applications to Seismic Imaging

Abstract Ambient Noise Imaging (ANI) of subsurface structures relies on seismic interferometry of diffuse seismic wavefields. However, the lack of effective methods to quantify and identify highly diffuse waves hampers applications of ANI, particularly in evaluating seismic attenuation and monitoring structural changes with high temporal resolution. Conventional ANI approaches require data normalization, which effectively suppresses the non‐diffuse component with large amplitude but also results in significant loss of amplitude and phase information in the continuous seismic records. In this study, we propose a frequency domain method to quantitatively evaluate the degree of diffuseness of seismic wavefields by analyzing their statistical characteristics of modal amplitudes for stationarity and randomness. Tests on synthetic waveform and field nodal records show that the proposed method can effectively distinguish between diffuse and non‐diffuse waveforms for either single‐ or three‐component data. As an application, we identify a 60‐s‐long diffuse coda of a local M 2.2 earthquake recorded by a dense nodal array on the San Jacinto Fault Zone, and successfully extract high‐quality dispersion curve andQ‐value without performing data normalization. These results are consistent with those obtained by conventional methods that assess the correlation between coherency and the Green's function, and by modeling ballistic waves generated by road traffic. Our proposed method can advance the imaging of subsurface velocity and attenuation structures as well as monitoring temporal changes for scientific studies and engineering applications.

Geochemistry & Geophysics↗

Techniques for improved statistical convergence in quantification of eddy diffusivity moments

While recent approaches, such as the macroscopic forcing method (MFM) or Green's function-based approaches, can be used to compute Reynolds-averaged Navier-Stokes closure operators using forced direct numerical simulations, MFM can also be used to directly compute moments of the effective nonlocal and anisotropic eddy diffusivities. The low-order spatial and temporal moments contain limited information about the eddy diffusivity but are often sufficient for quantification and modeling of nonlocal and anisotropic effects. However, when using MFM to compute eddy diffusivity moments, the statistical convergence can be slow for higher-order moments. In this work, we demonstrate that using the same direct numerical simulation (DNS) for all forced MFM simulations improves statistical convergence of the eddy diffusivity moments. We present its implementation in conjunction with a decomposition method that handles the MFM forcing semianalytically and allows for consistent boundary condition treatment, which we develop for both scalar and momentum transport. We demonstrate that for a two-dimensional Rayleigh-Taylor instability case study, using the same DNS for all forced MFM simulations results in convergence with 𝒪⁡(100) simulations rather than 𝒪⁡(1000) simulations. In conclusion, we then demonstrate the impacts of improved convergence on the quantification of the eddy diffusivity.

general physics↗

OmicsMLMentor: A Web Application for Guided Machine Learning Analysis of Omics Data

Expression-based omics technologies (e.g. proteomics, metabolomics, transcriptomics, etc.) increasingly rely on supervised and unsupervised machine learning (ML) models to find key biomolecules distinguishing conditions, identify natural groupings in biological data, or generate predictions for outcomes of interest. Fitting ML models to omics data presents several challenges, including handling missing data, selecting a normalization method, choosing a valid model, and optimizing hyperparameters, all requiring statistical programming skills to address these challenges. Thus, the open-source web application SLOPE was designed to lower the barrier to ML modeling for omics data. SLOPE supports the fitting of 15 ML models (10 supervised and 5 unsupervised) tailored to omics datasets, such as proteomics, metabolomics, lipidomics, and transcriptomics. SLOPE offers several omics-specific features, including methods for handling missingness (imputation, conversion, removal), normalization tests, ranking of models based on the structure of a user’s data and user input, and optimal hyperparameter selections using cross-validation splits. By streamlining ML workflows for omics analysis, SLOPE address critical gaps in existing online web tools, facilitating a broader adoption of these models for omics research. Here, SLOPE is applied to data from a lignin exposure study to highlight the workflow for fitting both supervised and unsupervised models to data.

lipidomics↗

Hierarchical Gaussian Random Field Sampling for Multilevel Markov Chain Monte Carlo: Coupling Stochastic Partial Differential Equation and the Karhunen–Loève Decomposition

This work introduces structure preserving hierarchical decompositions for sampling Gaussian random fields (GRFs) within the context of multilevel Bayesian inference in high-dimensional space. Existing scalable hierarchical sampling methods, such as those based on stochastic partial differential equations (SPDEs), often reduce the dimensionality of the sample space at the cost of accuracy of inference. Other approaches, such that those based on Karhunen-Loève (KL) expansions, offer sample space dimensionality reduction but sacrifice GRF representation accuracy and ergodicity of the Markov chain Monte Carlo (MCMC) sampler and are computationally expensive for high-dimensional problems. The proposed method integrates the dimensionality reduction capabilities of KL expansions with the scalability of SPDE-based sampling, thereby providing a robust, unified framework for high-dimensional uncertainty quantification (UQ) that is scalable and accurate, preserves ergodicity, and offers dimensionality reduction of the sample space. The hierarchy in our multilevel algorithm is derived from the geometric multigrid hierarchy. By constructing a hierarchical decomposition that maintains the covariance structure across the levels in the hierarchy, the approach enables efficient coarse-to-fine sampling while ensuring that all samples are drawn from the desired distribution. The effectiveness of the proposed method is demonstrated on a benchmark subsurface flow problem, demonstrating its effectiveness in improving computational efficiency and statistical accuracy. Furthermore, our proposed technique is more efficient and accurate and displays better convergence properties than existing methods for high-dimensional Bayesian inference problems.

Gaussian random fields↗

CoverM: read alignment statistics for metagenomics

SUMMARY: Genome-centric analysis of metagenomic samples is a powerful method for understanding the function of microbial communities. Calculating read coverage is a central part of analysis, enabling differential coverage binning for recovery of genomes and estimation of microbial community composition. Coverage is determined by processing read alignments to reference sequences of either contigs or genomes. Per-reference coverage is typically calculated in an ad-hoc manner, with each software package providing its own implementation and specific definition of coverage. Here we present a unified software package CoverM which calculates several coverage statistics for contigs and genomes in an ergonomic and flexible manner. It uses "Mosdepth arrays" for computational efficiency and avoids unnecessary I/O overhead by calculating coverage statistics from streamed read alignment results. AVAILABILITY AND IMPLEMENTATION: CoverM is free software available at https://github.com/wwood/coverm. CoverM is implemented in Rust, with Python (https://github.com/apcamargo/pycoverm) and Julia (https://github.com/JuliaBinaryWrappers/CoverM_jll.jl) interfaces.

Aroney, Samuel T N↗

Improvement of Drop‐Hammer Impact Testing for Safety Assessment of High Explosives Using 10‐mg Samples

Here, in this study, we established an improved method for drop-hammer impact testing of small quantities of high explosives (10 mg). We performed about seven hundred impact tests under various experimental conditions (e.g., sandpaper vs bare anvil, different sample masses, drop-weights, and striker diameters) to determine an optimal set of conditions and reaction detection methods (e.g., gas analysis, video, and sound recordings) that give the most statistically reliable results with 10 mg samples. We used both Frequentist and Bayesian statistical approaches to compare estimates of the drop height (DH50) that initiates a reaction 50% of the time, and to quantify the associated uncertainty. Gas analysis proved to be the most reliable reaction detection method, showing unambiguous rises in HE decomposition products (e.g., CO 2 ) even when the other indicators (e.g., sound, video) were inconclusive. The impact tests performed with a bare anvil showed much better reproducibility than those conducted with sandpaper, reducing the largest uncertainty observed in the data sets by a factor of 1.7. The DH 50 values obtained from three different sample masses (10, 20, and 35 mg) fell within the uncertainties of the measurements. We demonstrated the improved procedure (i.e., 10-mg samples, gas analysis, bare anvil, and Bayesian approach) on a variety of PETN samples having different surface areas and thermal histories.

PETN↗

Method-independent cusps for atomic orbitals in quantum Monte Carlo

Here, we present an approach for augmenting Gaussian atomic orbitals with correct nuclear cusps. Like the atomic orbital basis set itself and unlike previous cusp corrections, this approach is independent of the many-body method used to prepare wave functions for quantum Monte Carlo. Once the basis set and molecular geometry are specified, the cusp-corrected atomic orbitals are uniquely specified, regardless of which density functionals, quantum chemistry methods, or subsequent variational Monte Carlo optimizations are employed. We analyze the statistical improvement offered by these cusps in a number of molecules and find them to offer similar advantages as molecular-orbital-based approaches while remaining independent of the choice of many-body method.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Hydrodynamic fluctuations near a Hopf bifurcation: Stochastic onset of vortex shedding behind a circular cylinder

Here, we investigate hydrodynamic fluctuations in the flow past a circular cylinder near the critical Reynolds number Re c for the onset of vortex shedding. Starting from the fluctuating Navier-Stokes equations, we perform a perturbation expansion around Re c to derive analytical expressions for the statistics of the fluctuating lift force. Molecular-level simulations using the direct simulation Monte Carlo method support the theoretical predictions of the lift power spectrum and amplitude distribution. Notably, we have been able to collect sufficient statistics at distances Re ⁡/ Re c – 1 = O ⁡(10 –3 ) from the instability that confirm the appearance of non-Gaussian fluctuations, and we observe that they are associated with intermittent vortex shedding. These results emphasize how unavoidable thermal-noise-induced fluctuations become dramatically amplified in the vicinity of oscillatory flow instabilities and that their onset is fundamentally stochastic.

42 ENGINEERING↗

A novel framework for increasing research transparency: Exploring the connection between diversity and innovation

A split sample/dual method research protocol is demonstrated to increase transparency while reducing the probability of false discovery. We apply the protocol to examine whether diversity in ownership teams increases or decreases the likelihood of a firm reporting a novel innovation using data from the 2018 United States Census Bureau’s Annual Business Survey. Transparency is increased in three ways: 1) all specification testing and identifying potentially productive models is done in an exploratory subsample that 2) preserves the validity of hypothesis test statistics fromde novoestimation in the holdout confirmatory sample with 3) all findings publicly documented in an earlier registered report and in this journal publication. Bayesian estimation procedures that leverage information from the exploratory stage included in the confirmatory stage estimation replace traditional frequentist null hypothesis significance testing. In addition to increasing statistical power by using information from the full sample, Bayesian methods directly estimate a probability distribution for the magnitude of an effect, allowing much richer inference. Estimated magnitudes of diversity along academic discipline, race, ethnicity, and foreign-born status dimensions are positively associated with innovation. A maximally diverse ownership team on these dimensions would be roughly six times more likely to report new-to-market innovation than a homophilic team.

Science & Technology - Other Topics↗

Quantifying and simulating the weather forecast uncertainty for advanced building control

Weather forecast uncertainty is unavoidable despite technological advancements. Accurately quantifying and modelling this uncertainty is essential for developing and comparing advanced building controllers. In this study, we present a structured approach using a first-order autoregressive model (AR(1)) to model uncertainty in ambient temperature and global solar irradiation (GHI) forecasts. We analyzed weather data from four cities and employed Jensen–Shannon divergence (JSD) to evaluate the similarity between synthetic and actual forecast errors. The average JSD values for temperature are 0.027 (Berkeley), 0.021 (Leuven), 0.018 (Berlin), and 0.008 (Oslo), and for GHI, the average JSD values are 0.016 (Berkeley), 0.058 (Leuven), and 0.013 (Berlin). The low JSD values indicate a high similarity between the synthetic and real forecast error distributions. Further, our approach successfully generates synthetic weather forecasts that mirror the statistical properties of actual forecasts. The implementation of our method for uncertain forecast generation is being added to the BOPTEST framework.

54 ENVIRONMENTAL SCIENCES↗

Impact of composition and symmetry energy on the temperature of quasiprojectiles simulated with antisymmetrized molecular dynamics

The equation of state describes the emergent physical properties of matter. Experimental data is needed to help constrain the equation of state for nuclear matter. These constraints can help distinguish between an “asy-stiff” and an “asy-soft” equation of state, which has astrophysical implications. One path to help constrain the models is to analyze the nuclear caloric curve; some experiments have shown dependence on neutron excess, and may thus be sensitive to the asymmetry. A difference in the caloric curve based on the asymmetry of the reconstructed quasiprojectile (QP) had been observed using 70 Zn on 70 Zn at 35 MeV/nucleon taken with the NIMROD array. Antisymmetrized molecular dynamics calculations were performed for the same system and deexcited with gemini++. Both Gogny (asy-soft) and Gogny-as (asy-stiff) data sets were generated. The particles were then filtered based on detector geometric acceptance and thresholds. From the accepted particles, the excitation energy and temperature were calculated in the same way as for experimental data. Additionally, filter effects on the observed nuclear caloric curves were investigated. A tendency for the asy-stiff nuclear caloric curves to have higher temperatures than their asy-soft counterparts was observed for a number of probes. In addition, some probes may show sensitivity to the reconstructed composition of the QP, but this is inconclusive due to high statistical fluctuations and a large dependence on the exact method of event selection.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Lectures on statistical mechanics

Presented here is a transcription of the lecture notes from Professor Allan N. Kaufman’s graduate statistical mechanics course Physics 212A and 212B at the University of California Berkeley from the 1972–1973 academic year. 212A addressed equilibrium statistical mechanics with topics: fundamentals (micro-canonical and sub-canonical ensembles, adiabatic law and action conservation, fluctuations, pressure, and virial theorem), classical fluids and other systems (equation of state, deviations from ideality, virial coefficients and van der Waals potential, canonical ensemble and partition function, quasistatic evolution, grand-canonical ensemble and partition function, chemical potential, simple model of a phase transition, quantum virial expansion, numerical simulation of equations of state, and phase transition), chemical equilibrium (systems with multiple species and chemical reactions, law of mass action, Saha equation, chemical equilibrium including ionization and excited states), and long-range interactions (including Coulomb, dipole, and gravitational interactions, Debye–Hückel theory, and shielding). 212B addressed nonequilibrium statistical mechanics with topics: fundamentals (definitions: realizations, moments, characteristic function, and discrete variables), Brownian motion (Langevin equation, fluctuation–dissipation theorem, spatial diffusion, Boltzmann’s H-theorem), Liouville and Klimontovich equations, Landau equation (derivation, elaboration, and H-theorem, and irreversibility), Markov processes and Fokker–Planck equation (derivations of the Fokker–Planck equation and a master equation), linear response and transport theory (linear Boltzmann equation, linear response theory of Kubo and Mori, relation of entropy production to electrical conductivity, transport relations and coefficients, normal mode solutions of the transport equations, sketch of a generalized Langevin equation method for transport theory), and an introduction to nonequilibrium quantum statistical mechanics.

plasma dynamics↗

Improving neutrino energy estimation of charged-current interaction events with recurrent neural networks in MicroBooNE

We present a deep learning-based method for estimating the neutrino energy of charged-current neutrino-argon interactions. We employ a recurrent neural network (RNN) architecture for neutrino energy estimation in the MicroBooNE experiment, utilizing liquid argon time projection chamber (LArTPC) detector technology. Traditional energy estimation approaches in LArTPCs, which largely rely on reconstructing and summing visible energies, often experience sizable biases and resolution smearing because of the complex nature of neutrino interactions and the detector response. The estimation of neutrino energy can be improved after considering the kinematics information of reconstructed final-state particles. Utilizing kinematic information of reconstructed particles, the deep learning-based approach shows improved resolution and reduced bias for the muon neutrino Monte Carlo simulation sample compared to the traditional approach. In order to address the common concern about the effectiveness of this method on experimental data, the RNN-based energy estimator is further examined and validated with dedicated data-simulation consistency tests using MicroBooNE data. We also assess its potential impact on a neutrino oscillation study after accounting for all statistical and systematic uncertainties and show that it enhances physics sensitivity. This method has good potential to improve the performance of other physics analyses. Published by the American Physical Society 2024

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Fully quantum algorithm for mesoscale fluid simulations with application to partial differential equations

Fluid flow simulations marshal our most powerful computational resources. In many cases, even this is not enough. Quantum computers provide an opportunity to speed up traditional algorithms for flow simulations. We show that lattice-based mesoscale numerical methods can be executed as efficient quantum algorithms due to their statistical features. This approach revises a quantum algorithm for lattice gas automata to reduce classical computations and state preparation at every time step. For this, the algorithm approximates the qubit relative phases and subtracts them at the end of each time step. Phases are evaluated using the iterative phase estimation algorithm and subtracted using single-qubit rotation phase gates. Further, this method optimizes the quantum resource required and makes it more appropriate for near-term quantum hardware. We also demonstrate how the checkerboard deficiency that the D1Q2 scheme presents can be resolved using the D1Q3 scheme. The algorithm is validated by simulating two canonical partial differential equations: the diffusion and Burgers' equations on different quantum simulators. We find good agreement between quantum simulations and classical solutions for the presented algorithm.

97 MATHEMATICS AND COMPUTING↗

EMPDF : inferring the Milky Way mass with data-driven distribution function in phase space

We introduce the emPDF (empirical distribution function), a novel dynamical modelling method that infers the gravitational potential from kinematic tracers with optimal statistical efficiency under the minimal assumption of steady state. emPDF determines the best-fitting potential by maximizing the similarity between instantaneous kinematics and the time-averaged phase-space distribution function (DF), which is empirically constructed from observation upon the theoretical foundation of oPDF (Han et al. 2016). This approach eliminates the need for presumed functional forms of DFs or orbit libraries required by conventional DF- or orbit-based methods. emPDF stands out for its flexibility, efficiency, and capability in handling observational effects, making it preferable to the popular Jeans equation or other minimal assumption methods, especially for the Milky Way (MW) outer halo where tracers often have limited sample size and poor data quality. We apply emPDF to infer the MW mass profile using Gaia DR3 data of satellite galaxies and globular clusters, obtaining enclosed masses of M (,r) = 26±8, 46±8, 90±13⁠, and 149±40 x 10 10 M ⊙ at r = 30, 50, 100⁠, and 200 kpc, respectively. These are consistent with the updated constraints from simulation-informed DF fitting (Li et al. 2020). While the simulation-informed DF offers superior precision owing to the additional information extracted from simulations, emPDF is independent of such supplementary knowledge and applicable to general tracer populations. emPDF is currently implemented for tracers with complete 6D kinematics within spherical potentials, but it can potentially be extended to address more general problems.

Astrophysics of Galaxies (astro-ph.GA)↗

Isotopic gamma lines for identification of shielding materials

Identifying the constituting materials of concealed objects is crucial in a wide range of sectors, such as medical imaging, geophysics, nonproliferation, national security investigations, and so on. Existing methods face limitations, particularly when multiple materials are involved or when there are challenges posed by scattered radiation and large areal mass. Here we introduce a novel brute-force statistical approach for material identification using high spectral resolution detectors, such as HPGe. The method relies upon updated semianalytic formulae for computing uncollided flux from source of gamma radiation, shielded by a sequence of nested spherical or cylindrical materials. These semianalytical formulae make possible rapid flux estimation for material characterization via combinatorial search through all possible combinations of materials, using a high-resolution HPGe counting detector. An important prerequisite for the method is that the geometry of the objects is known (for example, from X-ray radiography). We demonstrate the viability of this material characterization technique in several use cases with both simulated and experimental data in spherical geometry.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Real-time neutron multiplicity and source localization for criticality safety during fuel debris removal

Advancing neutron detection and analysis techniques for complex radiation environments is an ongoing focus in nuclear instrumentation and monitoring. This proposal presents research and development of a generalized real-time neutron monitoring and analysis system, applicable to any detector capable of producing time-tagged neutron count data. While the work is demonstrated using the Neutron Multiplication Analysis Detector (NoMAD), a modular 15-tube helium-3 (He-3) array, due to its availability, spatial resolution, and flexible deployment, the methods developed are extensible to other systems, including organic scintillators and fast digital detectors. This research investigates two complementary analytical techniques for real-time characterization of neutron emitting sources: neutron multiplicity estimation based on the Hage-Cifarelli formalism and spatial localization using supervised machine learning applied to spatial count rate patterns. These methods are designed to operate under dynamic, evolving conditions such as fuel debris retrieval or reactor startup, where neutron-emitting material geometries may be partially unknown or changing over time. By integrating statistical neutron emission data with spatial localization, this research aims to develop and evaluate methods for real time neutron monitoring, source characterization, and material verification. Key contributions include implementation of a low-latency data pipeline for continuous neutron multiplicity analysis, development and validation of machine learning models for spatial inference, and experimental evaluation of system performance under variable measurement conditions. The outcomes are intended to support applications in nuclear safeguards, verification, emergency response, and reactor startup.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

The impact of urban configuration types on urban heat islands, air pollution, CO 2 emissions, and mortality in Europe: a data science approach

The world is becoming increasingly urbanized. As cities around the world continue to grow, it is important for urban planners and policymakers to understand how different urban configuration patterns affect the environment and human health. We aimed at identifying European urban configuration types, based on the Local Climate Zones categories and street design variables from Open Street Map, and evaluating their association with motorized traffic flows, Surface Urban Heat Island (SUHI) intensities, tropospheric nitrogen dioxide (NO 2 ), CO 2 per capita emissions and age-standardized mortality. We considered 946 European cities from 31 countries for the analysis defined in the 2018 Urban Audit database, of which 919 European cities were analysed. Data were collected at a 250 m × 250 m grid cell resolution. We divided all cities into five concentric rings based on the Burgess concentric urban planning model and calculated the mean values of all variables for each ring. First, to identify distinct urban configuration types, we applied the Uniform Manifold Approximation and Projection for Dimension Reduction method, followed by the k-means clustering algorithm. Next, statistical differences in exposures (including SUHI) and mortality between the resulting urban configuration types were evaluated using a Kruskal–Wallis test followed by a post-hoc Dunn's test. We identified four distinct urban configuration types characterising European cities: compact high density (n=246), open low-rise medium density (n=245), open low-rise low density (n=261), and green low density (n=167). Compact high density cities were a small size, had high population densities, and a low availability of natural areas. In contrast, green low-density cities were a large size, had low population densities, and a high availability of natural areas and cycleways. The open low-rise medium and low-density cities were a small to medium size with medium to low population densities and low to moderate availability of green areas. Motorised traffic flows and NO 2 exposure were significantly higher in compact high density and open low rise medium density cities when compared with green low density and open low-rise low density cities. Additionally, green low-density cities had a significantly lower SUHI effect compared with all other urban configuration types. Per person CO 2 emissions were significantly lower in compact high density cities compared with green low density cities. Lastly, green low density cities had significantly lower mortality rates when compared with all other urban configuration types. Our findings indicate that, although the compact city model is more sustainable, European compact cities still face challenges related to poor environmental quality and health. Our results have notable implications for urban and transport planning policies in Europe and contribute to the ongoing discussion on which city models can bring the greatest benefits for the environment, climate, and health.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗