Search NASA⌕ Search

SEARCH · Search NASA

Results for “STATISTICS”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

Object-Based Evaluation of Dynamical and Statistical Downscaled Precipitation Products over CONUS

High-resolution precipitation data, generated through dynamical downscaling (DD) or statistical downscaling (SD) of global climate model output, provide critical information for regional climate assessment and adaptation planning. Most downscaling development and validation have focused on accurate gridscale precipitation construction and ignored the spatial structure of precipitation across model grids and at the event scale. However, many applications, e.g., hydrologic modeling and the analysis using the downscaled precipitation, require a reasonable representation of the spatial structure of precipitation within watersheds. Therefore, a set of standard metrics to evaluate the representation of the spatial structure of individual storms across diverse downscaled precipitation products is desired. To address this need, we conducted an object-based evaluation of precipitation in decades-long DD and SD products over the contiguous United States (CONUS). Specifically, we evaluate their ability to reproduce various features of precipitation objects in the observations: total volume, precipitation area, peak intensity, and spatial structure. Multiple metrics (bias, Perkins score, and nonparametric statistical tests) are used to quantify model performance. Our evaluation reveals notable variations in performance among individual products across different climate zones and seasons, as well as between extreme and nonextreme events. In general, most DD products exhibit balanced performance across the four precipitation object features, while SD products vary more significantly in their performance across products. Based on this comprehensive evaluation, we provide guidance on choosing downscaled products for specific regions, seasons, and precipitation object features. These findings and recommendations can inform precipitation-relevant modeling and analysis over CONUS, guide future downscaling technique developments, and provide actionable information for climate impact assessment and adaptation.

Downscaling↗

Forced Component Estimation Statistical Method Intercomparison Project (ForceSMIP)

Anthropogenic climate change is unfolding rapidly, yet its regional manifestation can be obscured by internal variability. A primary goal of climate science is to identify the externally forced climate response from among the noise of internal variability. Separating the forced response from internal variability can be addressed in climate models by using a large ensemble to average over different possible realizations of internal variability. However, with only one realization of the real world, it is a major challenge to isolate the forced response directly in observations. In the Forced Component Estimation Statistical Method Intercomparison Project (ForceSMIP), contributors used existing and newly developed statistical and machine learning methods to estimate the forced response over 1950–2022 within individual realizations of the climate system. Participants used neural networks, linear inverse models, fingerprinting methods, and low-frequency component analysis, among other approaches. These methods were trained using large ensembles from multiple climate models and then applied to observations. Here, we evaluate method performance within large ensembles and investigate the estimates of the forced response in observations. Our results show that many different types of methods are skillful for estimating the forced response in climate models, though the relative skill of individual methods varies depending on the variable and evaluation metric. Methods with comparable skill in models can give a wide range of estimates of the forced response pattern in observations, illustrating the epistemic uncertainty in forced response estimates. ForceSMIP gives new insights into the forced response in observations, its uncertainty, and methods for its estimation.

Climate attribution↗

A statistical model for interpreting computerized dynamic posturography data

Computerized dynamic posturography (CDP) is widely used for assessment of altered balance control. CDP trials are quantified using the equilibrium score (ES), which ranges from zero to 100, as a decreasing function of peak sway angle. The problem of how best to model and analyze ESs from a controlled study is considered. The ES often exhibits a skewed distribution in repeated trials, which can lead to incorrect inference when applying standard regression or analysis of variance models. Furthermore, CDP trials are terminated when a patient loses balance. In these situations, the ES is not observable, but is assigned the lowest possible score--zero. As a result, the response variable has a mixed discrete-continuous distribution, further compromising inference obtained by standard statistical methods. Here, we develop alternative methodology for analyzing ESs under a stochastic model extending the ES to a continuous latent random variable that always exists, but is unobserved in the event of a fall. Loss of balance occurs conditionally, with probability depending on the realized latent ES. After fitting the model by a form of quasi-maximum-likelihood, one may perform statistical inference to assess the effects of explanatory variables. An example is provided, using data from the NIH/NIA Baltimore Longitudinal Study on Aging.

NASA Discipline Neuroscience↗

Statistical and linguistic features of DNA sequences

We present evidence supporting the idea that the DNA sequence in genes containing noncoding regions is correlated, and that the correlation is remarkably long range--indeed, base pairs thousands of base pairs distant are correlated. We do not find such a long-range correlation in the coding regions of the gene. We resolve the problem of the "non-stationary" feature of the sequence of base pairs by applying a new algorithm called Detrended Fluctuation Analysis (DFA). We address the claim of Voss that there is no difference in the statistical properties of coding and noncoding regions of DNA by systematically applying the DFA algorithm, as well as standard FFT analysis, to all eukaryotic DNA sequences (33 301 coding and 29 453 noncoding) in the entire GenBank database. We describe a simple model to account for the presence of long-range power-law correlations which is based upon a generalization of the classic Levy walk. Finally, we describe briefly some recent work showing that the noncoding sequences have certain statistical features in common with natural languages. Specifically, we adapt to DNA the Zipf approach to analyzing linguistic texts, and the Shannon approach to quantifying the "redundancy" of a linguistic text in terms of a measurable entropy function. We suggest that noncoding regions in plants and invertebrates may display a smaller entropy and larger redundancy than coding regions, further supporting the possibility that noncoding regions of DNA may carry biological information.

Non-NASA Center↗

A statistical model of the human core-temperature circadian rhythm

We formulate a statistical model of the human core-temperature circadian rhythm in which the circadian signal is modeled as a van der Pol oscillator, the thermoregulatory response is represented as a first-order autoregressive process, and the evoked effect of activity is modeled with a function specific for each circadian protocol. The new model directly links differential equation-based simulation models and harmonic regression analysis methods and permits statistical analysis of both static and dynamical properties of the circadian pacemaker from experimental data. We estimate the model parameters by using numerically efficient maximum likelihood algorithms and analyze human core-temperature data from forced desynchrony, free-run, and constant-routine protocols. By representing explicitly the dynamical effects of ambient light input to the human circadian pacemaker, the new model can estimate with high precision the correct intrinsic period of this oscillator ( approximately 24 h) from both free-run and forced desynchrony studies. Although the van der Pol model approximates well the dynamical features of the circadian pacemaker, the optimal dynamical model of the human biological clock may have a harmonic structure different from that of the van der Pol oscillator.

NASA Discipline Regulatory Physiology↗

Statistical Analysis of CFD Solutions From the Fifth AIAA Drag Prediction Workshop

A graphical framework is used for statistical analysis of the results from an extensive N-version test of a collection of Reynolds-averaged Navier-Stokes computational fluid dynamics codes. The solutions were obtained by code developers and users from North America, Europe, Asia, and South America using a common grid sequence and multiple turbulence models for the June 2012 fifth Drag Prediction Workshop sponsored by the AIAA Applied Aerodynamics Technical Committee. The aerodynamic configuration for this workshop was the Common Research Model subsonic transport wing-body previously used for the 4th Drag Prediction Workshop. This work continues the statistical analysis begun in the earlier workshops and compares the results from the grid convergence study of the most recent workshop with previous workshops.

Statistical analysis↗

Ensemble Statistical Post-Processing of the National Air Quality Forecast Capability: Enhancing Ozone Forecasts in Baltimore, Maryland

An ensemble statistical post-processor (ESP) is developed for the National Air Quality Forecast Capability (NAQFC) to address the unique challenges of forecasting surface ozone in Baltimore, MD. Air quality and meteorological data were collected from the eight monitors that constitute the Baltimore forecast region. These data were used to build the ESP using a moving-block bootstrap, regression tree models, and extreme-value theory. The ESP was evaluated using a 10-fold cross-validation to avoid evaluation with the same data used in the development process. Results indicate that the ESP is conditionally biased, likely due to slight overfitting while training the regression tree models. When viewed from the perspective of a decision-maker, the ESP provides a wealth of additional information previously not available through the NAQFC alone. The user is provided the freedom to tailor the forecast to the decision at hand by using decision-specific probability thresholds that define a forecast for an ozone exceedance. Taking advantage of the ESP, the user not only receives an increase in value over the NAQFC, but also receives value for An ensemble statistical post-processor (ESP) is developed for the National Air Quality Forecast Capability (NAQFC) to address the unique challenges of forecasting surface ozone in Baltimore, MD. Air quality and meteorological data were collected from the eight monitors that constitute the Baltimore forecast region. These data were used to build the ESP using a moving-block bootstrap, regression tree models, and extreme-value theory. The ESP was evaluated using a 10-fold cross-validation to avoid evaluation with the same data used in the development process. Results indicate that the ESP is conditionally biased, likely due to slight overfitting while training the regression tree models. When viewed from the perspective of a decision-maker, the ESP provides a wealth of additional information previously not available through the NAQFC alone. The user is provided the freedom to tailor the forecast to the decision at hand by using decision-specific probability thresholds that define a forecast for an ozone exceedance. Taking advantage of the ESP, the user not only receives an increase in value over the NAQFC, but also receives value for

ozone↗

Statistical Analysis of Instantaneous Frequency Scaling Factor as Derived From Optical Disdrometer Measurements At KQ Bands

The rain rate data and statistics of a location are often used in conjunction with models to predict rain attenuation. However, the true attenuation is a function not only of rain rate, but also of the drop size distribution (DSD). Generally, models utilize an average drop size distribution (Laws and Parsons or Marshall and Palmer [1]). However, individual rain events may deviate from these models significantly if their DSD is not well approximated by the average. Therefore, characterizing the relationship between the DSD and attenuation is valuable in improving modeled predictions of rain attenuation statistics. The DSD may also be used to derive the instantaneous frequency scaling factor and thus validate frequency scaling models. Since June of 2014, NASA Glenn Research Center (GRC) and the Politecnico di Milano (POLIMI) have jointly conducted a propagation study in Milan, Italy utilizing the 20 and 40 GHz beacon signals of the Alphasat TDP#5 Aldo Paraboni payload. The Ka- and Q-band beacon receivers provide a direct measurement of the signal attenuation while concurrent weather instrumentation provides measurements of the atmospheric conditions at the receiver. Among these instruments is a Thies Clima Laser Precipitation Monitor (optical disdrometer) which yields droplet size distributions (DSD); this DSD information can be used to derive a scaling factor that scales the measured 20 GHz data to expected 40 GHz attenuation. Given the capability to both predict and directly observe 40 GHz attenuation, this site is uniquely situated to assess and characterize such predictions. Previous work using this data has examined the relationship between the measured drop-size distribution and the measured attenuation of the link [2]. The focus of this paper now turns to a deeper analysis of the scaling factor, including the prediction error as a function of attenuation level, correlation between the scaling factor and the rain rate, and the temporal variability of the drop size distribution both within a given rain event and across different varieties of rain events. Index Terms-drop size distribution, frequency scaling, propagation losses, radiowave propagation.

Radio Attenuation↗

Statistical Analysis of Instantaneous Frequency Scaling Factor as Derived From Optical Disdrometer Measurements At KQ Bands

The rain rate data and statistics of a location are often used in conjunction with models to predict rain attenuation. However, the true attenuation is a function not only of rain rate, but also of the drop size distribution (DSD). Generally, models utilize an average drop size distribution (Laws and Parsons or Marshall and Palmer. However, individual rain events may deviate from these models significantly if their DSD is not well approximated by the average. Therefore, characterizing the relationship between the DSD and attenuation is valuable in improving modeled predictions of rain attenuation statistics. The DSD may also be used to derive the instantaneous frequency scaling factor and thus validate frequency scaling models. Since June of 2014, NASA Glenn Research Center (GRC) and the Politecnico di Milano (POLIMI) have jointly conducted a propagation study in Milan, Italy utilizing the 20 and 40 GHz beacon signals of the Alphasat TDP#5 Aldo Paraboni payload. The Ka- and Q-band beacon receivers provide a direct measurement of the signal attenuation while concurrent weather instrumentation provides measurements of the atmospheric conditions at the receiver. Among these instruments is a Thies Clima Laser Precipitation Monitor (optical disdrometer) which yields droplet size distributions (DSD); this DSD information can be used to derive a scaling factor that scales the measured 20 GHz data to expected 40 GHz attenuation. Given the capability to both predict and directly observe 40 GHz attenuation, this site is uniquely situated to assess and characterize such predictions. Previous work using this data has examined the relationship between the measured drop-size distribution and the measured attenuation of the link]. The focus of this paper now turns to a deeper analysis of the scaling factor, including the prediction error as a function of attenuation level, correlation between the scaling factor and the rain rate, and the temporal variability of the drop size distribution both within a given rain event and across different varieties of rain events. Index Terms-drop size distribution, frequency scaling, propagation losses, radiowave propagation.

Statistics↗

Statistical Analysis of CFD Solutions from the 6th AIAA CFD Drag Prediction Workshop

A graphical framework is used for statistical analysis of the results from an extensive N-version test of a collection of Reynolds-averaged Navier-Stokes computational fluid dynamics codes. The solutions were obtained by code developers and users from North America, Europe, Asia, and South America using both common and custom grid sequences as well as multiple turbulence models for the June 2016 6th AIAA CFD Drag Prediction Workshop sponsored by the AIAA Applied Aerodynamics Technical Committee. The aerodynamic configuration for this workshop was the Common Research Model subsonic transport wing-body previously used for both the 4th and 5th Drag Prediction Workshops. This work continues the statistical analysis begun in the earlier workshops and compares the results from the grid convergence study of the most recent workshop with previous workshops.

Statistical analysis↗

Relevance of the American Statistical Society's Warning on P-Values for Conjunction Assessment

On March 7, 2016, the American Statistical Association (ASA) issued an online editorial paper on the "context, process, and purpose of p-values." According to the paper, "the statement articulates in non-technical terms a few select principles that could improve the conductor interpretation of quantitative science, according to widespread consensus in the statistical community." This "consensus" statement was accompanied by 21 "commentaries" expressing a diversity of opinions among the panel of experts ASA convened. In the present work, we express our view that the ASA p-value warning has relevance to the space object Conjunction Assessment (CA) community.

p-values↗

Statistical Descriptors of Composite Fiber Aggregation

This study introduces a method of characterizing fiber aggregation and resin rich regions in composite microstructures. Microscale models of representative elements (RVE) need to be indicative of the extend of clustering (i.e. close fiber-to-fiber interaction) and resin rich “pools” which may impact the overall strength and performance of a composite structure. This algorithm was used to evaluate different unidirectional 2-D microstructure scans, which will be compared to their manufacturing method or any special treatment processes. These cluster and pool scan statistics can be used as criteria to judge statistical equivalency of artificially constructed microstructures.

Statistics↗

Artificial Generation of 2-D Fiber Reinforced Composite Microstructures with Statistically Equivalent Features

Fiber reinforced composites are used widely for their high strength and low weight advantages in various aerospace and automotive applications. While their use may be sought after, modeling of these material requires increasing fidelity at the lower scales to capture accurate material behavior under loading. The first steps in creating statistically equivalent models to real life cases is developing a method of rapid evaluation and artificial microstructure generation. The outlined work is capable of tracking microscale fiber positions and determining regions of localized volume fraction extrema (high and low end). Groupings of high and low volume fraction regions are called clusters and their geometry is used to characterize the microstructure. These cluster features can be evaluated for both artificial models and actual scans, allowing correlation to be established which can ultimately be used to regenerate statistically equivalent models. The results of this work show that if one feature is to be correlated, a model can be generated which matches almost exactly. But once more features are equally taken into account, the regeneration loses accuracy.

micromechanics↗

Enhancing Interpretability in Generative Modeling: Statistically Disentangled Latent Spaces Guided by Generative Factors in Scientific Datasets

This study addresses the challenge of statistically extracting generative factors from complex, high-dimensional datasets in unsupervised or semi-supervised settings. We investigate encoder-decoder-based generative models for nonlinear dimensionality reduction, focusing on disentangling low-dimensional latent variables corresponding to independent physical factors. Introducing Aux-VAE, a novel architecture within the classical Variational Autoencoder framework, we achieve disentanglement with minimal modifications to the standard VAE loss function by leveraging prior statistical knowledge through auxiliary variables. These variables guide the shaping of the latent space by aligning latent factors with learned auxiliary variables. We validate the efficacy of Aux-VAE through comparative assessments on multiple datasets, including astronomical simulations.

97 MATHEMATICS AND COMPUTING↗

A statistical and simulation-informed model for estimating permeability from pore size distribution in saturated geomaterials

Accurate permeability estimation is essential across subsurface engineering applications but remains challenging due to the complex pore structures of natural geomaterials. Traditional empirical methods and simplified theoretical models often inadequately capture the role of pore size distribution and connectivity. Here, this study develops a statistical and simulation-informed permeability model that collapses pore-scale complexity into a compact scaling of the form k = αϕμ d 2 , where ϕ is porosity, μ d is mean pore size, and α is a weakly varying coefficient. By combining pore network simulations with statistical analysis of unimodal and bimodal pore size distributions, we identify three key findings: (i) permeability is much more sensitive to mean pore size than to porosity; (ii) across extensive datasets, the ratio σ d /μ d (standard deviation to mean) clusters around a characteristic value ∼0.4, allowing the effects of the full pore size distribution to be represented by μ d and a narrowly varying α ≈ 0.05; and (iii) for bimodal systems, there exists a critical fraction of small pores ∼0.78 above which flow becomes small-pore dominated, enabling the definition of an effective flow-controlling pore population and facilitating simplified permeability estimation for such systems. The resulting model, which requires only porosity and a representative mean pore size as inputs, is validated against comprehensive experimental datasets (>1700 samples) spanning diverse soils and rocks and achieves good predictive accuracy. Overall, this work provides a physically grounded yet practically simple permeability estimator suitable for subsurface engineering, environmental protection, and resource management applications.

Permeability↗

Statistical distributions for transient transport

Here, this paper introduces the use of statistical distributions based on transport differential equations for clear distinction of transport modes within transient kinetic experiments. More specifically, novel techniques are developed for the transient data obtained through the Temporal Analysis of Products (TAP) reactor and are applicable to experiments where pulse response into a gas flow is used. The methodology allows distinguishing between two domains of diffusion transport in heterogeneous catalytic systems, i.e., Knudsen and non-Knudsen diffusion, using statistical fingerprints, and finding the transition domain. Two distribution parameters were obtained that directly result in coefficients that correspond to the concentration and the rate of transport. Using a linear relationship between the rate and concentration coefficients, Knudsen diffusion is revealed when the rate of transport is constant and non-Knudsen diffusion is confirmed when the rate of transport coefficient is a function of the concentration coefficient. As a result, accurate transport information can be extracted from experimental data even in the presence of comprising instrument drift or noise particularly when analyzing higher pressure pulse responses with complex transport. This enables more direct investigation of experiments influenced by gas-phase reactions.

TAP reactor↗

A high-throughput approach for statistical process optimization in Laser Powder Bed Fusion

Process variability is inherent in metal additive manufacturing (AM). However, it is often overlooked in process optimization frameworks, constraining the understanding of process uncertainties and their influence on parameter selection. To address this, we present an integrated framework that combines high-throughput single-track experiments, GAN-based melt pool geometry extraction, robust statistical and machine learning modeling, and uncertainty-quantified process mapping. Process variability is characterized through single-track melt pool behaviors, and its influence on defect formation is systematically quantified to enable statistically guided process parameter optimization. This approach is demonstrated on Laser Powder Bed Fusion (L-PBF) of stainless steel 316L, effectively capturing the interplay between process parameters, melt pool variability, and defect probability. By integrating uncertainty quantification into process optimization, this study provides a structured methodology for addressing variability challenges in AM quality control, ultimately contributing to enhanced manufacturing reliability.

Laser Powder Bed Fusion↗

Statistical White-Line Analysis in High-Throughput TXM-XANES for Chemical State Quantification

The transmission X-ray microscopy (TXM) based X-ray absorption near-edge structure (XANES) technique provides three-dimensional mapping of element-specific chemical states at nanometer-scale spatial resolution and micrometer-scale fields of view. However, compared to conventional volume-averaged XANES (VA-XANES) measurements, the inherently small voxel size in TXM-XANES leads to a lower signal-to-noise ratio, making full-spectrum analysis computationally demanding and less robust. Here, we present the structural and compositional conditions for a statistical white-line analysis framework under which chemical state information can be directly extracted from the white-line peak position in voxel spectra without the need for voxel-wise background subtraction or normalization, under well-defined structural and compositional conditions. The method is validated on layered oxide cathode materials, where low-order polynomial fitting accurately reproduces white-line features, and the extracted energy distributions correlate strongly with VA-XANES results. This statistical approach enables high-throughput, dose-efficient, and noise-robust chemical state quantification in TXM-XANES, offering broad applicability to functional materials requiring nanoscale oxidation-state mapping.

TXM↗