Search NASA⌕ Search

SEARCH · Search NASA

Results for “data statistics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Unlocking hidden information in sparse small-angle neutron scattering measurements

Hypothesis Small-Angle Neutron Scattering (SANS) is a powerful technique for studying soft matter systems such as colloids, polymers, and lyotropic phases, providing nanoscale structural insights. However, its effectiveness is limited by low neutron flux, leading to long acquisition times and noisy data. Here, we hypothesize that Bayesian statistical inference using Gaussian Process Regression (GPR) can reconstruct high-fidelity scattering data from sparse measurements by leveraging intensity smoothness and continuity. Experiments and Simulations The method was benchmarked computationally and validated through SANS experiments on various soft matter systems, including wormlike micelles, colloidal suspensions, polymeric structures, and lyotropic phases. GPR-based inference was applied to both experimental and synthetic data to evaluate its effectiveness in noise reduction and intensity reconstruction. Findings GPR significantly enhances SANS data quality and therefore reducing measurement times by up to two orders of magnitude. This cost-effective approach maximizes experimental efficiency, enabling high-throughput studies and real-time monitoring of dynamic systems. It is particularly beneficial for weakly scattering and time-sensitive studies. Beyond SANS, this framework applies to other low-SNR techniques, including laboratory-based small-angle X-ray scattering and various dynamical scattering methods. Furthermore, it offers transformative potential for compact neutron sources, enhancing their viability for structural analysis in resource-limited settings.

Small angle neutron scattering↗

Machine learning tools for epigenetics

The software provides machine learning analysis and visualization to detect patterns in epigenetic data, including conventional machine learning and statistical methods, and open-source packages like pyBigWig (https://github.com/deeptools/pyBigWig) for data processing. The software is written in python, it uses some python libraries.

Kim, Anastasiia↗

Sensitive Detection of Structural Differences using a Statistical Framework for Comparative Crystallography

Chemical and conformational changes underlie the functional cycles of proteins. Comparative crystallography can reveal these changes over time, over ligands, and over chemical and physical perturbations in atomic detail. A key difficulty, however, is that the resulting observations must be placed on the same scale by correcting for experimental factors. We recently introduced a Bayesian framework for correcting (scaling) X-ray diffraction data by combining deep learning with statistical priors informed by crystallographic theory. To scale comparative crystallography data, we here combine this framework with a multivariate statistical theory of comparative crystallography. By doing so, we find strong improvements in the detection of protein dynamics, element-specific anomalous signal, and the binding of drug fragments.

Hekstra, Doeke R. [Harvard Univ., Cambridge, MA (U↗

Quantifying Microstructure Variability in Laser Powder Bed Fusion 316 L Stainless Steel Microstructures with Spatial Statistics

Here, we have explored data-driven methods for material microstructure quantification that improve sensitivity to microstructural changes compared to traditional approaches. The methods integrate multiple microstructural properties, including grain morphology, crystallographic orientation, and material phase information. The simpler method employs maps of the Euclidean distance transformation metric to evaluate the morphology of grain boundary networks. The more intensive approach employs generalized spherical harmonic mapping for crystallographic orientations, per-pixel phase information, and a variational auto-encoder for dimensionality reduction and results in a multidimensional clustering of by microstructure similarity. Applied to an experimental dataset of additively manufactured steel, both methods detected slight variations in samples produced under nominally identical processing conditions. Both methods were able to distinguish between samples from multiple (nominally identical) builds, while the generalized spherical harmonics-based method could additionally cluster data samples rotated at two orientations on the build plate. The improved sensitivity of the methods, demonstrated through comparison with traditional microstructure characterization techniques, offers advantages for microstructure quantification and comparisons in advanced manufacturing applications.

SS316L↗

Transient anisotropic kernel for probabilistic learning on manifolds

PLoM (Probabilistic Learning on Manifolds) is a method introduced in 2016 for handling small training datasets by projecting an Itô equation from a stochastic dissipative Hamiltonian dynamical system, acting as the MCMC generator, for which the KDE-estimated probability measure with the training dataset is the invariant measure. PLoM performs a projection on a reduced-order vector basis related to the training dataset, using the diffusion maps (DMAPS) basis constructed with a time-independent isotropic kernel. In this paper, we propose a new ISDE projection vector basis built from a transient anisotropic kernel, providing an alternative to the DMAPS basis to improve statistical surrogates for stochastic manifolds with heterogeneous data. The construction ensures that for times near the initial time, the DMAPS basis coincides with the transient basis. For larger times, the differences between the two bases are characterized by the angle of their spanned vector subspaces. The optimal instant yielding the optimal transient basis is determined using an estimation of mutual information from Information Theory, which is normalized by the entropy estimation to account for the effects of the number of realizations used in the estimations. Consequently, this new vector basis better represents statistical dependencies in the learned probability measure for any dimension. Three applications with varying levels of statistical complexity and data heterogeneity validate the proposed theory, showing that the transient anisotropic kernel improves the learned probability measure.

Diffusion maps↗

Grid Reliability Statistics [SWR-25-45]

This codebase houses a suite of statistical and descriptive analyses of NERC GADS data of interest to grid modelers and planners. This repository is envisioned to house a collection of statistical and descriptive summaries of NERC GADS data, particularly summaries that on their own might be insufficient to warrant publication. In addition, it is intended to disseminate results rather than to enable reproduction as GADS is non-public.

Murphy, Sinnott [National Renewable Energy Laborat↗

The Past, Present and Future of Structural Health Monitoring: An Overview of Three Ages

This paper presents an overview of the discipline of structural health monitoring (SHM), organised in terms of three proposed ages. The first age is delineated by the prehistory of SHM and the period where nondestructing testing methods evolved into an organised set of principles built upon physics-based models; this age ended when the model-based approaches reached an impasse in terms of their ability to properly deal with real-world problems. The second age of SHM began with a transition to data-based methods based on statistical pattern recognition, which provided a holistic approach to SHM problems for the first time. This age arguably ended when the methods foundered in situations where the necessary training data were scarce. It is argued here that the third age began with the development of population-based SHM, which has been designed to overcome the problem of data scarcity. As there is very limited space in a single article to provide a comprehensive overview, an appendix has been provided here that gives a very systematic bibliography of SHM reviews—a meta-bibliography.

60 APPLIED LIFE SCIENCES↗

Misclassification in Workers’ Telecommuting Frequency Choices Using a Generalized Extreme Value Model

Telecommuting frequency is a response variable collected in travel surveys and is, therefore, prone to errors leading to mismeasurements or misclassification. Misclassification of explanatory variables is a common risk when using statistical modeling techniques. We define “misclassification” as a response reported or recorded in the wrong category; for example, a variable is recorded as a 1 when it should be 0. Here, in this context, this study aims to develop a statistical model to analyze telecommuting data which accounts for potential misclassification errors by building on existing literature in econometrics. The empirical analysis was undertaken using the 2017 National Household Travel Survey (NHTS) and the general extreme value (GEV) models available in the literature. Specifically, the frequency of telecommuting days was analyzed using the negative binomial (NB) model recast as the multinomial logit (MNL) model. By nature—and consistent with other studies—NHTS data are prone to errors that can be classified as intentional or unintentional misinformation provided by the person being interviewed. Ignoring these errors while modeling telecommuting frequencies using standard discrete count models can result in biased parameter estimates. The misclassification parameter was calculated for both over-reporting and under-reporting scenarios. The misclassification errors can be as high as 14% over-reported and 10% under-reported, particularly for the neighboring values. Statistical fit comparison between the models shows that models that ignore misclassification have worse data fit and biased parameter estimates with significant policy implications.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Statistically Resolved Planetary Boundary Layer Height Diurnal Variability Using Spaceborne Lidar Data

The Planetary Boundary Layer Height (PBLH) significantly impacts weather, climate, and air quality. Understanding the global diurnal variation of the PBLH is particularly challenging due to the necessity of extensive observations and suitable retrieval algorithms that can adapt to diverse thermodynamic and dynamic conditions. This study utilized data from the Cloud-Aerosol Transport System (CATS) to analyze the diurnal variation of PBLH in both continental and marine regions. By leveraging CATS data and a modified version of the Different Thermo-Dynamics Stability (DTDS) algorithm, along with machine learning denoising, the study determined the diurnal variation of the PBLH in continental mid-latitude and marine regions. The CATS DTDS-PBLH closely matches ground-based lidar and radiosonde measurements at the continental sites, with correlation coefficients above 0.6 and well-aligned diurnal variability, although slightly overestimated at nighttime. In contrast, PBLH at the marine site was consistently overestimated due to the viewing geometry of CATS and complex cloud structures. The study emphasizes the importance of integrating meteorological data with lidar signals for accurate and robust PBLH estimations, which are essential for effective boundary layer assessment from satellite observations.

54 ENVIRONMENTAL SCIENCES↗

Measurement of beam-recoil observables 𝐶 𝑥 and 𝐶 𝑧 for 𝐾 + ⁢Λ photoproduction

Exclusive photoproduction of 𝐾 + ⁢Λ final states off a proton target has been an important component in the search for missing nucleon resonances and our understanding of the production of final states containing strange quarks. Polarization observables have been instrumental in this effort. The current work is an extension of previously published CLAS results on the beam-recoil transferred polarization observables 𝐶 𝑥 and 𝐶 𝑧 . Here, we extend the kinematic range up to invariant mass 𝑊 = 3.33 GeV from the previous limit of 𝑊 = 2.5 GeV with significantly improved statistical precision in the region of overlap. These data will provide for tighter constraints on the reaction models used to unravel the spectrum of nucleon resonances and their properties by not only improving the statistical precision of the data within the resonance region, but also constraining 𝑡-channel processes that dominate at higher 𝑊 but extend into the resonance region.

Electrons↗

Machine Learning-Guided Identification of PET Hydrolases from Natural Diversity

The enzymatic depolymerization of poly(ethylene terephthalate) (PET) is emerging as a leading chemical recycling technology for waste polyester. As part of this endeavor, new candidate enzymes identified from natural diversity can serve as useful starting points for enzyme evolution and engineering. In this study, we improved upon HMM searches by applying an iterative machine learning strategy to identify 400 putative PET-degrading enzymes (PET hydrolases) from naturally occurring homologs. Using high-throughput (HTP) experimental techniques, we successfully expressed and purified >200 enzyme candidates and assayed them for PET hydrolysis activity as a function of pH, temperature, and substrate crystallinity. From this library, we discovered 91 previously unknown PET hydrolases, 35 of which retain activity at pH 4.5 on crystalline material, which are conditions relevant to developing more efficient commercial processes. Notably, four enzymes showed equal to or higher activity than LCC-ICCG, a benchmark PET hydrolase, at this challenging condition in our screening assay, and 11 of which have pH optima <7. Using these data, we identified regions of PETases statistically correlated to activity at lower pH. We additionally investigated the effect of condition-specific activity data on trained machine learning predictors and found a precision (putative hit rate) improvement of up to 30% compared to a Hidden Markov Model alone. Our findings show that by pointing enzyme discovery toward conditions of interest with multiple rounds of experimental and machine learning, we can discover large sets of active enzymes and explore factors associated with activity at those conditions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

GNSS-based Vegetation Optical Depth, Tree Sway, and Evapotranspiration data from the Niwot Ridge Subalpine Forest (US-NR1) AmeriFlux site

This data package contains data and information about Global Navigation Satellite System (GNSS)-based Vegetation Optical Depth (VOD), tree sway motion, and eddy-covariance evapotranspiration (ET) data collected at the Niwot Ridge Subalpine Forest AmeriFlux site (US-NR1). The raw GNSS data were collected between May 2022 and August 2023. Other processed datasets such as tree sway motion and ET data are also included. The goal was to study the water content within a subalpine forest and, more specifically, examine the canopy evaporation process. This data archive includes all data that were used within the following Biogeosciences discussion paper that further summarizes the research objectives and conclusions:Burns, S.P., V. Humphrey, E.D. Gutmann, M.S. Raleigh, D.R. Bowling, and P.D. Blanken, 2025: Using GNSS-based vegetation optical depth, tree sway motion, and eddy-covariance to examine evaporation of canopy-intercepted rainfall in a subalpine forest. EGUsphere [preprint],https://doi.org/10.5194/egusphere-2025-1755This data archive also supplements the 30-min Lawrence Berkeley National Laboratory (LBNL) AmeriFlux dataset for US-NR1 (i.e., https://doi.org/10.17190/AMF/1246088) and updates what was in the 2020 ESS-DIVE US-NR1 archive (https://doi.org/10.15485/1671825) to include data from the years 2020-2025. More specifically, the following updates are provided: (i) five-minute statistics (means, variances, covariances) of all data measured by the US-NR1 data system between Sep 2020 and Jun 2025 in netCDF format, (ii) the electronic logbook of US-NR1 site visits, (iii) a web calendar (in HTML format) documenting activity at the site (a replica of https://urquell.colorado.edu/calendar/), (iv) photos taken at the site between years 2020 and present day (Aug 2025), and (v) several auxiliary datasets, primary related to trees near the site, soil properties, soil moisture and soil temperature, and subcanopy radiation data. The data package is setup so that the web calendar, photos, and electronic logbook can be easily accessed on a local computer using a web browser. The provided data files are in either BINEX or SBF format (for the raw GNSS data), netCDF, CSV, ASCII, or MATLAB format. To obtain a better understanding about the archive, please start by reading the following PDF which is included within the data archive:README_ESS_DIVE_USNR1_2025_readme_first.pdf.

54 ENVIRONMENTAL SCIENCES↗

A physics informed bayesian optimization approach for material design: application to NiTi shape memory alloys

Abstract The design of materials and identification of optimal processing parameters constitute a complex and challenging task, necessitating efficient utilization of available data. Bayesian Optimization (BO) has gained popularity in materials design due to its ability to work with minimal data. However, many BO-based frameworks predominantly rely on statistical information, in the form of input-output data, and assume black-box objective functions. In practice, designers often possess knowledge of the underlying physical laws governing a material system, rendering the objective function not entirely black-box, as some information is partially observable. In this study, we propose a physics-informed BO approach that integrates physics-infused kernels to effectively leverage both statistical and physical information in the decision-making process. We demonstrate that this method significantly improves decision-making efficiency and enables more data-efficient BO. The applicability of this approach is showcased through the design of NiTi shape memory alloys, where the optimal processing parameters are identified to maximize the transformation temperature.

Chemistry↗

A More Precise Measurement of the Radius of PSR J0740+6620 Using Updated NICER Data

PSR J0740+6620 is the neutron star with the highest precisely determined mass, inferred from radio observations to be 2.08 ± 0.07 M ⊙ . Measurements of its radius therefore hold promise to constrain the properties of the cold, catalyzed, high-density matter in neutron star cores. Previously, Miller et al. and Riley et al. reported measurements of the radius of PSR J0740+6620 based on Neutron Star Interior Composition Explorer (NICER) observations accumulated through 2020 April 17, and an exploratory analysis utilizing NICER background estimates and a data set accumulated through 2021 December 28 was presented in Salmi et al. Here we report an updated radius measurement, derived by fitting models of X-ray emission from the neutron star surface to NICER data accumulated through 2022 April 21, totaling ~1.1 Ms additional exposure compared to the data set analyzed in Miller et al. and Riley et al., and to data from XMM-Newton observations. We find that the equatorial circumferential radius of PSR J0740+6620 is ${12.92}_{-1.13}^{+2.09}$ km (68% credibility), a fractional uncertainty ~83% the width of that reported in Miller et al., in line with statistical expectations given the additional data. If we were to require the radius to be less than 16 km, as was done in Salmi et al., then our 68% credible region would become $R={12.76}_{-1.02}^{+1.49}$ km, which is close to the headline result of Salmi et al. Our updated measurements, along with other laboratory and astrophysical constraints, imply a slightly softer equation of state than that inferred from our previous measurements.

79 ASTRONOMY AND ASTROPHYSICS↗

Cosmological constraints using Minkowski functionals from the first year data of the Hyper Suprime-Cam

We use Minkowski functionals to analyse weak lensing convergence maps from the first-year data release of the Subaru Hyper Suprime-Cam (HSC-Y1) survey. Minkowski functionals provide a description of the morphological properties of a field, capturing the non-Gaussian features of the Universe matter-density distribution. Using simulated catalogues that reproduce survey conditions and encode cosmological information, we emulate Minkowski functionals predictions across a range of cosmological parameters to derive the best-fit from the data. By applying multiple scales cuts, we rigorously mitigate systematic effects, including baryonic feedback and intrinsic alignments. From the analysis, combining constraints of the angular power spectrum and Minkowski functionals, we obtain S8≡σ8Ωm/0.3=0.808−0.046+0.033 and Ωm=0.293−0.043+0.157⁠. These results represent a 40 per cent improvement on the S8 constraints compared to using power spectrum only. Minkowski functionals results are consistent with other two-point, and higher order statistics constraints using the same data, being in agreement with CMB results from the Planck S8 measurements. Our study demonstrates the power of Minkowski functionals beyond two-point statistics to constrain and break the degeneracy between Ωm and σ8⁠.

79 ASTRONOMY AND ASTROPHYSICS↗

Dark energy survey year 3 results: likelihood-free, simulation-based w CDM inference with neural compression of weak-lensing map statistics

We present simulation-based cosmological wcold dark matter (wCDM) inference using dark energy survey year 3 weak-lensing maps, via neural data compression of weak-lensing map summary statistics: power spectra, peak counts, and direct map-level compression/inference with convolutional neural networks (CNN). Using simulation-based inference, also known as likelihood-free or implicit inference, we use forward-modelled mock data to estimate posterior probability distributions of unknown parameters. This approach allows all statistical assumptions and uncertainties to be propagated through the forward-modelled mock data; these include sky masks, non-Gaussian shape noise, shape measurement bias, source galaxy clustering, photometric redshift uncertainty, intrinsic galaxy alignments, non-Gaussian density fields, neutrinos, and non-linear summary statistics. We include a series of tests to validate our inference results. This paper also describes the Gower Street simulation suite: 791 full-sky pkdgrav3 dark matter simulations, with cosmological model parameters sampled with a mixed active-learning strategy, from which we construct over 3000 mock dark energy survey lensing data sets. For wCDM inference, for which we allow –1 < w < –$\frac{1}{3}$⁠, our most constraining result uses power spectra combined with map-level (CNN) inference. Using gravitational lensing data only, this map-level combination gives Ω m = 0.283$^{+0.020}_{–0.027}$⁠, S 8 = 0.804$^{+0.025}_{–0.017⁠}$, and w < –0.80 (with a 68 per cent credible interval); compared to the power spectrum inference, this is more than a factor of two improvement in dark energy parameter (Ω⁠ DE , w⁠) precision.

79 ASTRONOMY AND ASTROPHYSICS↗

Predictive Modeling of NOx Emissions from Lean Direct Injection of Hydrogen and Hydrogen/Natural Gas Blends Using Flame Imaging and Machine Learning

This research paper explores the use of machine learning to relate images of flame structure and luminosity to measured NOx emissions. Images of reactions produced by 16 aero-engine derived injectors for a ground-based turbine operated on a range of fuel compositions, air pressure drops, preheat temperatures and adiabatic flame temperatures were captured and postprocessed. The experimental investigations were conducted under atmospheric conditions, capturing CO, NO and NOx emissions data and OH* chemiluminescence images from 27 test conditions. The injector geometry and test conditions were based on a statistically designed test plan. These results were first analyzed using the traditional analysis approach of analysis of variance (ANOVA). The statistically based test plan yielded 432 data points, leading to a correlation for NOx emissions as a function of injector geometry, test conditions and imaging responses, with 70.2% accuracy. As an alternative approach to predicting emissions using imaging diagnostics as well as injector geometry and test conditions, a random forest machine learning algorithm was also applied to the data and was able to achieve an accuracy of 82.6%. This study offers insights into the factors influencing emissions in ground-based turbines while emphasizing the potential of machine learning algorithms in constructing predictive models for complex systems.

08 HYDROGEN↗