Search NASASearch

SEARCH · Search NASA

Results for “analysis and statistical methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

An R Shiny graphical user interface for analyzing, visualizing, and interpreting high precision mass spectrometric data

There is currently a lack of software that meets the needs for the analysis of raw data produced by modern isotope ratio mass spectrometers for both R&D and routine use at SRNL and other US national labs • Needs to accommodate multiple isotope systems, instruments, and manufacturers • Include modern statistical methods and handling/visualization of uncertainty • Flexible software with transparent (no “black box”) and reproducible methods • This project is inspired by existing discipline-specific data analysis software (e.g., Tripoli1 , ET_Redux2 , IsoplotR3) used in the geochemical community • Our goal is to build an open source data analysis software package that focuses on flexibility, transparency, and reproducibility

Labone, Elizabeth

An R shiny graphical user interface for highprecision mass spectrometric data analysis

• There is currently a lack of software that meets the needs for the analysis of raw data produced by modern isotope ratio mass spectrometers for both R&D and routine use at SRNL and other US national labs • Needs to accommodate multiple isotope systems, instruments, and manufacturers • Include modern statistical methods and handling/visualization of uncertainty • Flexible software with transparent (no “black box”) and reproducible methods • This project is inspired by existing discipline-specific data analysis software (e.g., Tripoli1 , ET_Redux2, IsoplotR3) used in the geochemical community • Our goal is to build an open source data analysis software package that focuses on flexibility, transparency, and reproducibility

LABONE, ELIZABETH

Probabilistic neural networks for improved analyses with phenomenological R -matrix

Here we present a method for measurement analyses based on probabilistic deep neural networks that provide several advantages over conventional analyses with phenomenological models. These include predicting physical quantities directly from data, the rapid generation of statistically robust uncertainties, and the ability to bypass some parameters that may induce ambiguities and complications in data analysis. As deep learning methods make predictions through “black boxes,” the uncertainty quantification is typically challenging. We use a probabilistic framework that provides thorough uncertainty quantification and is straightforward to follow in practice. With the network architecture based on the Transformer, we demonstrate the current method for predicting nuclear resonance parameters from scattering data using the phenomenological R-matrix model.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

FREDA: A Web Application for the Processing, Analysis, and Visualization of Fourier‐Transform Mass Spectrometry Data

The high-resolution measurement capability of Fourier-transform mass spectrometry (FT-MS) has made it a necessity for exploring the molecular composition of complex organic mixtures, like soil, plant, aquatic, and petroleum samples. This demand has driven a need for informatics tools to explore and analyze FT-MS data in a robust and reproducible manner. FREDA is an interactive web application developed to enable spectrometrists to format, process, and explore their FT-MS data without the need for statistical programming expertise. FREDA was built to explore outputs from a molecular identification tool, like CoreMS, and provide a suite of methods to filter data, compute chemical properties of peaks, statistically compare samples and groups of samples, conduct exploratory data analysis, and download the results with a report detailing all steps conducted. To demonstrate the utility of FREDA, an example analysis was conducted using FT-MS data from a soil microbiology study of samples collected in two different soil depths at the Sphagnum bog forest north of Grand Rapids, Minnesota. Differences between the two depths are observed using Kendrick, Gibbs free energy, and van Krevelen plots. G-tests are used to quantify a significant difference between the groups. All analyses and plotting are conducted using only the FREDA application. FREDA is an open-source and readily available web application that allows users to explore and make statistically valid conclusions about their FT-MS data. The application is available online (https://map.emsl.pnnl.gov/app/freda) with a tutorial web series (https://youtu.be/k5HLE2kNSBY?si=yB6sGoyvzxrFf5MP) and freely accessible code on Github (https://github.com/EMSL-Computing/FREDA).

47 OTHER INSTRUMENTATION

Determination of proton PDF uncertainties with Markov chain Monte Carlo

We present an analysis of parton distribution functions (PDFs) of the proton using Markov chain Monte Carlo (MCMC) methods. The MCMC approach naturally implements Bayes’ theorem and, thus, provides a means to directly sample the underlying probability distribution—in this case, the probability distribution of the PDF parameters. This allows for a straightforward propagation of the resulting uncertainties into any PDF-dependent observable, preserving their simple probabilistic interpretation. In our analysis we include a broad set of deep inelastic scattering data from HERA, BCDMS and NMC experiments along with the Drell-Yan, 𝑊 and 𝑍 boson data from LHC and Tevatron experiments, which combined with theoretical calculations at next-to-next-to-leading order in QCD allow for realistic determination of PDFs. The main focus of this analysis is to explore alternative methods for PDF uncertainty estimation that are more firmly grounded in statistical principles. We show that the flexibility of the Bayes framework, allowing one, e.g., to account for non-Gaussianity or inconsistencies of datasets, is crucial to extract realistic uncertainties when such assumptions are not fulfilled. We also demonstrate that MCMC allows one to determine the Δ⁢𝜒 2 value corresponding to a given confidence level in the sample, which can, in turn, be used as a statistically well-founded tolerance criterion used in the Hessian method, thus addressing one of its main long-standing drawbacks.

Risse, Peter Clemens [Universität Münster (Germany

Source Levels of In‐Cloud Air in Shallow Cumulus: Consistency Between Paluch Diagram and Lagrangian Particle Tracking

Abstract The Paluch diagram is a widely used tool for interpreting aircraft measurements of shallow cumulus clouds. A prior study conducted by Heus et al. (2008,https://doi.org/10.1175/2008jas2572.1) concluded that the source levels of in‐cloud air inferred from the Paluch diagram exhibit biases, sometimes of several hundred meters, in comparison to those derived from Lagrangian particle tracking. In this short study we revisit this comparison. The results indicate that the upper source levels of in‐cloud air determined from the Lagrangian Particle Tracking and the Paluch diagram are consistent, and the choice of statistical methods is crucial. The significance of this research lies in confirming the reliability of the Paluch analysis, enabling its confident application to aircraft data.

Meteorology & Atmospheric Sciences

Unraveling design principles of protein landscapes in photosynthetic membranes in plant chloroplasts

The supramolecular organization of proteins within photosynthetic membranes is crucial for energy conversion in plants. Here, we introduce an analytical and computational pipeline that integrates high-resolution cryo–scanning electron microscopy, biochemical quantification, advanced Monte Carlo computer simulations, and statistical methods to elucidate the elusive protein landscapes of grana membranes in intact Arabidopsis leaves. Our integrated analysis challenges the prevailing view that particles on the exoplasmic fracture faces in freeze-fracture samples represent photosystem II exclusively. Instead, these particles also include cytochrome b 6 f complexes. Furthermore, our steric clash analysis demonstrates that stacked membranes contain a mixture of larger PSII supercomplexes (C 2 S 2 M 2 and C 2 S 2 ) in addition to a smaller complex (C 2 ). This suggests that in vivo PSII supercomplexes exist in an equilibrium distribution of differing sizes. Furthermore, we discovered that, although size exclusion effects govern the global protein arrangement, local packing exhibits orientational order indicative of lateral attractive protein-protein interactions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Emulating ab initio computations of infinite nucleonic matter

We construct efficient emulators for the computation of the infinite nuclear matter equation of state. These emulators are based on the subspace-projected coupled-cluster method for which we here develop a new algorithm called small-batch voting to eliminate spurious states that might appear when emulating quantum many-body methods based on a non-Hermitian Hamiltonian. The efficiency and accuracy of these emulators facilitate a rigorous statistical analysis within which we explore nuclear matter predictions for > 10 6 different parametrizations of a chiral interaction model with explicit Δ -isobars at next-to-next-to leading order. Constrained by nucleon-nucleon scattering phase shifts and bound-state observables of light nuclei up to He 4 , we use history matching to identify nonimplausible domains for the low-energy coupling constants of the chiral interaction. Within these domains we perform a Bayesian analysis using sampling and importance resampling with different likelihood calibrations and study correlations between interaction parameters, calibration observables in light nuclei, and nuclear matter saturation properties. Published by the American Physical Society 2024

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

Constraining Galaxy-Halo connection using machine learning

We investigate the potential of machine learning (ML) methods to model small-scale galaxy clustering for constraining Halo Occupation Distribution (HOD) parameters. Our analysis reveals that while many ML algorithms report good statistical fits, they often yield likelihood contours that are significantly biased in both mean values and variances relative to the true model parameters. This highlights the importance of careful data processing and algorithm selection in ML applications for galaxy clustering, as even seemingly robust methods can lead to biased results if not applied correctly. ML tools offer a promising approach to exploring the HOD parameter space with significantly reduced computational costs compared to traditional brute-force methods if their robustness is established. Using our ANN-based pipeline, we successfully recreate some standard results from recent literature. Properly restricting the HOD parameter space, transforming the training data, and carefully selecting ML algorithms are essential for achieving unbiased and robust predictions. Among the methods tested, artificial neural networks (ANNs) outperform random forests (RF) and ridge regression in predicting clustering statistics, when the HOD prior space is appropriately restricted. We demonstrate these findings using the projected two-point correlation function (w p (r p )), angular multipoles of the correlation function (ξ ℓ (r)), and the void probability function (VPF) of Luminous Red Galaxies from Dark Energy Spectroscopic Instrument mocks. Our results show that while combining w p (r p ) and VPF improves parameter constraints, adding the multipoles ξ 0 , ξ 2 , and ξ 4 to w p (r p ) does not significantly improve the constraints.

cosmology

Catalog-level blinding on the bispectrum for DESI-like galaxy surveys

We evaluate the performance of the catalog-level blind analysis technique (blinding) presented in Brieden et al. (2020) in the context of a fixed template power spectrum and bispectrum analysis. This blinding scheme, which is tailored for galaxy redshift surveys similar to the Dark Energy Spectroscopic Instrument (DESI), has two components: the so-called “AP blinding” (concerning the dilation parameters α$_{∥}$, α$_{⊥}$) and “RSD blinding” (redshift space distortions, affecting the growth rate parameter f). Through extensive testing, including checks for the RSD part in cubic boxes, the impact of AP blinding on mocks with realistic survey sky coverage, and the implementation of a full AP+RSD blinding pipeline, our analysis demonstrates the effectiveness of the technique in preserving the integrity of cosmological parameter estimation when the analysis includes the bispectrum statistic. We emphasize the critical role of sophisticated — and difficult to accidentally unblind — blinding methods in precision cosmology.

79 ASTRONOMY AND ASTROPHYSICS

Structured illumination for surface-resolved grazing-incidence X-ray scattering

Grazing-incidence (GI) scattering techniques are widely used to characterize thin films, offering high surface sensitivity and insight into morphology and structure. However, these approaches typically provide statistical averaged information due to elongated footprint or limited spatial resolution due to beam size. Here we introduce a method that combines structured illumination with GI X-ray scattering and leverages our computational imaging approach to resolve local structural details. We demonstrate that our method captures local features of an organic semiconductor thin film without the need for sample rotation as in tomography. The method expands GI techniques from statistical averaging to high-resolution imaging, thereby providing the capability for detailed analysis of local material properties, such as domain shape, orientation and polymorphism, which are critical for advancing material design towards more efficient and tailored materials.

97 MATHEMATICS AND COMPUTING

Cluster spin glass correlations and dynamics in Zn 0.5⁢ Mn 0.5⁢ Te

Here, we present a combined magnetometry, muon spin-relaxation (𝜇⁢SR), and neutron-scattering study of the insulating spin glass Zn 0.5 ⁢Mn 0.5 ⁢Te, for which magnetic Mn 2+ and nonmagnetic Zn 2+ ions are randomly distributed on a face-centered cubic lattice. The magnetometry and 𝜇⁢SR results confirm a spin freezing transition around 𝑇 𝑓 ≈ 23 K, with the spin-fluctuation rate decreasing gradually and somewhat inhomogeneously through the sample volume as the temperature decreases toward 𝑇 𝑓 . Characteristic spin-correlation times well above 𝑇 𝑓 are on the order of 10 −10 s, much slower than typically observed in canonical spin glasses but in line with expectations for a cluster spin glass. Using magnetic pair distribution function (mPDF) analysis and reverse Monte Carlo (RMC) modeling of the magnetic diffuse neutron-scattering data, we show that the spin-glass ground state consists of clusters of spins exhibiting short-range-ordered type-III antiferromagnetic correlations with a locally ordered moment of 3.1⁢(1)⁢𝜇 B between nearest-neighbor spins. The type-III correlations decay exponentially as a function of spin separation distance with a correlation length of approximately 5 Å. The diffuse magnetic scattering and corresponding mPDF show no significant changes across 𝑇 𝑓 , indicating that the dynamically fluctuating short-range spin correlations in the paramagnetic state retain the same basic type-III configuration that characterizes the spin-glass state; the only change apparent from the neutron-scattering data is a gradual reduction of the correlation length and locally ordered moment with increasing temperature. Taken together, these results paint a unique and detailed picture of the local magnetic structure and dynamics in Zn 0.5 ⁢Mn 0.5⁢ Te and provide strong evidence that this material is best described as a cluster spin glass. In addition, this work showcases a statistical method for extracting diffuse scattering signals from neutron powder diffraction data, which we developed to facilitate the mPDF and RMC analysis of the neutron data. This method has the potential to be broadly useful for neutron powder diffraction experiments on a variety of materials with short-range atomic or magnetic order.

magnetism

Effect of adaptive cruise control on fuel consumption in real-world driving conditions

This paper presents a comprehensive analysis of the impact of adaptive cruise control on energy consumption in real-world driving conditions based on a natural experiment: a large-scale observational dataset of driving data from a diverse fleet of vehicles and drivers. The analysis is conducted at two different fidelity levels: (1) a macroscopic trip-level benefit estimate that compares trips with and without cruise control in a counterfactual way using statistical methods, and (2) a situation-based comparison achieved through the segmentation of trips into distinct driving situations such as acceleration, braking, cruising, and other maneuvers. The results of this research show that the effect of cruise control on energy consumption varies across different driving situations and levels of analysis. In a macroscopic trip-level analysis, cruise control engagement is associated with a slight increase in fuel consumption across the fleet. As revealed later by the situation-based analysis, this result can be attributed to the negative impact of cruise control on energy consumption in cruising mode, which is the most common driving situation. However, the situation-based comparison demonstrates that cruise control can provide fuel consumption benefits in situations involving acceleration and braking, particularly when a preceding vehicle is present. The study also emphasizes the importance of controlling for various factors that can influence both fuel consumption and the likelihood of cruise control engagement to properly evaluate its effects.

33 ADVANCED PROPULSION SYSTEMS

Algorithm to extract direction in 2D discrete distributions and a continuous Frobenius norm

In this study, we present a novel algorithm for determining directionality in 2D distributions of discrete data. We compare a reference dataset with a known direction to a measured dataset with an unknown direction by the Frobenius norm of the difference (FND) to find the unknown direction. To generalize this concept, we develop a continuous Frobenius norm of the difference (CFND) as a continuous analog of the FND and derive its analytical expression. By relating fitted and normalized 2D Gaussian distributions, we show that the CFND approximates the FND, and we validate this relationship with computer simulations. We find that a first-order approximation of the CFND between two similar Gaussian distributions takes the form of an absolute sine function, offering a simple analytical form with potential for specialized applications in segmented inverse beta decay (IBD) neutrino detectors, astronomy, machine learning, and more. Although this method may easily extend to 3D scalar fields, our focus here is on 2D real-valued fields as it directly applies to directionality. Our methodology consists of modeling a 2D Gaussian distribution, binning the data into a histogram, and encoding it as a square matrix. Rotating this matrix around its geometric center and comparing it to a measured dataset using the FND gives us rotational data that we fit with an absolute sine function. The location of the minimum of this fit is the angle closest to the true angle of the direction in the measured dataset. We present the derivation and discuss initial applications of the CFND in our novel algorithm, demonstrating its success in approximating directionality in 2D distributions.

Data Analysis, Statistics and Probability (physics

Dynamic data-driven multiscale modeling for predicting the degradation of a 316L stainless steel nuclear cladding material

Here, we have developed a long short-term memory stacked ensemble (LSTM-SE) surrogate modeling approach that can provide rapid predictions of microstructural evolution and the resultant mechanical properties of American Iron and Steel Institute (AISI) 316L series stainless steel (316LSS) fuel cladding under conditions of varying temperature and radiation dose rate. To acquire training data, we developed and implemented a kinetic Monte Carlo (KMC) model to simulate precipitation kinetics of M 23 C 6 , γ', and G phases within SS316L cladding. Experimentally reported precipitation kinetics of SS316L in literature were linked to the kinetic parameters of the simulated precipitation in our KMC model. The model was then used to simulate microstructure evolution under synthetically generated treatments of varying temperature and radiation dose rate, for periods of up to 3000 hours. Changes in volume fraction, number density, and particle size of precipitates were recorded, and particle area fractions were correlated using statistical methods to develop the surrogate model. Simultaneously, the mechanical properties of the simulated microstructures were evaluated using microstructure-based finite element method (FEM) analysis to determine the elastic modulus, yield stress, ultimate tensile strength, and elongation to failure of the aged microstructures. Using this approach, our surrogate model can predict precipitation behavior within 0.25% volume fraction and mechanical properties within 6% relative error from the values predicted by the KMC and FEM models using 50 training simulations as input. The trained recurrent neural network-based model can return estimations of precipitation kinetics and mechanical properties ~1000 times faster than the physics-based codes. This work demonstrates, as a proof of concept, that reactor material service lifetimes under variable service conditions can be predicted for a statistics-based model from a practicably obtainable dataset.

36 MATERIALS SCIENCE

Seeing is Believing: Autonomous Microscopy and the Data Revolution in Materials Science

This presentation explores the transformative potential of autonomous electron microscopy and artificial intelligence (AI) in accelerating materials science discovery, particularly for energy applications and materials operating in extreme environments. We discuss pioneering self-driving laboratories at NREL designed to intelligently probe material synthesis and degradation across multiple scales, aiming to rapidly bridge the gap between atomic-level understanding and the development of high-performance, reliable materials. Utilizing advanced machine learning techniques, such as few-shot learning and multimodal analysis integrating imaging and spectroscopy, we demonstrate methods to extract actionable descriptors for material behavior, quantify complex microstructural evolution, and statistically link synthesis parameters to defect populations. This AI-driven approach promises to accelerate the creation of predictive materials tailored for specific missions, enabling faster development cycles and enhanced material assurance.

36 MATERIALS SCIENCE

Seeing is Believing: Autonomous Microscopy and the Data Revolution in Materials Science

This presentation explores the transformative potential of autonomous electron microscopy and artificial intelligence (AI) in accelerating materials science discovery, particularly for energy applications and materials operating in extreme environments. We discuss pioneering self-driving laboratories at NREL designed to intelligently probe material synthesis and degradation across multiple scales, aiming to rapidly bridge the gap between atomic-level understanding and the development of high-performance, reliable materials. Utilizing advanced machine learning techniques, such as few-shot learning and multimodal analysis integrating imaging and spectroscopy, we demonstrate methods to extract actionable descriptors for material behavior, quantify complex microstructural evolution, and statistically link synthesis parameters to defect populations. This AI-driven approach promises to accelerate the creation of predictive materials tailored for specific missions, enabling faster development cycles and enhanced material assurance.

97 MATHEMATICS AND COMPUTING

A cross-dimensional analysis of data-driven short-term load forecasting methods with large-scale smart meter data

Electricity load forecasting is essential to utility operation and power grid stability. A wide spectrum of data-driven methods, ranging from linear regression models to more recent deep learning models have been adopted to forecast electric load over the years. However, there still lacks a holistic evaluation of the applicability of conventional statistical and machine learning based algorithms with respect to different temporal and spatial scopes, computational requirements, and sensitivity of model-tuning. Enabled by a large-scale electricity load profile dataset of over 40,000 residential customers in a utility region, we conducted a cross-dimensional analysis of data-driven load forecasting methods. Three regression-based and seven deep learning algorithms with different model configurations were evaluated in terms of their overall and peak load prediction accuracy, and training burdens, across spatial aggregation levels ranging from the transformer, feeder, substation, to neighborhood. We found, first, the load forecasting accuracy is constrained by a predictability boundary, influenced by the forecasting horizon and spatial aggregation level. Specifically, RandomForest, XGBoost, TFT, TSMixer, and TiDE models achieved less than 10 % prediction error for up to 96-h ahead forecasting for district, substation, and feeder levels, while other models struggle at long-horizon predictions; Second, for winter and summer peak load dates, most models were able to predict the peak demand timing within ± 1 h, but the prediction percentage error varied by models, with TFT and TiDE models being the top performers; Third, models with similar prediction accuracy can differ in training burden by an order of magnitude. Therefore, choosing model configurations that balance prediction performance and computational resource is an important practical consideration for large-scale deployment of the machine learning based load forecasting. The outcome of this study can guide researchers and practitioners to choose the proper load forecasting algorithms based on their problem scope, required accuracy, and available resources. The predictability boundary can serve as a benchmark for electricity load forecasting problems with new algorithms and datasets.

Li, Han