Search NASA⌕ Search

SEARCH · Search NASA

Results for “Statistical sampling techniques”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

86 records · Page 5

A novel conditional generative model for efficient ensemble forecasts of state variables in large-scale geological carbon storage

Integrating monitoring data to efficiently update reservoir pressure and CO 2 plume distribution forecasts presents a significant challenge in geological carbon storage (GCS) applications. Inverse modeling techniques are commonly used to fuse observational data and refine reservoir model parameters, thereby improving state variable forecasts. However, these techniques often rely on linear or Gaussian assumptions, which can limit their effectiveness in accurately predicting state variables. Moreover, simulating large-scale three-dimensional (3D) GCS problems is computationally expensive, making iterative runs in inverse problems prohibitive. To address these challenges, we propose a conditional generative model utilizing the score-based diffusion method for real-time 3D pressure and saturation field distribution predictions. Our approach involves solving the score function with a mini-batch-based Monte Carlo estimator to generate labeled data. This data is subsequently employed to train a fully connected neural network, enabling it to learn the conditional sample generator within a supervised learning framework. This method enables the rapid generation of a large ensemble of predictions, facilitating comprehensive uncertainty quantification of state variables. Here we applied our method to forecast the dynamic 3D distributions of pressure and saturation fields over a 30-year injection period. The statistical assessment with low root mean square error (RMSE) values demonstrates that our method can accurately predict the spatiotemporal distributions of both pressure and saturation fields. Moreover, the developed conditional generative model shows high computational efficiency by generating 100 ensemble forecasts of 3D state variables in less than 10 min. The consistency between ensemble averages and ground truth values further illustrates the model’s capability to capture state variable dynamics during the CO 2 plume injection process. Notably, the ground truth values fall within the ensemble forecasts, indicating that our uncertainty quantification effectively captures variability and potential noise in the observations. Thus, the developed conditional generative model proves to be a more efficient, accurate, and practical tool for GCS applications, facilitating timely risk analysis and informed decision-making.

58 GEOSCIENCES↗

A Machine Learning Framework for Modeling Ensemble Properties of Atomically Disordered Materials

Atomic disorder can strongly influence material properties such as charge transport, optical response, and catalytic activity. However, efficiently modeling these disorder effects remains challenging for first-principles methods due to the cost of sampling large configurational spaces and computing complex physical quantities. Recent advances of machine learning techniques, particularly graph neural networks (GNNs), has enabled the efficient and accurate predictions of complex material properties, offering promising tools for studying disordered systems. In this work, we present a general machine-learning-assisted computational framework that integrates equivariant GNNs with Monte Carlo simulations to compute the thermodynamic and ensemble-averaged functional properties of disordered materials. Using the surface-termination-disordered MXene monolayer Ti 3 C 2 T 2–x as a representative system, we find that electrical conductivity exhibits an emergent peak near the order–disorder phase transition temperature due to the interplay between electron scattering and doping. In contrast, optical conductivity remains largely insensitive to local atomic disorder and reflects the global surface chemical composition. These results highlight the role of atomic disorder in affecting material properties and demonstrate the potential of our approach for statistically modeling disorder effects in a wide range of materials such as high-entropy alloys and spin liquids.

MXene↗

Designing an Optimal Sensor Network via Minimizing Information Loss

Optimal experimental design is a classic topic in statistics, with many well-studied problems, applications, and solutions. The design problem we study is the placement of sensors to monitor spatiotemporal processes, explicitly accounting for the temporal dimension in our modeling and optimization. We observe that recent advancements in computational sciences often yield large datasets based on physics-based simulations, which are rarely leveraged in experimental design. We introduce a novel model-based sensor placement criterion, along with a highly-efficient optimization algorithm, which integrates physics-based simulations and Bayesian experimental design principles to identify sensor networks that “minimize information loss” from simulated data. Our technique relies on sparse variational inference and (separable) Gauss-Markov priors, and thus may adapt many techniques from Bayesian experimental design. We validate our method through a case study monitoring air temperature in Phoenix, Arizona, using state-of-the-art physics-based simulations. Our results show our framework to be superior to random or quasi-random sampling, particularly with a limited number of sensors. We conclude by discussing practical considerations and implications of our framework, including more complex modeling tools and real-world deployments.

54 ENVIRONMENTAL SCIENCES↗

The development and application of the stirred‐reactor coupon analysis (SRCA) test method

A new technique, termed the stirred‐reactor coupon analysis (SRCA) method, has been developed to measure the rate of glass dissolution in forward‐rate conditions. Monolithic glass coupons are partially masked with an inert material before placement in a large volume of well‐mixed solution with known chemistry and temperature for a predetermined duration. After the test, the mask is removed, and the difference in step height between the protected area and the exposed corroded portions of the sample coupon is measured to determine the extent of glass dissolution. The step height is converted to a rate measurement using the test duration and glass density. Test parameters such as sample surface preparation and test duration were evaluated to determine their effects on the measured rates. Additionally, results from an interlaboratory study (ILS) consisting of 12 laboratories from 11 different institutions are presented, where each laboratory performed 12 independent tests. When removing experimental outlier data, the 95% reproducibility limits for the SRCA method has no statistical difference with previously published standardized test methods used to determine the forward rate of glass dissolution. Overall, this paper describes steps necessary to perform the test method and provides the statistical calculations to evaluate test accuracy.

chemical durability↗

The DESI One-Percent Survey: Modelling the clustering and halo occupation of all four DESI tracers with U CHUU

We present results from a set of mock lightcones for the DESI One-Percent Survey, created from the UCHUU simulation. This 8 h −3 Gpc 3 N-body simulation comprises 2.1 trillion particles and provides high-resolution dark matter (sub)haloes in the framework of the Planck-based ΛCDM cosmology. Employing the subhalo abundance matching (SHAM) technique, we populated the UCHUU (sub)haloes with all four DESI tracers – Bright Galaxy Survey (BGS), luminous red galaxies (LRGs), emission line galaxies (ELGs), and quasars (QSOs) – to z = 2.1. Our method accounts for redshift evolution as well as the clustering dependence on luminosity and stellar mass. The two-point clustering statistics of the DESI One-Percent Survey generally agree with predictions from UCHUU across scales ranging from 0.3 h −1 Mpc to 100 h −1 Mpc for the BGS and across scales ranging from 5 h −1 Mpc to 100 h −1 Mpc for the other tracers. We observed some differences in clustering statistics that can be attributed to incompleteness of the massive end of the stellar mass function of LRGs, our use of a simplified galaxy-halo connection model for ELGs and QSOs, and cosmic variance. We find that at the high precision of UCHUU, the shape of the halo occupation distribution (HOD) of the BGS and LRG samples is smaller bias values, likely due to cosmic variance. The bias dependence on absolute magnitude, stellar mass, and redshift aligns with that of previous surveys. These results provide DESI with tools to generate high-fidelity lightcones for the remainder of the survey and enhance our understanding of the galaxy-halo connection.

cosmology↗

Probabilistic inference of the structure and orbit of Milky Way satellites with semi-analytic modelling

Semi-analytic modelling furnishes an efficient avenue for characterizing dark matter haloes associated with satellites of Milky Way-like systems, as it easily accounts for uncertainties arising from halo-to-halo variance, the orbital disruption of satellites, baryonic feedback, and the stellar-to-halo mass (SMHM) relation. We use the SatGen semi-analytic satellite generator, which incorporates both empirical models of the galaxy–halo connection as well as analytic prescriptions for the orbital evolution of these satellites after accretion onto a host to create large samples of Milky Way-like systems and their satellites. By selecting satellites in the sample that match observed properties of a particular dwarf galaxy, we can infer arbitrary properties of the satellite galaxy within the cold dark matter paradigm. For the Milky Way’s classical dwarfs, we provide inferred values (with associated uncertainties) for the maximum circular velocity v max and the radius r max at which it occurs, varying over two choices of baryonic feedback model and two prescriptions for the SMHM relation. While simple empirical scaling relations can recover the median inferred value for v max and r max , this approach provides realistic correlated uncertainties and aids interpretability. We also demonstrate how the internal properties of a satellite’s dark matter profile correlate with its orbit, and we show that it is difficult to reproduce observations of the Fornax dwarf without strong baryonic feedback. Furthermore, the technique developed in this work is flexible in its application of observational data and can leverage arbitrary information about the satellite galaxies to make inferences about their dark matter haloes and population statistics.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Robust Design Under Uncertainty in Quantum Error Mitigation

Error mitigation techniques are crucial to achieving near-term quantum advantage. Classical postprocessing of quantum computation outcomes is a popular approach for error mitigation, which includes methods, such as zero noise extrapolation, virtual distillation, and learning-based error mitigation. However, these techniques have limitations due to the propagation of uncertainty resulting from the finite shot number of a quantum measurement. In this work, we introduce general and unbiased methods for quantifying the uncertainty and error of error-mitigated observables based on the strategic sampling of error mitigation outcomes. We then extend our approach to demonstrate the optimization of performance and robustness of error mitigation under uncertainty. To illustrate our methods, we apply them to zero noise extrapolation and Clifford date regression in the ground state of the XY model simulated using depolarizing and International Business Machines Corporation (IBM) Toronto noise models, respectively. In particular, we optimize the choice of noise levels and the allocation of shots for zero noise extrapolation and the distribution of the training circuits for Clifford data regression. While our methods are readily applicable to any postprocessing-based error mitigation approach, in practice they must not be prohibitively expensive—even though they perform optimizations of the error mitigation hyperparameters requiring sampling of a statistical distribution of error mitigation outcomes. By leveraging surrogate-based optimization, we show that our methods can efficiently perform optimal design for a zero noise extrapolation implementation. We then further demonstrate the transferability of learned zero noise extrapolation hyperparameters to other similar circuits.

97 MATHEMATICS AND COMPUTING↗

Sampling Size Optimization for Bioburden Density Estimation in Planetary Protection

Planetary protection (PP) is a discipline that focuses on minimizing the biological contamination of spacecraft to ensure compliance with international policy. Precise estimation of bioburden - the total number of microbes in or on spacecraft hardware – and the bioburden density are of utmost importance for PP. Such estimation is the way concordance with requirements is demonstrated, and it is critical for quantifying the potential risk of inadvertently contaminating other planetary bodies. Although a suite of molecular techniques have been used to thoroughly characterize and profile the microbiome of various cleanroom environments and spacecraft, the gold standard remains the physical enumeration of microbes via culturing of samples directly taken from spacecraft and associated surfaces. However, due to technical, budgetary, and programmatic constraints, only a manageable portion (around 10%) of the entire spacecraft surface is directly sampled with cotton swabs or wipes. To generate the bioburden current best estimate (CBE) for components not directly verifiable, the accepted approach is to apply a NASA-defined bioburden estimate based on the components’ manufacturing or assembly environment. This approach utilizes a prespecified bioburden density estimation that applies a maximum value across the total surface area of the specified component. For hardware components that underwent similar assembly processes, an implied bioburden is adopted for all components, based on a direct verification of a representative component within the same lot. Once all components have a CBE, the bioburden estimates are generated. In previous publication [ 1], we have shown that statistical risks quantifying the accuracy of the estimates for sampled, prespecified, and implied components can be derived and ranked. For mean squared error (MSE) function, the risks are available analytically and hence a cost function can be obtained to optimize the risks with respect to the sampling area and sampling cost. Since the sampling area and sampling cost are two complimentary variables, their sum will have a well-defined minimum. This paper presents the multivariate optimization of the integrated risk of an empirical Bayes estimator to determine the optimal sampling schedule for a given number of components. It is assumed that given a number of components, N, the bioburden density for each component can either be sampled, implied, or prespecified. The multivariate optimization searches through different options to sample, imply or prespecify the bioburden density for a component, and account for the component’s surface area and cost of sampling. The idea of the optimization is based on the observation that the statistical risk of using an estimator is a monotonically decreasing function of the sampled area. The larger the sampled area, the lower the risk of using the estimator as the estimator becomes more and more accurate as the sampling area increases. On the other hand, the cost of sampling is monotonically increasing as the sampled surface grows. This makes the risk and total cost of sampling complimentary variables which can be counterbalanced to achieve an optimal overall value with respect to the sampled surface. In this paper, the integrated risk has been used to quantify the accuracy of the estimator. This risk has been selected because it depends on neither the true value of the parameter nor on the collected data. The cost of each sample was also available to obtain the total cost of sampling of N components. The paper will present the results based on computer-simulated data as well as the data collected during the InSight mission. The computer-simulated data have N components with randomly generated total areas and each component assigned to one of the three categories according to the method of estimating of bioburden density: sampled, implied, or prespecified. The cost of sampling is also available. The cost of sampling is estimated based on a cost model provided by the planetary protection group at JPL. For this paper, the overall cost was assumed to be a linear function of exposure. The optimization process finds the allocation of the components to the three categories that minimizes the tradeoff between integrated risk and total cost. For the InSight data, a set of components is selected representing all three categories, and optimization is performed to determine if the performed allocation was optimal or if a better allocation could have been obtained. To the best of our knowledge, this work is the first attempt not only perform an accurate estimation of bioburden density but also do it in an optimal way.

97 - MATHEMATICS AND COMPUTING↗

Evaluating User Errors and Temporal Trends in Marine Fish Communities Using 360-Degree Underwater Photography

The use of environmental DNA (eDNA) sampling has been proposed as a complementary method to monitor fish species in marine environments, offering a non-invasive and potentially more efficient approach to marine species observations. eDNA monitoring could be especially useful in and around sites targeted for marine energy generation as these regions need regular monitoring that would be impractical with traditional techniques. Before we can fully rely upon eDNA, we must first verify its accuracy against other proven methods, such as the use of underwater photography. In this study, I deployed a 360-degree camera in the tidal channel of Sequim Bay once a month during several hours overlapping slack tide. I investigated how having multiple people identify and count fish on underwater images could affect the overall results. Using chi square tests in R, I compared my fish identifications and counts to those made by another intern on the same images recorded in August. I found significant differences in the number of species identified and the total individual counts between the two different datasets. I also tested the statistical differences in both Shannon diversity and Pielou evenness indices between the August, September, and November camera deployments using a Hutcheson t-test. Only one significant difference was found in the Shannon index comparisons, and none were found between the Pielou evenness comparisons. These findings show that if multiple identifiers are used to process underwater images, quality control checks must be made to reduce the potential for error. This also points toward the possibility to leverage more advanced image analysis processes, such as automated image analysis software. The findings from this study also show that the dynamics of marine fish communities can vary over a few months; however, further analysis is needed to determine the extent of the seasonal changes in Sequim Bay.

59 BASIC BIOLOGICAL SCIENCES↗

An In Situ , Automated High-Explosives Aging Method Utilizing Two-Dimensional Gas Chromatography–Mass Spectrometry

Understanding chemical changes that occur in high explosives as they age is of great importance to the safe employment and storage of these compounds. Traditional methods of aging high explosives even under accelerated aging conditions are time intensive with durations on the order of months to years. The nature of traditional aging analyses reduces each sample to a snapshot data point often separated widely in time, requiring many assumptions as to how the degradation products develop. Further complicating matters, several analytical techniques are typically employed for each sample analysis in order to ascertain an entire picture of the decomposition pathways. To address these shortcomings with existing methods, a new method of accelerated aging of high explosives utilizing comprehensive two-dimensional gas chromatography coupled to high-resolution mass spectrometry (GC × GC-HRMS) was developed using 2,4,6,8,10,12-hexanitro-2,4,6,8,10,12-hexaazaisowurtzitane (CL-20) as a model compound for method development. This in situ automated method reduces the time scale of aging to a matter of hours using the inlet of the GC × GC as the aging vessel. GC × GC in combination with HRMS allowed for the collection of both evolved gases and other decomposition products produced during the entire aging process in real time with HRMS providing far greater certainty in identification of explosives aging products. Additionally, this method allowed for a higher throughput of samples with greatly simplified sample preparation. Chemometric analysis of the GC × GC-HRMS data set via the alteration analysis (ALA) enabled discovery of statistically significant chemical changes providing insight into the variation of decomposition pathways with varying aging temperatures.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Imaging Bragg Edge Analysis TooLs for Engineering Structures (iBeatles)

The Spallation Neutron Source (SNS) at Oak Ridge National Laboratory (ORNL) provides pulsed neutrons with energies varying from epithermal to cold. In preparation for VENUS, the neutron imaging beamline to be located at beam port 10, we have performed a series of experiments focused on wavelength-dependent radiography and computed tomography for a broad range of applications, from materials science to biological tissues.One of the time-of-flight (TOF) techniques that is of interest to the scientific community is the 2-dimensional mapping of phases and average crystalline plane orientation in samples both ex-situ and during applied stresses such as tensile loading and heating. This technique is known as Bragg edgeimaging and relies on the identification of changes of transmission values, fitting of the edge to measure its displacement, and thus identify the shift in lattice parameter due to stresses. One of the challenges of TOF imaging measurements is the amount of data and the inability to observe Bragg edge shifts in real time during an experiment. Thus, we have been focusing on creating a Python-based interface that allows fast data processing and instantaneous mapping and fitting of the Bragg edges, and their evolution through time. Python libraries and Jupyter notebooks have been implemented to facilitate decision making during an experiment. The advantage of the notebooks is the possibility to guide an experiment as they can quickly process and display Bragg edge data. These notebooks can be used independently, or can be combined in a Python Graphical User Interface (GUI) tool called iBeatles. This interface permits visualization and fitting of the Bragg edges, and ultimately back-projects the fitting results onto the radiographs to display a strain map. Assuming data collection has sufficient statistics, the strain mapping analysis can be performed on a pixel-by-pixel basis. This development is a step forward toward a better user experience at the future VENUS beamline in terms of live feedback and productivity. Analysis that used to take days of switching between different applications can now be done in minutes within the

Bilheux, JeanChristophe [Oak Ridge National Labora↗

The DESI Early Data Release white dwarf catalogue

The Early Data Release (EDR) of the Dark Energy Spectroscopic Instrument (DESI) comprises spectroscopy obtained from 2020 December 14 to 2021 June 10. White dwarfs were targeted by DESI both as calibration sources and as science targets and were selected based on Gaia photometry and astrometry. Here, we present the DESI EDR white dwarf catalogue, which includes 2706 spectroscopically confirmed white dwarfs of which approximately 60 per cent have been spectroscopically observed for the first time, as well as 66 white dwarf binary systems. We provide spectral classifications for all white dwarfs, and discuss their distribution within the Gaia Hertzsprung–Russell diagram. We provide atmospheric parameters derived from spectroscopic and photometric fits for white dwarfs with pure hydrogen or helium photospheres, a mixture of those two, and white dwarfs displaying carbon features in their spectra. We also discuss the less abundant systems in the sample, such as those with magnetic fields, and cataclysmic variables. The DESI EDR white dwarf sample is significantly less biased than the sample observed by the Sloan Digital Sky Survey, which is skewed to bluer and therefore hotter white dwarfs, making DESI more complete and suitable for performing statistical studies of white dwarfs.

79 ASTRONOMY AND ASTROPHYSICS↗

Deep-field analytical calibration

The next generation of imaging surveys, including the Vera C. Rubin Observatory Legacy Survey of Space and Time (LSST), Euclid, and the Nancy Grace Roman Space Telescope, will provide unprecedented constraints on cosmology using weak gravitational lensing. To fully exploit this statistical power, shear measurement methods must achieve sub- per cent accuracy while mitigating systematic biases from noise, the point-spread function (PSF), blending, and shear-dependent detection. The analytical calibration framework (AnaCal) has demonstrated such accuracy but requires adding noise to images, reducing effective depth. We introduce Deep-Field Analytical Calibration (DEEP-FIELD AnaCal), an extension of AnaCal that uses deep-field images to compute shear responses while preserving the statistical power of wide-field data. We validate DEEP-FIELD AnaCal on isolated and blended galaxy image simulations with LSST-like conditions, finding it meets the stringent requirement of multiplicative bias $|m| < 3\times 10^{-3}$ at 99.7 per cent confidence. Compared to standard AnaCal applied to wide-field images, DEEP-FIELD AnaCal increases the effective galaxy number density from 17 to 30 arcmin$^{-2}$ for simulated 10-yr LSST data. With deep fields $10\times$ longer than the wide field, we find pixel noise variance in shear estimation is reduced by 30 per cent and overall uncertainty by $\sim 25~{{\ \rm per\ cent}}$. Finally, using the LSST Deep Drilling Fields strategy, we assess sample variance and find an equivalent calibration uncertainty of $\lesssim 0.3~{{\ \rm per\ cent}}$. These results demonstrate that DEEP-FIELD AnaCal offers a promising path to achieve the required shear calibration for upcoming weak lensing surveys.

79 ASTRONOMY AND ASTROPHYSICS↗

Impact of survey spatial variability on galaxy redshift distributions and the cosmological 3 × 2-point statistics for the Rubin Legacy Survey of Space and Time (LSST)

We investigate the impact of spatial survey non-uniformity on the galaxy redshift distributions for forthcoming data releases of the Rubin Observatory Legacy Survey of Space and Time (LSST). Specifically, we construct a mock photometry data set degraded by the Rubin OpSim observing conditions, and estimate photometric redshifts of the sample using a template-fitting photo-z estimator, BPZ, and a machine learning method, FlexZBoost. We select the Gold sample, defined as $i\lt 25.3$ for 10 yr LSST data, with an adjusted magnitude cut for each year and divide it into five tomographic redshift bins for the weak lensing lens and source samples. We quantify the change in the number of objects, mean redshift, and width of each tomographic bin as a function of the coadd i-band depth for 1-yr (Y1), 3-yr (Y3), and 5-yr (Y5) data. In particular, Y3 and Y5 have large non-uniformity due to the rolling cadence of LSST, hence provide a worst-case scenario of the impact from non-uniformity. We find that these quantities typically increase with depth, and the variation can be $10\!-\!40~{{\rm per\ cent}}$ at extreme depth values. Using Y3 as an example, we propagate the variable depth effect to the weak lensing $3\times 2$ pt analysis, and assess the impact on cosmological parameters via a Fisher forecast. We find that galaxy clustering is most susceptible to variable depth, and non-uniformity needs to be mitigated below 3 per cent to recover unbiased cosmological constraints. There is little impact on galaxy–shear and shear–shear power spectra, given the expected LSST Y3 noise.

cosmology↗