Search NASA⌕ Search

SEARCH · Search NASA

Results for “statistical model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22

Snow Distribution Patterns Revisited: A Physics-Based and Machine Learning Hybrid Approach to Snow Distribution Mapping in the Sub-Arctic

Snowpack distribution in Arctic and alpine landscapes often occurs in repeating, year-to-year patterns due to local topographic, weather, and vegetation characteristics. Previous studies have suggested that with years of observational data, these snow distribution patterns can be statistically integrated into a snow process modeling workflow. Recent advances in snow hydrology and machine learning (ML) have increased our ability to predict snowpack distribution using in-situ observations, remote sensing data sets, and simple landscape characteristics that can be easily obtained for most environments. Here, we propose a hybrid approach to couple a ML snow distribution pattern (MLSDP) map with a physics-based, snow process model. We trained a random forest ML algorithm on tens of thousands of snow survey observations from a subarctic study area on the Seward Peninsula, Alaska, collected during peak snow water equivalent (SWE). We validated hybrid model outputs using in-situ snow depth and SWE observations, as well as a light detection and ranging data set and a distributed temperature profiling sensor data set. When the hybrid results were compared with the physics-based method, the hybrid method more accurately depicted the spatial patterns of the snowpack, areas of drifting snow, and years when no in-situ observations were used in the random forest ML training data set. The hybrid method also showed improvements in root mean squared error at 61% of locations where time-series estimations of snow depth were observed. These results can be applied to any physics-based model to improve the snow distribution patterning to reflect observed conditions in high latitude and high elevation cold region environments.

54 ENVIRONMENTAL SCIENCES↗

Data-based filtered dissipation rate modelling for multi-modal turbulent combustion: evaluating a priori model generalizability

Manifold-based models offer a computationally efficient alternative to directly transporting the thermochemical state in computational simulations of turbulent reacting flows, projecting the high-dimensional thermochemical state-space onto a low-dimensional manifold. Recent efforts have yielded a manifold-based model applicable to multi-modal combustion, enabling reconstruction of the thermochemical state from solutions to two-dimensional manifold equations in mixture fraction and generalized progress variable that are parameterised by three scalar dissipation rates. In coarse-grained simulations such as Large Eddy Simulation (LES), closure of the multi-modal manifold equations and subfilter variances/covariance requires closure of three filtered scalar dissipation rates. Here, the present work adopts a data-based approach, providing closure for the three filtered scalar dissipation rates via deep neural networks (DNNs). High-fidelity datasets corresponding to an autoigniting n-dodecane jet flame and a bluff body swirl-stabilized confined lifted spray flame of two aviation fuels (Jet-A and C1) with different ignition propensities are leveraged to generate training data that spans a diverse range of thermodynamic conditions and combustion modes, including low- and high-temperature ignition regimes in addition to premixed and nonpremixed behaviour. A final DNN model is trained to enforce inherent physical constraints by learning nonlinear functional transformations of the three filtered scalar dissipation rates. The generalizability of this constrained DNN model is demonstrated a priori via conditional statistics evaluated on the lifted spray flame with C1–a configuration that had not been included in the training data. Excellent DNN agreement with conditional DNS statistics is observed, and integrated gradients are computed to identify the most sensitive input variables. The similarity of the marginal PDFs of the most informative input variables and outputs across configurations are quantified via the Wasserstein metric, demonstrating that data-based models may successfully generalize to unseen parametric conditions so long as the most informative input variables share similar distributions across training and testing datasets.

Data-based modelling↗

Effect of the volume fraction gradient on the phase interaction force model for disperse two-phase flows

In this work, the effects of the particle volume fraction gradient on fluid-particle interactions are studied. The phase interaction force is decomposed into three terms. For the first term, namely the symmetrized force density, we present theoretical reasoning and numerical evidence to assume that it is independent of the particle volume fraction gradient. The second term is the particle volume fraction gradient times a newly introduced diffusion stress. The third term is the divergence of the particle-fluid-particle (PFP) stress. If this assumption of independence of the particle volume fraction gradient for the first term can be verified, to the first order of the ratio of the mean distance between particles to the macroscopic lengthscale, all three terms can be studied and modeled in flows with uniform particle distributions. Models thus obtained are applicable to statistically inhomogeneous flows, with the second and third terms accounting for statistical inhomogeneity. To verify this assumption, numerical simulations of flows passing fixed arrays of particles are performed. Both uniform and nonuniform particle volume fractions are studied and compared for disperse multiphase flows with the particle Reynolds numbers ranging from 1 to 100, and particle volume fraction ranging from 1% to 26% in statistically steady states. It is found that the symmetrized force (first) term can be well approximated by the drag force obtained from studies of uniform flows. The diffusion stress is positive along the flow direction and negative in the directions perpendicular to the flow. In the case of moving particles, this stress could potentially cause particle aggregation in the flow direction and dispersion in the directions perpendicular to the flow. Finally, the diffusion stress is only important when there is a volume fraction gradient, while the PFP stress can be important in inhomogeneous flows with either nonuniform particle concentrations or nonuniform average relative velocities between the phases.

42 ENGINEERING↗

Code Release for “Unlocking Extreme Space Weather through Advanced Modeling of Legacy Vela Spacecraft” ER

The software being developed for this project has two mains aims. First, a Bayesian Model Calibration (BMC) procedure is being developed to calibrate a spallation model that simulates protons hitting a spacecraft orbiting earth to real data. The procedure will be developed for general data (there is no data release requested as part of this code release). Second, an inverse physics modeling task is being undertaken to map the number of resulting neutrons observed from this process to the expected number of protons that hit the model. This second task is of statistical interest; to publish on it, the code will need to be open source.

Murph, Alexander↗

Grid Reliability Statistics [SWR-25-45]

This codebase houses a suite of statistical and descriptive analyses of NERC GADS data of interest to grid modelers and planners. This repository is envisioned to house a collection of statistical and descriptive summaries of NERC GADS data, particularly summaries that on their own might be insufficient to warrant publication. In addition, it is intended to disseminate results rather than to enable reproduction as GADS is non-public.

Murphy, Sinnott [National Renewable Energy Laborat↗

Generalized Tensor-on-Tensor Regression (GToTR)

SAND2026-23069O Generalized Tensor-on-Tensor Regression (GToTR) is a Python-based tool for conducting generalized tensor-on-tensor regression. It provides Canonical Polyadic (CP)-based generalized tensor regression models, support for generalized linear model-like families and links, alternating-optimization model fitting methods, and a standard statistics software interface. The tool supports tensor-valued responses and covariates using the open-source Python Tensor Toolbox (pyttb) software package. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Dunlavy, Daniel [Sandia National Lab. (SNL-CA), Li↗

Regulatory Testing of RPP-WTP HLW Glasses to Support Delisting Compliance, VSL-04R4780-1, Rev. 0 (Sep 2004)

The primary goal of the testing described in this report was to collect data to demonstrate compliance of the immobilized high-level waste (IHLW) glasses with delisting requirements. The collected data will be used to support a petition to delist the IHLW glasses destined for the national disposal facility. The Delisting Data Quality Objectives (DQO) (Cook and Blumenkranz 2003) identified a list of (16) inorganic constituents of potential concern (COPCs) and their associated limits for delisting. These COPCs can be divided into three groups, Cases 1, 2, and 3, based on their Toxicity Characteristic Leaching Procedure (TCLP) responses versus their respective delisting limits. To briefly summarize, Case 1 COPCs are those that, when loaded at their highest expected concentration in Waste Treatment Plant (WTP) glasses, are not expected to leach at their respective delisting limits when the glasses are exposed to the TCLP. Case 2 COPCs may reach the delisting limits in TCLP leachates of WTP glasses if loaded to concentrations near their maximum expected concentrations in glass. Finally, Case 3 COPCs are components that are likely to be present in concentrations sufficient to exceed their respective delisting limits in TCLP leachates of some possible glasses. The test objective was to show that, for (i) the expected range of inorganic contents in the waste feed to the Hanford WTP high level waste (HLW) vitrification facility, (ii) the expected range of glass product compositions, and (iii) the Case 1 and Case 2 COPCs identified by the DQO, the IHLW glasses meet all the relevant requirements for delisting. For Case 3 COPCs (i.e., Cd), the testing was to demonstrate the relationship between glass composition and TCLP cadmium (Cd) release, and then employ the results to develop TCLP-composition response models. During WTP operations, TCLP-composition models can be used to predict, within the required statistical uncertainties, TCLP responses of IHLW production glasses that are within the compositional region used to develop the model.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

First Sterile Neutrino Search at NOvA Experiment Using both Neutrino and Anti-Neutrino Data

NOvA is a long-baseline experiment with two functionally identical detectors with a Near Detector placed 1 km from the neutrino source at Fermilab and a Far Detector 810 km away at Ash River, Minnesota. NOvA uses a very high-intensity ($\sim$900 kW) beam of neutrinos and antineutrinos from Fermilab's NuMI beamline, together with two functionally identical detectors placed 14 mrad off the beam axis. NOvA is built to investigate the complex properties of neutrinos with an emphasis on neutrino oscillations. NOvA not only probes active neutrino mixing but also explores exotic oscillations, including sterile neutrino searches. This analysis uses the joint $\nu_{\mu}$ and neutral current (NC) disappearance channels to probe active-sterile mixing in the 3+1 model. Furthermore, this analysis leverages our statistics of 27$\times$10$^{20}$ protons on target (POT) in neutrino mode and 12.5$\times$10$^{20}$ protons on target (POT) in antineutrino mode. In this poster, we will present the first dual-baseline search for sterile neutrinos using both the neutrino and antineutrino datasets.

Chaudhary, S. [Indian Inst. Tech., Guwahati]↗

Implementing belowground controls on nutrient uptake in ELMv2-SPRUCE improves representation of a boreal peatland ecosystem

Boreal peatlands store 13 %–32 % of the global soil carbon (C) stock, a service dependent on plant-mycorrhizal fungi associations. In these nutrient poor systems, ectomycorrhizal and ericoid mycorrhizal fungi supply up to >80 % of the nutrient requirements of their plant hosts, partly with mined nitrogen (N) and phosphorus (P) from soil organic matter that are otherwise inaccessible to plants. Despite the ecological significance, mycorrhizal associations are only represented in a few land surface or ecosystem models. We modify the peatland branch of version 2 of the Energy Exascale Earth System Land Model (ELMv2-SPRUCE) to replace the default photosynthesis-driven inorganic N and P (NP) uptake process with a more realistic representation of the process via three pathways: (1) direct inorganic NP uptake by uncolonized fine roots, (2) indirect inorganic NP acquisition and (3) indirect NP acquisition from organic sources by mycorrhizal roots. We systematically evaluated the performance of the default and modified models with field observations from a whole ecosystem warming and carbon dioxide fertilization experimental site: Spruce and Peatland Responses Under Changing Environment (SPRUCE), in northern Minnesota, USA. The modified model reduces the underestimation of the growth response of shrubs in the default model to warming from 40 %–80 % to 17 %–35 % and reduces the overall relative absolute error on C fluxes from 1.61 to 1.54 in calibration. Improvements on modeled shrub growths and shrub-moss community net ecosystem exchanges are also seen in validation. The improved growth response of shrubs to warming is accompanied by several-fold increase in direct inorganic NP uptake and decrease in fungal colonization rate. The modified model simulates a smaller magnitude of transition of the ecosystem from C sink to C source under warming due to alleviation of plant nutrient limitation. Equifinality analysis shows the newly added parameters in the modified model can be constrained by the observed C fluxes. Sensitivity analysis shows the newly added parameters have stronger statistical interactions than the preexisting parameters in the default model. Overall, the modified model is an improvement over the default ELMv2-SPRUCE and will be a useful tool for understanding boreal peatland change.

Wang, Yaoping [Oak Ridge National Laboratory (ORNL↗

Constraining primordial non-Gaussianity from the large scale structure two-point and three-point correlation functions

Surveys of cosmological large-scale structure (LSS) are sensitive to the presence of local primordial non-Gaussianity (PNG), and may be used to constrain models of inflation. Local PNG, characterized by f NL ⁠, the amplitude of the quadratic correction to the potential of a Gaussian random field, is traditionally measured from LSS two-point and three-point clustering via the power spectrum and bi-spectrum. We propose a framework to measure f NL using the configuration space two-point correlation function (2pcf) monopole and three-point correlation function (3pcf) monopole of survey tracers. Our model estimates the effect of the scale-dependent bias induced by the presence of PNG on the 2pcf and 3pcf from the clustering of simulated dark matter haloes. We describe how this effect may be scaled to an arbitrary tracer of the cosmological matter density. The 2pcf and 3pcf of this tracer are measured to constrain the value of f NL ⁠. In LSS surveys, the effect of imaging systematics on two-point statistics is often degenerate with the PNG signal. Our proposed model employs three-point statistics primarily to break this degeneracy. Using simulations of luminous red galaxies observed by the Dark Energy Spectroscopic Instrument (DESI), we demonstrate the accuracy and constraining power of our method. Our forecast indicates the ability to constrain f NL to a precision of σf NL ≈ 22 with one year of DESI survey data, as well as the ability to constrain the imaging systematic weights in situ.

early Universe↗

Nonparametric Multiparticle Set Methods for Interpreting Environmental Samples

Collection and analysis of environmental samples is commonly used by a range of stakeholders in nuclear safeguards and security contexts. While the ubiquity of samples and their transport in the environment allow regular collection, developing and demonstrating methods for analyzing these samples is difficult. In this work, an environmental sample consists of a set of one or more individual particles. Recent advances in reactor simulation have allowed us to generate data that are more representative of real-world environmental samples, enabling statistically defensible method development and testing. The most notable of these advances is a drastic increase in the number of material depletion regions, which allows our simulations to capture the variation in isotopic composition seen at length scales consistent with environmental samples. Traditional approaches for handling multiparticle samples treat each particle in the sample individually, estimating the quantity of interest (e.g., core-average burnup) resulting from measurement and analysis of signatures (e.g., nuclide assays) from each individual particle. Individual estimates are then averaged to generate a single estimate of the quantity of interest over the entire sample. In this presentation, we introduce two novel approaches for interpreting environmental samples that comprise of multiple particles: (1) the Quantile-Quantile Comparator, which uses a multivariate generalization of quantile-quantile plots for comparing unknown statistical distributions, and (2) the Set Transformer, an attention-based neural network module designed to model interactions among elements (particles) in the input set (sample). Statistically representative sampling cannot be guaranteed as samples are passively collected and are beholden to what particles are available in the environment. These new analysis methods for set-input problems are expected to be more robust than traditional approaches to issues of sampling bias where particles are not uniformly distributed throughout regions of interest, as well as generally outperform traditional approaches by jointly considering all elements in the set. We will present results comparing the performance of traditional single particle approaches and the novel Quantile-Quantile Comparator and Set Transformer for interpretation of simulated environmental samples.

Phathanapirom, Birdy↗

Surrogate model evaluation and building energy benchmarking for commercial buildings

Building energy consumption benchmarking involves challenges associated with various energy patterns for different building types; heating, ventilating, and air-conditioning (HVAC) system types; and climates. Given significant variation in energy use patterns, accurate prediction of long-term energy use using surrogate models remains challenging. Multiple linear regression (MLR) is commonly used for building energy benchmarking because of its simple structure; however, it lacks accuracy compared to other black-box models. Although many studies have compared surrogate models and offer guidance on model selection based on metrics, they do not provide detailed analysis on improving the surrogate model accuracy. In this paper, we implement a surrogate model using polynomial ridge regression (i.e., MLR with interaction terms combined with ridge regularization) for small office and retail strip mall buildings across six HVAC system types and all climate zones, for electricity and natural gas in baseline and proposed scenarios. A simulation workflow is developed using OpenStudio TM /EnergyPlus TM to generate simulation data using measures over a wide range of efficiency inputs. Enhancements based on statistical insights are used for improving the model accuracy using filters, input transformations, and change points. Surrogate models achieved average coefficient of variation of the root mean squared error (CVRMSE) values of 2.17, 1.06, 2.05, and 3.26 for proposed electricity, proposed natural gas, baseline electricity, and baseline natural gas, respectively, with enhancements reducing CVRMSE by an average of 14.9% across all combinations. We provide model interpretation via Shapley additive explanations to determine which input variables most influence energy consumption and provide supportive arguments for enhancements.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Nonlinearities in Magnetic Confinement, Ionospheric Physics, and Population Explosion Leading to Profile Resilience Нелінійності в магнетному утриманні, фізиці іоносфери та процесі демографічного вибуху, які приводять до стійкості профілю

Nonlinearities play an important role in many fields. In the field of thermonuclear fusion, they are involved in questions such as profile resilience and fluid closure. A nonlinear phenomenon common to both fusion and astrophysical planets is the generation of zonal flows. These flows play a significant role in determining the level of turbulence and fluid closure in fusion. The effects of resonance broadening and nonlinearities are investigated, specifically focusing on the case of nonlinear instability that has appeared in drift waves. Similarities and differences between our systems are discussed, with population explosion and the dynamics of nonlinear systems for drift waves by different states in profile resilience described with great precision. The aim of our study is to put our fluid model for drift waves in tokamaks within the wider framework of statistical physics principles. This reinforces our belief in the broad application of our drift wave model, which encompasses current tokamaks, ITER, and the fusion pilot plant.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Ranking Biological Features in Soil-Based Microbial Multi-Omics Data with Integration Modeling

Distinguishing the most important features (e.g. proteins, metabolites, etc.) per group (e.g. control and treatment) is a critical challenge in feature-rich multi-omics experiments, especially in soil data. Traditional feature identification and ranking approaches, such as differential expression, are based on single omics and thus not directly translatable to multi-omics experiments. Here, 5 multi-omics integration models (DIABLO, JACA, MOFA, MultiMLP, and SLIDE) that were not explicitly built for soil data applications were tested using a soil-based multi-omics experiment. The data were obtained from an experimental setup of an autoclaved soil system inoculated with 8 bacteria and using chitin as the carbon source and including samples collected at 0- (control), 4-, 8-, and 12-weeks post-inoculation. The omics data included metaproteomics, 16S rRNA sequencing, and LC-MS/MS metabolomics (in positive and negative mode). Each multi-omics integration model was implemented, and top features were compared to differential univariate statistics per omic type, demonstrating that integration approaches cut the potential number of top features from 2957 identified by differential statistics to 13-224 (a 99.6% to 92.4% reduction). Interestingly, most top features across integration models were not shared; though, scaling and averaging ranks across models shared similar patterns. This work highlights the usefulness of multi-omics integration models in soil-based microbial studies and the power of using multiple integration models together to interpret results.

54 ENVIRONMENTAL SCIENCES↗

Statistical Uncertainty of Inhalation Dose Coefficients: Impact of Particle Deposition in ICRP 66 Human Respiratory Tract Model

Inhaled radioactive materials can pose a long-term health concern, as the material can be incorporated into the body’s metabolic pathways and remain in organs and tissues for extended durations. During the retention period, the radioactive material may localize in a source organ and irradiate adjacent target organs and tissues. Distribution of these materials changes over time, requiring biokinetic modeling to evaluate their movement through various tissues and organs. The evolving distribution depends on multiple inputs characterizing the inhaled material, such as particle size and size distribution, particle density, aspect ratio, specific radionuclide, the chemical form, and solubility. In addition, biological parameters such as breathing rate, breathing type (nasal or nasal/oral), respiratory system morphometry, tidal volume, functional residual capacity, and anatomical dead space all influence material transport. These aerosol properties and physiological characteristics of the respiratory tract jointly define a range of initial conditions that influence the time-dependent distribution of radioactive material. To evaluate both uncertainty in the initial conditions of inhalation exposure and the final output (committed effective dose) from biokinetic models, a Python-based software tool, Radiological Exposure Dose Calculator (REDCAL), was developed to propagate uncertainty within the human respiratory tract model. Focusing on deposition fraction uncertainty, the primary objective was to characterize the initial activity distribution across respiratory regions as a function of anticipated particle sizes and distributions. The impact of the deposition fraction uncertainty was propagated to committed effective dose coefficients for selected radionuclides in a companion publication. For each particle size, a lognormal distribution, characterized by its geometric mean as defined within ICRP Publication 66, serves as the basis for introducing uncertainty into the physical processes governing deposition in various lung regions. Finally, this study addresses the deposition process and examines how uncertainty in deposition mechanisms affects activity distribution in the airways, ultimately presenting the expected range and standard deviation of deposited activity as a function of particle size.

International Commission on Radiological Protectio↗

Stochastic Modeling of the Joint Neutron Number-Cumulative Fission Fragment Kinetic Energy Deposition Distribution and its Statistical Moments [Slides]

We investigate the joint distribution of the neutron number and cumulative fission-fragment kinetic energy (FKE) deposition, with a specific focus on low-order statistical moments: the mean, variance, and correlation. Starting from a point-kinetic framework, we derive a forward Master equation (FME) for the joint distribution and develop the corresponding moment equations.

42 ENGINEERING↗

Analyzing historical snow trends in interior Alaska

Study region The Chena River watershed in Interior Alaska, USA Study focus This study examines 40 years (water years 1982–2021) of snowpack characteristics to consider its hydrological implications in the 5350 km² Chena River basin. Using observations and a fine-scale physics model, we analyzed trends of snow water equivalent (SWE), snow onset and disappearance, and snow cover duration (SCD). New hydrological insights for the region Results indicate a decline in SWE across the modeled domain, averaging a decrease of 3 mm per decade, with larger decreases (up to 10 mm per decade) at lower elevations. While domain-averaged SWE trends were not statistically significant, observed SCD showed statistically significant decreases: −5.2, −5.0, and −4.4 days per decade at Teuchet Creek, Fairbanks F.O., and Little Chena Ridge, respectively. Notably, observations at SNOTEL stations and modeling revealed no statistically significant change in domain-averaged Rain-on-Snow (ROS) events over the 40-year period, contrasting some regional future estimates of increased ROS frequency. Peak streamflow did not consistently correlate with peak SWE levels, suggesting that other environmental factors such as ROS events and rapid temperature increases (e.g., a 10°C spike observed in 1992) are key drivers of hydrological outcomes. These findings improve understanding of complex subarctic hydrological processes impacting permafrost and highlight the need for adaptive water resource management to mitigate multi-factor risks like flooding and wildfire, requiring proactive planning.

54 ENVIRONMENTAL SCIENCES↗

ELM2.1-XGBfire1.0: improving wildfire prediction by integrating a machine learning fire model in a land surface model

Wildfires have shown increasing trends in both frequency and severity across the contiguous United States (CONUS). However, process-based fire models have difficulties in accurately simulating the burned area over the CONUS due to a simplification of the physical process and cannot capture the interplay among fire, ignition, climate, and human activities. The deficiency of burned area simulation deteriorates the description of fire impact on energy balance, water budget, and carbon fluxes in the Earth system models (ESMs). Alternatively, fire models based on machine learning (ML), which capture statistical relationships between the burned area and environmental factors, have shown promising burned area predictions and corresponding fire impact simulation. We develop a hybrid framework (ELM2.1-XGBFire1.0) that integrates an eXtreme Gradient Boosting (XGBoost) wildfire model with the Energy Exascale Earth System Model (E3SM) land model (ELM) version 2.1. A Fortran–C–Python deep learning bridge is adapted to support online communication between ELM and the ML fire model. Specifically, the burned area predicted by the ML-based wildfire model is directly passed to ELM to adjust the carbon pool and vegetation dynamics after disturbance, which are then used as predictors in the ML-based fire model in the next time step. Evaluated against the historical burned area from Global Fire Emissions Database 5 from 2001–2019, the ELM2.1-XGBFire1.0 outperforms process-based fire models in terms of spatial distribution and seasonal variations. The ELM2.1-XGBFire1.0 has proven to be a new tool for studying vegetation–fire interactions and, more importantly, enables seamless exploration of climate–fire feedback, working as an active component of E3SM.

54 ENVIRONMENTAL SCIENCES↗