Search NASA⌕ Search

SEARCH · Search NASA

Results for “Models, Statistical”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20

Predicting the Evolution of Shallow Cumulus Clouds With a Lotka‐Volterra Like Model

Abstract In numerical weather prediction and climate models, boundary‐layer clouds are controlled by a wide range of subgrid‐scale processes. However, understanding the nature of these processes and their role in the evolution of the cloud size distribution as a whole has been elusive. To address this issue, we adopt a novel empirical framework from the field of population dynamics to model the evolution of cloud size statistics by using the shallow cumulus properties obtained from a large‐eddy simulation (LES). Our approach involves representing the cloud size distribution and the total cloud area using a revised Lotka‐Volterra model and ridge linear model, respectively. The physical interpretation of the total cloud area and coefficients obtained from the optimization of the models reveals three stages probably interpreted by dominant processes: the formation of new clouds, the growth of single clouds, and a steady state with organized transitions involving the growth and decay of multiple clouds. Furthermore, we showcase the potential of this framework to serve as a component of scale‐aware parameterizations of shallow‐convective clouds in atmospheric models.

54 ENVIRONMENTAL SCIENCES↗

CMIP6-based Multi-model Streamflow Projections over the Conterminous US, Version 1.1

This dataset presents an ensemble of streamflow projections covering the conterminous United States (CONUS), developed to support the SECURE Water Act Section 9505 Assessment for the US Department of Energy (DOE) Water Power Technologies Office (WPTO). Multiple Coupled Models Intercomparison Project phase 6 (CMIP6) Global Climate Models (GCMs) were downscaled using either statistical (DBCCA) or dynamical (RegCM) downscaling methods, based on two meteorological reference datasets (Daymet and Livneh). Subsequently, the downscaled precipitation, temperature, and wind speed data were used to drive two calibrated hydrologic models (VIC and PRMS), with total runoff routed through the Routing Application for Parallel computatIon of Discharge (RAPID) routing model, producing an ensemble of streamflow projections across 2.7 million NHDPlusV2 stream reaches across the CONUS. Each ensemble member covers the 1980-2019 baseline and 2020-2059 near-future periods under the high-end (SSP585) emission scenario. Additionally, using only DBCCA and Daymet, the projections extend to the 2060-2099 far-future period and encompass three additional emission scenarios (SSP370, SSP245, and SSP126). This dataset is designed to support the SECURE Water Act Section 9505 Assessment for the US Department of Energy (DOE) Water Power Technologies Office (WPTO). For further details, refer to Kao et al. (2022), Rastogi et al. (2022), and Ghimire et al. (2023).

13 HYDRO ENERGY↗

Applying Transfer Learning for Street-Scale Nuisance Flood Forecasting in Coastal-Urban Cities

An important challenge with Machine Learning (ML) is its transferability; that is, whether a ML model trained on one set of data can be applied to a second set of data without requiring a full re-training of the model. Transfer Learning (TL) addresses this challenge by transferring knowledge learned in the source domain (the data it was trained on) to the target domain (a second set of data that is statistically different but related, which the model was not trained on). This study investigates the use of TL for street-scale nuisance flood forecasting by exploring whether a ML model trained on data collected for one set of streets can effectively forecast flooding for another set of streets in the same city using TL. The envisioned use case is a city deploying a new flood depth monitoring sensor on a street and using TL to apply a ML model, trained on sensor data from an existing flood depth sensor network, to this new street. Eventually, the new flood depth sensor will have a sufficient dataset for training its own ML model, but TL can be used to fill the gap in time while this new dataset is being generated. This method is explored using a Long Short-Term Memory (LSTM) model trained on data for the flood-prone streets of Norfolk City, Virginia. The data used for training includes environmental time series (rainfall, tide), topographic features (Digital Elevation Model (DEM), Topographic Wetness Index (TWI), Depth To Water (DTW)), and street-scale flood depth time series obtained from a high-fidelity physics-based model, acting as a synthetic street-scale stream depth sensor dataset since actual stream depth sensor data is generally unavailable for most cities. A set of 180 flood-prone streets was used to train a base model, while another set of 180 flood-prone streets was used to re-train that model using different TL strategies. The results show that full-weight re-training proved most effective and minimal re-training of only the output layer was insufficient. The advantage of TL was most pronounced when target data was limited, meaning data collected at the new water depth sensor location included generally less than 18 flood events. As target data increased beyond 18 flood events, the benefit of TL diminished relative to training a ML model directly on the local flood events. These findings can assist cities as they implement street-scale flood sensing systems to create accurate forecasts for new sensing locations that do not yet have sufficient data records to train a local ML model.

Roy, Binata [Univ. of Virginia, Charlottesville, V↗

Anomalous electroweak physics unraveled via evidential deep learning

The ever-growing ecosystem of beyond standard model (BSM) calculations and parametrizations has motivated the development of systematic methods for making quantitative cross-comparisons over the wide range of possible models, especially with controllable uncertainties. In this setting, the language of uncertainty quantification (UQ) furnishes useful metrics for assessing statistical overlaps and discrepancies among BSM and related models. In this study, we leverage recent machine learning (ML) developments in evidential deep learning (EDL) for UQ to separate data (aleatoric) and knowledge (epistemic) uncertainties in a model-discrimination setting. We construct several potentially BSM-motivated scenarios for the anomalous electroweak interaction (AEWI) of neutrinos with nucleons in deep inelastic scattering ( v DIS). These scenarios are then quantitatively mapped, as a demonstration, alongside Monte Carlo replicas of the CT18 PDFs used to calculate the $\varDelta \chi ^{2}$ statistic for a typical multi-GeV v DIS experiment, CDHSW. Our framework effectively highlights areas of model agreement and provides a classification of out-of-distribution (OOD) samples. By offering the opportunity to quantitatively understand model overlaps, the approach presented in this work can help facilitate efficient BSM model exploration and exclusion for future New Physics searches.

AI↗

Assessing the Impact of a Forest Canopy on Near-Surface Wind Statistics

Representing the forest canopy in atmospheric numerical models should improve simulated winds within and above the canopy up to a few hundred meters above the ground. Here, in this study, we implement a forest canopy parameterization into the Weather Research and Forecasting (WRF) Model in a large-eddy simulation (LES) mode by applying drag forces across multiple layers within the canopy height. We use unique observations from the Lidar Experiments for Assessing Flow over Forests (LEAFF) field campaign at the Wind River Experimental Forest (WREF) in the U.S. Pacific Northwest to evaluate model performance. In a 2-day case study, the canopy parameterization improved wind predictions both within and above the canopy, particularly during the daytime and at finer grid resolution. Without it, winds were frequently overpredicted above the canopy. Similarly, derived quantities such as the wind shear index also yielded estimates closer to observations with the canopy parameterization implemented. These findings suggest that representing the canopy using drag forces alone can improve simulated mean winds up to 200 m above the surface. Furthermore, second-order statistical moments of wind were more sensitive to canopy density than first-order moments, especially during the daytime. This increased sensitivity and the improved daytime performance in wind speed—evidenced by the lowest bias from observations (3% compared to 20% over diurnal cycle)—imply that winds above the canopy layer are strongly influenced by how well turbulence above the canopy is modeled. The results of this study can serve as a foundation for parameterizing forest canopy effects in coarser weather forecast models.

Energy - Wind↗

Meeting Global Health Needs via Infectious Disease Forecasting: Development of a Reliable Data-Driven Framework

Infectious diseases (IDs) have a significant detrimental impact on global health. Timely and accurate ID forecasting can result in more informed implementation of control measures and prevention policies. To meet the operational decision-making needs of real-world circumstances, we aimed to build a standardized, reliable, and trustworthy ID forecasting pipeline and visualization dashboard that is generalizable across a wide range of modeling techniques, IDs, and global locations. We forecasted 6 diverse, zoonotic diseases (brucellosis, campylobacteriosis, Middle East respiratory syndrome, Q fever, tick-borne encephalitis, and tularemia) across 4 continents and 8 countries. We included a wide range of statistical, machine learning, and deep learning models (n=9) and trained them on a multitude of features (average n=2326) within the One Health landscape, including demography, landscape, climate, and socioeconomic factors. The pipeline and dashboard were created in consideration of crucial operational metrics—prediction accuracy, computational efficiency, spatiotemporal generalizability, uncertainty quantification, and interpretability—which are essential to strategic data-driven decisions. While no single best model was suitable for all disease, region, and country combinations, our ensemble technique selects the best-performing model for each given scenario to achieve the closest prediction. For new or emerging diseases in a region, the ensemble model can predict how the disease may behave in the new region using a pretrained model from a similar region with a history of that disease. The data visualization dashboard provides a clean interface of important analytical metrics, such as ID temporal patterns, forecasts, prediction uncertainties, and model feature importance across all geographic locations and disease combinations. As the need for real-time, operational ID forecasting capabilities increases, this standardized and automated platform for data collection, analysis, and reporting is a major step forward in enabling evidence-based public health decisions and policies for the prevention and mitigation of future ID outbreaks.

60 APPLIED LIFE SCIENCES↗

Statistical inference of anomalous thermal transport with uncertainty quantification for interpretive 2D SOL models

The critical task of inferring anomalous cross-field transport coefficients is addressed in simulations of boundary plasmas with fluid models. A workflow for parameter inference in the UEDGE fluid code is developed using Bayesian optimization with parallelized sampling and integrated uncertainty quantification. In this workflow, transport coefficients are inferred by maximizing their posterior probability distribution, which is generally multidimensional and non-Gaussian. Uncertainty quantification is integrated throughout the optimization within the Bayesian framework that combines diagnostic uncertainties and model limitations. As a concrete example, we infer the anomalous electron thermal diffusivity $\chi_\perp$ from an interpretive 2D model describing electron heat transport in the conduction-limited region with radiative power loss. The workflow is first benchmarked against synthetic data and then tested on H-, L-, and I-mode discharges to match their midplane temperature and divertor heat flux profiles. We demonstrate that the workflow efficiently infers diffusivity and its associated uncertainty, generating 2D profiles that match 1D measurements. Future efforts will focus on incorporating more complicated fluid models and analyzing transport coefficients inferred from a large database of experimental results.

Bayesian optimization↗

Amplified Mesoscale and Submesoscale Variability and Increased Concentration of Precipitation under Global Warming over Western North America

Abstract Cold-season precipitation statistics in simulations from the storm-resolving WRF Model at 6-km and 1-h resolution over western North America are analyzed. Pseudo–global warming future simulations for the 2041–80 period, constrained by GCMs under the RCP8.5 scenario, are compared to the 1981–2020 historical simulation. The analysis focuses on the dynamical properties of precipitation time series at subdaily scales and on the morphology of storms. The statistical distribution of precipitation intensities in each pixel of the simulation domain is characterized through nonparametric statistical indicators: frequency of wet hours, mean wet-hour precipitation intensity, and Gini coefficient as a measure of the temporal concentration of the precipitation volume. Additionally, the temporal and spatial Fourier power spectra of precipitation time series and precipitation fields are analyzed. The half-power period (HPP) and half-power wavelength (HPW) are defined as spectral measures of the characteristic scales of precipitation’s temporal and spatial patterns. The results show statistically significant increases in the mean wet-hour precipitation intensity and in the Gini coefficient in 99% of the pixels, indicating that the seasonal precipitation volume becomes more concentrated within a smaller number of hours with higher precipitation intensity. The statistics of change in the frequency of wet hours are more contrasted across the simulation domain. The changes are also reflected in the power spectra, which show the spatial and temporal variability increasing proportionally more with finer spatial and temporal scales and the HPW and HPP decreasing. These projected changes are expected to have consequences, not only in terms of hydrologic impacts but also in terms of the predictability of precipitation patterns. Significance Statement The precipitation characteristics of winter storms over the western United States and southwestern Canada are analyzed in future climate simulations for the 2041–80 period. As compared to present-day climate, the most intense parts of the storms are projected to produce a higher rainfall volume, with increased concentration over smaller areas and shorter time intervals. The propensity of rainfall intensity to vary rapidly over time will be enhanced in the future according to the simulations. These model predictions imply an increased risk of rapid flooding in small basins. They also suggest that predicting several hours ahead the time and location at which a storm will produce maximum rainfall may become more challenging in the future.

Climate change↗

PCMDI Metrics Package

The Program for Climate Model Diagnosis & Intercomparison (PCMDI) Metrics Package (PMP) is used to provide "quick-look" objective comparisons of Earth System Models (ESMs) with one another and available observations. The PMP provides a diverse suite of analysis utilities each of which produce summary statistics that gauge the consistency between climate model simulations and available observations. The primary application of the PMP is to evaluate simulations from the Coupled Model Intercomparison Project (CMIP). It can also be used to provide objective performance summaries during the model development process as well as selected research purposes.

Ullrich, PaulA [Lawrence Livermore National Labora↗

An Advanced Microscopic Energy Consumption Model for Automated Vehicle:Development, Calibration, Verification

The automated vehicle (AV) equipped with the Adaptive Cruise Control (ACC) system is expected to reduce the fuel consumption for the intelligent transportation system. This paper presents the Advanced ACC-Micro (AA-Micro) model, a new energy consumption model based on micro trajectory data, calibrated and verified by empirical data. Utilizing a commercial AV equipped with the ACC system as the test platform, experiments were conducted at the Columbus 151 Speedway, capturing data from multiple ACC and Human-Driven (HV) test runs. The calibrated AA-Micro model integrates features from traditional energy consumption models and demonstrates superior goodness of fit, achieving an impressive 90% accuracy in predicting ACC system energy consumption without overfitting. A comprehensive statistical evaluation of the AA-Micro model's applicability and adaptability in predicting energy consumption and vehicle trajectories indicated strong model consistency and reliability for ACC vehicles, evidenced by minimal variance in RMSE values and uniform RSS distributions. Conversely, significant discrepancies were observed when applying the model to HV data, underscoring the necessity for specialized models to accurately predict energy consumption for HV and ACC systems, potentially due to their distinct energy consumption characteristics.

Ma, Ke↗

Mining Product Reviews for Important Product Features of Refurbished iPhones

Problem: Remanufacturers want to increase consumer interest in refurbished products, which motivates the need to understand which product features are important to buyers of refurbished products such as mobile phones. Research Questions: This study addresses two questions. First, which product features are most important for buyers of refurbished iPhones? Second, how do those preferences differ from the preferences of buyers of new iPhones? Methods: Online reviews of iPhones are obtained and converted into a document–term matrix. Using this text model, three subsets of features are identified using statistical analysis of frequency of mention: most frequent, average, and least frequent. A logistic regression (LR) model is then used to identify which features are most predictive of whether a review is for a new or refurbished phone. Results: Buyers of refurbished phones mention battery health, screen/display, shell condition, and brand significantly more often than other features. Directly contrasting reviews of refurbished versus new phones shows that shell condition, brand, speaker, and charger are found to be the most predictive product features indicated in reviews for refurbished phones. Of those, the shell condition is significantly more predictive than the others. Implications: The results identify product features that remanufacturers of iPhones can emphasize to increase customer demand.

Anisi, Atefeh↗

DESI 2024: reconstructing dark energy using crossing statistics with DESI DR1 BAO data

Here, we implement Crossing Statistics to reconstruct in a model-agnostic manner the expansion history of the universe and properties of dark energy, using DESI Data Release 1 (DR1) BAO data in combination with one of three different supernova compilations (PantheonPlus, Union3, and DES-SN5YR) and Planck CMB observations. Our results hint towards an evolving and emergent dark energy behaviour, with negligible presence of dark energy at z ≳ 1, at varying significance depending on data sets combined. In all these reconstructions, the cosmological constant lies outside the 95% confidence intervals for some redshift ranges. This dark energy behaviour, reconstructed using Crossing Statistics, is in agreement with results from the conventional w 0 –w a dark energy equation of state parametrization reported in the DESI Key cosmology paper. Our results add an extensive class of model-agnostic reconstructions with acceptable fits to the data, including models where cosmic acceleration slows down at low redshifts. We also report constraints on H 0 r d from our model-agnostic analysis, independent of the pre-recombination physics.

79 ASTRONOMY AND ASTROPHYSICS↗

A Probabilistic Model for Global EMIC Wave Activity Using Van Allen Probes Observations

Electromagnetic ion cyclotron (EMIC) waves play a key role in radiation belt dynamics through resonant interactions. However, their low occurrence probability, high variability, and spatial intermittency pose challenges for accurate modeling. In this study, we present a machine learning (ML)-based global EMIC wave model built on the entire data set from the Van Allen Probes mission. To capture the distinct statistical characteristics of wave occurrence and amplitude, the model is separated into two modules: an occurrence model trained using ML techniques, and a wave amplitude model sampled from observed probability distributions. The input parameters are limited to real-time or predictable variables to ensure practical applicability. Our model shows strong performance across the entire test set and demonstrates improved predictive capability over a baseline random occurrence model, particularly during quiet geomagnetic conditions. Evaluation during both quiet and active periods confirms the model's ability to represent the clustered and intermittent nature of EMIC wave activity. Furthermore, the model provides global estimates of wave power, enabling integration with radiation belt electron data and showing signatures consistent with wave-induced scattering. We found a good correlation between the global wave activity from the model and relativistic electron observation by Van Allen Probes, regardless of the availability of in situ wave observations. The modular structure of the model also allows for straightforward expansion for additional wave properties, such as wave frequency, which can be modeled independently. This flexible, event-sensitive approach offers a promising framework for data-driven radiation belt simulations and space weather applications.

79 ASTRONOMY AND ASTROPHYSICS↗

A new metrics framework for quantifying and intercomparing atmospheric rivers in observations, reanalyses, and climate models

We present a new atmospheric river (AR) analysis and benchmarking tool, namely Atmospheric River Metrics Package (ARMP). It includes a suite of new AR metrics that are designed for quick analysis of AR characteristics via statistics in gridded climate datasets such as model output and reanalysis. This package can be used for climate model evaluation in comparison with reanalysis and observational products. Integrated metrics such as mean bias and spatial pattern correlation are efficient for diagnosing systematic AR biases in climate models. For example, the package identifies the fact that, in CMIP5 and CMIP6 (Coupled Model Intercomparison Project Phases 5 and 6) models, AR tracks in the South Atlantic are positioned farther poleward compared to ERA5 reanalysis, while in the South Pacific, tracks are generally biased towards the Equator. For the landfalling AR peak season, we find that most climate models simulate a completely opposite seasonal cycle over western Africa. This tool can also be used for identifying and characterizing structural differences among different AR detectors (ARDTs). For example, ARs detected with the Mundhenk algorithm exhibit systematically larger size, width, and length compared to the TempestExtremes (TE) method. The AR metrics developed from this work can be routinely applied for model benchmarking and during the development cycle to trace performance evolution across model versions or generations and set objective targets for the improvement of models. They can also be used by operational centers to perform near-real-time climate and extreme event impact assessments as part of their forecast cycle.

58 GEOSCIENCES↗

Constraining the phase shift of relativistic species in DESI BAOs

In the early Universe, neutrinos decouple quickly from the primordial plasma and propagate without further interactions. The impact of free-streaming neutrinos is to create a temporal shift in the gravitational potential that impacts the acoustic waves known as baryon acoustic oscillations (BAOs), resulting in a non-linear spatial shift in the Fourier-space BAO signal. In this work, we make use of and extend upon an existing methodology to measure the phase shift amplitude $\beta _{\phi }$ and apply it to the Dark Energy Spectroscopic Instrument (DESI) Data Release 1 (DR1) BAOs with an anisotropic BAO fitting pipeline. We validate the fitting methodology by testing the pipeline with two publicly available fitting codes applied to highly precise cubic box simulations and realistic simulations representative of the DESI DR1 data. We find further study towards the methods used in fitting the BAO signal will be necessary to ensure accurate constraints on $\beta _{\phi }$ in future DESI data releases. Using DESI DR1, we present individual measurements of the anisotropic BAO distortion parameters and the $\beta _{\phi }$ for the different tracers, and additionally a combined fit to $\beta _{\phi }$ resulting in $\beta _{\phi } = 2.7 \pm 1.7$. After including a prior on the distortion parameters from constraints using Planck we find $\beta _{\phi } = 2.7^{+0.60}_{-0.67}$ suggesting $\beta _{\phi } > 0$ at 4.3$\sigma$ significance. This result may hint at a phase shift that is not purely sourced from the standard model expectation for $N_{\rm {eff}}$ or could be a upwards statistical fluctuation in the measured $\beta _{\phi }$; this result relaxes in models with additional freedom beyond Lambda-cold dark matter.

79 ASTRONOMY AND ASTROPHYSICS↗

Probabilistic Error Bounds for Low-Rank Tensor Decompositions Used in Large-Scale Data Analysis Applications (LDRD Final Report)

This report documents a research project on analyzing low-rank tensor models for data analysis that took place at Sandia National Laboratories from October 2023–September 2025. The focus of this work was to extend theoretical frameworks from statistics and probability theory for use with models for scalar, vector, and matrix data to models with tensor, or general multi-dimensional array, data. Through this work, we have provided a new set of tools for bounding errors on low-rank tensor models of both complete and sampled data. The remainder of this report is organized as follows. In Section 1, we describe the proposed work at the start of the project. Section 2 describes the research advances made as part of the project. Other research contributions in the form of conference presentations and software development is provided in Section 3. Workforce development at Sandia and Florida Atlantic University (via a subcontract on this project) is provided in Section 4.

97 MATHEMATICS AND COMPUTING↗

Comparison of Machine Learning-Based Predictive Models of the Nutrient Loads Delivered from the Mississippi/Atchafalaya River Basin to the Gulf of Mexico

Predicting nutrient loads is essential to understanding and managing one of the environmental issues faced by the northern Gulf of Mexico hypoxic zone, which poses a severe threat to the Gulf’s healthy ecosystem and economy. The development of hypoxia in the Gulf of Mexico is strongly associated with the eutrophication process initiated by excessive nutrient loads. Due to the complexities in the excessive nutrient loads to the Gulf of Mexico, it is challenging to understand and predict the underlying temporal variation of nutrient loads. The study was aimed at identifying an optimal predictive machine learning model to capture and predict nonlinear behavior of the nutrient loads delivered from the Mississippi/Atchafalaya River Basin (MARB) to the Gulf of Mexico. For this purpose, monthly nutrient loads (N and P) in tons were collected from US Geological Survey (USGS) monitoring station 07373420 from 1980 to 2020. Machine learning models—including autoregressive integrated moving average (ARIMA), gaussian process regression (GPR), single-layer multilayer perceptron (MLP), and a long short-term memory (LSTM) with the single hidden layer—were developed to predict the monthly nutrient loads, and model performances were evaluated by standard assessment metrics—Root Mean Square Error (RMSE) and Correlation Coefficient (R). The residuals of predictive models were examined by the Durbin–Watson statistic. The results showed that MLP and LSTM persistently achieved better accuracy in predicting monthly TN and TP loads compared to GPR and ARIMA. In addition, GPR models achieved slightly better test RMSE score than ARIMA models while their correlation coefficients are much lower than ARIMA models. Moreover, MLP performed slightly better than LSTM in predicting monthly TP loads while LSTM slightly outperformed for TN loads. Furthermore, it was found that the optimizer and number of inputs didn’t show effects on the LSTM performance while they exhibited impacts on MLP outcomes. This study explores the capability of machine learning models to accurately predict nonlinearly fluctuating nutrient loads delivered to the Gulf of Mexico. Further efforts focus on improving the accuracy of forecasting using hybrid models which combine several machine learning models with superior predictive performance for nutrient fluxes throughout the MARB.

54 ENVIRONMENTAL SCIENCES↗

“Which Projections Do I Use?” Strategies for Climate Model Ensemble Subset Selection Based on Regional Stakeholder Needs

Climate model (or earth system model) projections are increasingly used for climate adaptation planning and impact assessments. As part of this process, many end‐users evaluate a subset of downscaled climate projections without being aware of the implications of downscaling methodology for statistics or event outcomes. Approaches for determining a subset of global climate models to use often focus on values from the raw models, rather than from their downscaled counterparts, in other words assuming that the statistical distribution of the multi‐model ensemble does not change post downscaling. This study demonstrates that a downscaled ensemble will typically retain the change distribution as a raw ensemble, but individual models can differ dramatically post‐downscaling. We recommend that subset‐selection methods account for this possibility and that decision‐relevant downscaled climate projections provide proper descriptions of fitness‐for‐purpose and essential caveats, so that non‐specialists can interpret the results with an appropriate level of confidence.

54 ENVIRONMENTAL SCIENCES↗