Search NASA⌕ Search

SEARCH · Search NASA

Results for “statistical model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Development of an Unbiased Future Solar Dataset for Solar Resource Adequacy Research Over CONUS

A high-resolution, long-term solar dataset is essential for capturing the variability of solar energy resources and informing strategies to ensure grid reliability and resilience in systems with high levels of solar energy integration. This study focuses on generating unbiased, high-resolution projections of solar irradiance through a statistical downscaling framework, using Earth system model (ESM) simulations obtained from the North American Coordinated Regional Climate Downscaling Experiment (NA-CORDEX). The National Solar Radiation Database (NSRDB) is used to calibrate statistical downscaling models. The newly developed dataset provides solar irradiance, surface air temperature, and surface wind speed at 4-km and hourly resolutions across the contiguous United States (CONUS), based on two future scenarios (RCP4.5 and RCP8.5). This study outlines key steps in developing the high-resolution future solar dataset, including (1) regridding ESM data to a common 20-km resolution grid, (2) correcting ESM biases using the NSRDB, and (3) applying temporal and spatial downscaling methods to generate high-resolution (4-km, hourly) solar projections. Preliminary results indicate that downscaled projections (4-km) captured reasonable spatial patterns when compared to observations across CONUS for four variables. On average across all pixels, 4-km daily-total GHI and DNI projections showed normalized bias (nBias) less than 1% and 6% for GHI and DNI against NSRDB, respectively (nBias less than 1% and 5% for daily-average surface air temperature and surface wind speed). In terms of long-term trend for GHI and DNI, there was no strong increasing or decreasing trend (when compared to surface air temperature), but it showed a very weak decreasing trend.

14 SOLAR ENERGY↗

Hourly PM 2.5 Estimates across California from 2018 to 2023

This study presents a new data set of hourly PM 2.5 concentrations across California from 2018 to 2023 at a three-kilometer resolution. This data set was developed by assimilating observations from PurpleAir and the U.S. EPA Air Quality System monitors into wildfire smoke forecasts from the High-Resolution Rapid Refresh Smoke (HRRR-Smoke) model using the Gridpoint Statistical Interpolation (GSI) three-dimensional variational data assimilation framework. Archived forecasts of modeled wildfire smoke PM 2.5 from HRRR-Smoke create the background field for assimilation, which is then corrected using surface observations of total PM 2.5 . The resulting reanalysis from GSI provides an estimate of total PM 2.5 that minimizes error from both the observational and the model data. Validation results indicate strong performance, with monthly R 2 values ranging from 0.73 to 0.91 across the six-year data set, comparable to other PM 2.5 data sets. Case studies are presented for three major fire events, the 2018 Camp Fire, 2019 Kincade Fire, and 2020 Lightning Complex Fires to demonstrate the data set’s fidelity in resolving plume dynamics and local exposure patterns. Root-mean-squared error averaged over each month scales with average PM 2.5 concentrations, resulting in a low error under typical conditions but higher absolute errors during extreme smoke events. This is the first long-term, hourly PM 2.5 data set of its kind for California and enables the generation of subdaily exposure metrics, such as peak hourly concentrations, exceedance durations, and time-of-day exposure peaks. The novelty and strong validation of this data set make it a compelling resource for future studies on the impact and significance of subdaily PM 2.5 exposure.

PM2.5↗

Generalized master equation for particle transport in binary random media with renewal statistics

Particle transport in binary stochastic mixtures is classically modeled assuming Markovian or exponential mixing statistics but in many applications material memory invalidates the Markov assumption. For non-Markovian mixing characterized by alternating renewal processes, a transport-theoretic framework is presented that provides an exact description of transport in nonscattering random binary media with general non-exponential statistics. Our approach is to Markovianize the problem by augmenting the {material type, particle flux} state space with the age or distance from the last interface. A Chapman-Kolmogorov equation is formulated for the joint probability density of the material type, particle flux, and age, and subsequently reduced to a generalized Master equation (GME) in differential form. This constitutes the primary result of this work. A state-updating Monte Carlo algorithm consistent with the GME is developed and benchmarked against analytical solutions for multiple chord-length laws. For purely absorbing renewal statistical media, the GME reproduces analytical benchmarks for the equilibrium age distribution, interior mean/variance of material-conditioned fluxes, and boundary transmittance. Simulations further demonstrate that a Markov (exponential) approximation of non-exponential statistics can introduce large errors in transmittance and interior flux profiles. Lastly, the reintroduction of memory due to scattering is briefly addressed through heuristic considerations.

Fluctuations & noise↗

A Novel Standard Gibbs Energy of Formation Model for High-Enthalpy Water Systems

This work provides an advanced standard molar Gibbs energy of formation model, informed by molecular statistical thermodynamics (MST), and validated by mineral solubility and ion association reactions. This new model aligns with experimental data within uncertainties and highlights the role of specific MST interactions around the critical point. This work will provide a reliable tool for the optimization of industrial processes operating at the edges of the supercritical domain.

Hall, Derek↗

Development of a 95-Year Solar Dataset for Resource Adequacy Studies

Long-term high-resolution solar data provides enhanced understanding of variability of solar generation and enhances our ability to develop strategies for a resilient and reliable electric grid under high deployment of solar energy. Therefore, it is important to develop long-term synthetic datasets that can provide multiple occurrences of various severe weather scenarios that are expected to test the limits of resource adequacy under scenarios contain various energy generation sources. Examples of such scenarios could be long periods of high temperatures when demand for electricity is high or periods where high winds could lead to a shut-down of transmission lines for long periods of time to ensure fire safety. NREL has developed the first version of such a dataset covering a 95-year period covering 2006-2100 at a 4km hourly resolution. This dataset contains all variables necessary to calculate solar generation. During development of this dataset, we focused on creating unbiased, high-resolution solar irradiance through statistical downscaling methods, using Regional Climate Model (RCM) simulations from the North American Coordinated Regional Climate Downscaling Experiment (NA-CORDEX) as input. The National Solar Radiation Database (NSRDB) containing over 25 years of observations was used to calibrate the statistical downscaling models. This presentation will outline the primary steps in developing this dataset, including (1) regridding RCM data to a common grid at 20-km resolution, (2) correcting RCM biases with NSRDB, (3) applying temporal and spatial downscaling methods to generate high-resolution (4-km, hourly) solar and ancillary data. Additionally, we will present an evaluation of the downscaled data against the NSRDB across various zones in the CONUS. Lastly, we will present a user guide for accessing the datasets.

14 SOLAR ENERGY↗

Dynamic data-driven multiscale modeling for predicting the degradation of a 316L stainless steel nuclear cladding material

Here, we have developed a long short-term memory stacked ensemble (LSTM-SE) surrogate modeling approach that can provide rapid predictions of microstructural evolution and the resultant mechanical properties of American Iron and Steel Institute (AISI) 316L series stainless steel (316LSS) fuel cladding under conditions of varying temperature and radiation dose rate. To acquire training data, we developed and implemented a kinetic Monte Carlo (KMC) model to simulate precipitation kinetics of M 23 C 6 , γ', and G phases within SS316L cladding. Experimentally reported precipitation kinetics of SS316L in literature were linked to the kinetic parameters of the simulated precipitation in our KMC model. The model was then used to simulate microstructure evolution under synthetically generated treatments of varying temperature and radiation dose rate, for periods of up to 3000 hours. Changes in volume fraction, number density, and particle size of precipitates were recorded, and particle area fractions were correlated using statistical methods to develop the surrogate model. Simultaneously, the mechanical properties of the simulated microstructures were evaluated using microstructure-based finite element method (FEM) analysis to determine the elastic modulus, yield stress, ultimate tensile strength, and elongation to failure of the aged microstructures. Using this approach, our surrogate model can predict precipitation behavior within 0.25% volume fraction and mechanical properties within 6% relative error from the values predicted by the KMC and FEM models using 50 training simulations as input. The trained recurrent neural network-based model can return estimations of precipitation kinetics and mechanical properties ~1000 times faster than the physics-based codes. This work demonstrates, as a proof of concept, that reactor material service lifetimes under variable service conditions can be predicted for a statistics-based model from a practicably obtainable dataset.

36 MATERIALS SCIENCE↗

Bias Correction and Statistical Downscaling of Future Solar Irradiance Projections Using the NSRDB

Assessing renewable energy resources under future climate scenarios has been highlighted to understand potential impacts of future climate change in renewable generation on the power sector. Climate model projection has been recognized by the renewable energy community as a useful data set to analyze the impacts of future climate change on renewable resources. However, future climate projections generated from general circulation models (GCMs) contain inherent biases that need to be corrected for accurate analysis of future projections of climate variables. In addition, the coarse spatiotemporal resolution of GCMs needs to be improved for regional climate studies. In this work, we develop statistical methods to downscale future projections of global horizontal irradiance (GHI) in a computationally efficient way. Our approach builds statistical downscaling models that correct bias of climate projection of GHI and downscale the future GHI projection from daily-scale to hourly-scale. The National Solar Radiation Database (NSRDB) is used to calibrate the statistical models and validate the downscaled GHI projections across the contiguous United State (CONUS). Preliminary results show that the statistical approach efficiently downscales climate projections of GHI with a nBIAS of 3%, nMAE of 34 % and nRMSE of 46% calculated against NSRDB for CONUS. This study describes the implemented methodology and initial results as well as future research to create high-resolution climate data sets for solar energy applications.

analytical models↗

AI Model Benchmarking for Nonproliferation Applications: Steel Thread Benchmarking Task Force Technical Report (Rev. 2)

Steel Thread is a NA-22 venture that seeks to build trustworthy, reliable AI models that can be used in a wide variety of nonproliferation tasks. A key aspect of building these models is developing appropriate benchmarks and evaluation methods, which will enable the venture to identify and adapt models to provide the most value in the nonproliferation domain. Benchmarks must be relevant to key tasks in this domain, such as question answering, information retrieval, document summarization and classification, consensus analysis, and image and data analysis. This report 1) provides an overview of benchmark design, evaluation, and challenges; 2) reviews a variety of open benchmarks, with a focus on language models and tasks; and 3) identifies benchmarks that are most relevant to Steel Thread. This report is intended to serve as a basis for further efforts to classify and evaluate benchmarks and their correlation with success on nonproliferation-specific tasks. The Steel Thread venture has defined benchmarks to be a particular combination of a dataset (or datasets) and a metric (or metrics) conceptualized as representing one or more specific tasks or sets of abilities for a specific modality. It is adopted by a research community as a shared framework for comparing methods.1 It includes 1) Data: Labeled (a designated subset not used for training, which could be all the data), 2) Metric: A way to quantify performance, 3) Task/Ability: What the benchmark is testing, 4) Protocol: A structured and repeatable evaluation process, 5) Baseline/Reference Model: For comparison; could be statistical, rule-based, SME-derived, or another model, and 6) Maintenance Plan: to update with new information over time; important for long-term utility. For further clarity, the definition includes what a benchmark, in this context, is not. It is not a corpus of training data, specific to a model (it is intended to apply to a range of models), a universal evaluation of performance, a guarantee that the ‘top’ model on the leaderboard will be the best fit for every specific use case, an all-encompassing proof of a model’s universal quality, nor is it a one-size-fits-all measure of success. It does not cover every real-world constraint (like operational, ethical, or cost considerations), a systems integration test, or a unit test. This definition was inspired by and resulted from discussions within the Steel Thread Benchmarking Task Force. This group was formed to define what we would mean as a benchmark within Steel Thread but persisted as the need to develop a thorough understanding of the large and expanding existing benchmarking space. This technical report is a result of the group’s divide and conquer approach to exploring this space. The release of benchmarks might not be progressing as quickly as model development, but it is moving very fast, as many benchmarks quickly become saturated, when state-of-the-art models score so close to the benchmark’s ceiling that their results are virtually indistinguishable. At that point, the test no longer differentiates between new systems, so researchers usually stop reporting scores as the benchmark no longer informs about improvements from the next generation of models. In the OpenAI announcement of GPT-5, they reported results on six flagship public benchmarks (AIME 2025, SWE-bench Verified, Aider Polyglot, MMMU, HealthBench Hard, GPQA) but the full system-card covers roughly thirty-five separate evaluations, comprising hundreds of test task items in total. There have been some efforts to summarize benchmarks in specific fields, like for text-to-image generation, but these surveys have had a narrow methodology scope. Therefore, a comprehensive survey of all benchmarks or even all benchmarks that could be relevant to Steel Thread is outside of the scope of this report. We chose some specific benchmarks to investigate in detail.

97 MATHEMATICS AND COMPUTING↗

elm-diagnostics

elm-diagnostics is a Python package for computing diagnostic analyses and visualizations for the E3SM Land Model (ELM) component and is meant to support new feature development in ELM. The tool reads model history files and performs quantitative analyses including budget-closure checking, variable transformations, temporal aggregations, and statistical summaries to support model evaluation, validation, and scientific interpretation. The framework is designed for extensibility, with modular architecture enabling straightforward addition of new diagnostic methods, derived variables, analysis types, visualization approaches, and model-specific adaptations

Hoffman, Matt [Los Alamos National Laboratory]↗

Primordial nucleosynthesis with non-extensive statistics

The conventional Big Bang model successfully anticipates the initial abundances of 2 H(D), 3 He, and 4 He, aligning remarkably well with observational data. However, a persistent challenge arises in the case of 7 Li, where the predicted abundance exceeds observations by a factor of approximately three. Despite numerous efforts employing traditional nuclear physics to address this incongruity over the years, the enigma surrounding the lithium anomaly endures. In this context, we embark on an exploration of Big Bang nucleosynthesis (BBN) of light element abundances with the application of Tsallis non-extensive statistics. A comparison is made between the outcomes obtained by varying the non-extensive parameter q away from its unity value and both observational data and abundance predictions derived from the conventional big bang model. Here, a good agreement is found for the abundances of 4 He, 3 He and 7 Li, implying that the lithium abundance puzzle might be due to a subtle fine-tuning of the physics ingredients used to determine the BBN. However, the deuterium abundance deviates from observations.

Bertulani, Carlos A.↗

Hierarchical semi-Markov models with duration-aware dynamics for activity sequences

Residential electricity demand at granular scales is driven by what people do and for how long. Accurately forecasting this demand for applications like microgrid management and demand response therefore requires generative models for activities that can produce realistic daily activity sequences, capturing both the timing and duration of human behavior. This paper develops a generative model of human activity sequences using nationally representative time-use diaries at a 10-min resolution. We use this model to quantify which demographic factors are most critical for improving predictive performance. We propose a hierarchical semi-Markov framework that addresses two key modeling challenges. First, a time-inhomogeneous Markov router learns the patterns of “which activity comes next.” Second, a semi-Markov hazard component explicitly models activity durations, capturing “how long” activities realistically last. To ensure statistical stability when data are sparse, the model pools information across related demographic groups and time blocks. The entire framework is trained and evaluated using survey design weights to ensure our findings are representative of the U.S. population. On a held-out test set, we demonstrate that explicitly modeling durations with the hazard component provides a substantial and statistically significant improvement over purely Markovian models. Furthermore, our analysis reveals a clear hierarchy of demographic factors: Sex, Day-Type, and Household Size provide the largest predictive gains, while Region and Season, though important for energy calculations, contribute little to predicting the activity sequence itself. The result is an interpretable and robust generator of synthetic activity traces, providing a high-fidelity foundation for downstream energy systems modeling.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Microstructure Clones

Background: A material’s microstructure drives its material performance. Contemporary crystal plasticity experiments compare full-field strain measurements of polycrystal specimens to models. Because each specimen is unique, it is impossible to know which features of the observed deformation are deterministic vs statistical; thus, differences between model and experiment may or may not be significant. Objective: This paper introduces the invention of microstructure clones. Microstructure clones are 2D oligocrystal specimens that have nearly identical microstructures to remedy the aforementioned experimental limitations. Having specimens with nearly identical microstructures will allow for multiple destructive tests of a microstructure (either as repeats or intentionally different experiments), an ability to “see the future” by providing insight into how a specimen will deform, variability quantification, and experimental investigations of response to small microstructural changes. Methods: This work introduces microstructure clones. Repeatability of these clones is demonstrated in tensile bars of pure nickel. Local strain measurements from digital image correlation are compared between clone specimens and compared to results from a crystal plasticity finite element model. Results: Two sets of microstructure clones were tested in this study and displayed very consistent deformation responses within each clone set. Small observed differences in deformation invite investigation into microstructure stochasticity and the effect of small microstructural and loading differences. Conclusions: Microstructure clones represent a significant shift in understanding structure–property relationships. This work reshapes experimental crystal plasticity to allow for experiments that control for specific variables, quantification of microstructural stochasticity (and other sources of stochasticity), and opportunities for replicating experiments.

Crystal Plasticity↗

Thermodynamically informed priors for uncertainty propagation in first-principles statistical mechanics

Here, this work demonstrates how first-principles statistical mechanics approaches within a Bayesian framework can quantify and propagate uncertainties to downstream thermodynamic calculations. To address the issue of Bayesian prior selection, knowledge of 0 K ground states in the material system of interest is incorporated into the prior. The effectiveness of this framework is shown by creating a phase diagram for the fcc zirconium nitride system, including confidence intervals on order-disorder transition temperatures.

Bayesian methods↗

Downscaled CMIP5 projections of physical fire risk understate historical trends

Reliable projections of wildfire risk are important for multi-sector impacts analysis. Statistically downscaled and bias-corrected Earth system model ensemble products are routinely used to analyze regional physical wildfire risk, but evaluations of historical observed trends and variability are lacking. Here, we evaluate physical fire risk over the western United States using the Canadian Forest Fire Weather Index (FWI) by comparing model outputs from the Coupled Model Intercomparison Project Phase 5 (CMIP5), statistically downscaled via the Multivariate Adaptive Constructed Analogs (MACA) approach, against the observational target dataset gridMET, a gridded high-resolution surface meteorological product. We analyze multidecadal trends and interannual variability in seasonal average FWI for the historical period and future projections under two emissions scenarios, and we compare MACA-CMIP5 ensemble results with a simple time series model that generates historical and future projections of seasonal FWI based on bootstrapping observed historical trends and variability. Our findings indicate that MACA-CMIP5 accurately captures the magnitude and spatial patterns of seasonally averaged FWI but tends to underestimate historical decadal trends. We show that future increases in fire risk may be underestimated relative to the simple time series model that projects historical variability into the future. We also highlight that model biases in relative humidity contribute significantly to model-data differences. Our results underscore the importance of historical hindcasting exercises for informing broader multi-sector applications.

FWI↗

SUNBIRD : a simulation-based model for full-shape density-split clustering

Combining galaxy clustering information from regions of different environmental densities can help break cosmological parameter degeneracies and access non-Gaussian information from the density field that is not readily captured by the standard two-point correlation function (2PCF) analyses. However, modelling these density-dependent statistics down to the non-linear regime has so far remained challenging. We present a simulation-based model that is able to capture the cosmological dependence of the full shape of the density-split clustering (DSC) statistics down to intra-halo scales. Our models are based on neural-network emulators that are trained on high-fidelity mock galaxy catalogues within an extended-ΛCDM framework, incorporating the effects of redshift-space, Alcock–Paczynski distortions, and models of the halo–galaxy connection. Our models reach sub-percent level accuracy down to $1 \, h^{-1}\text{Mpc}$ and are robust against different choices of galaxy–halo connection modelling. When combined with the galaxy 2PCF, DSC can tighten the constraints on ω cdm , σ 8 , and n s by factors of 2.9, 1.9, and 2.1, respectively, compared to a 2PCF-only analysis. DSC additionally puts strong constraints on environment-based assembly bias parameters.

79 ASTRONOMY AND ASTROPHYSICS↗

Modeling distributed energy resource aggregations in security constrained unit commitment and economic dispatch

The Federal Energy Regulatory Commission (FERC) recently issued Order 2222, which requires all wholesale electricity markets in the US to allow distributed energy resources (DERs) to participate in the market as aggregated resources. These DER aggregations may be composed of many individual resources that are offered and dispatched by the market as a single entity. We present here a model of a distributed energy resource aggregator (DERA) that is scheduled by a market operator’s security constrained unit commitment (SCUC) and security constrained economic dispatch (SCED). The DERA model includes constraints for battery energy storage systems (BESSs), demand response resources (DRRs), and a simple distributed energy resource (DER). This paper describes a model for each resource type and presents two methods for the DERA to generate market offer curves: a profit-maximizing optimization to compute cost curves and a direct cost algorithm to determine dispatch costs for each resource and combine into cost curves. Once all participating DERAs are scheduled in SCUC/SCED, the model is then modified to dispatch individual DERs to maximize profit or minimize schedule deviation of the DERAs. A simulation of a representative day illustrates the DERA offers, the scheduled generation, and the DERA dispatch. Findings show the potential for unavoidable schedule deviations due to internal DER constraints and due to economic incentives to deviate from the SCUC/SCED schedules. This highlights the importance of DERA offer construction on market efficiency and system reliability. Novel aspects of our approach include: (1) We consider the asymmetry of price incentives impacting DERAs from the wholesale market compared to those impacting consumers from the retail market, as imposed by current regulations and laws. (2) We model aggregate consumer response through statistically parameterizable utility functions rather than a potentially impractical approach of modeling each individual consumer. (3) We show how to use the DERA operational dispatch model to create offers into the wholesale electricity market. (4) We show how DERAs may fail to meet their scheduled dispatch because the market offer format may not permit them to fully express their operational features such as intertemporal costs and constraints to the market.

aggregations↗

Conserved Macromolecular Architecture of Poplar Secondary Cell Walls Revealed by ssNMR and Atomistic Modeling

The macromolecular architecture of plant secondary cell walls governs wood's mechanical and biochemical properties, yet its natural intra-species variability remains poorly characterized. Here, we combined 13C solid-state NMR (ssNMR), multivariate statistical analysis, and molecular modeling to profile nanoscale structure across 13 genetically diverse Populus trichocarpa genotypes grown in 13C-enriched atmospheres. SsNMR-derived phenotypes spanning composition, structure, mobility, and inter-polymer proximities reveal a conserved architecture, with a subtle yet coordinated variation organizing into dominant structural and secondary mobility axes. A representative atomistic model captures these features and reproduces experimental metrics. Molecular dynamics simulations support a weak but consistent positive correlation between cellulose abundance and crystalline-like order, with interior cellulose chains enriched in tg (trans-gauche) conformations without expanding crystalline cores. Together, experiment and simulation reveal a genetically buffered, broadly conserved nanoscale architecture across genotypes, where subtle fine-tuning of cellulose bundling and matrix packing balances mechanical performance with biological function.

09 BIOMASS FUELS↗

A high-throughput approach for statistical process optimization in Laser Powder Bed Fusion

Process variability is inherent in metal additive manufacturing (AM). However, it is often overlooked in process optimization frameworks, constraining the understanding of process uncertainties and their influence on parameter selection. To address this, we present an integrated framework that combines high-throughput single-track experiments, GAN-based melt pool geometry extraction, robust statistical and machine learning modeling, and uncertainty-quantified process mapping. Process variability is characterized through single-track melt pool behaviors, and its influence on defect formation is systematically quantified to enable statistically guided process parameter optimization. This approach is demonstrated on Laser Powder Bed Fusion (L-PBF) of stainless steel 316L, effectively capturing the interplay between process parameters, melt pool variability, and defect probability. By integrating uncertainty quantification into process optimization, this study provides a structured methodology for addressing variability challenges in AM quality control, ultimately contributing to enhanced manufacturing reliability.

Laser Powder Bed Fusion↗