Search NASASearch

SEARCH · Search NASA

Results for “Statistical techniques”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Partnership Center for High-Fidelity Boundary Plasma Simulation (Final Report)

Within the Partnership Center for High-Fidelity Boundary Plasma Simulation (HBPS), work at UT-Austin was aimed at improved verification, validation, and uncertainty quantification (VVUQ) for edge plasma simulations and on performing gyrokinetics simulations of pedestal instabilities and turbulence in order to expand foundational understanding of pedestal transport. Regarding VVUQ, the accomplishments can be summarized as follows. First, it was shown that the Moment Preserving Constrained Resampling technique, when applied periodically in particle-in-cell simulations in the XGC code, can dramatically improve the accuracy of the simulation at essentially equivalent computational cost. Second, a technique for estimating model correlations, which are required to solve the model selection and sample allocation problem in multifidelity UQ techniques, without sampling the highest fidelity, most computationally expensive model, was developed and demonstrated. Third, previously developed methods for estimating statistical and discretization errors were applied to numerical methods relevant to edge plasma simulations, namely in particle-in-cell-based approaches, and shown to work. Finally, benchmark studies for comparing gyrokinetic codes were developed and performed, leading to reasonable agreement between four commonly used codes. Regarding physics studies, gyrokinetic simulations to investigate microtearing modes in the DIII-D pedestal were performed using the GENE code.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

Measurement of muon antineutrino charged current - 0 meson scattering, using the NOvA Near Detector

Antineutrino interaction cross sections are, at present, poorly constrained, particularly regarding the role of multi-nucleon processes such as 2-particle 2-hole (2p2h) interactions. The associated crosssection systematic uncertainties represent a significant challenge for precision oscillation measurements, especially for the next generation of neutrino experiments such as DUNE. We present a new measurement of the muon antineutrino charged-current cross section without mesons in the final state, using the high-statistics data set of the NOvA Near Detector. The analysis employs a cut-based selection enhanced by machine learning techniques to isolate a high-purity sample dominated by quasielastic (QE) and 2p2h interactions. We present the cross section as a function of the kinetic energy and scattering angle of the outgoing muon. We also present measurements of more model-dependent kinematic variables such as the neutrino energy and momentum transfer, to better probe the underlying nuclear physics. The results are compared against various neutrino event generators to test the robustness of current interaction models.

Vockerodt, Kevin John [Ohio State U.; Queen Mary,

Exascale granular microstructure reconstruction in 3D volumes of arbitrary geometries with generative learning

Reconstructing 3D granular microstructures within volumes of arbitrary geometries from limited 2D image data is crucial for predicting the material properties, as well as performances of structural components accounting for material microstructural effects. We present a novel generative learning framework that enables exascale reconstruction of granular microstructures within complex 3D geometric volumes. Building upon existing transfer learning techniques using pre-trained convolutional neural networks (CNN), we introduce several key innovations to overcome the difficulties inherent in arbitrary geometries. Our framework incorporates periodic boundary conditions using circular padding techniques, ensuring continuity and representativeness of the reconstructed microstructures. We also introduce a novel seamless transition reconstruction (STR) method that creates statistically equivalent transition zones to integrate multiple pre-existing 3D microstructure volumes. Based on STR, we propose a cost-effective strategy for reconstructing microstructures within complex geometric volumes, minimizing computational waste. Validation through numerical experiments using kinetic Monte Carlo simulations demonstrates accurate reproduction of grain statistics, including grain size distributions and morphology. A case study involving the reconstruction of a 4-blade propeller microstructure illustrates the method’s capability to efficiently handle complex geometries. In conclusion, the proposed framework significantly reduces computational demands while maintaining high reconstruction quality, paving the way for scalable microstructure reconstruction in materials design and analysis.

36 MATERIALS SCIENCE

Numerical simulation of involute-plate research reactor flow behavior using RANS, LES and DNS

This paper investigates the flow behavior of involute-plate research reactors by performing Reynolds-Averaged Navier Stokes simulation (RANS), Large Eddy Simulation (LES) and Direct Numerical Simulation (DNS) of the channel flow between fuel plates. By modeling turbulence with different numerical approaches, this study provides data with three levels of fidelity. For the RANS simulation, three widely used turbulence models, i.e., k-ε, k-ω, Reynolds Stress Turbulence model (RST) are applied by using the commercial CFD code STAR-CCM +. For LES and DNS, the open-source CFD code, Nek5000, is used given its outstanding scalability on High Performance Computer (HPC) and high-order technique. The results from RANS simulations are compared with that from LES and DNS for benchmarking. Both macroscale parameters and turbulence statistics, such as velocity magnitude, lateral velocity and turbulence kinetic energy, are presented and analyzed. The results from RANS simulation achieve good agreement with LES and DNS on velocity and turbulence kinetic energy prediction. The RST turbulence model predicts the most similar flow pattern of lateral velocity as compared to LES and DNS. The Lambda-2 (λ2) criterion with a reasonable threshold is used to demonstrate the instantaneous vortices distribution in the involute channel from both LES and DNS calculation. The DNS simulation captures more detailed turbulence especially near the corner, which explains the discrepancy between LES and DNS results near the corner. The normalized RMS error are defined and calculated to assess the performance of those turbulence models. The RST model captures the anisotropic feature of turbulence, which enable it to outperform other turbulence models for predicting the flow behavior in an involute channel. Although some discrepancies are found between LES and DNS results in the corner, the overall deviations between LES and DNS are found to be small. In conclusion, given that the computational cost of DNS calculation is an order of magnitude higher, using LES data for benchmarking RANS model is a cost-effective approach.

DNS

Information-theoretic astrophysical uncertainties in the effective theory of dark matter direct detection

The impact of astrophysical uncertainties in direct detection searches can vary significantly across particle dark matter models and detector targets, due to the different velocity and momentum dependencies of the scattering cross section. We address these uncertainties for all operators of the nonrelativistic effective field theory of dark-matter/nucleon interactions, making use of the Kullback-Leibler (KL) information divergence to measure the deviation of the true dark matter velocity distribution from the Maxwell-Boltzmann form. This approach quantifies how astrophysical uncertainties affect each operator in the effective theory, without assuming any specific functional form for the velocity distribution. While for some operators the uncertainties are smaller than 1 order of magnitude for entropically motivated deviations from the Maxwell-Boltzmann form, for other operators, these uncertainties can be as large as three orders of magnitude near threshold. Furthermore, we identify the dependence of the scattering rate for various operators of the effective theory with different velocity-weighted moments of the velocity distribution, functionally analogous to the mean, variance, or skewness. This provides new analytic insight into which features of the velocity distribution are most relevant to detect a given particle dark matter model. Our technique is general and could be applied to a broader class of physics problems where a physical observable depends on the statistical moments of an uncertain theoretical distribution.

Herrera, Gonzalo [MIT, MKI; Harvard U.; Virginia T

Automating galaxy morphology classification using k -nearest neighbours and non-parametric statistics

ABSTRACT Morphology is a fundamental property of any galaxy population. It is a major indicator of the physical processes that drive galaxy evolution and in turn the evolution of the entire Universe. Historically, galaxy images were visually classified by trained experts. However, in the era of big data, more efficient techniques are required. In this work, we present a k-nearest neighbours based approach that utilizes non-parametric morphological quantities to classify galaxy morphology in Sloan Digital Sky Survey images. Most previous studies used only a handful of morphological parameters to identify galaxy types. In contrast, we explore 1023 morphological spaces (defined by up to 10 non-parametric statistics) to find the best combination of morphological parameters. Additionally, while most previous studies broadly classified galaxies into early types and late types or ellipticals, spirals, and irregular galaxies, we classify galaxies into 11 morphological types with an average accuracy of ${\sim} 80\!-\!90 \, {{\rm per\, cent}}$ per T-type. Our method is simple, easy to implement, and is robust to varying sizes and compositions of the training and test samples. Preliminary results on the performance of our technique on deeper images from the Hyper Suprime-Cam Subaru Strategic Survey reveal that an extension of our method to modern surveys with better imaging capabilities might be possible.

Mukundan, Kavya

Advanced Laboratory and Field Arrays (ALFA)/Lab Collaboration Project (LCP) for Marine Energy (Final Scientific/Technical Report)

The objective of the Advanced Laboratory and Field Arrays (ALFA) project was to reduce the Levelized Cost of Energy (LCOE) of Marine and Hydrokinetic (MHK) energy by leveraging research, development, and testing capabilities at Oregon State University, University of Washington, and the University of Alaska, Fairbanks. ALFA is a project within the Pacific Marine Energy Center (PMEC; formerly NNMREC), a multi-institution entity with a diverse funding base that focuses on research and development for marine renewables. The ALFA project aimed to accelerate the development of next-generation arrays of wave energy conversion (WEC) and tidal energy conversion (TEC) devices through a suite of field-focused R&D activities spanning a broad range of strategic opportunity areas identified in the Funding Opportunity Announcement: • Device and/or array operation and maintenance (O&M) logistics development; • High-fidelity resource characterization and/or modeling technique development and validation; • Array-specific component technology development (e.g. moorings and foundations, transmission, and other offshore grid components); • Array performance testing and evaluation; and • Novel cost-effective environmental monitoring techniques and instrumentation testing and evaluation. The objective of the Lab Collaboration Project (LCP) was to accelerate the development of next-generation marine energy conversion systems. The LCP aimed to achieve these project objectives in collaboration with the national laboratories by: • Developing concept generation and assessment tools; • Improving access to existing testing resources; • Validating collision risk models between fish and turbines; and • Advancing analysis and simulation capabilities for wave-WEC interactions and PTO analysis in nonlinear ocean waves. The ALFA portion of the project was comprised of six overarching technical tasks: • Task 1: Debris Modeling, Detection and Mitigation; • Task 2: Autonomous Monitoring & Intervention; • Task 3: Resource Characterization for Extreme Conditions; • Task 4: Robust Models for Design of Offshore Anchoring and Mooring Systems; • Task 5: Performance Enhancement for Marine Energy Converter (MEC) Arrays; and • Task 6: Evaluating Sampling Techniques for MHK Biological Monitoring. The LCP was divided into four overarching technical tasks: • Task 7: Project Management and Reporting • Task 8: Novel Design and Assessment Methodologies for Wave Energy Converter Design (Wave- SPARC) • Task 9: Testing Access for Commercial Marine Renewable Energy Technology Developers • Task 10: Quantifying Collision Risk for Fish and Turbines • Task 11: Nonlinear Ocean Waves and PTO Control Strategy Each ALFA/LCP task listed above functioned as a separate and discreet project. A final Technical Report was written for each individual task and these reports were uploaded to OSTI, after receiving DOE approval. The following document is a compilation of each of these final, approved reports arranged as individual chapters.

13 HYDRO ENERGY

Transfer functions for Q A /Q B international regulatory limits for the safe transport of radioactive materials

This paper presents a proposed revision of the International Atomic Energy Agency transport regulations, related to the A 1 and A 2 limit values used to determine the radioactive transport classification. Based on the 'Q system', a novel methodology was introduced to derive Q A and Q B values related to scenarios involving external exposure from a distant source. These values are key parameters that respectively represent the total effective dose and total equivalent dose to the skin, from all primary and secondary particles contributing to radiation exposure. The International Working Group (WG A 1 /A 2 ) is established and associated with the TRANSSC Technical Expert Group on Radiation Protection. A review of the A 1 and A 2 values is performed in response to identified limitations within the existing Q system. The followed approach is based on Monte Carlo simulations that enabled the development of transfer functions aimed at reducing computational time and increasing the flexibility of dose evaluations for any radionuclide with known particle emission spectra. This method allows updating the Q A and Q B values to account for future data evolutions (decay data, fluence-to-dose conversion coefficients) and standardizing the calculation of regulation limits across all referenced radionuclides and scenarios related to external exposure. The transfer functions are established using three Monte Carlo simulation codes—FLUKA, Geant4, and MCNP—and address the previous limitations of the 'Q system', reflecting the latest International Commission for Radiation Protection recommendations and improvements in calculation techniques. The results of the WG show consistent agreement across the codes, with minor discrepancies observed at low primary energies due to statistical uncertainties and different handling of stopping power for electrons/positrons in the codes. This revised approach aligns with current standards and recommendations, ensuring that the radiological consequences of transport accidents are acceptable for the new A 1 and A 2 limits from a radiological protection perspective.

61 RADIATION PROTECTION AND DOSIMETRY

Cross sections for the formation of Rb84m,g, Rb83, and Rb82m in Sr86(d,x) reactions up to deuteron energies of 49 MeV: Competition between α-particle and multinucleon emission processes

Cross sections of Sr86(d,x) reactions leading to the products Rb84m,g, Rb83, and Rb82m were measured by the stacked-sample activation technique up to deuteron energies of 49 MeV. Nuclear model calculations were performed using the codes talys and empire, which combine the statistical, precompound, and direct interaction components. In all cases, the empire results were much higher than the talys calculation. Fairly good agreement was obtained between measured data and the talys calculation after some optimization of the input model parameters. Insight into competition between α-particle and multinucleon emission in the Y88 compound-nucleus system was also gained.

59 ≤ A ≤ 89

Precision Measurement of the Neutron Magnetic Form Factor via the Ratio Method at Jefferson Lab Hall A

Protons and neutrons, collectively known as nucleons, are composed of quarks and gluons. The Sachs electromagnetic form factors encode information about the spatial distributions of charge and magnetization in the nucleon, particularly at low momentum transfer. In particular, the neutron magnetic form factor (GMn) provides crucial information about the distribution of magnetization inside the neutron and helps constrain theoretical models of nucleon structure. Quasi-elastic electron scattering from deuterium was measured up to Q^2=13.5 GeV^2 using the Super BigBite Spectrometer in Hall A at Jefferson Lab. In this work, the neutron magnetic form factor GMn was extracted at Q^2 = 3.0 GeV^2 and Q^2=4.5 GeV^2 using the Ratio Method. These results represent a subset of the full dataset collected in this experiment, which extended to significantly higher Q^2. The extracted GMn values agree with the existing global fit within approximately two standard deviations at Q^2=3.0 and show excellent agreement at Q^2=4.5. The measurements achieved systematic uncertainties of about 2% and statistical uncertainties below 0.5%, among the most precise determinations of GMn at these kinematics. These results demonstrate the robustness of the experimental technique and provide an important validation point for future extractions at higher Q^2, where data remain scarce. In addition, the GRINCH heavy gas Cherenkov detector—a key component of the experimental apparatus—was commissioned and achieved an electron detection efficiency of approximately 97%, supporting reliable particle identification. Together, the analysis presented here advances both our understanding of nucleon structure and the validation of the experimental methods and instrumentation used to access it.

Satnik, Maria [College of William and Mary, Willia

Searching for Neutrino Tridents in the NOvA Near Detector

This dissertation presents a search for neutrino trident production in the NOvA near detector through the coherent ``dimuon" channel: $\nu_\mu +\hspace{1pt}\text{X} \rightarrow \nu_\mu + \mu^- + \mu^+ +\hspace{1pt}\text{X}$. Trident production is a rare, purely electroweak process with sensitivity to physics beyond the Standard Model. The theoretical background, motivation for studying the process, and previous experimental measurements are reviewed. The analysis uses data collected by the NOvA near detector (ND) from Fermilab's Neutrinos at the Main Injector (NuMI) beam between November 2014 and February 2024, corresponding to an exposure of $25.5\times 10^{20}$ protons on target. The ND is a segmented tracking calorimeter located 800~m from the beam target, receiving neutrinos with a mean energy of 2~GeV. A multi-pass background reduction strategy is implemented, including the development of a novel dimuon-specific tracking technique. Trident candidates are identified using a boost ed decision tree classifier trained on simulated signal and background events. Limited background Monte Carlo statistics necessitate the use of functional fits to sideband data, which are extrapolated to estimate backgrounds in the signal region. The unblinded data contain 9 trident-like events, with an estimated background of 5.66 $\pm$ 5.15 events. This yields a best fit estimate of 3.34 tridents compared to the Standard Model prediction of 4.66. A profiled Feldman-Cousins method is used to determine a 90\% confidence interval of [0,9.1] on the number of signal events, corresponding to an upper limit of 1.95$\times$ the Standard Model prediction. This result represents the lowest energy search for trident events to date, and the first experimental contribution to the process in 27 years.

Bowles, Reed Scott [Indiana U.]

Searches for Light Dark Matter and Evidence of Coherent Elastic Neutrino-Nucleus Scattering of Solar Neutrinos with the LUX-ZEPLIN (LZ) Experiment

We present searches for light dark matter (DM) with masses 3–9 GeV/𝑐 2 in the presence of coherent elastic neutrino-nucleus scattering (CE⁢𝜈⁢NS) from 8 B solar neutrinos with the LUX-ZEPLIN experiment. This analysis uses a 5.7 tonne-yr exposure with data collected between March 2023 and April 2025. In an energy range spanning 1–6 keV, we report no significant excess of events attributable to dark matter nuclear recoils, but we observe a significant signal from 8 B CE ⁢𝜈 ⁢NS interactions that is consistent with expectation. We set world-leading limits on spin-independent and spin-dependent-neutron DM-nucleon interactions for masses down to 5 GeV/𝑐 2 . In the no-dark-matter scenario, we observe a signal consistent with 8 B CE⁢ 𝜈 ⁢NS events, corresponding to a 4.5⁢𝜎 statistical significance. This is the most significant evidence of 8 B CE 𝜈 ⁢NS interactions and is enabled by robust background modeling and mitigation techniques. This demonstrates LZ’s ability to detect rare signals at keV-scale energies.

Dark matter detectors

Topological Interpretability for Deep Learning

With the growing adoption of AI-based systems across everyday life, the need to understand their decision-making mechanisms is correspondingly increasing. The level at which we can trust the statistical inferences made from AI-based decision systems is an increasing concern, especially in high-risk systems such as criminal justice or medical diagnosis, where incorrect inferences may have tragic consequences. Despite their successes in providing solutions to problems involving real-world data, deep learning (DL) models cannot quantify the certainty of their predictions. These models are frequently quite confident, even when their solutions are incorrect. This work presents a method to infer prominent features in two DL classification models trained on clinical and non-clinical text by employing techniques from topological and geometric data analysis. We create a graph of a model's feature space and cluster the inputs into the graph's vertices by the similarity of features and prediction statistics. We then extract subgraphs demonstrating high-predictive accuracy for a given label. These subgraphs contain a wealth of information about features that the DL model has recognized as relevant to its decisions. We infer these features for a given label using a distance metric between probability measures, and demonstrate the stability of our method compared to the LIME and SHAP interpretability methods. This work establishes that we may gain insights into the decision mechanism of a DL model. This method allows us to ascertain if the model is making its decisions based on information germane to the problem or identifies extraneous patterns within the data.

Spannaus, Adam

Insights into distorted lamellar phases with small-angle scattering and machine learning

Lamellar phases are essential in various soft matter systems, with topological defects significantly influencing their mechanical properties. In this report, we present a machine-learning approach for quantitatively analyzing the structure and dynamics of distorted lamellar phases using scattering techniques. By leveraging the mathematical framework of Kolmogorov–Arnold networks, we demonstrate that the conformations of these distorted phases – expressed as superpositions of complex waves – can be reconstructed from small-angle scattering intensities. Through the contour analysis of wave field phase singularities, we obtain the statistics of the spatial distribution of topological defects. Furthermore, we establish that the temporal evolution of these defects can be derived from the time-dependent traveling wave field, informed by the dispersion relation of spectral components. This method opens new avenues for investigating the dynamics of distorted lamellar phases using various dynamic scattering techniques such as neutron spin echo and X-ray photon correlation spectroscopy. These findings enhance our microscopic understanding of how defects influence the physical properties of lamellar materials, with implications for both equilibrium and non-equilibrium states in general lamellar systems.

36 MATERIALS SCIENCE

Top-down proteomics

Proteoforms arising from posttranslational modifications, genetic polymorphisms, and RNA splice variants, play a pivotal role as the key drivers in biology. Thus, a comprehensive understanding of proteoforms is essential for unraveling the intricacies of biological systems and bridging the gap between genotype and phenotype. By analyzing whole proteins without digestion, top-down proteomics (TDP) provides a holistic view of the proteome and presents a next-generation approach for deciphering protein function, uncovering disease mechanisms, and advancing precision medicine. This Primer embarks on a journey into the world of TDP by encapsulating its historical context, underlying principles, recent advances, and an outlook on the future of TDP. The experimental section navigates instrumentation, sample preparation, intact protein separation, tandem mass spectrometry techniques, and data collection. Results decipher raw data, visualize intact protein spectra, unravel data analysis, and explain proteoform identification, characterization, and quantitation, as well as statistical analysis. Various applications of TDP spanning the human proteoform project, biomedical, biopharmaceutical, and clinical applications are described. These are complemented by discussions on measurement reproducibility, limitations, and a forward-looking perspective outlining uncharted waters where the field can advance, and potential exciting future applications of TDP.

Roberts, David S.

Statistical fracture behavior of doped UO 2 using a ball-on-ring equibiaxial flexure test method

Metal oxide dopants, such as titanium and chromium oxides, have garnered considerable attention for their potential to increase grain size (≥ 30 µm) in UO 2 fuel, purportedly enhancing fission gas retention during reactor operation. Fuel performance is significantly impacted by fuel fracture behavior, so it is important to understand the effects of enhanced grain size and dopant content on UO 2 fuel fracture. UO 2 pellets were doped with 0.1 wt% TiO 2 and 0.3 wt% Cr 2 O 3 to alter density and grain size. Inductively coupled plasma mass spectroscopy measured dopant levels pre- and post-sintering. X-ray diffraction revealed lattice changes and microstrain via Rietveld refinement. Field emission scanning electron microscopy determined grain sizes of approximately 30 µm for TiO 2 doping and 7 µm for Cr 2 O 3 doping. Transverse rupture strength tests were performed on over 30 samples per dataset to obtain characteristic strength and Weibull modulus. Results indicate no statistical difference in fracture strength between 0.1 wt% TiO 2 doped UO 2 and undoped UO 2 , while 0.3 wt% Cr 2 O 3 doped UO 2 exhibited a 20% decrease in fracture strength. Doped UO 2 samples also showed reduced Weibull modulus compared to undoped UO 2 , suggesting increased scatter in fracture strength. This study's findings suggest that titanium and chromium oxide doping in UO 2 , regardless of grain size, induce residual stresses, decreasing fracture strength and increasing variability in fracture behavior.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS

Meeting Global Health Needs via Infectious Disease Forecasting: Development of a Reliable Data-Driven Framework

Infectious diseases (IDs) have a significant detrimental impact on global health. Timely and accurate ID forecasting can result in more informed implementation of control measures and prevention policies. To meet the operational decision-making needs of real-world circumstances, we aimed to build a standardized, reliable, and trustworthy ID forecasting pipeline and visualization dashboard that is generalizable across a wide range of modeling techniques, IDs, and global locations. We forecasted 6 diverse, zoonotic diseases (brucellosis, campylobacteriosis, Middle East respiratory syndrome, Q fever, tick-borne encephalitis, and tularemia) across 4 continents and 8 countries. We included a wide range of statistical, machine learning, and deep learning models (n=9) and trained them on a multitude of features (average n=2326) within the One Health landscape, including demography, landscape, climate, and socioeconomic factors. The pipeline and dashboard were created in consideration of crucial operational metrics—prediction accuracy, computational efficiency, spatiotemporal generalizability, uncertainty quantification, and interpretability—which are essential to strategic data-driven decisions. While no single best model was suitable for all disease, region, and country combinations, our ensemble technique selects the best-performing model for each given scenario to achieve the closest prediction. For new or emerging diseases in a region, the ensemble model can predict how the disease may behave in the new region using a pretrained model from a similar region with a history of that disease. The data visualization dashboard provides a clean interface of important analytical metrics, such as ID temporal patterns, forecasts, prediction uncertainties, and model feature importance across all geographic locations and disease combinations. As the need for real-time, operational ID forecasting capabilities increases, this standardized and automated platform for data collection, analysis, and reporting is a major step forward in enabling evidence-based public health decisions and policies for the prevention and mitigation of future ID outbreaks.

60 APPLIED LIFE SCIENCES

Mountain Basin Controls on the Snow-to-Streamflow Signal: An AIC-Weighted Multiple Linear Regression Framework

A regression-based analysis quantifies how basin characteristics modulate the snow-to-streamflow signal. First, we use the ERA5-Land reanalysis gridded product (European Centre for Medium Range Weather Forecasts reanalysis 5 -Land component) for 4,655 hydrologic unit code - 10 (HUC10) mountain basins across the western United States (US) for water years 1987–2024. Linear regressions are performed for peak snow water equivalent (SWE) and annual streamflow for each mountain basin. Models use ordinary least squares in Python’s statsmodels package. After which, an Akaike Information Criterion (AIC)–weighted ensemble multiple linear regression (MLR) framework with 47 watershed traits is used to predict the linear regression coefficient of determination (r-squared) defining the ability of peak SWE to predict annual streamflow across all mountain basin. Predictor sets are constrained to avoid multicollinearity by excluding models with variance inflation factors (VIF) greater than 5. Mountain basin traits included in the MLR include seasonal climate, topography, vegetation type and structure, and bedrock geology. Accepted models are considered if their AIC is within 2.0 of the model with the minimum AIC, or best model. To compare predictor influence across acceptable models, we computed standardized regression coefficients. To evaluate structural redundancy among models, we constructed binary inclusion vectors for each acceptable model, denoting whether a predictor was present (1) or absent (0). Core predictor variables are defined as occurring in at least 67% of the acceptable models. For this regional analysis, only one model was found acceptable, with higher snow-to-streamflow translation (higher r-squared) occurring in colder mountain basins with higher relative winter precipitation, more snow accumulation and a lower fraction of annual precipitation that falls in the spring and summer. The second component of the data package uses previously published, high-resolution output from an integrated hydrological model of the East River watershed using the U.S. Geological Survey Groundwater and Surface water Flow model (GSFLOW, doi:10.15485/1998576). East River MLR expands upon the approach described above to explore the response of five streamflow metrics—annual streamflow, runoff efficiency, 7-day minimum flow, low-flow duration, and non-perennial stream fraction to snow system indicators including peak SWE, snow-covered area, snow disappearance date, and the fraction of basin area characterized by low-to-no snow, as well as seasonal precipitation and temperature, and annual hydrologic variables representing soil moisture, evapotranspiration (ET), the partitioning of incoming precipitation to evapotranspiration (ET/P), groundwater storage, and groundwater inflow to streams. MLR was done on all water years (P0: 1987-2024) and for each period as determined in the split analysis using pooled regression techniques (P1: 1987-2011 and P2: 2012-2024) to evaluate shifting predictor variable emphasis on streamflow generation. Results indicate that since 2012, peak SWE has lost statistical strength in its prediction of annual streamflow and runoff efficiency, and the indirect influence of spring temperature has emerged as critically important. Low-flow metrics remain largely influenced by soil moisture, vegetation water use and groundwater inflows with summer precipitation becoming a direct influence on minimum summer flow. Together, these data and Python-based analysis tools provide a framework for identifying the key watershed characteristics that control how streamflow responds to snow from year to year. The package also helps quantify uncertainty in statistical models and assess how snow–streamflow relationships vary across regions and over time. This dataset contains comma-separated values files (.csv), text files (.txt), python code files (.py), figure files (.png), and shapefiles (.cpg, .dbf, .prj, .sbn, .sbx, .shp, .xml). Further details on file contents and MLR execution can be found in the readme file and the FLMD files. Work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

54 ENVIRONMENTAL SCIENCES