Search NASASearch

SEARCH · Search NASA

Results for “quantile regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Unveiling the drivers contributing to global wheat yield shocks through quantile regression

Sudden reductions in crop yield (i.e., yield shocks) severely disrupt the food supply, intensify food insecurity, depress farmers' welfare, and worsen a country's economic conditions. Here, we study the spatiotemporal patterns of wheat yield shocks, quantified by the lower quantiles of yield fluctuations, in 86 countries over 30 years. Furthermore, we assess the relationships between shocks and their key ecological and socioeconomic drivers using quantile regression based on statistical (linear quantile mixed model) and machine learning (quantile random forest) models. Using a panel dataset that captures spatiotemporal patterns of yield shocks and possible drivers in 86 countries, we find that the severity of yield shocks has been increasing globally since 1997. Moreover, our cross-validation exercise shows that quantile random forest outperforms the linear quantile regression model. Despite this performance difference, both models consistently reveal that the severity of shocks is associated with higher weather stress, nitrogen fertilizer application rate, and gross domestic product (GDP) per capita (a typical indicator for economic and technological advancement in a country). While the unexpected negative association between more severe wheat yield shocks and higher fertilizer application rate and GDP per capita does not imply a direct causal effect, they indicate that the advancement in wheat production has been primarily on achieving higher yields and less on lowering the possibility and magnitude of sharp yield reductions. Hence, in the context of growing extreme weather stress, there is a critical need to enhance the technology and management practices that mitigate yield shocks to improve the resilience of the world food systems.

60 APPLIED LIFE SCIENCES

Enhancing Solar Power Forecasting with Regularized Constrained Quantile Regression Averaging and Bootstrapping Techniques

Probabilistic solar power forecasting (SPF) plays an essential role in optimizing power-grid operations by quantifying the forecast uncertainty. To improve the accuracy and robustness of probabilistic SPF, this paper introduces the regularized constrained quantile regression averaging (rCQRA) method to combine outputs from multiple PSPF models. In addition, a bootstrapping method was used to quantify model uncertainty, providing insights into the reliability and significance of each ensemble component. To evaluate its efficacy, the proposed rCQRA method is used to integrate four PSPF methods. The resulting SPF models are trained and validated using a real-world six-year dataset from a rooftop solar plant in the USA. The performance of the proposed rCQRA method is evaluated and compared with two benchmark methods under three categories of weather conditions. It is shown that the rCQRA method has superior performance in its forecast reliability, sharpness, and accuracy.

Ensemble learning, probabilistic solar power forec

Quantifying mean, variability, and uncertainty in indoor radon exposure in Pennsylvania using random forest and quantile regression forest models

Radon is a naturally occurring radioactive gas that poses a serious health risk as the primary cause of lung cancer in non-smokers. Despite the well-known adverse association with health outcomes, current radon exposure assessments are limited to county-level or average-level estimates, which fail to capture regional variability. This study uses Machine Learning models, including Random Forest (RF) and Quantile Regression Forest (QRF), to estimate the indoor radon concentrations at the ZCTA (Zip code tabulation area)-level and characterize uncertainties in model estimates. Incorporating geological, meteorological, and building-specific data, the models aim to improve radon risk assessment by capturing mean exposure, variability, and extreme concentration levels. Processed radon test data (n = 718,111) were analyzed using average, variability, and quantile prediction methods. Models that estimate the average radon exposure at the ZCTA-level can yield promising model-fit results, but they do not capture the underlying variability of indoor radon exposure within a ZCTA. We utilize volatility analyses to identify characteristics indicative of high variability of indoor radon exposure. We also show that a QRF model can be used to estimate upper quantiles of residential radon exposure, thereby uncovering localized areas of elevated exposure that were not apparent in mean estimates. The results highlighted the need for a deep characterization of exposure risk and show that regions with moderate average exposure levels could still harbor extreme outliers with implications for evaluating health risks. Utilizing multiple radon exposure models allows for a deeper characterization of radon risk within a geographic area and can better identify high-risk areas. The results from this study provide a foundation for developing mitigation strategies and examining associations between radon exposure and health outcomes at fine scales. Future research should extend the geographic scope and incorporate additional environmental risk factors to establish a comprehensive framework for risk assessment.

Lee, Heechan [ORNL]

Generalized Bayesian MARS: Tools for Stochastic Computer Model Emulation

The multivariate adaptive regression spline (MARS) approach of Friedman and its Bayesian counterpart are effective approaches for the emulation of computer models. The traditional assumption of Gaussian errors limits the usefulness of MARS, and many popular alternatives, when dealing with stochastic computer models. Here, we propose a generalized Bayesian MARS (GBMARS) framework which admits the broad class of generalized hyperbolic distributions as the induced likelihood function. This allows us to develop tools for the emulation of stochastic simulators which are parsimonious, scalable, and interpretable and require minimal tuning, while providing powerful predictive and uncertainty quantification capabilities. GBMARS is capable of robust regression with t distributions, quantile regression with asymmetric Laplace distributions, and a general form of “Normal-Wald” regression in which the shape of the error distribution and the structure of the mean function are learned simultaneously. We demonstrate the effectiveness of GBMARS on various stochastic computer models, and we show that it compares favorably to several popular alternatives.

97 MATHEMATICS AND COMPUTING

Joint Modeling of Wind Speed and Wind Direction Through a Conditional Approach

Atmospheric near surface wind speed and wind direction play an important role in many applications, ranging from air quality modeling, building design, wind turbine placement to climate change research. It is therefore crucial to accurately estimate the joint probability distribution of wind speed and direction. In this work, we develop a conditional approach to model these two variables, where the joint distribution is decomposed into the product of the marginal distribution of wind direction and the conditional distribution of wind speed given wind direction. To accommodate the circular nature of wind direction, a von Mises mixture model is used; the conditional wind speed distribution is modeled as a directional dependent Weibull distribution via a two-stage estimation procedure, consisting of a directional binned Weibull parameter estimation, followed by a harmonic regression to estimate the dependence of the Weibull parameters on wind direction. A Monte Carlo simulation study indicates that our method outperforms two other approaches in estimation efficiency: one that utilizes periodic spline quantile regression and another that generates data from the commonly used Abe-Ley distribution for cylindrical data. We illustrate our method by using the output from a regional climate model to investigate how the joint distribution of wind speed and direction may change under some future climate scenarios. Our method indicates significant changes in the variation of wind speed with respect to some directions.

17 WIND ENERGY

A Data-Driven Exploration of the Impact of Renewable Energy on Inter-Area Oscillations in the U.S. Eastern Interconnection

As increasing amounts of renewable energy (RE) resources are incorporated into the bulk-power grid, power system oscillations are expected to change. This work investigates how RE generation impacts the frequency and damping ratio (DR) of two dominant inter-area modes in the U.S. Eastern Interconnection (EI) using regularly updated estimates collected over a 12-month period. Quantile regression is used to derive the correlation between operating conditions and mode properties, and a bootstrap method is used to quantify the uncertainty associated with the correlation estimates. Results show that with an increase in system load, the frequency of a mode decreases and DR increases. Evidence that increasing RE generation results in an increase in frequency and decline in DR was found for one of the two modes studied. This work shows that increasing RE levels will impact the properties of inter-area oscillations in the EI, but it does not indicate the presence of immediate threats to grid stability. The outlined approach can be used to periodically assess changing mode properties as RE levels continue to grow and flag stability concerns before they become serious reliability threats.

Inter-area oscillation, mode meters, quantile regr

Quantifying spatial and vertical variations in soil C:N relationships in permafrost-affected landscapes

Permafrost regions are experiencing rapid changes that affect carbon (C) and nitrogen (N) cycles, with implications for vegetation dynamics and gas exchanges with the atmosphere. Soil C:N ratio is a key indicator of organic matter quality, yet spatial estimates of N stocks and C:N ratios lag behind those for C. We used quantile regression forests to compare direct and indirect digital soil mapping approaches for predicting soil C:N ratios at 0–30, 30–60, and 60–100 cm depths across a latitudinal transect in Alaska. The indirect approach – deriving C:N from separately predicted C and N stocks – outperformed direct mapping for the surface layer (0–30 cm), while direct mapping was marginally better at greater depths. However, prediction accuracy decreased with depth for both methods. Temperature and topography were the most important predictors. Both approaches overestimated low and underestimated high C:N ratios, with direct mapping showing greater bias. Our results underscore the challenges of modeling C:N ratios in heterogeneous, data-sparse permafrost soils, but also suggest that indirect mapping holds promise if supported by more extensive datasets.

54 ENVIRONMENTAL SCIENCES

What Shapes Transportation Charging Infrastructure Availability? Evidence from Tennessee

This study examines how community, travel, and freight characteristics relate to public charging infrastructure availability across Tennessee ZIP codes. We link Alternative Fuels Data Center station locations with traffic, socioeconomic, demographic, commuting, and freight employment data to build a ZIP code-level dataset. Ordinary least squares regression captures variation in chargers per 10,000 residents (R2=0.311). Quantile regressions at the 25th, 50th, and 75th percentiles, with pseudo R2 values up to 0.099, show that the determinants of infrastructure availability differ across low-, medium-, and high-availability areas. Percent female, percent car commuters, average household size, and median age are negatively associated with charging availability across much of the distribution. Truck traffic is positively associated only in lower-availability ZIP codes, while vehicle miles traveled shifts from a negative association at the lower end of the distribution to a positive association at the upper end. The results provide insight into how public charging deployment aligns with community characteristics, mobility demand, and freight activity across Tennessee. Future work can distinguish charger types and power levels, incorporate land-use and temporal rollout patterns, and examine how charging infrastructure needs differ across urban and rural contexts.

Calderón, Oriana [University of Tennessee, Knoxvil

Uncertainty quantification for neural network potential foundation models

Abstract For neural network potentials (NNPs) to gain widespread use, researchers must be able to trust model outputs. However, the blackbox nature of neural networks and their inherent stochasticity are often deterrents, especially for foundation models trained over broad swaths of chemical space. Uncertainty information provided at the time of prediction can help reduce aversion to NNPs. In this work, we detail two uncertainty quantification (UQ) methods. Readout ensembling, by finetuning the readout layers of an ensemble of foundation models, provides information about model uncertainty, while quantile regression, by replacing point predictions with distributional predictions, provides information about uncertainty within the underlying training data. We demonstrate our approach with the MACE-MP-0 model, applying UQ to the foundation model and a series of finetuned models. The uncertainties produced by the readout ensemble and quantile methods are demonstrated to be distinct measures by which the quality of the NNP output can be judged.

36 MATERIALS SCIENCE

Beyond Point Estimates: Benchmarking Uncertainty Quantification Methods on the AION-1 Astronomical Foundation Model

Foundation models for astronomical surveys offer powerful learned representations that can be transferred to downstream regression tasks such as galaxy property estimation. However, point predictions alone are insufficient for scientific inference; reliable uncertainty quantification (UQ) is essential. We compare seven UQ methods on galaxy property regression using frozen AION-1 foundation-model embeddings, predicting redshift, stellar mass, stellar-population age, gas-phase metallicity, and specific star-formation rate, from Legacy Survey photometry/imaging and DESI spectra, with PROVABGS-derived labels. Distribution-free conformal methods achieve marginal coverage within $\sim$1 pp of the nominal 90% across all properties, while non-conformal baselines (Deep Ensembles, MC~Dropout) fail to calibrate reliably. Among conformal approaches, Conformalized Quantile Regression (CQR) delivers the best coverage in the bin with the poorest model predictions. More importantly, only the Locally Valid and Discriminative (LVD) framework -- particularly when operating on AION-1 embeddings -- also provides finite-sample \emph{local validity}, producing intervals that adapt to each galaxy's local prediction difficulty rather than relying on marginal guarantees alone. These results establish conformal prediction, and LVD in particular, as the preferred UQ framework for uncertainty-aware inference on foundation-model embeddings in astrophysics.

Tame-Narvaez, Karla [Fermilab] (ORCID:000000022249

Methane fluxes in tidal marshes of the conterminous United States

Abstract Methane (CH 4 ) is a potent greenhouse gas (GHG) with atmospheric concentrations that have nearly tripled since pre‐industrial times. Wetlands account for a large share of global CH 4 emissions, yet the magnitude and factors controlling CH 4 fluxes in tidal wetlands remain uncertain. We synthesized CH 4 flux data from 100 chamber and 9 eddy covariance (EC) sites across tidal marshes in the conterminous United States to assess controlling factors and improve predictions of CH 4 emissions. This effort included creating an open‐source database of chamber‐based GHG fluxes ( https://doi.org/10.25573/serc.14227085 ). Annual fluxes across chamber and EC sites averaged 26 ± 53 g CH 4 m −2 year −1 , with a median of 3.9 g CH 4 m −2 year −1 , and only 25% of sites exceeding 18 g CH 4 m −2 year −1 . The highest fluxes were observed at fresh‐oligohaline sites with daily maximum temperature normals (MATmax) above 25.6°C. These were followed by frequently inundated low and mid‐fresh‐oligohaline marshes with MATmax ≤25.6°C, and mesohaline sites with MATmax >19°C. Quantile regressions of paired chamber CH 4 flux and porewater biogeochemistry revealed that the 90th percentile of fluxes fell below 5 ± 3 nmol m −2 s −1 at sulfate concentrations >4.7 ± 0.6 mM, porewater salinity >21 ± 2 psu, or surface water salinity >15 ± 3 psu. Across sites, salinity was the dominant predictor of annual CH 4 fluxes, while within sites, temperature, gross primary productivity (GPP), and tidal height controlled variability at diel and seasonal scales. At the diel scale, GPP preceded temperature in importance for predicting CH 4 flux changes, while the opposite was observed at the seasonal scale. Water levels influenced the timing and pathway of diel CH 4 fluxes, with pulsed releases of stored CH 4 at low to rising tide. This study provides data and methods to improve tidal marsh CH 4 emission estimates, support blue carbon assessments, and refine national and global GHG inventories.

54 ENVIRONMENTAL SCIENCES

Advancing Multiscale Simulation of Plasma-Surface Interfaces

We report the development of an atomistic-informed, surface-state-dependent predictive model for particle exchange in a carbon-tungsten plasma-surface interface. The predictive model uses machine learning (ML) techniques to learn the energy and angular distributions for particle exchange and rate functions for surface state evolution from molecular dynamics simulations of cumulative bombardment of tungsten by energetic carbon ions. Each predictive component is sensitive to the energy and trajectory of incident plasma species and the surface state. The surface state is represented by a set of surface state descriptors, which were derived from the atomistic surface state for each independent carbon bombardment event. These descriptors are representative of the composition and degree of amorphization of the outermost angstrom of surface material and were chosen to optimize predictive performance for particle exchange at the interface. The distributions for particle exchange (reflection/sputtering) are demonstrated to vary with each surface state descriptor, motivating the development of surface-state-dependent particle exchange models for plasma simulations. The performance of various ML methods was compared, including polynomial quantile regression, artificial neural networks, k-nearest neighbors, and random forest algorithms, with polynomial regression performing the best for interpolation and extrapolation of learned relationships. In addition to the particle exchange model, a neutral network was developed and used to identify data sufficiency throughout surface descriptor space, which will enable real-time feedback during future data production to ensure data is produced where it is most needed, and we provide commentary on improvements to the data production workflow for future endeavors.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

Examples of Mission-driven Data Science from Jefferson Lab and ACES

This presentation details mission-driven data science initiatives at Jefferson Lab and the Joint Institute for Advanced Computing on Environmental Studies (ACES). JLab, a U.S. Department of Energy Office of Science national laboratory, operates the Continuous Electron Beam Accelerator Facility (CEBAF), and is the lead institute for the new High Performance Data Facility (HPDF) Hub. The Joint Institute for ACES brings together interdisciplinary teams in health informatics, climate modeling, computer science, and physics to address environmental challenges, including flood modeling. The Hampton Roads region, particularly Norfolk and Virginia Beach, faces increasing flood risks, motivating the need for rapid, reliable, and risk-aware decision support. ACES’s flooding work has a focus on uncertainty quantification (UQ) and machine learning (ML) for coastal flood management. The work is motivated by the increasing vulnerability of communities such as Norfolk and Virginia Beach, Virginia, to frequent coastal flooding events, and the need for rapid, reliable decision support. The research develops computationally efficient ML surrogate models to forecast water levels and flooding risk. A central theme is the quantification and calibration of predictive uncertainty, especially for out-of-distribution (OOD) scenarios, using techniques such as Monte Carlo Dropout, Deep Ensembles, Gaussian Processes, and Deep Quantile Regression (DQR). The study demonstrates that distance-aware UQ is critical for reliable scientific AI, particularly in high-dimensional, safety-critical, and real-time applications.

McSpadden, Diana [Thomas Jefferson National Accele

UNCERTAINTY-AWARE DEEP LEARNING FRAMEWORK FOR FORECASTING COASTAL WATER LEVEL IN VIRGINIA BEACH

Coastal areas like Virginia Beach, USA, are increasingly vulnerable to flooding. To mitigate the impact of flooding, it is crucial for the City of Virginia Beach to have reliable 72-hour-ahead (3 days) forecasts of water levels at key gauge locations. To support this effort, several sensors have been installed throughout the city to monitor water levels and other environmental parameters such as wind speed, precipitation, and atmospheric pressure. Leveraging sensor data from one of these locations, we developed an uncertainty-aware deep learning model to forecast water levels. We employed deep quantile regression (DQR) to quantify variability in the predictions and examined the performance of three different model architectures. In addition to exclusively including historical data, we investigated the improvement wind forecasts provide to the accuracy of 72-hour-ahead water level predictions. The results show a twelvefold improvement in the flood forecast for a real flooding event.

Hasan, Mahmud [Thomas Jefferson National Accelerat

Beyond pinball loss: Quantile methods for calibrated uncertainty quantification

Amongthemanywaysofquantifying uncertainty in a regression setting, specifying the full quantile function is attractive, as quantiles are amenable to interpretation and evaluation. A model that predicts the true conditional quantiles for each input, at all quantile levels, presents a correct and efficient representation of the underlying uncertainty. To achieve this, many current quantile-based methods focus on optimizing the pinball loss. However, this loss restricts the scope of applicable regression models, limits the ability to target many desirable properties (e.g. calibration, sharpness, centered intervals), and may produce poor conditional quantiles. In this work, we develop new quantile methods that address these shortcomings. In particular, we propose methods that can apply to any class of regression model, select an explicit balance between calibration and sharpness, optimize for calibration of centered intervals, and produce more accurate conditional quantiles. We provide a thorough experimental evaluation of our methods, which includes a high dimensional uncertainty quantification task in nuclear fusion.

97 MATHEMATICS AND COMPUTING

Data and scripts associated with “Allometric scaling of hyporheic respiration across basins in the Pacific Northwest USA"

This data package is associated with the publication “Allometric scaling of hyporheic respiration across basins in the Pacific Northwest USA” submitted to JGR-Biogeosciences (Regier et al. 2025).This study used reach-scale modeled estimates of hyporheic aerobic respiration made by the River Corridor Model (Fang et al. 2020) and watershed characteristics across the Willamette and Yakima River basins to explore potential allometric scaling (i.e., power-law relationships between size and function) of cumulative hyporheic respiration across catchment-to-basin scales. Scaling was explored quantitatively via the R2, slope, and y-intercept of relationships between cumulative hyporheic respiration and watershed area, divided into hyporheic exchange flux (HEF) quantiles. We also explored relationships between allometric scaling and other watershed characteristics through linear regression, spatial patterns, and mutual information analyses. Our results also suggest variability of hyporheic respiration allometry for middle exchange flux quantiles, and in relation to land-cover. Our findings provide initial evidence that allometric scaling may be useful for predicting hyporheic biogeochemical dynamics across watersheds from reach to basin scales. This data package is associated with the GitHub repository found at https://github.com/peterregier/rc_wrb_yrb_scaling. The data package is organized into several key directories. The “data” folder contains multiple CSV files, including landscape heterogeneity, scaling analysis, and watershed boundary data. The “figures” folder has all figure files in both PDF and PNG formats. Core analysis scripts and figure generation scripts are in the “scripts” directory, systematically numbered for sequential execution. The root directory includes essential project files; please see the file ending in “flmd.csv” for a list and description of all files contained in this data package and the file ending in “dd.csv” for data dictionaries used to describe tabular column headers.

54 ENVIRONMENTAL SCIENCES

wa-hls4ml and lui-gnn: A benchmark and GNN-based surrogate model for hls4ml resource and latency estimation

As machine learning (ML) increasingly serves as a tool for addressing real-time challenges in scientific applications, the development of advanced tooling has significantly reduced the time required to iterate on various designs. These advancements have solved major obstacles, but also exposed new challenges. For example, processes that were not previously considered bottlenecks, such as model synthesis, are now becoming limiting factors in the rapid iteration of designs. To reduce these emerging constraints, multiple efforts are being launched toward designing an ML-based surrogate model that estimates resource usage of synthesized accelerator architectures. This model would reduce the design iteration time, especially when designing within a set of given hardware constraints. This approach shows considerable potential, but as it stands, the effort is early and would benefit from coordination and standardization to assist future work as it emerges. We introduce wa-hls4ml, a benchmark for ML accelerator resource and latency estimation, and its corresponding initial dataset of more than 100,000 fully connected neural networks, all synthesized using hls4ml and targeting Xilinx FPGAs. In addition to the resource utilization and latency data provided, the dataset includes generated artifacts and log files for many of the synthesized neural networks, in order to support future research in ML-based code generation. The benchmark evaluates the performance of resource and latency predictors against several common ML model architectures, primarily originating from scientific domains, as exemplar models, as well as the average performance across a subset of the dataset. We measure the performance of a given predictor model through multiple metrics, including $R^2$ score and SMAPE on regression tasks, as well as inference time to further characterize the estimator under test. Additionally, we introduce the latency/utilization inference graph neural network (lui-gnn), a surrogate model that uses a graph neural network to represent input architectures in the form of a directed graph. This graph representation allows for a diverse set of model architectures to all be effectively handled by a surrogate model. We present the architecture and performance of the model, as evaluated by the new proposed benchmark, including SMAPE, $R^2$ score, and inference times, and find that lui-gnn generally predicts latency and utilization for the 75\% quantile within several percent of the synthesized resources on the synthetic test dataset, indicating that this approach of estimating resource and latency via a surrogate models has promise and warrants further research.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS