Search NASASearch

SEARCH · Search NASA

Results for “quantile regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

26 records · Page 2

Statistical Approaches for the Definition of Landslide Rainfall Thresholds and their Uncertainty Using Rain Gauge and Satellite Data

Models for forecasting rainfall-induced landslides are mostly based on the identification of empirical rainfall thresholds obtained exploiting rain gauge data. Despite their increased availability, satellite rainfall estimates are scarcely used for this purpose. Satellite data should be useful in ungauged and remote areas, or should provide a significant spatial and temporal reference in gauged areas. In this paper, the analysis of the reliability of rainfall thresholds based on rainfall remote sensed and rain gauge data for the prediction of landslide occurrence is carried out. To date, the estimation of the uncertainty associated with the empirical rainfall thresholds is mostly based on a bootstrap resampling of the rainfall duration and the cumulated event rainfall pairs (D,E) characterizing rainfall events responsible for past failures. This estimation does not consider the measurement uncertainty associated with D and E. In the paper, we propose (i) a new automated procedure to reconstruct ED conditions responsible for the landslide triggering and their uncertainties, and (ii) three new methods to identify rainfall threshold for the possible landslide occurrence, exploiting rain gauge and satellite data. In particular, the proposed methods are based on Least Square (LS), Quantile Regression (QR) and Nonlinear Least Square (NLS) statistical approaches. We applied the new procedure and methods to define empirical rainfall thresholds and their associated uncertainties in the Umbria region (central Italy) using both rain-gauge measurements and satellite estimates. We finally validated the thresholds and tested the effectiveness of the different threshold definition methods with independent landslide information. The NLS method among the others performed better in calculating thresholds in the full range of rainfall durations. We found that the thresholds obtained from satellite data are lower than those obtained from rain gauge measurements. This is in agreement with the literature, where satellite rainfall data underestimate the 'ground' rainfall registered by rain gauges.

landslide prediction

Tropical Tropospheric Ozone Trends (1990 to 2022): A Re-evaluation Based on SHADOZ and IAGOS Profiles and TOMS/OMI Columns

Changes in tropical tropospheric ozone (TTO) are of importance because this region spans roughly a third of the Earth and portions of it are experiencing variability in trends of ozone precursors (CO, NO x , CH 4 and nonmethane hydrocarbons) associated with economic growth and fires. In addition to ozone changes affecting radiative forcing, tropical ozone is an important source of the OH radical and thus, the oxidizing capacity of the planet (Thompson, 1992). Recent studies examining TTO trends satellite and in-situ observations over the past ~25 years include: Thompson et al., JGR, 2021; Gaudel et al., ACP, 2023; Stauffer et al., ACP, 2023. The results show considerable regional and seasonal variability in TTO trends and sensitivity to data selection, frequency, and statistical method used. The satellite data vary most widely in method, time period and reliability. Here we revisit trends for the 1990-2022 period with the best-characterized buv-based satellite products that span that period (derived from TOMS and OMI, Ziemke et al., 2019). The satellite-based trends are compared to trends based on in-situ data from the Southern Hemisphere Additional Ozonesondes (SHADOZ) network (1998-2022), measurements from selected pre-SHADOZ and IAGOS commercial aircraft data (Gaudel et al., 2023). Among sensitivities examined are the dependence of ozone trends on start and end years, impacts of ENSO events and the COVID-19 perturbation to emissions. Two statistical methods are used, quantile regression (QR) and multiple linear regression (MLR). Trends of total TTO, ozone segments in the boundary layer (to ~700 hPa), and free troposphere (700-300 hPa) are compared.

ozone, OMI, tropospheric ozone, SHADOZ

Beyond pinball loss: Quantile methods for calibrated uncertainty quantification

Amongthemanywaysofquantifying uncertainty in a regression setting, specifying the full quantile function is attractive, as quantiles are amenable to interpretation and evaluation. A model that predicts the true conditional quantiles for each input, at all quantile levels, presents a correct and efficient representation of the underlying uncertainty. To achieve this, many current quantile-based methods focus on optimizing the pinball loss. However, this loss restricts the scope of applicable regression models, limits the ability to target many desirable properties (e.g. calibration, sharpness, centered intervals), and may produce poor conditional quantiles. In this work, we develop new quantile methods that address these shortcomings. In particular, we propose methods that can apply to any class of regression model, select an explicit balance between calibration and sharpness, optimize for calibration of centered intervals, and produce more accurate conditional quantiles. We provide a thorough experimental evaluation of our methods, which includes a high dimensional uncertainty quantification task in nuclear fusion.

97 MATHEMATICS AND COMPUTING

Developing a Hydrological Monitoring and Sub-Seasonal to Seasonal Forecasting System for South and Southeast Asian River Basins

South and Southeast Asia is subject to significant hydrometeorological extremes, including drought. Under rising temperatures, growing populations, and an apparent weakening of the South Asian monsoon in recent decades, concerns regarding drought and its potential impacts on water and food security are on the rise. Reliable sub-seasonal to seasonal (S2S) hydrological forecasts could, in principle, help governments and international organizations to better assess risk and act in the face of an oncoming drought. Here, we leverage recent improvements in S2S meteorological forecasts and the growing power of Earth Observations to provide more accurate monitoring of hydrological states for forecast initialization. Information from both sources is merged in a South and Southeast Asia sub-seasonal to seasonal hydrological forecasting system (SAHFS-S2S), developed collaboratively with the NASA SERVIR program and end-users across the region. This system applies the Noah-MultiParameterization (NoahMP) Land Surface Model (LSM) in the NASA Land Information System (LIS), driven by downscaled meteorological fields from the Global Data Assimilation System (GDAS) and Climate Hazards InfraRed Precipitation products (CHIRP and CHIRPS) to optimize initial conditions. The NASA Goddard Earth Observing System Model - sub-seasonal to seasonal (GEOS-S2S) forecasts, downscaled using the National Center for Atmospheric Research (NCAR) General Analog Regression Downscaling (GARD) tool and quantile mapping, are then applied to drive 5-km resolution hydrological forecasts to a 9-month forecast time horizon. Results show that the skillful predictions of root zone soil moisture can be made one to two months in advance for forecasts initialized in rainy seasons and up to 8 months when initialized in dry seasons. The memory of accurate initial conditions can positively contribute to forecast skills throughout the entire 9-month prediction period in areas with limited precipitation. This SAHFS-S2S has been operationalized at the International Centre for Integrated Mountain Development (ICIMOD) to support drought monitoring and warning needs in the region.

Yifan Zhou

Tolerance bounds for log gamma regression models

The present procedure for finding lower confidence bounds for the quantiles of Weibull populations, on the basis of the solution of a quadratic equation, is more accurate than current Monte Carlo tables and extends to any location-scale family. It is shown that this method is accurate for all members of the log gamma(K) family, where K = 1/2 to infinity, and works well for censored data, while also extending to regression data. An even more accurate procedure involving an approximation to the Lawless (1982) conditional procedure, with numerical integrations whose tables are independent of the data, is also presented. These methods are applied to the case of failure strengths of ceramic specimens from each of three billets of Si3N4, which have undergone flexural strength testing.

Jones, R. A.

Data and scripts associated with “Allometric scaling of hyporheic respiration across basins in the Pacific Northwest USA"

This data package is associated with the publication “Allometric scaling of hyporheic respiration across basins in the Pacific Northwest USA” submitted to JGR-Biogeosciences (Regier et al. 2025).This study used reach-scale modeled estimates of hyporheic aerobic respiration made by the River Corridor Model (Fang et al. 2020) and watershed characteristics across the Willamette and Yakima River basins to explore potential allometric scaling (i.e., power-law relationships between size and function) of cumulative hyporheic respiration across catchment-to-basin scales. Scaling was explored quantitatively via the R2, slope, and y-intercept of relationships between cumulative hyporheic respiration and watershed area, divided into hyporheic exchange flux (HEF) quantiles. We also explored relationships between allometric scaling and other watershed characteristics through linear regression, spatial patterns, and mutual information analyses. Our results also suggest variability of hyporheic respiration allometry for middle exchange flux quantiles, and in relation to land-cover. Our findings provide initial evidence that allometric scaling may be useful for predicting hyporheic biogeochemical dynamics across watersheds from reach to basin scales. This data package is associated with the GitHub repository found at https://github.com/peterregier/rc_wrb_yrb_scaling. The data package is organized into several key directories. The “data” folder contains multiple CSV files, including landscape heterogeneity, scaling analysis, and watershed boundary data. The “figures” folder has all figure files in both PDF and PNG formats. Core analysis scripts and figure generation scripts are in the “scripts” directory, systematically numbered for sequential execution. The root directory includes essential project files; please see the file ending in “flmd.csv” for a list and description of all files contained in this data package and the file ending in “dd.csv” for data dictionaries used to describe tabular column headers.

54 ENVIRONMENTAL SCIENCES

Quantiles, parametric-select density estimation, and bi-information parameter estimators

A quantile-based approach to statistical analysis and probability modeling of data is presented which formulates statistical inference problems as functional inference problems in which the parameters to be estimated are density functions. Density estimators can be non-parametric (computed independently of model identified) or parametric-select (approximated by finite parametric models that can provide standard models whose fit can be tested). Exponential models and autoregressive models are approximating densities which can be justified as maximum entropy for respectively the entropy of a probability density and the entropy of a quantile density. Applications of these ideas are outlined to the problems of modeling: (1) univariate data; (2) bivariate data and tests for independence; and (3) two samples and likelihood ratios. It is proposed that bi-information estimation of a density function can be developed by analogy to the problem of identification of regression models.

Parzen, E.

wa-hls4ml and lui-gnn: A benchmark and GNN-based surrogate model for hls4ml resource and latency estimation

As machine learning (ML) increasingly serves as a tool for addressing real-time challenges in scientific applications, the development of advanced tooling has significantly reduced the time required to iterate on various designs. These advancements have solved major obstacles, but also exposed new challenges. For example, processes that were not previously considered bottlenecks, such as model synthesis, are now becoming limiting factors in the rapid iteration of designs. To reduce these emerging constraints, multiple efforts are being launched toward designing an ML-based surrogate model that estimates resource usage of synthesized accelerator architectures. This model would reduce the design iteration time, especially when designing within a set of given hardware constraints. This approach shows considerable potential, but as it stands, the effort is early and would benefit from coordination and standardization to assist future work as it emerges. We introduce wa-hls4ml, a benchmark for ML accelerator resource and latency estimation, and its corresponding initial dataset of more than 100,000 fully connected neural networks, all synthesized using hls4ml and targeting Xilinx FPGAs. In addition to the resource utilization and latency data provided, the dataset includes generated artifacts and log files for many of the synthesized neural networks, in order to support future research in ML-based code generation. The benchmark evaluates the performance of resource and latency predictors against several common ML model architectures, primarily originating from scientific domains, as exemplar models, as well as the average performance across a subset of the dataset. We measure the performance of a given predictor model through multiple metrics, including $R^2$ score and SMAPE on regression tasks, as well as inference time to further characterize the estimator under test. Additionally, we introduce the latency/utilization inference graph neural network (lui-gnn), a surrogate model that uses a graph neural network to represent input architectures in the form of a directed graph. This graph representation allows for a diverse set of model architectures to all be effectively handled by a surrogate model. We present the architecture and performance of the model, as evaluated by the new proposed benchmark, including SMAPE, $R^2$ score, and inference times, and find that lui-gnn generally predicts latency and utilization for the 75\% quantile within several percent of the synthesized resources on the synthetic test dataset, indicating that this approach of estimating resource and latency via a surrogate models has promise and warrants further research.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS