Search NASA⌕ Search

SEARCH · Search NASA

Results for “Pareto distribution”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Measuring the giant radio galaxy length distribution with the LoTSS

Many massive galaxies launch jets from the accretion disk of their central black hole, but only ~10 3 instances are known in which the associated outflows form giant radio galaxies (GRGs, or giants): luminous structures of megaparsec extent that consist of atomic nuclei, relativistic electrons, and magnetic fields. Large samples are imperative to understanding the enigmatic growth of giants, and recent systematic searches in homogeneous surveys constitute a promising development. For the first time, it is possible to perform meaningful precision statistics with GRG lengths, but a framework to do so is missing. We measured the intrinsic GRG length distribution by combining a novel statistical framework with a LOFAR Two-metre Sky Survey (LoTSS) sample of freshly discovered giants. In turn, this allowed us to answer an array of questions on giants. For example, we can now assess how rare a 5 Mpc giant is compared with one of 1 Mpc, and how much larger – given a projected length – the corresponding intrinsic length is expected to be. Notably, we can now also infer the GRG number density in the Local Universe. We assumed the intrinsic GRG length distribution to be Paretian (i.e. of power-law form) with tail index ξ, and predicted the observed distribution by modelling projection and selection effects. To infer ξ, we also systematically searched the LoTSS for hitherto unknown giants and compiled the largest catalogue of giants to date. We show that if intrinsic GRG lengths are Pareto distributed with index ξ, then projected GRG lengths are also Pareto distributed with index ξ. Selection effects induce curvature in the observed projected GRG length distribution: angular length selection flattens it towards the lower end, while surface brightness selection steepens it towards the higher end. We explicitly derived a GRG’s posterior over intrinsic lengths given its projected length, laying bare the ξ dependence. We also discovered 2060 giants within LoTSS DR2 pipeline products; our sample more than doubles the known population. Spectacular discoveries include the largest, second-largest, and fourth-largest GRG known (l p = 5.1 Mpc, l p = 5.0 Mpc, and l p = 4.8 Mpc), the largest GRG known hosted by a spiral galaxy (l p = 2.5 Mpc), and the largest secure GRG known beyond redshift 1 (l p = 3.9 Mpc). We increase the number of known giants whose angular length exceeds that of the Moon from 10 to 23; among the discoveries is the angularly largest known radio galaxy in the Northern Sky, which is also the angularly largest known GRG ($\phi$ = 2°). Combining theory and data, we determined that intrinsic GRG lengths are well described by a Pareto distribution, and measured the index ξ = –3.5 ± 0.5. This implies that, given its projected length, a GRG’s intrinsic length is expected to be just 15% larger. Finally, we determined the comoving number density of giants in the Local Universe to be n GRG = 5 ± 2(100 Mpc) –3 . We developed a practical mathematical framework that elucidates the statistics of giant radio galaxy lengths. Through a LoTSS search, we also discovered 2060 new giants. By combining both advances, we determined that intrinsic GRG lengths are well described by a Pareto distribution with index ξ = –3.5 ± 0.5, and that giants are truly rare in a cosmological sense: most clusters and filaments of the Cosmic Web are not currently home to a giant. Thus, our work yields new observational constraints for analytical models and simulations featuring radio galaxy growth.

79 ASTRONOMY AND ASTROPHYSICS↗

Detecting Permafrost Active Layer Thickness Change From Nonlinear Baseflow Recession

We report permafrost underlies about one fifth of the global land area and affects ground stability, freshwater runoff, soil chemistry, and surface-atmosphere gas exchange. The depth of thawed ground overlying permafrost (active layer thickness) has broadly increased across the Arctic in recent decades, coincident with a period of increased streamflow, especially the lowest flows (baseflow). Mechanistic links between active layer thickness and baseflow have recently been explored using linear reservoir theory, but most watersheds behave as nonlinear reservoirs. We derive theoretical nonlinear relationships between long-term average saturated soil thickness $\bar{η}$(proxy for active layer thickness) and long-term average baseflow. When applied to 38 years of daily streamflow data for the Kuparuk River basin on the North Slope of Alaska, the theory predicts $\bar{η}$ increased 0.17 ± 0.22[2σ] cm a -1 between 1983 and 2020 (6.4 ± cm total). The rate of increase nearly doubled to 0.29 ± 0.31 cm a -1 between 1990 and 2020, during which time local field measurements from Circumpolar Active Layer Monitoring sites indicate the active layer increased 0.31 ± 0.22 cm a -1 . The predicted rate of increase more than doubled again between 2002 and 2020, outpacing a near doubling of observed active layer thickening, consistent with trends in terrestrial water storage inferred from Gravity Recovery and Climate Experiment satellite gravimetry and Modern-Era Retrospective Analysis for Research and Applications climate reanalysis. Overall, hydrologic change is accelerating in the Kuparuk River basin, and we provide a theoretical framework for estimating basin-scale changes in active layer water storage from streamflow measurements.

54 ENVIRONMENTAL SCIENCES↗

baseflow: a MATLAB and GNU Octave package for baseflow recession analysis

baseflow is a MATLAB® toolbox designed for baseflow recession analysis, a technique used in hydrologic science to infer aquifer properties from streamflow. By leveraging widely available streamflow data, baseflow can be used to estimate aquifer properties such as hydraulic conductivity and drainable porosity over the modern instrumental stream gage record. The toolbox is intended for analysis of measured streamflow values recorded on a daily timestep, and is tailored for shallow, unconfined riparian aquifers that discharge groundwater laterally into adjacent streams. Additionally, baseflow can analyze the collective behavior of individual hillslope aquifers constituting hydrologic catchments, known as “watersheds”, from a nonlinear dynamical systems perspective. The toolbox incorporates recent advances in baseflow recession analysis to enable objective estimations of aquifer properties, and their sensitivity to methodological decisions, at both hillslope and catchment scales.

97 MATHEMATICS AND COMPUTING↗

Seeing through noise in power laws

Despite widespread claims of power laws across the natural and social sciences, evidence in data is often equivocal. Modern data and statistical methods reject even classic power laws such as Pareto’s law of wealth and the Gutenberg–Richter law for earthquake magnitudes. We show that the maximum-likelihood estimators and Kolmogorov–Smirnov (K-S) statistics in widespread use are unexpectedly sensitive to ubiquitous errors in data such as measurement noise, quantization noise, heaping and censorship of small values. This sensitivity causes spurious rejection of power laws and biases parameter estimates even in arbitrarily large samples, which explains inconsistencies between theory and data. We show that logarithmic binning by powers of λ > 1 attenuates these errors in a manner analogous to noise averaging in normal statistics and that λ thereby tunes a trade-off between accuracy and precision in estimation. Binning also removes potentially misleading within-scale information while preserving information about the shape of a distribution over powers of λ, and we show that some amount of binning can improve sensitivity and specificity of K-S tests without any cost, while more extreme binning tunes a trade-off between sensitivity and specificity. We therefore advocate logarithmic binning as a simple essential step in power-law inference.

97 MATHEMATICS AND COMPUTING↗

Seeing through noise in power laws

Despite widespread claims of power laws across the natural and social sciences, evidence in data is often equivocal. Modern data and statistical methods reject even classic power laws such as Pareto’s law of wealth and the Gutenberg–Richter law for earthquake magnitudes. We show that the maximum-likelihood estimators and Kolmogorov–Smirnov (K-S) statistics in widespread use are unexpectedly sensitive to ubiquitous errors in data such as measurement noise, quantization noise, heaping and censorship of small values. This sensitivity causes spurious rejection of power laws and biases parameter estimates even in arbitrarily large samples, which explains inconsistencies between theory and data. We show that logarithmic binning by powers of λ > 1 attenuates these errors in a manner analogous to noise averaging in normal statistics and that λ thereby tunes a trade-off between accuracy and precision in estimation. Binning also removes potentially misleading within-scale information while preserving information about the shape of a distribution over powers of λ, and we show that some amount of binning can improve sensitivity and specificity of K-S tests without any cost, while more extreme binning tunes a trade-off between sensitivity and specificity. We therefore advocate logarithmic binning as a simple essential step in power-law inference.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Long-Term Impacts of Constrained Transmission Deployment on the Cost-Reliability Tradeoff

Traditional Resource Adequacy (RA) frameworks in the U.S. undervalue the contributions of inter-regional transmission to resource adequacy during stress periods, focusing on the availability of nameplate capacity instead. However, availability of nameplate capacity does not always translate into electricity delivery, especially during tail events. Moreover, the rapid deployment of energy-limited resources and increasing electricity demand challenge existing resource adequacy frameworks and couple regional electricity demand and availability of supply via transmission. We propose a two-stage framework that goes beyond the existing capacity-centered approaches to reveal the RA contributions of transmission. In the first stage we introduce a multi-objective optimization framework to quantify the merits of transmission expansion via Pareto Frontiers under alternative futures of no transmission investment, primary energy resources availability and demand growth. The second stage focuses on tail events and leverages the results of the first stage to characterize the risk profile of regional consumers across the U.S. under the alternative energy futures. We find that no new transmission can lead to a more expensive and less reliable national grid across scenarios, however, the impact on regional RA can vary. The probabilistic analysis reveals that transmission investments can alleviate the tail risk of consumers, however, the availability of fuel resources does not always alleviate regional tail risks. Our findings inform policymakers and utilities on the prioritization of transmission investments to mitigate the risk of widespread outages, also for tail events, and ensure reliable and affordable electricity delivery to all.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Low Space Harmonic Content Windings (LSHWs) Applied to Improve the Pareto Front in Design Optimization of Electric Machines

In design optimization of high performance electric machines where the power density is increasingly a key metric, it is often desirable to maximize the average torque output within a given volume. However, if a smooth output is also required, the torque ripple may become a hurdle as it tends to increase with the average torque for conventional windings. A direct benefit of low space harmonic content windings (LSHWs), featuring the suppression or elimination of targeted non-working orders in the winding magnetomotive force (MMF), is the torque ripple reduction. This paper provides a detailed analysis of torque ripple pairs and the space-time harmonic order mapping, with concise results that applies to both distributed windings and fractionalslot concentrated windings (FSCWs) of any slot/pole combination. Metamodel-based optimizations are performed to compare two LSHWs with a conventional winding. Improved Pareto fronts are validated, which confirms higher torque producing capabilities at various torque ripple thresholds for LSHWs.

33 ADVANCED PROPULSION SYSTEMS↗

Quantifying uncertainty in Pareto estimates of global lake area

This software contains the code for Bayesian uncertainty analysis of global lake area. Computed uncertainties are compared against more conventional estimation using lake size-abundance distributions. Routines are included to explore sensitivity to observational errors and ad-hoc "censoring" strategies.

Stachelek, Jemma↗

Active learning enables generation of molecules that advance the known Pareto front

Although generative models hold promise for discovering molecules with optimized desired properties, they often fail to suggest synthesizable molecules that improve upon the properties of the structures represented in the training distribution. We find that this limitation arises not only from the molecule generation process itself, but also from the poor generalization capabilities of molecular property predictors. We address this challenge by creating a closed-loop molecule generation pipeline with iterative retraining on new quantum chemical simulation data. Compared against static, single-pass generative modeling approaches, only our closed-loop iterative workflow generates molecules with properties extending beyond the training distribution (up to 0.44 standard deviations beyond the original range) and achieves a 79% improvement in out-of-distribution molecule classification accuracy. Furthermore, by conditioning molecular generation on thermodynamic stability data obtained during the iterative loop, the proportion of stable and hence potentially synthesizable molecules generated is 3.5x higher than the next-best model.

Chemistry↗

Decomposition and Algorithmic Approaches for Solving Large-Scale Process Family Design Problems

Our most recent work expands the water desalination case study from 76 variants to 10,897 variants using the equation-oriented model built in Pyomo as part of the PARETO project. Using the discretization formulation presented in Stinchfield (2024a), rather than solving for all 10,897 variants simultaneously, we decompose the formulation into subproblems containing subsets of variants from the process family. We solve the overall problem with Progressive Hedging (PH) deployed in parallel on a distributed HPC cluster using the open-source Python package mpi-sppy (Knueven et al., 2023). This approach allowed us to solve this process family design problem to ~1.5% relative optimality gap in about 5 hours; in comparison, Gurobi reached ~50% relative optimality gap in about 6 hours (Stinchfield et al., 2024b). However, this approach still requires discretization of the common unit module design ranges; additionally, PH acts as a heuristic for MILP’s with gap-closing capabilities. Ideally, we would not have to use ML surrogates or discretization to solve this problem, instead solving the process family design problem with the equation-oriented model directly to achieve the most accurate results. However, recall that we did not consider solving the MINLP directly due to complexity and size. In this work, we aim to decompose and solve this large-scale MINLP using a Structured Nonlinear Global Optimization algorithm presented by Cao and Zavala (2019).

Stinchfield, Georgia↗

Quantifying uncertainty in Pareto estimates of global lake area

Abstract Size is a critical factor determining the rate and occurrence of specific lake processes such as carbon sequestration and greenhouse gas emissions and emerging evidence suggests that small lakes in particular have particularly large CO 2 flux rates. Because we do not have a complete census of all lakes, upscaling estimates of such processes to small lakes at broad spatial scales requires the use of lake size‐abundance distributions rather than empirical measurements of area. Existing lake census efforts are incomplete such that as lakes become smaller, they are more likely to be omitted either because they are too small to be resolved from remote sensing products or because of limited ground surveying effort (i.e., “censoring” of small lakes relative to large lakes). The present study explores one potential shortcoming of prior approaches estimating global lake area using lake size‐abundance distributions. Namely, that these prior approaches rely on frequentist curve fitting techniques combined with an ad‐hoc cutoff determination strategy (visual inspection to determine a likely censoring point). This yields an over‐exact lake area estimate that is typically reported with no uncertainty bounds. I show how these shortcomings can be addressed with a Bayesian model that produces larger estimates of lake area uncertainty relative to the typical approach. When used as part of a sensitivity analysis, such an approach has the potential to enable more robust intercomparisons among studies of aquatic processes upscaling.

54 ENVIRONMENTAL SCIENCES↗

Microgrid Design Toolkit (MDT) User Guide: Software v1.4

The MDT is a decision support software that can provide the information needed to identify optimal microgrid designs in the early stages of the design process. MDT searches the trade space of alternative microgrid designs in terms of user-defined objectives, such as performance, reliability, and cost. It produces a Pareto frontier of solutions embodying the efficient tradeoffs amongst multiple user-defined objectives.

24 POWER TRANSMISSION AND DISTRIBUTION↗

MaTableGPT: GPT‐Based Table Data Extractor from Materials Science Literature

Abstract Efficiently extracting data from tables in the scientific literature is pivotal for building large‐scale databases. However, the tables reported in materials science papers exist in highly diverse forms; thus, rule‐based extractions are an ineffective approach. To overcome this challenge, the study presents MaTableGPT, which is a GPT‐based table data extractor from the materials science literature. MaTableGPT features key strategies of table data representation and table splitting for better GPT comprehension and filtering hallucinated information through follow‐up questions. When applied to a vast volume of water splitting catalysis literature, MaTableGPT achieves an extraction accuracy (total F1 score) of up to 96.8%. Through comprehensive evaluations of the GPT usage cost, labeling cost, and extraction accuracy for the learning methods of zero‐shot, few‐shot, and fine‐tuning, the study presents a Pareto‐front mapping where the few‐shot learning method is found to be the most balanced solution owing to both its high extraction accuracy (total F1 score >95%) and low cost (GPT usage cost of 5.97 US dollars and labeling cost of 10 I/O paired examples). The statistical analyses conducted on the database generated by MaTableGPT revealed valuable insights into the distribution of the overpotential and elemental utilization across the reported catalysts in the water splitting literature.

Yi, Gyeong Hoon [Computational Science Research Ce↗

Optimizing fluvial flood mitigation strategies: A multi-objective approach for cost-effective and socially-aware infrastructure feasibility analysis

Effective levee planning must balance capital cost, risk reduction, and community priorities. These objectives are rarely optimized together. This study presents a feasibility phase, simulationin-the-loop framework that couples terrain-based flood modeling with a socially aware multiobjective optimizer. Flood risk is measured as Expected Annual Exposed Population (EAEP), obtained by integrating exposure over Annual Exceedance Probability (AEP) nodes, mirroring the Hydrologic Engineering Center's Flood Damage Reduction Analysis (HEC-FDA) expected-annual formulation but with people rather than dollars. Exposure per scenario is computed by overlaying binary inundation masks with a population surface at the tract level. Distributional fairness is encoded through a Group Benefit Share (GBS) constraint that requires high-SVI tracts to receive at least a baseline share of annualized benefits. Capital cost is represented by a height-dependent unit-cost model suitable for screening. This study addresses the two-objective problem, minimize cost and expected annual exposure subject to the GBS constraint, using Non-Dominated Sorting Genetic Algorithm II (NSGA-II) and leveraging Pareto front for feasibility phase decision making. Implemented with terrain-based flood modeling, GeoFlood, for rapid scenario evaluation, the framework is demonstrated in Southeast Texas. The results reveal clear trade-offs among cost, risk, and social benefits and identify non-dominated levee height configurations that satisfy the benefit-share floor. The contributions are a scalable decision support method that operationalizes expected annual population-based risk, embeds enforceable benefit-sharing guarantees, and uses lightweight simulation to explore large design spaces before higher fidelity design stages.

Flood mitigation↗

Selecting Critical Scenarios of DER Adoption in Distribution Grids Using Bayesian Optimization

We develop a new methodology to select scenarios of DER adoption most critical for distribution grids. Anticipating risks of future voltage and line flow violations due to additional PV adopters is central for utility investment planning but continues to rely on deterministic or ad hoc scenario selection. We propose a highly efficient search framework based on multi-objective Bayesian Optimization. We treat underlying grid stress metrics as computationally expensive black-box functions, approximated via Gaussian Process surrogates and design an acquisition function based on probability of scenarios being Pareto-critical across a collection of line- and bus-based violation objectives. Our approach provides a statistical guarantee and offers an order of magnitude speed-up relative to a conservative exhaustive search. Case studies on realistic feeders with 200-400 buses demonstrate the effectiveness and accuracy of our approach.

Mulkin, Olivier↗

Tula: Optimizing Time, Cost, and Generalization in Distributed Large-Batch Training

Distributed training increases the number of batches processed per iteration either by scaling-out (adding more nodes) or scaling-up (increasing the batch-size). However, the largest configuration does not necessarily yield the best performance. Horizontal scaling introduces additional communication overhead, while vertical scaling is constrained by computation cost and device memory limits. Thus, simply increasing the batch-size leads to diminishing returns: training time and cost decrease initially but eventually plateaus, creating a knee-point in the time/cost vs. batch-size pareto curve. The optimal batch-size therefore depends on the underlying model, data and available compute resources. Large batches also suffer from worse model quality due to the well-known “generalization gap”. In this paper, we present Tula, an online service that automatically optimizes time, cost, and convergence quality for large-batch training of convolutional models. It combines parallel-systems modeling with statistical performance prediction to identify the optimal batchsize. Tula predicts training time and cost within 7.5−14% error across multiple models, and achieves up to 20× overall speedup and improves test accuracy by ≈9% on average over standard large-batch training on various vision tasks, thus successfully mitigating the generalization gap and accelerating training at the same time.

Tyagi, Sahil [ORNL] (ORCID:0009000783144745)↗

Optimisation of the Kaplan hydropower system via PID 2 and digital twin

Here, this paper proposes a proportional–integral-double–derivative (PID 2 ) optimisation method for the Kaplan hydropower system by building a digital twin. The study first uses one multilayer perceptron (MLP) to model the hydroturbine dynamic and then adopts three connected MLPs to model the generator dynamic, both in an open-loop fashion. Inspired by stochastic distribution control (SDC) theory, we regard the training of the turbine's neural network model as a process control problem, and we propose minimising entropy loss to update the network parameters. The next step is to build the digital twin by connecting the neural network models with a PID 2 controller and a lead-lag exciter and run the whole model in a closed-loop fashion. After that, a binary search approach is applied to optimise the PID 2 parameters based on the obtained digital twin model. The simulation results show that the proposed method can reduce the mean square tracking error by more than 90%. Furthermore, the method is extended to jointly optimise the PID 2 controller and excitation system gains through multiobjective optimisation, leveraging Pareto frontier analysis to balance active power and voltage tracking performance. Simulation results confirm the effectiveness of the proposed method, achieving a 83.46% reduction in relative mean square error of active power, a 47.13% reduction in terminal voltage tracking error, and an 82.78% improvement in the overall scalarized objective.

Hydropower system↗