Search NASA⌕ Search

SEARCH · Search NASA

Results for “Pareto distribution”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Measuring the giant radio galaxy length distribution with the LoTSS

Many massive galaxies launch jets from the accretion disk of their central black hole, but only ~10 3 instances are known in which the associated outflows form giant radio galaxies (GRGs, or giants): luminous structures of megaparsec extent that consist of atomic nuclei, relativistic electrons, and magnetic fields. Large samples are imperative to understanding the enigmatic growth of giants, and recent systematic searches in homogeneous surveys constitute a promising development. For the first time, it is possible to perform meaningful precision statistics with GRG lengths, but a framework to do so is missing. We measured the intrinsic GRG length distribution by combining a novel statistical framework with a LOFAR Two-metre Sky Survey (LoTSS) sample of freshly discovered giants. In turn, this allowed us to answer an array of questions on giants. For example, we can now assess how rare a 5 Mpc giant is compared with one of 1 Mpc, and how much larger – given a projected length – the corresponding intrinsic length is expected to be. Notably, we can now also infer the GRG number density in the Local Universe. We assumed the intrinsic GRG length distribution to be Paretian (i.e. of power-law form) with tail index ξ, and predicted the observed distribution by modelling projection and selection effects. To infer ξ, we also systematically searched the LoTSS for hitherto unknown giants and compiled the largest catalogue of giants to date. We show that if intrinsic GRG lengths are Pareto distributed with index ξ, then projected GRG lengths are also Pareto distributed with index ξ. Selection effects induce curvature in the observed projected GRG length distribution: angular length selection flattens it towards the lower end, while surface brightness selection steepens it towards the higher end. We explicitly derived a GRG’s posterior over intrinsic lengths given its projected length, laying bare the ξ dependence. We also discovered 2060 giants within LoTSS DR2 pipeline products; our sample more than doubles the known population. Spectacular discoveries include the largest, second-largest, and fourth-largest GRG known (l p = 5.1 Mpc, l p = 5.0 Mpc, and l p = 4.8 Mpc), the largest GRG known hosted by a spiral galaxy (l p = 2.5 Mpc), and the largest secure GRG known beyond redshift 1 (l p = 3.9 Mpc). We increase the number of known giants whose angular length exceeds that of the Moon from 10 to 23; among the discoveries is the angularly largest known radio galaxy in the Northern Sky, which is also the angularly largest known GRG ($\phi$ = 2°). Combining theory and data, we determined that intrinsic GRG lengths are well described by a Pareto distribution, and measured the index ξ = –3.5 ± 0.5. This implies that, given its projected length, a GRG’s intrinsic length is expected to be just 15% larger. Finally, we determined the comoving number density of giants in the Local Universe to be n GRG = 5 ± 2(100 Mpc) –3 . We developed a practical mathematical framework that elucidates the statistics of giant radio galaxy lengths. Through a LoTSS search, we also discovered 2060 new giants. By combining both advances, we determined that intrinsic GRG lengths are well described by a Pareto distribution with index ξ = –3.5 ± 0.5, and that giants are truly rare in a cosmological sense: most clusters and filaments of the Cosmic Web are not currently home to a giant. Thus, our work yields new observational constraints for analytical models and simulations featuring radio galaxy growth.

79 ASTRONOMY AND ASTROPHYSICS↗

Tuning Monotonic Basin Hopping: Improving the Efficiency of Stochastic Search as Applied to Low-Thrust Trajectory Optimization

Trajectory optimization methods using MBH have become well developed during the past decade. An essential component of MBH is a controlled random search through the multi-dimensional space of possible solutions. Historically, the randomness has been generated by drawing RVs from a uniform probability distribution. Here, we investigate the generating the randomness by drawing the RVs from Cauchy and Pareto distributions, chosen because of their characteristic long tails. We demonstrate that using Cauchy distributions (as first suggested by Englander significantly improves MBH performance, and that Pareto distributions provide even greater improvements. Improved performance is defined in terms of efficiency and robustness, where efficiency is finding better solutions in less time, and robustness is efficiency that is undiminished by (a) the boundary conditions and internal constraints of the optimization problem being solved, and (b) by variations in the parameters of the probability distribution. Robustness is important for achieving performance improvements that are not problem specific. In this work we show that the performance improvements are the result of how these long-tailed distributions enable MBH to search the solution space faster and more thoroughly. In developing this explanation, we use the concepts of sub-diffusive, normally-diffusive, and super-diffusive RWs originally developed in the field of statistical physics.

mission design↗

Tuning Monotonic Basin Hopping: Improving the Efficiency of Stochastic Search as Applied to Low-Thrust Trajectory Optimization

Trajectory optimization methods using monotonic basin hopping (MBH) have become well developed during the past decade [1, 2, 3, 4, 5, 6]. An essential component of MBH is a controlled random search through the multi-dimensional space of possible solutions. Historically, the randomness has been generated by drawing random variable (RV)s from a uniform probability distribution. Here, we investigate the generating the randomness by drawing the RVs from Cauchy and Pareto distributions, chosen because of their characteristic long tails. We demonstrate that using Cauchy distributions (as first suggested by J. Englander [3, 6]) significantly improves monotonic basin hopping (MBH) performance, and that Pareto distributions provide even greater improvements. Improved performance is defined in terms of efficiency and robustness. Efficiency is finding better solutions in less time. Robustness is efficiency that is undiminished by (a) the boundary conditions and internal constraints of the optimization problem being solved, and (b) by variations in the parameters of the probability distribution. Robustness is important for achieving performance improvements that are not problem specific. In this work we show that the performance improvements are the result of how these long-tailed distributions enable MBH to search the solution space faster and more thoroughly. In developing this explanation, we use the concepts of sub-diffusive, normally-diffusive, and super-diffusive random walks (RWs) originally developed in the field of statistical physics.

autonomous↗

Detecting Permafrost Active Layer Thickness Change From Nonlinear Baseflow Recession

We report permafrost underlies about one fifth of the global land area and affects ground stability, freshwater runoff, soil chemistry, and surface-atmosphere gas exchange. The depth of thawed ground overlying permafrost (active layer thickness) has broadly increased across the Arctic in recent decades, coincident with a period of increased streamflow, especially the lowest flows (baseflow). Mechanistic links between active layer thickness and baseflow have recently been explored using linear reservoir theory, but most watersheds behave as nonlinear reservoirs. We derive theoretical nonlinear relationships between long-term average saturated soil thickness $\bar{η}$(proxy for active layer thickness) and long-term average baseflow. When applied to 38 years of daily streamflow data for the Kuparuk River basin on the North Slope of Alaska, the theory predicts $\bar{η}$ increased 0.17 ± 0.22[2σ] cm a -1 between 1983 and 2020 (6.4 ± cm total). The rate of increase nearly doubled to 0.29 ± 0.31 cm a -1 between 1990 and 2020, during which time local field measurements from Circumpolar Active Layer Monitoring sites indicate the active layer increased 0.31 ± 0.22 cm a -1 . The predicted rate of increase more than doubled again between 2002 and 2020, outpacing a near doubling of observed active layer thickening, consistent with trends in terrestrial water storage inferred from Gravity Recovery and Climate Experiment satellite gravimetry and Modern-Era Retrospective Analysis for Research and Applications climate reanalysis. Overall, hydrologic change is accelerating in the Kuparuk River basin, and we provide a theoretical framework for estimating basin-scale changes in active layer water storage from streamflow measurements.

54 ENVIRONMENTAL SCIENCES↗

Hopping with an Adaptive Hop Probability Distribution

Monotonic Basin Hopping (MBH) is a stochastic global search technique that may be used to design complex interplanetary trajectories. Previous research on MBH empirically demonstrated that a bi-polar Pareto distribution is an effective way to generate search “hops.” However, no analytical foundation exists to explain this performance. In this work we provide the beginning of that analytical foundation and also introduce a new variant of MBH that uses an adaptive probability distribution to generate the hops. The technique is demonstrated on historical interplanetary trajectory design problems.

A C Englander↗

Evaluation of NASA's MERRA Precipitation Product in Reproducing the Observed Trend and Distribution of Extreme Precipitation Events in the United States

This study evaluates the performance of NASA's Modern-Era Retrospective Analysis for Research and Applications (MERRA) precipitation product in reproducing the trend and distribution of extreme precipitation events. Utilizing the extreme value theory, time-invariant and time-variant extreme value distributions are developed to model the trends and changes in the patterns of extreme precipitation events over the contiguous United States during 1979-2010. The Climate Prediction Center (CPC) U.S.Unified gridded observation data are used as the observational dataset. The CPC analysis shows that the eastern and western parts of the United States are experiencing positive and negative trends in annual maxima, respectively. The continental-scale patterns of change found in MERRA seem to reasonably mirror the observed patterns of change found in CPC. This is not previously expected, given the difficulty in constraining precipitation in reanalysis products. MERRA tends to overestimate the frequency at which the 99th percentile of precipitation is exceeded because this threshold tends to be lower in MERRA, making it easier to be exceeded. This feature is dominant during the summer months. MERRA tends to reproduce spatial patterns of the scale and location parameters of the generalized extreme value and generalized Pareto distributions. However, MERRA underestimates these parameters, particularly over the Gulf Coast states, leading to lower magnitudes in extreme precipitation events. Two issues in MERRA are identified: 1) MERRA shows a spurious negative trend in Nebraska and Kansas, which is most likely related to the changes in the satellite observing system over time that has apparently affected the water cycle in the central United States, and 2) the patterns of positive trend over the Gulf Coast states and along the East Coast seem to be correlated with the tropical cyclones in these regions. The analysis of the trends in the seasonal precipitation extremes indicates that the hurricane and winter seasons are contributing the most to these trend patterns in the southeastern United States. In addition, the increasing annual trend simulated by MERRA in the Gulf Coast region is due to an incorrect trend in winter precipitation extremes.

MERRA↗

baseflow: a MATLAB and GNU Octave package for baseflow recession analysis

baseflow is a MATLAB® toolbox designed for baseflow recession analysis, a technique used in hydrologic science to infer aquifer properties from streamflow. By leveraging widely available streamflow data, baseflow can be used to estimate aquifer properties such as hydraulic conductivity and drainable porosity over the modern instrumental stream gage record. The toolbox is intended for analysis of measured streamflow values recorded on a daily timestep, and is tailored for shallow, unconfined riparian aquifers that discharge groundwater laterally into adjacent streams. Additionally, baseflow can analyze the collective behavior of individual hillslope aquifers constituting hydrologic catchments, known as “watersheds”, from a nonlinear dynamical systems perspective. The toolbox incorporates recent advances in baseflow recession analysis to enable objective estimations of aquifer properties, and their sensitivity to methodological decisions, at both hillslope and catchment scales.

97 MATHEMATICS AND COMPUTING↗

Security Vulnerability Profiles of NASA Mission Software: Empirical Analysis of Security Related Bug Reports

NASA develops, runs, and maintains software systems for which security is of vital importance. Therefore, it is becoming an imperative to develop secure systems and extend the current software assurance capabilities to cover information assurance and cybersecurity concerns of NASA missions. The results presented in this report are based on the information provided in the issue tracking systems of one ground mission and one flight mission. The extracted data were used to create three datasets: Ground mission IVV issues, Flight mission IVV issues, and Flight mission Developers issues. In each dataset, we identified the software bugs that are security related and classified them in specific security classes. This information was then used to create the security vulnerability profiles (i.e., to determine how, why, where, and when the security vulnerabilities were introduced) and explore the existence of common trends. The main findings of our work include:- Code related security issues dominated both the Ground and Flight mission IVV security issues, with 95 and 92, respectively. Therefore, enforcing secure coding practices and verification and validation focused on coding errors would be cost effective ways to improve mission's security. (Flight mission Developers issues dataset did not contain data in the Issue Category.)- In both the Ground and Flight mission IVV issues datasets, the majority of security issues (i.e., 91 and 85, respectively) were introduced in the Implementation phase. In most cases, the phase in which the issues were found was the same as the phase in which they were introduced. The most security related issues of the Flight mission Developers issues dataset were found during Code Implementation, Build Integration, and Build Verification; the data on the phase in which these issues were introduced were not available for this dataset.- The location of security related issues, as the location of software issues in general, followed the Pareto principle. Specifically, for all three datasets, from 86 to 88 the security related issues were located in two to four subsystems.- The severity levels of most security issues were moderate, in all three datasets.- Out of 21 primary security classes, five dominated: Exception Management, Memory Access, Other, Risky Values, and Unused Entities. Together, these classes contributed from around 80 to 90 of all security issues in each dataset. This again proves the Pareto principle of uneven distribution of security issues, in this case across CWE classes, and supports the fact that addressing these dominant security classes provides the most cost efficient way to improve missions' security. The findings presented in this report uncovered the security vulnerability profiles and identified the common trends and dominant classes of security issues, which in turn can be used to select the most efficient secure design and coding best practices compiled by the part of the SARP project team associated with the NASA's Johnson Space Center. In addition, these findings provide valuable input to the NASA IVV initiative aimed at identification of the two 25 CWEs of ground and flight missions.

vulnerability↗

Seeing through noise in power laws

Despite widespread claims of power laws across the natural and social sciences, evidence in data is often equivocal. Modern data and statistical methods reject even classic power laws such as Pareto’s law of wealth and the Gutenberg–Richter law for earthquake magnitudes. We show that the maximum-likelihood estimators and Kolmogorov–Smirnov (K-S) statistics in widespread use are unexpectedly sensitive to ubiquitous errors in data such as measurement noise, quantization noise, heaping and censorship of small values. This sensitivity causes spurious rejection of power laws and biases parameter estimates even in arbitrarily large samples, which explains inconsistencies between theory and data. We show that logarithmic binning by powers of λ > 1 attenuates these errors in a manner analogous to noise averaging in normal statistics and that λ thereby tunes a trade-off between accuracy and precision in estimation. Binning also removes potentially misleading within-scale information while preserving information about the shape of a distribution over powers of λ, and we show that some amount of binning can improve sensitivity and specificity of K-S tests without any cost, while more extreme binning tunes a trade-off between sensitivity and specificity. We therefore advocate logarithmic binning as a simple essential step in power-law inference.

97 MATHEMATICS AND COMPUTING↗

Seeing through noise in power laws

Despite widespread claims of power laws across the natural and social sciences, evidence in data is often equivocal. Modern data and statistical methods reject even classic power laws such as Pareto’s law of wealth and the Gutenberg–Richter law for earthquake magnitudes. We show that the maximum-likelihood estimators and Kolmogorov–Smirnov (K-S) statistics in widespread use are unexpectedly sensitive to ubiquitous errors in data such as measurement noise, quantization noise, heaping and censorship of small values. This sensitivity causes spurious rejection of power laws and biases parameter estimates even in arbitrarily large samples, which explains inconsistencies between theory and data. We show that logarithmic binning by powers of λ > 1 attenuates these errors in a manner analogous to noise averaging in normal statistics and that λ thereby tunes a trade-off between accuracy and precision in estimation. Binning also removes potentially misleading within-scale information while preserving information about the shape of a distribution over powers of λ, and we show that some amount of binning can improve sensitivity and specificity of K-S tests without any cost, while more extreme binning tunes a trade-off between sensitivity and specificity. We therefore advocate logarithmic binning as a simple essential step in power-law inference.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Evolutionary Computing for Low-thrust Navigation

The development of new mission concepts requires efficient methodologies to analyze, design and simulate the concepts before implementation. New mission concepts are increasingly considering the use of ion thrusters for fuel-efficient navigation in deep space. This paper presents parallel, evolutionary computing methods to design trajectories of spacecraft propelled by ion thrusters and to assess the trade-off between delivered payload mass and required flight time. The developed methods utilize a distributed computing environment in order to speed up computation, and use evolutionary algorithms to find globally Pareto-optimal solutions. The methods are coupled with two main traditional trajectory design approaches, which are called direct and indirect. In the direct approach, thrust control is discretized in either arc time or arc length, and the resulting discrete thrust vectors are optimized. In the indirect approach, a thrust control problem is transformed into a costate control problem, and the initial values of the costate vector are optimized. The developed methods are applied to two problems: 1) an orbit transfer around the Earth and 2) a transfer between two distance retrograde orbits around Europa, the closest to Jupiter of the icy Galilean moons. The optimal solutions found with the present methods are comparable to other state-of-the-art trajectory optimizers and to analytical approximations for optimal transfers, while the required computational time is several orders of magnitude shorter than other optimizers thanks to an intelligent design of control vector discretization, advanced algorithmic parameterization, and parallel computing.

optimization↗

A Simulation Based Approach to Optimize Berth Throughput Under Uncertainty at Marine Container Terminals

Berth scheduling is a critical function at marine container terminals and determining the best berth schedule depends on several factors including the type and function of the port, size of the port, location, nearby competition, and type of contractual agreement between the terminal and the carriers. In this paper we formulate the berth scheduling problem as a bi-objective mixed-integer problem with the objective to maximize customer satisfaction and reliability of the berth schedule under the assumption that vessel handling times are stochastic parameters following a discrete and known probability distribution. A combination of an exact algorithm, a Genetic Algorithms based heuristic and a simulation post-Pareto analysis is proposed as the solution approach to the resulting problem. Based on a number of experiments it is concluded that the proposed berth scheduling policy outperforms the berth scheduling policy where reliability is not considered.

Golias, Mihalis M.↗

Long-Term Impacts of Constrained Transmission Deployment on the Cost-Reliability Tradeoff

Traditional Resource Adequacy (RA) frameworks in the U.S. undervalue the contributions of inter-regional transmission to resource adequacy during stress periods, focusing on the availability of nameplate capacity instead. However, availability of nameplate capacity does not always translate into electricity delivery, especially during tail events. Moreover, the rapid deployment of energy-limited resources and increasing electricity demand challenge existing resource adequacy frameworks and couple regional electricity demand and availability of supply via transmission. We propose a two-stage framework that goes beyond the existing capacity-centered approaches to reveal the RA contributions of transmission. In the first stage we introduce a multi-objective optimization framework to quantify the merits of transmission expansion via Pareto Frontiers under alternative futures of no transmission investment, primary energy resources availability and demand growth. The second stage focuses on tail events and leverages the results of the first stage to characterize the risk profile of regional consumers across the U.S. under the alternative energy futures. We find that no new transmission can lead to a more expensive and less reliable national grid across scenarios, however, the impact on regional RA can vary. The probabilistic analysis reveals that transmission investments can alleviate the tail risk of consumers, however, the availability of fuel resources does not always alleviate regional tail risks. Our findings inform policymakers and utilities on the prioritization of transmission investments to mitigate the risk of widespread outages, also for tail events, and ensure reliable and affordable electricity delivery to all.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Quantifying uncertainty in Pareto estimates of global lake area

This software contains the code for Bayesian uncertainty analysis of global lake area. Computed uncertainties are compared against more conventional estimation using lake size-abundance distributions. Routines are included to explore sensitivity to observational errors and ad-hoc "censoring" strategies.

Stachelek, Jemma↗

Active learning enables generation of molecules that advance the known Pareto front

Although generative models hold promise for discovering molecules with optimized desired properties, they often fail to suggest synthesizable molecules that improve upon the properties of the structures represented in the training distribution. We find that this limitation arises not only from the molecule generation process itself, but also from the poor generalization capabilities of molecular property predictors. We address this challenge by creating a closed-loop molecule generation pipeline with iterative retraining on new quantum chemical simulation data. Compared against static, single-pass generative modeling approaches, only our closed-loop iterative workflow generates molecules with properties extending beyond the training distribution (up to 0.44 standard deviations beyond the original range) and achieves a 79% improvement in out-of-distribution molecule classification accuracy. Furthermore, by conditioning molecular generation on thermodynamic stability data obtained during the iterative loop, the proportion of stable and hence potentially synthesizable molecules generated is 3.5x higher than the next-best model.

Chemistry↗

Decomposition and Algorithmic Approaches for Solving Large-Scale Process Family Design Problems

Our most recent work expands the water desalination case study from 76 variants to 10,897 variants using the equation-oriented model built in Pyomo as part of the PARETO project. Using the discretization formulation presented in Stinchfield (2024a), rather than solving for all 10,897 variants simultaneously, we decompose the formulation into subproblems containing subsets of variants from the process family. We solve the overall problem with Progressive Hedging (PH) deployed in parallel on a distributed HPC cluster using the open-source Python package mpi-sppy (Knueven et al., 2023). This approach allowed us to solve this process family design problem to ~1.5% relative optimality gap in about 5 hours; in comparison, Gurobi reached ~50% relative optimality gap in about 6 hours (Stinchfield et al., 2024b). However, this approach still requires discretization of the common unit module design ranges; additionally, PH acts as a heuristic for MILP’s with gap-closing capabilities. Ideally, we would not have to use ML surrogates or discretization to solve this problem, instead solving the process family design problem with the equation-oriented model directly to achieve the most accurate results. However, recall that we did not consider solving the MINLP directly due to complexity and size. In this work, we aim to decompose and solve this large-scale MINLP using a Structured Nonlinear Global Optimization algorithm presented by Cao and Zavala (2019).

Stinchfield, Georgia↗

Reducing the Volume of NASA Earth-Science Data

A computer program reduces data generated by NASA Earth-science missions into representative clusters characterized by centroids and membership information, thereby reducing the large volume of data to a level more amenable to analysis. The program effects an autonomous data-reduction/clustering process to produce a representative distribution and joint relationships of the data, without assuming a specific type of distribution and relationship and without resorting to domain-specific knowledge about the data. The program implements a combination of a data-reduction algorithm known as the entropy-constrained vector quantization (ECVQ) and an optimization algorithm known as the differential evolution (DE). The combination of algorithms generates the Pareto front of clustering solutions that presents the compromise between the quality of the reduced data and the degree of reduction. Similar prior data-reduction computer programs utilize only a clustering algorithm, the parameters of which are tuned manually by users. In the present program, autonomous optimization of the parameters by means of the DE supplants the manual tuning of the parameters. Thus, the program determines the best set of clustering solutions without human intervention.

Lee, Seungwon↗

Quantifying uncertainty in Pareto estimates of global lake area

Abstract Size is a critical factor determining the rate and occurrence of specific lake processes such as carbon sequestration and greenhouse gas emissions and emerging evidence suggests that small lakes in particular have particularly large CO 2 flux rates. Because we do not have a complete census of all lakes, upscaling estimates of such processes to small lakes at broad spatial scales requires the use of lake size‐abundance distributions rather than empirical measurements of area. Existing lake census efforts are incomplete such that as lakes become smaller, they are more likely to be omitted either because they are too small to be resolved from remote sensing products or because of limited ground surveying effort (i.e., “censoring” of small lakes relative to large lakes). The present study explores one potential shortcoming of prior approaches estimating global lake area using lake size‐abundance distributions. Namely, that these prior approaches rely on frequentist curve fitting techniques combined with an ad‐hoc cutoff determination strategy (visual inspection to determine a likely censoring point). This yields an over‐exact lake area estimate that is typically reported with no uncertainty bounds. I show how these shortcomings can be addressed with a Bayesian model that produces larger estimates of lake area uncertainty relative to the typical approach. When used as part of a sensitivity analysis, such an approach has the potential to enable more robust intercomparisons among studies of aquatic processes upscaling.

54 ENVIRONMENTAL SCIENCES↗