Search NASA⌕ Search

SEARCH · Search NASA

Results for “hierarchical sampling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Iterative self-organizing SCEne-LEvel sampling (ISOSCELES) for large-scale building extraction

Convolutional neural networks (CNN) provide state-of-the-art performance in many computer vision tasks, including those related to remote-sensing image analysis. Successfully training a CNN to generalize well to unseen data, however, requires training on samples that represent the full distribution of variation of both the target classes and their surrounding contexts. With remote sensing data, acquiring a sufficiently representative training set is a challenge due to both the inherent multi-modal variability of satellite or aerial imagery and the general high cost of labeling data. To address this challenge, we have developed ISOSCELES, an Iterative Self-Organizing SCEne LEvel Sampling method for hierarchical sampling of large image sets. Using affinity propagation, ISOSCELES automates the selection of highly representative training images. Compared to random sampling or using available reference data, the distribution of the training is principally data driven, reducing the chance of oversampling uninformative areas or undersampling informative ones. In comparison to manual sample selection by an analyst, ISOSCELES exploits descriptive features, spectral and/or textural, and eliminates human bias in sample selection. Using a hierarchical sampling approach, ISOSCELES can obtain a training set that reflects both between-scene variability, such as in viewing angle and time of day, and within-scene variability at the level of individual training samples. We verify the method by demonstrating its superiority to stratified random sampling in the challenging task of adapting a pre-trained model to a new image and spatial domain for country-scale building extraction. Using a pair of hand-labeled training sets comprising 1,987 sample image chips, a total of 496,000,000 individually labeled pixels, we show, across three distinct model architectures, an increase in accuracy, as measured by F1-score, of 2.2–4.2%.

42 ENGINEERING↗

Enabling machine learning-ready HPC ensembles with Merlin

With the growing complexity of computational and experimental facilities, many scientific researchers are turning to machine learning (ML) techniques to analyze large scale ensemble data. With complexities such as multi-component workflows, heterogeneous machine architectures, parallel file systems, and batch scheduling, care must be taken to facilitate this analysis in a high performance computing (HPC) environment. Here, we present Merlin, a workflow framework to enable large ML-friendly ensembles of scientific HPC simulations. By augmenting traditional HPC with distributed compute technologies, Merlin aims to lower the barrier for scientific subject matter experts to incorporate ML into their analysis. As a producer–consumer workflow model, Merlin enables multi-machine, cross-batch job, dynamically allocated yet persistent workflows capable of utilizing surge-compute resources. Key features of Merlin are a flexible HPC-centric interface, low per-task overhead, multi-tiered fault recovery, and a hierarchical sampling algorithm that allows for $\mathscr{O}$(N) task execution and $\mathscr{O}$(N ln N) task queuing to ensembles of millions of tasks. In addition to Merlin’s design, we test the algorithm’s performance in an HPC center and demonstrate the ability to enqueue 40 million simulations in 100 s, with a 30 millisecond per-task overhead that is independent of ensemble size. Finally, we describe some example applications that Merlin has enabled on leadership-class HPC resources, such as the ML-augmented optimization of nuclear fusion experiments and the calibration of infectious disease models to study the progression of and possible mitigation strategies for COVID-19.

97 MATHEMATICS AND COMPUTING↗

Micro on a macroscale: relating microbial-scale soil processes to global ecosystem function

ABSTRACT Soil microorganisms play a key role in driving major biogeochemical cycles and in global responses to climate change. However, understanding and predicting the behavior and function of these microorganisms remains a grand challenge for soil ecology due in part to the microscale complexity of soils. It is becoming increasingly clear that understanding the microbial perspective is vital to accurately predicting global processes. Here, we discuss the microbial perspective including the microbial habitat as it relates to measurement and modeling of ecosystem processes. We argue that clearly defining and quantifying the size, distribution and sphere of influence of microhabitats is crucial to managing microbial activity at the ecosystem scale. This can be achieved using controlled and hierarchical sampling designs. Model microbial systems can provide key data needed to integrate microhabitats into ecosystem models, while adapting soil sampling schemes and statistical methods can allow us to collect microbially-focused data. Quantifying soil processes, like biogeochemical cycles, from a microbial perspective will allow us to more accurately predict soil functions and address long-standing unknowns in soil ecology.

59 BASIC BIOLOGICAL SCIENCES↗

Hierarchical Gaussian Random Field Sampling for Multilevel Markov Chain Monte Carlo: Coupling Stochastic Partial Differential Equation and the Karhunen–Loève Decomposition

This work introduces structure preserving hierarchical decompositions for sampling Gaussian random fields (GRFs) within the context of multilevel Bayesian inference in high-dimensional space. Existing scalable hierarchical sampling methods, such as those based on stochastic partial differential equations (SPDEs), often reduce the dimensionality of the sample space at the cost of accuracy of inference. Other approaches, such that those based on Karhunen-Loève (KL) expansions, offer sample space dimensionality reduction but sacrifice GRF representation accuracy and ergodicity of the Markov chain Monte Carlo (MCMC) sampler and are computationally expensive for high-dimensional problems. The proposed method integrates the dimensionality reduction capabilities of KL expansions with the scalability of SPDE-based sampling, thereby providing a robust, unified framework for high-dimensional uncertainty quantification (UQ) that is scalable and accurate, preserves ergodicity, and offers dimensionality reduction of the sample space. The hierarchy in our multilevel algorithm is derived from the geometric multigrid hierarchy. By constructing a hierarchical decomposition that maintains the covariance structure across the levels in the hierarchy, the approach enables efficient coarse-to-fine sampling while ensuring that all samples are drawn from the desired distribution. The effectiveness of the proposed method is demonstrated on a benchmark subsurface flow problem, demonstrating its effectiveness in improving computational efficiency and statistical accuracy. Furthermore, our proposed technique is more efficient and accurate and displays better convergence properties than existing methods for high-dimensional Bayesian inference problems.

Gaussian random fields↗

Stochastic nonlinear analysis of unidirectional fiber composites using image-based microstructural uncertainty quantification

Here, we present a data-driven nonlinear uncertainty quantification and propagation framework to study the microstructure-induced stochastic performance of unidirectional (UD) carbon fiber reinforced polymer (CFRP) composites. The proposed approach integrates (1) microscopic image characterization, (2) stochastic microstructure reconstruction, and (3) efficient multiscale finite element simulations enabled by self-consistent clustering (SCA) analysis. To model the complex microstructural variability, the proposed UQ methods take the non-Gaussian uncertainty sources into account through a distribution-free sampling approach leveraging nonparametric and asymptotic statistical tools. A hierarchical conditional sampling strategy enables the simultaneous sampling of multiple sources of uncertainties. Our approach provides insights into the impact of microstructural variabilities, which are shown to have an increasing impact on the nonlinear responses of UD CFRP parts under progressive compression loading and ultimately on the failure rate over time. We discover that before CFRP parts start to fail, a characteristic time period emerges with distinctive uncertainty distributions specific to the microstructure variability and the probability of failure. Identifying the failure time period is crucial to the reliability prediction, which is an essential component of CFRP design.

36 MATERIALS SCIENCE↗

hiposa v0.1.0

Fast construction of hierarchical N-dimensional experimental design. The software allows construction of sampling strategies that are self-avoiding, tunable in resolution in N dimensions. This code is a drop-in replacement for seeding optimization runs often performed using uniform random sampling. The hierarchical approach allows for a rapid, top down method for construction the most effective design that drives towards a global optimum.

Zwart, PetrusH [Lawrence Berkeley National Laborat↗

Using low volume eDNA methods to sample pelagic marine animal assemblages

Environmental DNA (eDNA) is an increasingly useful method for detecting pelagic animals in the ocean but typically requires large water volumes to sample diverse assemblages. Ship-based pelagic sampling programs that could implement eDNA methods generally have restrictive water budgets. Studies that quantify how eDNA methods perform on low water volumes in the ocean are limited, especially in deep-sea habitats with low animal biomass and poorly described species assemblages. Using 12S rRNA and COI gene primers, we quantified assemblages comprised of micronekton, coastal forage fishes, and zooplankton from low volume eDNA seawater samples (n = 436, 380–1800 mL) collected at depths of 0–2200 m in the southern California Current. We compared diversity in eDNA samples to concurrently collected pelagic trawl samples (n = 27), detecting a higher diversity of vertebrate and invertebrate groups in the eDNA samples. Differences in assemblage composition could be explained by variability in size-selectivity among methods and DNA primer suitability across taxonomic groups. The number of reads and amplicon sequences variants (ASVs) did not vary substantially among shallow (<200 m) and deep samples (>600 m), but the proportion of invertebrate ASVs that could be assigned a species-level identification decreased with sampling depth. Using hierarchical clustering, we resolved horizontal and vertical variability in marine animal assemblages from samples characterized by a relatively low diversity of ecologically important species. Low volume eDNA samples will quantify greater taxonomic diversity as reference libraries, especially for deep-dwelling invertebrate species, continue to expand.

59 BASIC BIOLOGICAL SCIENCES↗

Hierarchical Inference of the Lensing Convergence from Photometric Catalogs with Bayesian Graph Neural Networks

Abstract We present a Bayesian graph neural network (BGNN) that can estimate the weak lensing convergence ( κ ) from photometric measurements of galaxies along a given line of sight (LOS). The method is of particular interest in strong gravitational time-delay cosmography (TDC), where characterizing the “external convergence” ( κ ext ) from the lens environment and LOS is necessary for precise Hubble constant ( H 0 ) inference. Starting from a large-scale simulation with a κ resolution of ∼1′, we introduce fluctuations on galaxy–galaxy lensing scales of ∼1″ and extract random sight lines to train our BGNN. We then evaluate the model on test sets with varying degrees of overlap with the training distribution. For each test set of 1000 sight lines, the BGNN infers the individual κ posteriors, which we combine in a hierarchical Bayesian model to yield constraints on the hyperparameters governing the population. For a test field well sampled by the training set, the BGNN recovers the population mean of κ precisely and without bias (within the 2 σ credible interval), resulting in a contribution to the H 0 error budget well under 1%. In the tails of the training set with sparse samples, the BGNN, which can ingest all available information about each sight line, extracts a stronger κ signal compared to a simplified version of the traditional method based on matching galaxy number counts, which is limited by sample variance. Our hierarchical inference pipeline using BGNNs promises to improve the κ ext characterization for precision TDC. The code is available as a public Python package, Node to Joy ⏬ .

79 ASTRONOMY AND ASTROPHYSICS↗

Accelerating multilevel Markov Chain Monte Carlo using machine learning models

Here, this work presents an efficient approach for accelerating multilevel Markov Chain Monte Carlo (MCMC) sampling for large-scale problems using low-fidelity machine learning models. While conventional techniques for large-scale Bayesian inference often substitute computationally expensive high-fidelity models with machine learning models, thereby introducing approximation errors, our approach offers a computationally efficient alternative by augmenting high-fidelity models with low-fidelity ones within a hierarchical framework. The multilevel approach utilizes the low-fidelity machine learning model (MLM) for inexpensive evaluation of proposed samples thereby improving the acceptance of samples by the high-fidelity model. The hierarchy in our multilevel algorithm is derived from geometric multigrid hierarchy. We utilize an MLM to accelerate the coarse level sampling. Training machine learning model for the coarsest level significantly reduces the computational cost associated with generating training data and training the model. We present an MCMC algorithm to accelerate the coarsest level sampling using MLM and account for the approximation error introduced. We provide theoretical proofs of detailed balance and demonstrate that our multilevel approach constitutes a consistent MCMC algorithm. Additionally, we derive the expression for cost reduction due to machine learning model to facilitate cost analysis of the hierarchical sampling algorithm. Our technique is demonstrated on a standard benchmark inference problem in groundwater flow, where we estimate the probability density of a quantity of interest using a four-level MCMC algorithm. Our proposed algorithm accelerates multilevel sampling by a factor of two while achieving similar accuracy compared to sampling using the standard multilevel algorithm.

97 MATHEMATICS AND COMPUTING↗

3D printed Ni-Mn-Fe Bi-Functional Catalyst for Metal Air Batteries

Oxygen evolution catalyst with porous structures on the submillimeter and millimeter ranges have been shown to have high catalytic activity. Often, a limiting factor in the performance of a catalyst can be its specific surface area as this can limit the availability of active sites which encourage the desired reaction. As such, the inclusion of a hierarchical structure that include both meso-scale and nano-scale porosity can offer a means to greatly increase the performance of electrochemical catalysts. We have developed a method for manufacturing samples with such a hierarchical structure by combining techniques of direct ink writing (DIW), thermal sintering and free corrosion based dealloying. This allows us to prepare three-dimensionally structured samples with open-cell porosity at the macro-, micro- and nano-scales. Each level of porosity offers a different advantage to the overall functionality and applicability of this technology.

25 ENERGY STORAGE↗

A compendium of bacterial and archaeal single-cell amplified genomes from oxygen deficient marine waters

Oxygen-deficient marine waters referred to as oxygen minimum zones (OMZs) or anoxic marine zones (AMZs) are common oceanographic features. They host both cosmopolitan and endemic microorganisms adapted to low oxygen conditions. Microbial metabolic interactions within OMZs and AMZs drive coupled biogeochemical cycles resulting in nitrogen loss and climate active trace gas production and consumption. Global warming is causing oxygen-deficient waters to expand and intensify. Therefore, studies focused on microbial communities inhabiting oxygen-deficient regions are necessary to both monitor and model the impacts of climate change on marine ecosystem functions and services. Here we present a compendium of 5,129 single-cell amplified genomes (SAGs) from marine environments encompassing representative OMZ and AMZ geochemical profiles. Of these, 3,570 SAGs have been sequenced to different levels of completion, providing a strain-resolved perspective on the genomic content and potential metabolic interactions within OMZ and AMZ microbiomes. Hierarchical clustering confirmed that samples from similar oxygen concentrations and geographic regions also had analogous taxonomic compositions, providing a coherent framework for comparative community analysis.

59 BASIC BIOLOGICAL SCIENCES↗

Coalescence-Induced Spontaneous Shedding of Microdroplets on Superhydrophobic Surfaces Featuring Enclosed Micropillars with Hierarchical Roughness

This study investigated water vapor condensation on superhydrophobic surfaces (SHSs) featuring micropillars enclosed by wall lattices and having three-tier hierarchical roughness. A total of five samples were created with three (NW-J, W200-J and W400-J samples) having large micropillar depth (~6 μm) and two (NW-S and W200-S samples) having small micropillar depth (~1 μm). Two distinct condensate removal modes were observed during condensation: coalescence-induced jumping on samples with large micropillar depth and coalescence-induced shedding on samples with small micropillar depth. The results showed that the diameter of the shedding droplet on the W200-S sample having small micropillar depth could be as small as 107 μm, as compared to the theoretical critical diameter of 267 μm for gravitational shedding on the same sample. The enhanced functionality of the three-tier nanotextures on the W200-S sample could effectively suppress localized pinning of the three-phase contact line and Wenzel neck formation during the growth of condensate droplets. Consequently, during multidroplet coalescence, the released surface energy easily overcomes the solid–liquid adhesion, leading to spontaneous shedding of merged droplets. The inclusion of the wall lattice aids condensate growth by the droplet self-alignment along the walls and promoting coalescence. As a result, the W200-S sample exhibited the highest condensate collection as well. In conclusion, the proposed surface design has great potential for scaling up and implementation in heating, ventilation, and air-conditioning equipment due to the simplicity of the surface morphology and the facile spray-coating method used to achieve hierarchical roughness.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Evaluating the Benefits of Bayesian Hierarchical Methods for Analyzing Heterogeneous Environmental Datasets: A Case Study of Marine Organic Carbon Fluxes

Large compilations of heterogeneous environmental observations are increasingly available as public databases, allowing researchers to test hypotheses across datasets. Statistical complexities arise when analyzing compiled data due to unbalanced spatial sampling, variable environmental context, mixed measurement techniques, and other reasons. Hierarchical Bayesian modeling is increasingly used in environmental science to describe these complexities, however few studies explicitly compare the utility of hierarchical Bayesian models to simpler and more commonly applied methods. Here we demonstrate the utility of the hierarchical Bayesian approach with application to a large compiled environmental dataset consisting of 5,741 marine vertical organic carbon flux observations from 407 sampling locations spanning eight biomes across the global ocean. We fit a global scale Bayesian hierarchical model that describes the vertical profile of organic carbon flux with depth. Profile parameters within a particular biome are assumed to share a common deviation from the global mean profile. Individual station-level parameters are then modeled as deviations from the common biome-level profile. The hierarchical approach is shown to have several benefits over simpler and more common data aggregation methods. First, the hierarchical approach avoids statistical complexities introduced due to unbalanced sampling and allows for flexible incorporation of spatial heterogeneitites in model parameters. Second, the hierarchical approach uses the whole dataset simultaneously to fit the model parameters which shares information across datasets and reduces the uncertainty up to 95% in individual profiles. Third, the Bayesian approach incorporates prior scientific information about model parameters; for example, the non-negativity of chemical concentrations or mass-balance, which we apply here. We explicitly quantify each of these properties in turn. We emphasize the generality of the hierarchical Bayesian approach for diverse environmental applications and its increasing feasibility for large datasets due to recent developments in Markov Chain Monte Carlo algorithms and easy-to-use high-level software implementations.

54 ENVIRONMENTAL SCIENCES↗

The hierarchical growth of bright central galaxies and intracluster light as traced by the magnitude gap

Using a sample of 2800 galaxy clusters identified in the Dark Energy Survey across the redshift range 0.20 < z < 0.60, we characterize the hierarchical assembly of bright central galaxies (BCGs) and the surrounding intracluster light (ICL). To quantify hierarchical formation we use the stellar mass–halo mass (SMHM) relation, comparing the halo mass, estimated via the mass–richness relation, to the stellar mass within the BCG + ICL system. Moreover, we incorporate the magnitude gap (M14), the difference in brightness between the BCG (measured within 30 kpc) and fourth brightest cluster member galaxy within 0.5 $R_{200,c}$, as a third parameter in this linear relation. The inclusion of M14, which traces BCG hierarchical growth, increases the slope and decreases the intrinsic scatter, highlighting that it is a latent variable within the BCG + ICL SMHM relation. Moreover, the correlation with M14 decreases at large radii. However, the stellar light within the BCG + ICL transition region (30 –80 kpc) most strongly correlates with halo mass and has a statistically significant correlation with M14. Since the transition region and M14 are independent measurements, the transition region may grow due to the BCG’s hierarchical formation. Additionally, as M14 and ICL result from hierarchical growth, we use a stacked sample and find that clusters with large M14 values are characterized by larger ICL and BCG + ICL fractions, which illustrates that the merger processes that build the BCG stellar mass also grow the ICL. Furthermore, this may suggest that M14 combined with the ICL fraction can identify dynamically relaxed clusters.

79 ASTRONOMY AND ASTROPHYSICS↗

TDCOSMO - XVI. Measurement of the Hubble constant from the lensed quasar WGD 2038–4008

Time-delay cosmography is a powerful technique to constrain cosmological parameters, particularly the Hubble constant (H0). The TDCOSMO Collaboration is performing an ongoing analysis of lensed quasars to constrain cosmology using this method. In this work, we obtain constraints from the lensed quasar WGD 2038−4008 using new time-delay measurements and previous mass models by TDCOSMO. This is the first TDCOSMO lens to incorporate multiple lens modeling codes and the full time-delay covariance matrix into the cosmological inference. The models are fixed before the time delay is measured, and the analysis is performed blinded with respect to the cosmological parameters to prevent unconscious experimenter bias. We obtain DΔ t = 1.68−0.38+0.40 Gpc using two families of mass models, a power-law describing the total mass distribution, and a composite model of baryons and dark matter, although the composite model is disfavored due to kinematics constraints. In a flat ΛCDM cosmology, we constrain the Hubble constant to be H0 = 65−14+23 km s−1 Mpc−1. The dominant source of uncertainty comes from the time delays, due to the low variability of the quasar. Future long-term monitoring, especially in the era of the Vera C. Rubin Observatory’s Legacy Survey of Space and Time, could catch stronger quasar variability and further reduce the uncertainties. This system will be incorporated into an upcoming hierarchical analysis of the entire TDCOSMO sample, and improved time delays and spatially-resolved stellar kinematics could strengthen the constraints from this system in the future.Key words: gravitational lensing: strong / cosmological parameters / distance scale⋆ Corresponding author; kcwong19@gmail.com.⋆⋆ NHFP Einstein fellow.

79 ASTRONOMY AND ASTROPHYSICS↗

Haar-Like Wavelets on Hierarchical Trees

Here, discrete wavelet methods, originally formulated in the setting of regularly sampled signals, can be adapted to data defined on a point cloud if some multiresolution structure is imposed on the cloud. A wide variety of hierarchical clustering algorithms can be used for this purpose, and the multiresolution structure obtained can be encoded by a hierarchical tree of subsets of the cloud. Prior work introduced the use of Haar-like bases defined with respect to such trees for approximation and learning tasks on unstructured data. This paper builds on that work in two directions. First, we present an algorithm for constructing Haar-like bases on general discrete hierarchical trees. Second, with an eye towards data compression, we present thresholding techniques for data defined on a point cloud with error controlled in the $L$ $\infty$ norm and in a Hölder-type norm. In a concluding trio of numerical examples, we apply our methods to compress a point cloud dataset, study the tightness of the $L$ $\infty$ error bound, and use thresholding to identify MNIST classifiers with good generalizability.

97 MATHEMATICS AND COMPUTING↗