Search NASA⌕ Search

SEARCH · Search NASA

Results for “hierarchical sampling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Iterative self-organizing SCEne-LEvel sampling (ISOSCELES) for large-scale building extraction

Convolutional neural networks (CNN) provide state-of-the-art performance in many computer vision tasks, including those related to remote-sensing image analysis. Successfully training a CNN to generalize well to unseen data, however, requires training on samples that represent the full distribution of variation of both the target classes and their surrounding contexts. With remote sensing data, acquiring a sufficiently representative training set is a challenge due to both the inherent multi-modal variability of satellite or aerial imagery and the general high cost of labeling data. To address this challenge, we have developed ISOSCELES, an Iterative Self-Organizing SCEne LEvel Sampling method for hierarchical sampling of large image sets. Using affinity propagation, ISOSCELES automates the selection of highly representative training images. Compared to random sampling or using available reference data, the distribution of the training is principally data driven, reducing the chance of oversampling uninformative areas or undersampling informative ones. In comparison to manual sample selection by an analyst, ISOSCELES exploits descriptive features, spectral and/or textural, and eliminates human bias in sample selection. Using a hierarchical sampling approach, ISOSCELES can obtain a training set that reflects both between-scene variability, such as in viewing angle and time of day, and within-scene variability at the level of individual training samples. We verify the method by demonstrating its superiority to stratified random sampling in the challenging task of adapting a pre-trained model to a new image and spatial domain for country-scale building extraction. Using a pair of hand-labeled training sets comprising 1,987 sample image chips, a total of 496,000,000 individually labeled pixels, we show, across three distinct model architectures, an increase in accuracy, as measured by F1-score, of 2.2–4.2%.

42 ENGINEERING↗

Enabling machine learning-ready HPC ensembles with Merlin

With the growing complexity of computational and experimental facilities, many scientific researchers are turning to machine learning (ML) techniques to analyze large scale ensemble data. With complexities such as multi-component workflows, heterogeneous machine architectures, parallel file systems, and batch scheduling, care must be taken to facilitate this analysis in a high performance computing (HPC) environment. Here, we present Merlin, a workflow framework to enable large ML-friendly ensembles of scientific HPC simulations. By augmenting traditional HPC with distributed compute technologies, Merlin aims to lower the barrier for scientific subject matter experts to incorporate ML into their analysis. As a producer–consumer workflow model, Merlin enables multi-machine, cross-batch job, dynamically allocated yet persistent workflows capable of utilizing surge-compute resources. Key features of Merlin are a flexible HPC-centric interface, low per-task overhead, multi-tiered fault recovery, and a hierarchical sampling algorithm that allows for $\mathscr{O}$(N) task execution and $\mathscr{O}$(N ln N) task queuing to ensembles of millions of tasks. In addition to Merlin’s design, we test the algorithm’s performance in an HPC center and demonstrate the ability to enqueue 40 million simulations in 100 s, with a 30 millisecond per-task overhead that is independent of ensemble size. Finally, we describe some example applications that Merlin has enabled on leadership-class HPC resources, such as the ML-augmented optimization of nuclear fusion experiments and the calibration of infectious disease models to study the progression of and possible mitigation strategies for COVID-19.

97 MATHEMATICS AND COMPUTING↗

Micro on a macroscale: relating microbial-scale soil processes to global ecosystem function

ABSTRACT Soil microorganisms play a key role in driving major biogeochemical cycles and in global responses to climate change. However, understanding and predicting the behavior and function of these microorganisms remains a grand challenge for soil ecology due in part to the microscale complexity of soils. It is becoming increasingly clear that understanding the microbial perspective is vital to accurately predicting global processes. Here, we discuss the microbial perspective including the microbial habitat as it relates to measurement and modeling of ecosystem processes. We argue that clearly defining and quantifying the size, distribution and sphere of influence of microhabitats is crucial to managing microbial activity at the ecosystem scale. This can be achieved using controlled and hierarchical sampling designs. Model microbial systems can provide key data needed to integrate microhabitats into ecosystem models, while adapting soil sampling schemes and statistical methods can allow us to collect microbially-focused data. Quantifying soil processes, like biogeochemical cycles, from a microbial perspective will allow us to more accurately predict soil functions and address long-standing unknowns in soil ecology.

59 BASIC BIOLOGICAL SCIENCES↗

Hierarchical Gaussian Random Field Sampling for Multilevel Markov Chain Monte Carlo: Coupling Stochastic Partial Differential Equation and the Karhunen–Loève Decomposition

This work introduces structure preserving hierarchical decompositions for sampling Gaussian random fields (GRFs) within the context of multilevel Bayesian inference in high-dimensional space. Existing scalable hierarchical sampling methods, such as those based on stochastic partial differential equations (SPDEs), often reduce the dimensionality of the sample space at the cost of accuracy of inference. Other approaches, such that those based on Karhunen-Loève (KL) expansions, offer sample space dimensionality reduction but sacrifice GRF representation accuracy and ergodicity of the Markov chain Monte Carlo (MCMC) sampler and are computationally expensive for high-dimensional problems. The proposed method integrates the dimensionality reduction capabilities of KL expansions with the scalability of SPDE-based sampling, thereby providing a robust, unified framework for high-dimensional uncertainty quantification (UQ) that is scalable and accurate, preserves ergodicity, and offers dimensionality reduction of the sample space. The hierarchy in our multilevel algorithm is derived from the geometric multigrid hierarchy. By constructing a hierarchical decomposition that maintains the covariance structure across the levels in the hierarchy, the approach enables efficient coarse-to-fine sampling while ensuring that all samples are drawn from the desired distribution. The effectiveness of the proposed method is demonstrated on a benchmark subsurface flow problem, demonstrating its effectiveness in improving computational efficiency and statistical accuracy. Furthermore, our proposed technique is more efficient and accurate and displays better convergence properties than existing methods for high-dimensional Bayesian inference problems.

Gaussian random fields↗

Stochastic nonlinear analysis of unidirectional fiber composites using image-based microstructural uncertainty quantification

Here, we present a data-driven nonlinear uncertainty quantification and propagation framework to study the microstructure-induced stochastic performance of unidirectional (UD) carbon fiber reinforced polymer (CFRP) composites. The proposed approach integrates (1) microscopic image characterization, (2) stochastic microstructure reconstruction, and (3) efficient multiscale finite element simulations enabled by self-consistent clustering (SCA) analysis. To model the complex microstructural variability, the proposed UQ methods take the non-Gaussian uncertainty sources into account through a distribution-free sampling approach leveraging nonparametric and asymptotic statistical tools. A hierarchical conditional sampling strategy enables the simultaneous sampling of multiple sources of uncertainties. Our approach provides insights into the impact of microstructural variabilities, which are shown to have an increasing impact on the nonlinear responses of UD CFRP parts under progressive compression loading and ultimately on the failure rate over time. We discover that before CFRP parts start to fail, a characteristic time period emerges with distinctive uncertainty distributions specific to the microstructure variability and the probability of failure. Identifying the failure time period is crucial to the reliability prediction, which is an essential component of CFRP design.

36 MATERIALS SCIENCE↗

hiposa v0.1.0

Fast construction of hierarchical N-dimensional experimental design. The software allows construction of sampling strategies that are self-avoiding, tunable in resolution in N dimensions. This code is a drop-in replacement for seeding optimization runs often performed using uniform random sampling. The hierarchical approach allows for a rapid, top down method for construction the most effective design that drives towards a global optimum.

Zwart, PetrusH [Lawrence Berkeley National Laborat↗

Integrated flight/propulsion control system design based on a decentralized, hierarchical approach

A sample integrated flight/propulsion control system design is presented for the piloted longitiudinal landing task with a modern, statistically unstable fighter aircraft. The design procedure is summarized, the vehicle model used in the sample study is described, and the procedure for partitioning the integrated system is presented along with a description of the subsystems. The high-level airframe performance specifications and control design are presented and the control performance is evaluated. The generation of the low-level (engine) subsystem specifications from the airframe requirements are discussed, and the engine performance specifications are presented along with the subsystem control design. A compensator to accommodate the influence of airframe outputs on the engine subsystem is also considered. Finally, the entire closed loop system performance and stability characteristics are examined.

Mattern, Duane↗

Integrated flight/propulsion control system design based on a decentralized, hierarchical approach

A sample integrated flight/propulsion control system design is presented for the piloted longitudinal landing task with a modern, statistically unstable fighter aircraft. The design procedure is summarized. The vehicle model used in the sample study is described, and the procedure for partitioning the integrated system is presented along with a description of the subsystems. The high-level airframe performance specifications and control design are presented and the control performance is evaluated. The generation of the low-level (engine) subsystem specifications from the airframe requirements are discussed, and the engine performance specifications are presented along with the subsystem control design. A compensator to accommodate the influence of airframe outputs on the engine subsystem is also considered. Finally, the entire closed loop system performance and stability characteristics are examined.

Mattern, Duane↗

Using low volume eDNA methods to sample pelagic marine animal assemblages

Environmental DNA (eDNA) is an increasingly useful method for detecting pelagic animals in the ocean but typically requires large water volumes to sample diverse assemblages. Ship-based pelagic sampling programs that could implement eDNA methods generally have restrictive water budgets. Studies that quantify how eDNA methods perform on low water volumes in the ocean are limited, especially in deep-sea habitats with low animal biomass and poorly described species assemblages. Using 12S rRNA and COI gene primers, we quantified assemblages comprised of micronekton, coastal forage fishes, and zooplankton from low volume eDNA seawater samples (n = 436, 380–1800 mL) collected at depths of 0–2200 m in the southern California Current. We compared diversity in eDNA samples to concurrently collected pelagic trawl samples (n = 27), detecting a higher diversity of vertebrate and invertebrate groups in the eDNA samples. Differences in assemblage composition could be explained by variability in size-selectivity among methods and DNA primer suitability across taxonomic groups. The number of reads and amplicon sequences variants (ASVs) did not vary substantially among shallow (<200 m) and deep samples (>600 m), but the proportion of invertebrate ASVs that could be assigned a species-level identification decreased with sampling depth. Using hierarchical clustering, we resolved horizontal and vertical variability in marine animal assemblages from samples characterized by a relatively low diversity of ecologically important species. Low volume eDNA samples will quantify greater taxonomic diversity as reference libraries, especially for deep-dwelling invertebrate species, continue to expand.

59 BASIC BIOLOGICAL SCIENCES↗

Hierarchical Inference of the Lensing Convergence from Photometric Catalogs with Bayesian Graph Neural Networks

Abstract We present a Bayesian graph neural network (BGNN) that can estimate the weak lensing convergence ( κ ) from photometric measurements of galaxies along a given line of sight (LOS). The method is of particular interest in strong gravitational time-delay cosmography (TDC), where characterizing the “external convergence” ( κ ext ) from the lens environment and LOS is necessary for precise Hubble constant ( H 0 ) inference. Starting from a large-scale simulation with a κ resolution of ∼1′, we introduce fluctuations on galaxy–galaxy lensing scales of ∼1″ and extract random sight lines to train our BGNN. We then evaluate the model on test sets with varying degrees of overlap with the training distribution. For each test set of 1000 sight lines, the BGNN infers the individual κ posteriors, which we combine in a hierarchical Bayesian model to yield constraints on the hyperparameters governing the population. For a test field well sampled by the training set, the BGNN recovers the population mean of κ precisely and without bias (within the 2 σ credible interval), resulting in a contribution to the H 0 error budget well under 1%. In the tails of the training set with sparse samples, the BGNN, which can ingest all available information about each sight line, extracts a stronger κ signal compared to a simplified version of the traditional method based on matching galaxy number counts, which is limited by sample variance. Our hierarchical inference pipeline using BGNNs promises to improve the κ ext characterization for precision TDC. The code is available as a public Python package, Node to Joy ⏬ .

79 ASTRONOMY AND ASTROPHYSICS↗

Accelerating multilevel Markov Chain Monte Carlo using machine learning models

Here, this work presents an efficient approach for accelerating multilevel Markov Chain Monte Carlo (MCMC) sampling for large-scale problems using low-fidelity machine learning models. While conventional techniques for large-scale Bayesian inference often substitute computationally expensive high-fidelity models with machine learning models, thereby introducing approximation errors, our approach offers a computationally efficient alternative by augmenting high-fidelity models with low-fidelity ones within a hierarchical framework. The multilevel approach utilizes the low-fidelity machine learning model (MLM) for inexpensive evaluation of proposed samples thereby improving the acceptance of samples by the high-fidelity model. The hierarchy in our multilevel algorithm is derived from geometric multigrid hierarchy. We utilize an MLM to accelerate the coarse level sampling. Training machine learning model for the coarsest level significantly reduces the computational cost associated with generating training data and training the model. We present an MCMC algorithm to accelerate the coarsest level sampling using MLM and account for the approximation error introduced. We provide theoretical proofs of detailed balance and demonstrate that our multilevel approach constitutes a consistent MCMC algorithm. Additionally, we derive the expression for cost reduction due to machine learning model to facilitate cost analysis of the hierarchical sampling algorithm. Our technique is demonstrated on a standard benchmark inference problem in groundwater flow, where we estimate the probability density of a quantity of interest using a four-level MCMC algorithm. Our proposed algorithm accelerates multilevel sampling by a factor of two while achieving similar accuracy compared to sampling using the standard multilevel algorithm.

97 MATHEMATICS AND COMPUTING↗

3D printed Ni-Mn-Fe Bi-Functional Catalyst for Metal Air Batteries

Oxygen evolution catalyst with porous structures on the submillimeter and millimeter ranges have been shown to have high catalytic activity. Often, a limiting factor in the performance of a catalyst can be its specific surface area as this can limit the availability of active sites which encourage the desired reaction. As such, the inclusion of a hierarchical structure that include both meso-scale and nano-scale porosity can offer a means to greatly increase the performance of electrochemical catalysts. We have developed a method for manufacturing samples with such a hierarchical structure by combining techniques of direct ink writing (DIW), thermal sintering and free corrosion based dealloying. This allows us to prepare three-dimensionally structured samples with open-cell porosity at the macro-, micro- and nano-scales. Each level of porosity offers a different advantage to the overall functionality and applicability of this technology.

25 ENERGY STORAGE↗

A compendium of bacterial and archaeal single-cell amplified genomes from oxygen deficient marine waters

Oxygen-deficient marine waters referred to as oxygen minimum zones (OMZs) or anoxic marine zones (AMZs) are common oceanographic features. They host both cosmopolitan and endemic microorganisms adapted to low oxygen conditions. Microbial metabolic interactions within OMZs and AMZs drive coupled biogeochemical cycles resulting in nitrogen loss and climate active trace gas production and consumption. Global warming is causing oxygen-deficient waters to expand and intensify. Therefore, studies focused on microbial communities inhabiting oxygen-deficient regions are necessary to both monitor and model the impacts of climate change on marine ecosystem functions and services. Here we present a compendium of 5,129 single-cell amplified genomes (SAGs) from marine environments encompassing representative OMZ and AMZ geochemical profiles. Of these, 3,570 SAGs have been sequenced to different levels of completion, providing a strain-resolved perspective on the genomic content and potential metabolic interactions within OMZ and AMZ microbiomes. Hierarchical clustering confirmed that samples from similar oxygen concentrations and geographic regions also had analogous taxonomic compositions, providing a coherent framework for comparative community analysis.

59 BASIC BIOLOGICAL SCIENCES↗

Automated Classification of Thermal Infrared Spectra Using Self-organizing Maps

Existing and planned space missions to a variety of planetary and satellite surfaces produce an ever increasing volume of spectral data. Understanding the scientific informational content in this large data volume is a daunting task. Fortunately various statistical approaches are available to assess such data sets. Here we discuss an automated classification scheme based on Kohonen Self-organizing maps (SOM) we have developed. The SUM process produces an output layer were spectra having similar properties lie in close proximity to each other. One major effort is partitioning this output layer into appropriate regions. This is prefonned by defining dosed regions based upon the strength of the boundaries between adjacent cells in the SOM output layer. We use the Davies-Bouldin index as a measure of the inter-class similarities and intra-class dissimilarities that determines the optimum partition of the output layer, and hence number of SOM clusters. This allows us to identify the natural number of clusters formed from the spectral data. Mineral spectral libraries prepared at Arizona State University (ASU) and John Hopkins University (JHU) are used to test and evaluate the classification scheme. We label the library sample spectra in a hierarchical scheme with class, subclass, and mineral group names. We use a portion of the spectra to train the SOM, i.e. produce the output layer, while the remaining spectra are used to test the SOM. The test spectra are presented to the SOM output layer and assigned membership to the appropriate cluster. We then evaluate these assignments to assess the scientific meaning and accuracy of the derived SOM classes as they relate to the labels. We demonstrate that unsupervised classification by SOMs can be a useful component in autonomous systems designed to identify mineral species from reflectance and emissivity spectra in the therrnal IR.

Roush, Ted L.↗

Voids and constraints on nonlinear clustering of galaxies

Void statistics of the galaxy distribution in the Center for Astrophysics Redshift Survey provide strong constraints on galaxy clustering in the nonlinear regime, i.e., on scales R equal to or less than 10/h Mpc. Computation of high-order moments of the galaxy distribution requires a sample that (1) densely traces the large-scale structure and (2) covers sufficient volume to obtain good statistics. The CfA redshift survey densely samples structure on scales equal to or less than 10/h Mpc and has sufficient depth and angular coverage to approach a fair sample on these scales. In the nonlinear regime, the void probability function (VPF) for CfA samples exhibits apparent agreement with hierarchical scaling (such scaling implies that the N-point correlation functions for N greater than 2 depend only on pairwise products of the two-point function xi(r)) However, simulations of cosmological models show that this scaling in redshift space does not necessarily imply such scaling in real space, even in the nonlinear regime; peculiar velocities cause distortions which can yield erroneous agreement with hierarchical scaling. The underdensity probability measures the frequency of 'voids' with density rho less than 0.2 -/rho. This statistic reveals a paucity of very bright galaxies (L greater than L asterisk) in the 'voids.' Underdensities are equal to or greater than 2 sigma more frequent in bright galaxy samples than in samples that include fainter galaxies. Comparison of void statistics of CfA samples with simulations of a range of cosmological models favors models with Gaussian primordial fluctuations and Cold Dark Matter (CDM)-like initial power spectra. Biased models tend to produce voids that are too empty. We also compare these data with three specific models of the Cold Dark Matter cosmogony: an unbiased, open universe CDM model (omega = 0.4, h = 0.5) provides a good match to the VPF of the CfA samples. Biasing of the galaxy distribution in the 'standard' CDM model (omega = 1, b = 1.5; see below for definitions) and nonzero cosmological constant CDM model (omega = 0.4, h = 0.6 lambda(sub 0) = 0.6, b = 1.3) produce voids that are too empty. All three simulations match the observed VPF and underdensity probability for samples of very bright (M less than M asterisk = -19.2) galaxies, but produce voids that are too empty when compared with samples that include fainter galaxies.

Vogeley, Michael S.↗

Void statistics of the CfA redshift survey

Clustering properties of two samples from the CfA redshift survey, each containing about 2500 galaxies, are studied. A comparison of the velocity distributions via a K-S test reveals structure on scales comparable with the extent of the survey. The void probability function (VPF) is employed for these samples to examine the structure and to test for scaling relations in the galaxy distribution. The galaxy correlation function is calculated via moments of galaxy counts. The shape and amplitude of the correlation function roughly agree with previous determinations. The VPFs for distance-limited samples of the CfA survey do not match the scaling relation predicted by the hierarchical clustering models. On scales not greater than 10/h Mpc, the VPFs for these samples roughly follow the hierarchical pattern. A variant of the VPF which uses nearly all the data in magnitude-limited samples is introduced; it accounts for the variation of the sampling density with velocity in a magnitude-limited survey.

Vogeley, Michael S.↗