Search NASA⌕ Search

SEARCH · Search NASA

Results for “demographic inference”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Inferring demographic and selective histories from population genomic data using a 2-step approach in species with coding-sparse genomes: an application to human data

Abstract The demographic history of a population, and the distribution of fitness effects (DFE) of newly arising mutations in functional genomic regions, are fundamental factors dictating both genetic variation and evolutionary trajectories. Although both demographic and DFE inference has been performed extensively in humans, these approaches have generally either been limited to simple demographic models involving a single population, or, where a complex population history has been inferred, without accounting for the potentially confounding effects of selection at linked sites. Taking advantage of the coding-sparse nature of the genome, we propose a 2-step approach in which coalescent simulations are first used to infer a complex multi-population demographic model, utilizing large non-functional regions that are likely free from the effects of background selection. We then use forward-in-time simulations to perform DFE inference in functional regions, conditional on the complex demography inferred and utilizing expected background selection effects in the estimation procedure. Throughout, recombination and mutation rate maps were used to account for the underlying empirical rate heterogeneity across the human genome. Importantly, within this framework it is possible to utilize and fit multiple aspects of the data, and this inference scheme represents a generalized approach for such large-scale inference in species with coding-sparse genomes.

Soni, Vivak (ORCID:0000000294969562)↗

Developing an Evolutionary Baseline Model for Humans: Jointly Inferring Purifying Selection with Population History

Building evolutionarily appropriate baseline models for natural populations is not only important for answering fundamental questions in population genetics—including quantifying the relative contributions of adaptive versus nonadaptive processes—but also essential for identifying candidate loci experiencing relatively rare and episodic forms of selection (e.g., positive or balancing selection). Here, a baseline model was developed for a human population of West African ancestry, the Yoruba, comprising processes constantly operating on the genome (i.e., purifying and background selection, population size changes, recombination rate heterogeneity, and gene conversion). Specifically, to perform joint inference of selective effects with demography, an approximate Bayesian approach was employed that utilizes the decay of background selection effects around functional elements, taking into account genomic architecture. This approach inferred a recent 6-fold population growth together with a distribution of fitness effects that is skewed towards effectively neutral mutations. Importantly, these results further suggest that, although strong and/or frequent recurrent positive selection is inconsistent with observed data, weak to moderate positive selection is consistent but unidentifiable if rare.

59 BASIC BIOLOGICAL SCIENCES↗

Inferring Stochastic Rates from Heterogeneous Snapshots of Particle Positions

Many imaging techniques for biological systems—like fixation of cells coupled with fluorescence microscopy—provide sharp spatial resolution in reporting locations of individuals at a single moment in time but also destroy the dynamics they intend to capture. In this study, these snapshot observations contain no information about individual trajectories, but still encode information about movement and demographic dynamics, especially when combined with a well-motivated biophysical model. The relationship between spatially evolving populations and single-moment representations of their collective locations is well-established with partial differential equations (PDEs) and their inverse problems. However, experimental data is commonly a set of locations whose number is insufficient to approximate a continuous-in-space PDE solution. Here, motivated by popular subcellular imaging data of gene expression, we embrace the stochastic nature of the data and investigate the mathematical foundations of parametrically inferring demographic rates from snapshots of particles undergoing birth, diffusion, and death in a nuclear or cellular domain. Toward inference, we rigorously derive a connection between individual particle paths and their presentation as a Poisson spatial process. Using this framework, we investigate the properties of the resulting inverse problem and study factors that affect quality of inference. One pervasive feature of this experimental regime is the presence of cell-to-cell heterogeneity. Rather than being a hindrance, we show that cell-to-cell geometric heterogeneity can increase the quality of inference on dynamics for certain parameter regimes. Altogether, the results serve as a basis for more detailed investigations of subcellular spatial patterns of RNA molecules and other stochastically evolving populations that can only be observed for single instants in their time evolution.

59 BASIC BIOLOGICAL SCIENCES↗

Intra- and inter-subtype HIV diversity between 1994 and 2018 in southern Uganda: a longitudinal population-based study

There is limited data on human immunodeficiency virus (HIV) evolutionary trends in African populations. We evaluated changes in HIV viral diversity and genetic divergence in southern Uganda over a 24-year period spanning the introduction and scale-up of HIV prevention and treatment programs using HIV sequence and survey data from the Rakai Community Cohort Study, an open longitudinal population-based HIV surveillance cohort. Gag (p24) and env (gp41) HIV data were generated from people living with HIV (PLHIV) in 31 inland semi-urban trading and agrarian communities (1994–2018) and four hyperendemic Lake Victoria fishing communities (2011–2018) under continuous surveillance. HIV subtype was assigned using the Recombination Identification Program with phylogenetic confirmation. Inter-subtype diversity was evaluated using the Shannon diversity index, and intra-subtype diversity with the nucleotide diversity and pairwise TN93 genetic distance. Genetic divergence was measured using root-to-tip distance and pairwise TN93 genetic distance analyses. Demographic history of HIV was inferred using a coalescent-based Bayesian Skygrid model. Evolutionary dynamics were assessed among demographic and behavioral population subgroups, including by migration status. 9931 HIV sequences were available from 4999 PLHIV, including 3060 and 1939 persons residing in inland and fishing communities, respectively. In inland communities, subtype A1 viruses proportionately increased from 14.3% in 1995 to 25.9% in 2017 (P < .001), while those of subtype D declined from 73.2% in 1995 to 28.2% in 2017 (P < .001). The proportion of viruses classified as recombinants significantly increased by nearly four-fold from 12.2% in 1995 to 44.8% in 2017. Inter-subtype HIV diversity has generally increased. While intra-subtype p24 genetic diversity and divergence leveled off after 2014, intra-subtype gp41 diversity, effective population size, and divergence increased through 2017. Intra- and inter-subtype viral diversity increased across all demographic and behavioral population subgroups, including among individuals with no recent migration history or extra-community sexual partners. This study provides insights into population-level HIV evolutionary dynamics following the scale-up of HIV prevention and treatment programs. Continued molecular surveillance may provide a better understanding of the dynamics driving population HIV evolution and yield important insights for epidemic control and vaccine development.

60 APPLIED LIFE SCIENCES↗

An integrated integral projection model ( IPM 2 ) to disentangle size‐structured harvest and natural mortality

Abstract Body size is one of the most important traits governing individual‐level demographic rates and modulating population‐level processes. Multiple size‐dependent demographic rates can simultaneously change population structure, so distinguishing their individual contributions to overall population dynamics remains a challenge. Disentangling size‐dependent harvest rates from other demographic rates is critical for assessing the impact of removal on populations of invasive species. Inference about invasive populations can be difficult, however, as observations are often collected opportunistically as part of removal programs, rather than experimentally designed. Yet accurate inference is essential for understanding the feasibility of population suppression and optimising management decisions. We develop an integrated integral projection model (IPM 2 ) that leverages the strengths of the integrated population model and integral projection model to enable inference about complex, size‐structured demographic rates from imperfect observations. We apply the IPM 2 in the context of invasive European green crab ( Carcinus maenas ), a species for which individual body size strongly regulates both the observation‐generating process and latent, population dynamics. The IPM 2 facilitates the distinct estimation of green crab size‐structured harvest and natural mortality rates, parameters for which no explicit data is collected and that are unidentifiable in component datasets of the integrated population model. The model represents how the green crab population changes over time, providing the first estimates of size‐structured abundance of this high‐priority species. By forecasting the stable size distribution and equilibrium population size under varying removal efforts, we demonstrate that extremely high levels of removal effort can reduce the equilibrium green crab population size. Yet these high mortality rates also shift the stable size distribution and increase the equilibrium abundance of smaller crabs, since size‐selective removal alters intraspecific interactions. The ecological outcome of this shift in size structure will be variable, as green crab size modulates only some of its interactions with other species. These results highlight the value of the IPM 2 framework for inferring complex population dynamics with information needs that outpace information in individual observational datasets, providing a path forward for accurate assessment of conservation programs.

Keller, Abigail G. [Department of Environment Scie↗

pyFLANK, a graph neural network based null distribution inference model for F ST outlier detection

Detecting genomic regions under selection is essential for understanding how populations adapt to different environments, yet it remains challenging due to the confounding effects of demographic history and linkage disequilibrium (LD). Fixation index (F ST ) is a widely used statistic to identify genomic regions under adaptation. However, identifying genes under selection by defining F ST outliers often remains challenging, owing to confounding effects of underlying demographic history. Traditional methods assume independence among loci and rely on simple demographic models, while newer models perform much better but are computationally expensive and not easily scalable. Here, we present pyFLANK, an open-source and automated Python implementation which detects F ST outliers using a null distribution inferred from quasi-independent loci. Our tool integrates three approaches to identify loci obeying a null distribution: graph neural network (GNN) inference, linkage disequilibrium (LD)-based inference, and user-defined input. Because pyFLANK uses GNN-based inference of quasi-independent loci, it yields a more accurate null model with less need for user parameter input. In simulation experiments, pyFLANK achieved lower false positive rates than current methods while maintaining comparable detection power, indicating that its refined null model better distinguishes true adaptive loci from background variation. The GNN-based model, in particular, detected additional loci associated with phenotypic variance that were not identified by existing methods. Assessments of simulation and real data from different species demonstrate that pyFLANK achieves lower false positive rates compared with other commonly used F ST outlier detectors, while maintaining comparable detection power and excellent computational performance, providing a robust and user-friendly tool for identifying loci under divergent selection. It extends existing F ST outlier frameworks by incorporating explicit LD-aware strategies for null model calibration. The method is intended as a practical and scalable complement to existing genome scan approaches.

FST↗

The SEDs and Host Galaxies of the Dustiest GRB Afterglows

The afterglows and host galaxies of long gamma-ray bursts (GRBs) offer unique opportunities to study star-forming galaxies in the high-z Universe, Until recently, however. the information inferred from GRB follow-up observations was mostly limited to optically bright afterglows. biasing all demographic studies against sight-lines that contain large amounts of dust. Aims. Here we present afterglow and host observations for a sample of bursts that are exemplary of previously missed ones because of high visual extinction (A(sub v) (Sup GRB) approx > 1 mag) along the sight-line. This facilitates an investigation of the properties, geometry and location of the absorbing dust of these poorly-explored host galaxies. and a comparison to hosts from optically-selected samples. Methods. This work is based on GROND optical/NIR and Swift/XRT X-ray observations of the afterglows, and multi-color imaging for eight GRB hosts. The afterglow and galaxy spectral energy distributions yield detailed insight into physical properties such as the dust and metal content along the GRB sight-line as well as galaxy-integrated characteristics like the host's stellar mass, luminosity. color-excess and star-formation rate. Results. For the eight afterglows considered in this study we report for the first time the redshift of GRBs 081109 (z = 0.97S7 +/- 0.0005). and the visual extinction towards GRBs 0801109 (A(sub v) (Sup GRB) = 3.4(sup +0.4) (sub -0.3) mag) and l00621A (A(sub v) (Sup GRB) = 3.8 +/- 0.2 mag), which are among the largest ever derived for GRB afterglows. Combined with non-extinguished GRBs. there is a strong anti-correlation between the afterglow's metals-to-dust ratio and visual extinction. The hosts of the dustiest afterglows are diverse in their properties, but on average redder(((R - K)(sub AB)) approximates 1.6 mag), more luminous ( approximates 0.9 L (sup *)) and massive ((log M(sup *) [M(solar]) approximates 9.8) than the hosts of optically-bright events. We hence probe a different galaxy population. suggesting that previous host samples miss most of the massive. chemically-evolved and metal-rich members. This also indicates that the dust along the sight-line is often related to host properties, and thus probably located in the diffuse ISM or interstellar clouds and not in the immediate GRB environment. Some of the hosts in our sample. are blue, young or of small stellar mass illustrating that even apparently non-extinguished galaxies possess very dusty sight-lines due to a patchy dust distribution. Conclusions. The afterglows and host galaxies of the dustiest GRBs provide evidence for a complex dust geometry in star-forming galaxies. In addition, they establish a population of luminous. massive and correspondingly chemically-evolved GRB hosts. This suggests that GRBs trace the global star-formation rate better than studies based on optically-selected host samples indicate, and the previously-claimed deficiency of high-mass host galaxies was at least partially a selection effect.

Kruhler, T.↗

The Use of Artificial Intelligence in Head and Neck Cancers: A Multidisciplinary Survey

Artificial intelligence (AI) approaches have been introduced in various disciplines but remain rather unused in head and neck (H&N) cancers. This survey aimed to infer the current applications of and attitudes toward AI in the multidisciplinary care of H&N cancers. From November 2020 to June 2022, a web-based questionnaire examining the relationship between AI usage and professionals’ demographics and attitudes was delivered to different professionals involved in H&N cancers through social media and mailing lists. A total of 139 professionals completed the questionnaire. Only 49.7% of the respondents reported having experience with AI. The most frequent AI users were radiologists (66.2%). Significant predictors of AI use were primary specialty (V = 0.455; p < 0.001), academic qualification and age. AI’s potential was seen in the improvement of diagnostic accuracy (72%), surgical planning (64.7%), treatment selection (57.6%), risk assessment (50.4%) and the prediction of complications (45.3%). Among participants, 42.7% had significant concerns over AI use, with the most frequent being the ‘loss of control’ (27.6%) and ‘diagnostic errors’ (57.0%). This survey reveals limited engagement with AI in multidisciplinary H&N cancer care, highlighting the need for broader implementation and further studies to explore its acceptance and benefits.

60 APPLIED LIFE SCIENCES↗

Exploring how urban form and demographics are linked with pedestrian and bicycle safety

With pedestrian and bicycle safety as the focus, this study investigates the role of urban form, burdened communities (BCs), and demographics at the national level. Urban form can contribute to segregation, limiting access to crucial resources such as safe infrastructure, essential services, and economic opportunities. Leveraging recent data, this research applies six key indicators to identify BCs based on various socioeconomic and environmental factors. Here, the study creates a unique database combining 10 years of pedestrian-bicycle-involved fatal crashes with data for the 71,729 census tracts with burden indicators and census data. The data are analyzed using descriptives and rigorous zero-hurdle negative binomial models, which account for excessive zeros observed in the data. The inference-based analysis results reveal a positive correlation between burden indicators and pedestrian-bicycle-involved fatal crash occurrences, alongside a heightened risk in areas with high-intensity development. Higher Black, American Indian, or Alaska Native populations are associated with more fatal crashes. The study offers novel insights into safety dynamics across different contexts characterized by urban forms, BCs, and demographics. The study underscores the importance of targeted interventions to enhance pedestrian and bicycle safety.

bicycle crashes↗

Testing AGN outflow and accretion models with C IV and He II emission line demographics in z ≈ 2 quasars

Using ≈190 000 spectra from the 17th data release of the Sloan Digital Sky Survey (SDSS), we investigate the ultraviolet emission line properties in z ≈ 2 quasars. Specifically, we quantify how the shape of C IV λ1549 and the equivalent width (EW) of He II λ1640 depend on the black hole mass and Eddington ratio inferred from Mg II λ2800. Above L/L Edd ≳ 0.2, there is a strong mass dependence in both C IV blueshift and He II EW. Large C IV blueshifts are observed only in regions with both high mass and high accretion rate. Including X-ray measurements for a subsample of 5000 objects, we interpret our observations in the context of AGN accretion and outflow mechanisms. The observed trends in He II and 2 keV strength are broadly consistent with theoretical qsosed models of AGN spectral energy distributions (SEDs) for low spin black holes, where the ionizing SED depends on the accretion disc temperature and the strength of the soft excess. High spin models are not consistent with observations, suggesting SDSS quasars at z ≈ 2 may in general have low spins. We find a dramatic switch in behaviour at L/L Edd ≲ 0.1: the ultraviolet emission properties show much weaker trends, and no longer agree with qsosed predictions, hinting at changes in the structure of the broad line region. Overall, the observed emission line trends are generally consistent with predictions for radiation line driving where quasar outflows are governed by the SED, which itself results from the accretion flow and hence depends on both the SMBH mass and accretion rate.

79 ASTRONOMY AND ASTROPHYSICS↗

Initial Mobility Analysis for ORNL VA-EDH Synthetic Populations

Travel burdens are a major barrier to healthcare access among US Veteran patient populations, particularly those residing in rural areas. Spatial accessibility to points of care for US Veteran populations is commonly assessed in two ways. The first approach uses open data from the US Census to represent collective travel burdens, for example the distance between population-weighted census tract centroids and VHA points of care. The second approach uses restricted-access VHA patient data to measure travel costs (e.g., distance, time) for accessing points of care with respect to geolocated patient addresses and real or approximated transportation networks. While the advantage of the open data approach lies in its reproducibility, it has notable limitations in its tendency to infer individual travel behavior from aggregate population characteristics, a problem known as ecological fallacy. Conversely, while the patient data approach is able to account for individual travel behavior, its ability to account for localized access disparities (e.g., a neighborhood with exceptionally high transportation costs) and patient demographics is limited as protecting individual patient data requires their storage in closed systems with limited capacity for adequately modeling real-world travel patterns or for supplementing patient attributes. Additionally, the patient data approach cannot account for veterans who are not enrolled in the VHA system but who may be eligible for care. These challenges limit the ability to perform “what if” analyses on the effects of place-specific interventions on veteran populations with high access barriers to healthcare. To address these challenges, we explore the application of realistic synthetic populations to examine travel burdens and spatial accessibility issues among veteran patient populations. Synthetic populations provide a virtual, individually-resolved and cross-sectional representation of the veteran patient population that enables investigation of spatial access to points of care in ways in which aggregate data and patient data do not. First, synthetic populations allow one to directly assess how individuals access points of care, from synthesized residential locations to outpatient facilities on real-world transportation networks. Modeling access to points of care at the individual scale addresses the ecological fallacy problem associated with using aggregated census data to represent veteran populations and patterns of movement. Second, synthetic populations provide a means of completely representing an area’s veteran population using only publicly available, anonymized census microdata from the American Community Survey (ACS) to ensure the privacy of real-world individuals. Generating synthetic populations from the ACS also expands descriptive characteristics beyond what patient data typically offers to include socio-demographic, economic, housing, and mobility attributes. More detailed profiles of both VHA patient populations and veterans not enrolled in the VA system will provide a comprehensive picture of groups that may benefit from interventions or outreach. As an initial exercise for using synthetic populations to measure veteran travel burdens to VA care, we apply Oak Ridge National Laboratory’s (ORNL) UrbanPop capability to generate a series of synthetic VHA patient populations for 9 Veterans Integrated Services Networks (VISN) market areas in 9 Census Divisions across the continental United States, which are listed in Table 1. We use UrbanPop to produce synthetic populations for the VISN markets selected for each US Census Division, then assign VA outpatient clinic destinations to synthetic VHA patients based on travel about each VISN market’s road network. To demonstrate using the synthetic populations to evaluate healthcare travel burdens, we compare the time-based impedance between simulated home locations and VA outpatient clinics in each VISN market. We then perform validation exercises on the synthetic populations with respect to neighborhood (block group) demographic composition as well as patient mobility, comparing aggregate origin-destination statistics for the synthetic population to outpatient visits available in restricted patient data from the VA’s Corporate Data Warehouse (CDW) database.

97 MATHEMATICS AND COMPUTING↗

A Process Model of Interpersonal Relationship Formation in Isolated and Confined Environments

Future space crews will face several challenges such as living and working in a confined environment, isolated from others. These circumstances increase the importance of interpersonal compatibility, teamwork, and team performance. The interpersonal compatibility of space crews has been and continues to be of interest to both NASA and the Institute of Biomedical Problems (IBMP), whose research informs operations for Roscosmos. In NASA-sponsored research, interpersonal compatibility has been examined in terms of how the combination of team members’ personality traits, values, and demographics shape team member relationships and team performance overtime. IBMP-sponsored research mostly has moved away from trait-based approaches toward an idiographic (in-depth, heavily descriptive) approach to researching crew interpersonal relations. This research uses software such as Personal Self-Perception and Attitudes (PSPA), network approaches to team member relations, and content analysis of interactions to assess interpersonal compatibility, and infer states and team dynamics. Our research program integrated these ideas. Our primary research aim was to develop and empirically test a process model of interpersonal relationship formation in isolated and confinement environments. We created a model, collected data from teams in an isolated and confined environment, and applied a novel analytical strategy that combines trait, state, and interaction data (i.e., relational events). We focused specifically on the formation of strained relationships in isolation and confinement—or with whom fellow crewmembers find it difficult to work.

S T Bell↗

Climatic and Demographic Consequences of the Massive Volcanic Eruption of 1258

Somewhere in the tropics, a volcano exploded violently during the year 1258, producing a massive stratospheric aerosol veil that eventually blanketed the globe. Arctic and Antarctic ice cores suggest that this was the world's largest volcanic eruption of the past millennium. According to contemporary chronicles, the stratospheric dry fog possibly manifested itself in Europe as a persistently cloudy aspect of the sky and also through an apparently total darkening of the eclipsed Moon. Based on a sudden temperature drop for several months in England, the eruption's initiation date can be inferred to have been probably January 1258. The frequent cold and rain that year led to severe crop damage and famine throughout much of Europe. Pestilence repeatedly broke out in 1258 and 1259; it occurred also in the Middle East, reportedly there as plague. Another very cold winter followed in 1260-1261. The troubled period's wars, famines, pestilences, and earthquakes appear to have contributed in part to the rise of the European flagellant movement of 1260, one of the most bizarre social phenomena of the Middle Ages. Analogies can be drawn with the climatic aftereffects and European social unrest following another great tropical eruption, Tambora in 1815. Some generalizations about the climatic impacts of tropical eruptions are made from these and other data.

Stothers, Richard B.↗

Population Structure of White Sturgeon (Acipenser transmontanus) in the Columbia River Inferred from Single-Nucleotide Polymorphisms

White sturgeon (Acipenser transmontanus) are the largest freshwater fish in North America, with reproducing populations in the Sacramento-San Joaquin, Fraser, and Columbia River Basins. Of these, the Columbia River is the largest, but it is also highly fragmented by hydroelectric dams, and many segments are characterized by declining abundance and persistent recruitment failure. Efforts to conserve and supplement these fish requires an understanding of their spatial genetic structure. Here, we assembled a large set of samples from throughout the Columbia River Basin, along with representative collections from adjacent basins, and genotyped them using a panel of 325 single-nucleotide markers. Results from individual- and group-based analyses of these data indicate that white sturgeon in the uppermost Columbia River Basin, in the Kootenai and upper Snake Rivers, are the most distinct, while the remaining populations downstream in the basin can be described as a genetic gradient consistent with an isolation-by-distance effect. Notably, the population in the lowest reaches of the Columbia River is more distinct from the middle or upper reaches than from outside basins, and suggests historically a higher or more recent gene exchange through coastal routes than with populations in the interior Columbia Basin. Nonetheless, proximal reaches were generally only marginally or non-significantly divergent, suggesting that transplanting larvae or juveniles from nearby sources poses relatively little risk of outbreeding depression. Indeed, we inferred examples of dispersal between reaches via close-kin mark-recapture and genetic mark-recapture that indicate movement between nearby reaches is not unusual. Samples from the Kootenai and upper Snake Rivers exhibited notably lower genetic diversity than the remaining samples as a result of population bottlenecks, genetic drift, and/or historical divergence. Conservation actions, such as supplementation, are underway to maintain population viability and will require balanced efforts to increase demographic abundance while maintaining genetic diversity.

Willis, Stuart C.↗

Directionally supervised cellular automaton for the initial peopling of Sahul

Reconstructing the patterns of Homo sapiens expansion out of Africa and across the globe has been advanced using demographic and travel-cost models. However, modelled routes are ipso facto influenced by migration rates, and vice versa. We combined movement ‘superhighways’ with a demographic cellular automaton to predict one of the world's earliest peopling events — Sahul between 75000 and 50000 years ago. Novel outcomes from the superhighways-weighted model include (i) an approximate doubling of the predicted time to continental saturation (~10,000 years) compared to that based on the directionally unsupervised model (~5000 years), suggesting that rates of migration need to account for topographical constraints in addition to rate of saturation; (ii) a previously undetected movement corridor south through the centre of Sahul early in the expansion wave based on the scenarios assuming two dominant entry points into Sahul; and (iii) a better fit to the spatially de-biased, Signor-Lipps-corrected layer of initial arrival inferred from dated archaeological material. Our combined model infrastructure provides a data-driven means to examine how people initially moved through, settled, and abandoned different regions of the globe.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Producing High-fidelity Synthetic Population Ensembles at Scale

Used within social simulations, synthetic population ensembles enable uncertainty quantification (UQ) methods for obtaining more robust model inference and prediction. A synthetic population ensemble is a series of plausible virtual reconstructions of an area’s population at the granularity of people and residences, generated stochastically to preserve privacy of the source population survey’s respondents. In this paper, we demonstrate the production of large synthetic population ensembles for the US via Oak Ridge National Laboratory’s UrbanPop framework to support modeling of high spatial resolution energy affordability metrics from nationwide social surveys in collaboration with the fusionACS project. Our initial task involves creating ensembles for 17 US metropolitan areas, each consisting of 41 population instances (a base realization and 40 replicates). To accomplish this task at scale, we configured an integrated system comprised of a research cloud, virtual containerization, GPU-enhanced functionality, and a dual API/CLI to interact with UrbanPop’s maturing Likeness Python ecosystem. We observe a reduction in theoretical execution time while maintaining high-fidelity approximations of residential totals by metropolitan area and the demographic characteristics of neighborhoods. We discuss expansion of our approach to produce synthetic population ensembles for the entire US, particularly plans to establish automated workflows for job orchestration to increase computational efficiency, as well as provide outlook for broadening applications of the ensembles.

Gaboardi, James [ORNL] (ORCID:0000000247766826)↗

Scalable Generation of High-fidelity Synthetic Population Ensembles

Used within social simulations, synthetic population ensembles enable uncertainty quantification (UQ) methods for obtaining more robust model inference and prediction. A synthetic population ensemble is a series of plausible virtual reconstructions of an area’s population at the granularity of people and residences, generated stochastically to preserve privacy of the source population survey’s respondents. In this paper, we demonstrate the production of large synthetic population ensembles for the U.S. via Oak Ridge National Laboratory’s UrbanPop framework to support modeling of high spatial resolution energy affordability metrics from nationwide social surveys in collaboration with the fusionACS project. The study involves two scenarios: creating ensembles for (1) 17 U.S. metropolitan areas in 2019 and (2) full U.S. Census Divisions in 2023, with each scenario consisting of 41 population instances (a base realization and 40 replicates). To accomplish this task at scale, we configured an integrated system within a research cloud, comprised of virtual containerizations, GPU-enhanced functionality, and orchestrated deployments of UrbanPop’s maturing Likeness Python ecosystem. Results demonstrate we maintained high-fidelity approximations of residential totals by areas of interest and the demographic characteristics of neighborhoods while reducing manual workflow burdens. Finally, we discuss plans to fine-tune and further develop our automated workflows for truly distributed job orchestration to increase computational efficiency, as well as provide an outlook for broadening applications of the ensembles.

Cluster computing↗