Big Microstructure Datasets for Materials Informatics: Using Statistically Conditioned Generative Models to Curate Big Datasets
Explore the source record for details and available documents.
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Bottom-up urban energy models are crucial for understanding current energy use patterns and informing design strategies. However, accurately characterizing these models to represent different communities remains a challenge due to the extensive data needed for simulating existing energy use behavior. This data includes information related to human activities and building characteristics, all of which correlate with socioeconomic factors. To overcome this challenge, we developed an automated framework that utilizes both top-down and bottom-up data, to predict unknown building and occupant characteristics that are needed for more accurate and equitable modeling and analytics. Our framework, integrated into the URBANopt district energy modeling platform, uses statistical data models from ResStock. URBANopt models co-located buildings and neighborhoods. At this scale there are data gaps in building characteristic data, such as materials, insulation, occupancy, income, and energy usage of the buildings. To address this data gap, we use ResStock data, representative at the census tract scale, and develop machine-learning and deeplearning techniques to disaggregate it to individual buildings. By mapping unique occupant, building and economic properties to URBANopt energy models, we gain detailed insights into the variability of building energy use across different neighborhoods. This insight helps deploy technologies for co-located buildings and supports targeted upgrades for communities with unique economic and demographic characteristics, ensuring energy equity. Accurate characterization of energy models allows us to develop equitable strategies tailored to diverse neighborhoods, whether underserved or affluent. Our automated framework streamlines energy modeling and provides a reliable tool for building energy characterization.
Assessing local climate change impacts often requires downscaling coarse global climate model (GCM) output to finer resolution. Two main approaches exist: dynamical downscaling using high-resolution regional climate models, and statistical downscaling based on historical relationships between large-scale and local variables. In a recent analysis of five dynamically downscaled simulations over the western United States, Koszuta et al. (2024, https://doi.org/10.1029/2023gl107298) found that warming weakens orographic influence on winter precipitation, damping increases on windward slopes and amplifying them in rain-shadowed regions. Here we show that this effect is robust across seasons and multiple dynamically downscaled ensembles, and is more pronounced at higher model resolutions. However, it is absent in projections from a widely used statistical model (LOCA2), even when trained on high-resolution future simulations (LOCA2-Hybrid). This highlights a key limitation of many statistical downscaling methods: their preservation of parent GCM trends, which usually fail to capture emergent changes in orographic precipitation patterns.
The National Climate Database (NCDB) is a high resolution, bias-corrected climate dataset consisting of the three most widely used variables of solar radiation- global horizontal (GHI), direct normal (DNI), and diffuse horizontal irradiance (DHI)- as well as other meteorological data. The goal of the NCDB is to provide unbiased high temporal and spatial resolution climate data needed for renewable energy modeling. The NCDB is modeled using a statistical downscaling approach with Regional Climate Model (RCM)-based climate projections obtained from the North American Coordinated Regional Climate Downscaling Experiment (NA-CORDEX; linked below). Daily climate projections simulated by the Canadian Regional Climate Model 4 (CanRCM4) forced by the second-generation Canadian Earth System Model (CanESM2) for two Representative Concentration Pathways (RCP4.5 or moderate emissions scenario and RCP8.5 or highest baseline emission scenario) are selected as inputs to the statistical downscaling models. The National Solar Radiation Database (NSRDB) is used to build and calibrate statistical models.
We propose physically reasonable systems capable of avoiding ergodicity at infinite time in the thermodynamic limit, even with generic perturbations and when coupled to a heat bath. In two dimensions, the rainbow loop soup has (stretched) exponentially numerous absolutely stable nonergodic states with diverging energy but vanishing energy density. In three dimensions the rainbow membrane soup has (stretched) exponentially numerous nonergodic states with diverging energy barriers, leading to infinite-time robust ergodicity breaking that even survives coupling to a nonzero temperature heat bath. We describe our results in the language of exact emergent symmetries and demonstrate how the systems avoid common instabilities. Furthermore, our construction naturally connects to quantum dimer models, topologically ordered systems, the group word construction, and Hamiltonians whose low-energy eigenstates exhibit anomalous entanglement entropy.
This study presents a comprehensive framework for flood susceptibility mapping by integrating geospatial factors with both statistical and machine learning models. Thirteen Flood-related factors, including DEM, slope, TWI, NDVI, etc., are extracted as features of models, and historical flood data derived from Sentinel-1 SAR from 2018 to 2023 are used as the target variables of the models. These datasets are analyzed using a frequency-based statistical model and three machine learning models, including Random Forest, XGBoost, and CNN, to generate flood susceptibility maps. The performance of each model is evaluated through AUC; and SHAP scores are separately generated for Machine learning (ML) models to explain each feature contribution in the ML model. The generated susceptibility maps are validated by high-flood-risk locations monitored by flood sensors, BLE inundation models, and flood-prone areas suggested by the Local Community Task Force. The results indicate that the XGBoost model outperforms all other models, with an AUC of 0.92 and demonstrates the highest alignment with recommended high-flood-risk locations, while the frequency-based statistical model showed the weakest performance with an AUC of 0.65. SHAP value graphs highlight the elevation, slope, and TWI as the most influential features across all models. The susceptibility maps generated by the machine learning model show strong agreement with the BLE map and high-flood-risk areas identified by the local Community Task Force.
Direct statistical simulation (DSS) of nonlinear dynamical systems bypasses the traditional route of accumulating statistics by lengthy direct numerical simulations by solving the equations that govern the statistics themselves. DSS suffers, however, from the curse of dimensionality as the statistics (such as correlations) generally have higher dimensions than the underlying dynamical variables. Here we investigate two approaches to reduce the dimensionality of DSS, illustrating each method with numerical experiments with the Lorenz96 dynamical system. The forms of DSS chosen here involve approximate closures at second and third order in the equal-time cumulants. We demonstrate significant reduction in computational effort that can be achieved without sacrificing the accuracy of DSS. The methods developed here can be applied to turbulent fluid and magnetohydrodynamical systems. Published by the American Physical Society 2025
Abstract Background The findings of the 2023 AAPM Grand Challenge on Deep Generative Modeling for Learning Medical Image Statistics are reported in this Special Report. Purpose The goal of this challenge was to promote the development of deep generative models for medical imaging and to emphasize the need for their domain‐relevant assessments via the analysis of relevant image statistics. Methods As part of this Grand Challenge, a common training dataset and an evaluation procedure was developed for benchmarking deep generative models for medical image synthesis. To create the training dataset, an established 3D virtual breast phantom was adapted. The resulting dataset comprised about 108 000 images of size 512 512. For the evaluation of submissions to the Challenge, an ensemble of 10 000 DGM‐generated images from each submission was employed. The evaluation procedure consisted of two stages. In the first stage, a preliminary check for memorization and image quality (via the Fréchet Inception Distance [FID]) was performed. Submissions that passed the first stage were then evaluated for the reproducibility of image statistics corresponding to several feature families including texture, morphology, image moments, fractal statistics, and skeleton statistics. A summary measure in this feature space was employed to rank the submissions. Additional analyses of submissions was performed to assess DGM performance specific to individual feature families, the four classes in the training data, and also to identify various artifacts. Results Fifty‐eight submissions from 12 unique users were received for this Challenge. Out of these 12 submissions, 9 submissions passed the first stage of evaluation and were eligible for ranking. The top‐ranked submission employed a conditional latent diffusion model, whereas the joint runners‐up employed a generative adversarial network, followed by another network for image superresolution. In general, we observed that the overall ranking of the top 9 submissions according to our evaluation method (i) did not match the FID‐based ranking, and (ii) differed with respect to individual feature families. Another important finding from our additional analyses was that different DGMs demonstrated similar kinds of artifacts. Conclusions This Grand Challenge highlighted the need for domain‐specific evaluation to further DGM design as well as deployment. It also demonstrated that the specification of a DGM may differ depending on its intended use.
Crucial to many measurements at the LHC is the use of correlated multi-dimensional information to distinguish rare processes from large backgrounds, which is complicated by the poor modeling of many of the crucial backgrounds in Monte Carlo simulations. In this work, we introduce HI-SIGMA, a method to perform unbinned high-dimensional statistical inference with data-driven background distributions. In contradistinction to many applications of Simulation Based Inference in High Energy Physics, HI-SIGMA relies on generative ML models, rather than classifiers, to learn the signal and background distributions in the high-dimensional space. These ML models allow for interpretable inference while also incorporating model errors and other sources of systematic uncertainties. We showcase this methodology on a simplified version of a di-Higgs measurement in the bbγγ final state, where the di-photon resonance allows for background interpolation from sidebands into the signal region. We demonstrate that HI-SIGMA provides improved sensitivity as compared to standard classifier-based methods, and that systematic uncertainties can be straightforwardly incorporated by extending methods which have been used for histogram based analyses.
Quantifying parametric uncertainty using observations from individual sites provides a critical foundation for Earth system modeling, serving as a necessary first step before scaling up to regional or global applications. This study introduces a novel computational framework designed to enhance model predictability by reducing parametric uncertainty and assessing site and observable generalizability using various observational constraints. The framework integrates five components: Model Simulation, Statistical Emulation, Global Sensitivity Analysis (GSA), Model Calibration, and Model Prediction. Using the E3SM land model, we simulated site-level land-atmosphere carbon and energy fluxes from 2003 to 2007 across five evergreen needleleaf FLUXNET sites, perturbing 26 vegetation-related model parameters. Gaussian process emulators were employed to expedite GSA and model calibration. Four critical parameters that strongly influence selected land-atmosphere fluxes were identified by GSA. Bayesian approaches were used to infer parameter probability distributions leveraging synthetic data and FLUXNET observations. The results reveal that posterior parameter distributions vary significantly across different sites and observables within the same plant functional type. Probabilistic predictions indicate that parameters calibrated at one site can enhance predictive accuracy at other sites, although site heterogeneity may sometimes outweigh parametric uncertainty. Additionally, the probabilistic predictions demonstrate that calibration for one variable can also improve predictability for other variables, thereby maximizing predictive capabilities with limited observations. This framework provides a powerful approach for reducing parametric uncertainty in Earth system models and deepening our understanding of carbon dynamics and energy cycles. Its adaptability makes it a valuable tool for broader applications in Earth system modeling.
The characteristics of the hadron-to-quark first-order phase transition differ depending on whether charge neutrality is locally or globally fulfilled. In 𝛽-equilibrated matter, these two possibilities correspond to the Maxwell and Gibbs constructions. Recently, we presented a new framework in which a continuously varying parameter allows one to describe a first-order phase transition in intermediate scenarios to the two extremes of fully local and fully global charge neutrality. In this work, we extend the previous framework to finite temperatures and out-of-𝛽 equilibrium conditions, making it available for simulations of core-collapse supernovae and binary neutron star mergers. We investigate its impact on key thermodynamic quantities across a range of baryon densities, temperatures, and electron fractions. We find that when matter is not in 𝛽 equilibrium, the pressure in the mixed phase is not constant even for the case of fully local charge neutrality. Moreover, we compute the thermal index using three different approaches, demonstrating that the finite-temperature extension of an equation of state using a constant thermal index can be ill defined when applied to the mixed phase.
We calculate the magnetic dipole $\gamma$-ray strength functions in a chain of even-mass neodymium isotopes $^{144-152}$Nd in the framework of the configuration-interaction (CI) shell model. We infer the strength function by applying the maximum entropy method (MEM) to the exact imaginary-time response function calculated with the shell-model Monte Carlo (SMMC) method. The success of the MEM depends on the choice of a good strength function as a prior distribution. We investigate two choices for the prior strength function: the static path approximation (SPA) and the quasiparticle random-phase approximation (QRPA). We find that the QRPA is a better approximation at low temperatures (i.e., near the ground state), while the SPA is a better choice at finite temperatures. We identify a low-energy enhancement (LEE) in the MEM deexcitation $M1$ strength functions of the even-mass neodymium isotopes and compare with recent experimental results for the total deexcitation $\gamma$-ray strength functions. The LEE is already seen in the SPA strength function but not in the QRPA strength function, indicating the importance of large-amplitude static fluctuations around the mean field in reproducing the LEE. Our method is currently the only one which can reproduce LEE in heavy open-shell nuclei where conventional CI shell model calculations are prohibited. With the onset of deformation as number of neutrons increases along the chain of neodymium isotopes, we observe that some of the LEE strength transfers to a low-energy excitation, which we interpret as a finite-temperature ``scissors'' mode. Here, we also observe a finite-temperature spin-flip mode.
We have achieved the proposed goal to experimentally probe the AB + CD and AB + C types of reactions with state-to-state resolution, which we also compared to advanced theoretical calculations to help elucidate the role of quantum mechanics in the processes of bond breakage and formation. Our approach uses reactants that are prepared at ultracold temperatures (< 1µK) such that the quantum effects of translational motion are an important factor. Specific example reactions, including the potassium-rubidium metathesis reaction KRb + KRb → K 2 + Rb 2 as well as the atom exchange reaction Rb + KRb → Rb2 + K, are chosen because the technology of quantum internal and motional state control of these types of molecules is particularly advanced. The results for the entire funding period are fruitful. For the majority of this grant, we have constructed a one-of-the-kind quantum degenerate gas apparatus that integrates ion detection and velocity map imaging capabilities, allowing us to explore the KRb + KRb → K 2 + Rb 2 bimolecular reaction in detail. Specifically, we first verified such a reaction indeed proceed at ultracold temperatures by direct detection of reaction products. We then mapped out the complete product state distribution, which was compared to a state-counting model based on statistical theory. Our results show an overall agreement with the statistical state counting model, but also reveal several deviating state-pairs. An exact quantum calculation for molecule-molecule collisions, that is needed to understand these deviations, is however beyond the current state-of-the-art. Beside scrutinizing the reaction products, we also directly observe the reaction intermediate complex, which was quite a surprise to us. The intermediate complexes are long-lived and can interact with the inferred light that we use to trap the ultracold gas. After molecule-molecule collisions, we then explored the more theoretically tractable Rb + KRb reaction, which is endothermic. Surprisingly, we observed an exceedingly long-lived KRb$^*_2$ collisional complexes, with our experimentally measured complex lifetime deviating from conventional theoretical calculations by five orders of magnitude. This discrepancy has motivated many explorations of possible underlying causes, though no model yet captures this phenomenon completely. In the final year and the work that continues today, we extend upon these atom-molecule collision experiments to explore the origin of the long-lived KRb$^*_2$ complex lifetime and develop means to control the outcome of the reaction complex. The 5-year funded work advanced our understanding of chemical reactions at the lowest possible temperatures and at the same time opened up many new questions that are beyond our initial imaginations.
Laboratory plasma production almost always preferentially heats either the ions or electrons, leading to a two-temperature state. In this state, density functional theory molecular dynamic simulation is the state of the art for modeling bulk material properties. We construct a statistical mechanics model for the two temperature limit that is theoretically consistent with the molecular dynamics method. We proceed to derive the electron-ion multi-temperature quantum Ornstein-Zernike equations for the first time. This allows the construction of a two-temperature two-component plasma model using the average atom from which we can compute bulk material properties at a fraction of the computation time of the two-temperature density functional theory simulation. The accuracy of the model is benchmarked against ion pair correlation and self-diffusion results from ab initio simulation. Here, we proceed to compute the viscosity and ion thermal conductivity as a function of both ion and electron temperature.
In this letter, measurements of (anti)alpha production in central (0–10%) Pb–Pb collisions at a center-of-mass energy per nucleon–nucleon pair of $\sqrt{s_{NN}} = 5.02$ TeV are presented, including the first measurement of an antialpha transverse-momentum spectrum. Owing to its large mass, the production of (anti)alpha is expected to be sensitive to different particle production models. The production yields and transverse-momentum spectra of nuclei are of particular interest because they provide a stringent test of these models. The averaged antialpha and alpha spectrum is compared to the spectra of lighter particles, by including it into a common blast-wave fit capturing the hydrodynamic-like flow of all particles. This fit is indicating that the (anti)alpha also participates in the collective expansion of the medium created in the collision. A blast-wave fit including only protons, (anti)alpha, and other light nuclei results in a similar flow velocity as the fit that includes all particles. A similar flow velocity, but a significantly larger kinetic freeze-out temperature is obtained when only protons and light nuclei are included in the fit. The coalescence parameter B 4 is well described by calculations from a statistical hadronization model but significantly underestimated by calculations assuming nucleus formation via coalescence of nucleons. Similarly, the (anti)alpha-to-proton ratio is well described by the statistical hadronization model. On the other hand, coalescence calculations including approaches with different implementations of the (anti)alpha substructure tend to underestimate the data.
This research aims to develop a framework for establishing the correlation between in-situ monitoring data, process parameters, and microstructure evolution in blown-powder laser-directed energy deposition (DED) additive manufacturing (AM). To achieve this, a comprehensive manufacturing framework has been developed, spanning from in-situ data acquisition, melt-pool simulation, microstructure modeling, and statistical microstructure quantification. A machine learning-based surrogate model is constructed to predict melt pool geometry directly from in-situ coaxial camera data. The surrogate model is trained using outputs from a high-fidelity melt pool simulation, which provides accurate melt pool dimension data under varying process conditions. The predicted melt pool geometry is then used as input to a microstructure model to predict microstructural features. To rigorously compare and analyze microstructures, the project introduces statistical metrics that quantify differences based on key features such as morphology and texture. Microstructures are represented using advanced statistical descriptors including angular chord length distribution, two-point spatial statistics, orientation distribution function, and global spherical harmonic. These representations are used to compute four distinct “dissimilarity scores” that quantitatively capture differences in texture and morphology. This framework is demonstrated to enable automated calibration of simulation parameters by minimizing discrepancies between simulated and target microstructures. The technology developed in this project enables direct correlation between in-situ monitoring data and resulting microstructure, paving the way for adaptive microstructure control in metal AM. This capability strengthens the connection between process parameters and final material properties, facilitating more precise and reliable material design.
Not provided.
Explore the source record for details and available documents.