Search NASA⌕ Search

SEARCH · Search NASA

Results for “data model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

Modeling aerosol bolus inhalations in the human lung with the multiple path particle deposition model: Comparison with experimental data

Existing one-dimensional (1D) models of aerosol dosimetry often ignore mixing mechanisms of inhaled aerosols during their transport in the lung. This mixing or aerosol dispersion results from different physical mechanisms in different regions of the lung. It is a higher order effect, which cannot be directly captured in 1D modeling approaches, and thus is sometimes modeled as a diffusive process. Here, in this study, we improved our recently developed alveolar mixing module incorporated in the multiple path particle dosimetry model (MPPD) to account for flow irreversibility and particle trapping in the alveolar spaces, as well as mixing occurring in the tracheobronchial region. This new version of MPPD was coupled with CFPD-based predictions of aerosol bolus dispersion in the oral airway. The model was used to predict the deposition, dispersion, and mode shift of aerosol bolus inhaled at different penetration depths within the lung for breathing patterns and particle size matching those used in a previous experimental study (Darquenne et al., 2016). Even though a quite simplified approach was used, the computations appear to describe subject-specific and test-specific experimental data reasonably well. The proposed combined dispersion-deposition model can be a useful tool for targeted drug delivery and also for exposure health risk assessment.

MPPD↗

Global tuning of hadronic interaction models with accelerator-based and astroparticle data

In high-energy and astroparticle physics, event generators play an essential role, even in the simplest data analyses. As analysis techniques become more sophisticated, e.g. based on deep neural networks, their correct description of the observed event characteristics becomes even more important. Physical processes occurring in hadronic collisions are simulated within a Monte Carlo framework. A major challenge is the modeling of hadron dynamics at low momentum transfer, which includes the initial and final phases of every hadronic collision. QCD-inspired phenomenological models used for these phases cannot guarantee completeness or correctness over the full phase space. These models usually include parameters which must be tuned to suitable experimental data. Until now, event generators have been developed and tuned mainly on the basis of data from high-energy physics experiments at accelerators. The wealth of data available from the latest generation of astroparticle experiments has not yet been fully exploited, and in many cases is not satisfactorily described. Both kinds of data sets are complementary as astroparticle experiments provide sensitivity especially to hadrons produced nearly parallel to the collision axis and cover center-of-mass energies up to several hundred TeV, well beyond those reached at colliders so far. In this report, we provide an overview of state-of-the-art event generators and their tuning, including the most relevant inputs from high-energy accelerator and astroparticle experiments. We present a road map that shows, for the first time, how the unified tuning of event generators with accelerator-based and astroparticle data can be performed.

Albrecht, J. [Ruhr U., Bochum, RAPP Ctr.; Ruhr U.,↗

ARM Data-Oriented Metrics and Diagnostics Package for Climate Model Evaluation

A Python-based metrics and diagnostics package is currently being developed by the U.S. Department of Energy (DOE) Atmospheric Radiation Measurement (ARM) Infrastructure Team at Lawrence Livermore National Laboratory (LLNL) to facilitate the use of long-term, high-frequency measurements from the ARM Facility in evaluating the regional climate simulation of clouds, radiation, and precipitation. This metrics and diagnostics package computes climatological means of targeted climate model simulation and generates tables and plots for comparing the model simulation with ARM observational data. The Coupled Model Intercomparison Project (CMIP) model data sets are also included in the package to enable model intercomparison as demonstrated in Zhang et al. (2017). The mean of the CMIP model can serve as a reference for individual models. Basic performance metrics are computed to measure the accuracy of mean state and variability of climate models. The evaluated physical quantities include cloud fraction, temperature, relative humidity, cloud liquid water path, total column water vapor, precipitation, sensible and latent heat fluxes, and radiative fluxes, with plan to extend to more fields, such as aerosol and microphysics properties. Process-oriented diagnostics focusing on individual cloud- and precipitation-related phenomena are also being developed for the evaluation and development of specific model physical parameterizations. The version 1.0 package is designed based on data collected at ARM’s Southern Great Plains (SGP) Research Facility, with the plan to extend to other ARM sites. The metrics and diagnostics package is currently built upon standard Python libraries and additional Python packages developed by DOE (such as CDMS and CDAT). The ARM metrics and diagnostic package is available publicly with the hope that it can serve as an easy entry point for climate modelers to compare their models with ARM data. In this report, we first present the input data, which constitutes the core content of the metrics and diagnostics package in section 2, and a user's guide documenting the workflow/structure of the version 1.0 codes, and including step-by-step instruction for running the package in section 3.

54 ENVIRONMENTAL SCIENCES↗

Confronting Large‐Eddy Simulations With Stereo Camera Data by Means of Reconstructed Hemispheric Cloud Size Distributions

High-resolution hemispheric camera images at a meteorological site in western Germany are used to analyze the multi-dimensional spatial characteristics of continental cumulus cloud fields, and to evaluate Large-Eddy Simulations on this aspect. Traditional non-hemispheric cloud-detecting instruments provide additional reference data. The main model-observation comparison focuses on cloud size distributions (CSDs), employing two methods: (a) directly using three-dimensional model fields, direct CSDs, and (b) using rendered hemispheric images of the model fields as produced by a camera simulator based on path-tracing. In the latter method, both the real and rendered images are used to three-dimensionally reconstruct the cloud fields, yielding hemispheric CSDs. Advantages of hemispheric comparisons over more classic approaches include (a) fair comparisons between model and data, and (b) full use of the enhanced resolutions and hemispheric spatial coverage of the camera imagery. Basic evaluation of the simulations demonstrates good agreement on thermodynamic structure and its diurnal cycle. Cloud heights and cloud cover are intercompared between the model, camera data and other instrumentation, providing insight into their structural differences. A consistent alignment is found between the hemispheric CSDs from both the model and the cameras. Power law fits reveal structurally lower exponents in hemispheric CSDs compared to non-hemispheric CSDs, which particularly caution against directly comparing hemispheric CSDs to non-hemispheric distributions. This result is robust for sample size and fitting method. These findings inform future use of hemispheric camera systems for studying cumulus cloud field morphology and model evaluation.

54 ENVIRONMENTAL SCIENCES↗

Development of a data-driven neural network model for electron thermal transport in NSTX

A data-driven electron thermal transport neural network (ETT-NN) model, trained on TRANSP interpretative analysis results of National Spherical Torus Experiment (NSTX), was developed to enable faster and more accurate ETT computation for spherical tokamaks (STs). The model incorporates both convolutional NNs and recurrent NNs, allowing it to simultaneously account for the spatial and temporal non-localities and multi-scale features of turbulent transport, which have been considered only in a limited manner in conventional models. The model was validated through interpretative analysis and predictive simulations using Tokamak Reactor Integrated Automated Suite for Simulation and Computation, demonstrating relatively high accuracy. Additionally, parameter scans were performed on test discharges known to exhibit specific turbulent modes, such as microtearing mode, trapped electron mode, kinetic ballooning mode, and electron temperature gradient mode. The scanning results revealed that the ETT-NN model exhibits the same trends as those observed in conventional gyrokinetic simulations or theories, while also capturing the global nature of turbulent transport, indicating that the data-driven model accurately reflects the underlying physical characteristics. Furthermore, due to the dimensionless nature of the model, we can feasibly expand its applicability by incorporating data from other devices and uncovering the characteristics of ETT in STs in the future.

NSTX↗

Machine learning approaches for influenza A virus risk assessment identifies predictive correlates using ferret model in vivo data

In vivo assessments of influenza A virus (IAV) pathogenicity and transmissibility in ferrets represent a crucial component of many pandemic risk assessment rubrics, but few systematic efforts to identify which data from in vivo experimentation are most useful for predicting pathogenesis and transmission outcomes have been conducted. To this aim, we aggregated viral and molecular data from 125 contemporary IAV (H1, H2, H3, H5, H7, and H9 subtypes) evaluated in ferrets under a consistent protocol. Three overarching predictive classification outcomes (lethality, morbidity, transmissibility) were constructed using machine learning (ML) techniques, employing datasets emphasizing virological and clinical parameters from inoculated ferrets, limited to viral sequence-based information, or combining both data types. Among 11 different ML algorithms tested and assessed, gradient boosting machines and random forest algorithms yielded the highest performance, with models for lethality and transmission consistently better performing than models predicting morbidity. Comparisons of feature selection among models was performed, and highest performing models were validated with results from external risk assessment studies. Our findings show that ML algorithms can be used to summarize complex in vivo experimental work into succinct summaries that inform and enhance risk assessment criteria for pandemic preparedness that take in vivo data into account.

59 BASIC BIOLOGICAL SCIENCES↗

Dynamic data-driven multiscale modeling for predicting the degradation of a 316L stainless steel nuclear cladding material

Here, we have developed a long short-term memory stacked ensemble (LSTM-SE) surrogate modeling approach that can provide rapid predictions of microstructural evolution and the resultant mechanical properties of American Iron and Steel Institute (AISI) 316L series stainless steel (316LSS) fuel cladding under conditions of varying temperature and radiation dose rate. To acquire training data, we developed and implemented a kinetic Monte Carlo (KMC) model to simulate precipitation kinetics of M 23 C 6 , γ', and G phases within SS316L cladding. Experimentally reported precipitation kinetics of SS316L in literature were linked to the kinetic parameters of the simulated precipitation in our KMC model. The model was then used to simulate microstructure evolution under synthetically generated treatments of varying temperature and radiation dose rate, for periods of up to 3000 hours. Changes in volume fraction, number density, and particle size of precipitates were recorded, and particle area fractions were correlated using statistical methods to develop the surrogate model. Simultaneously, the mechanical properties of the simulated microstructures were evaluated using microstructure-based finite element method (FEM) analysis to determine the elastic modulus, yield stress, ultimate tensile strength, and elongation to failure of the aged microstructures. Using this approach, our surrogate model can predict precipitation behavior within 0.25% volume fraction and mechanical properties within 6% relative error from the values predicted by the KMC and FEM models using 50 training simulations as input. The trained recurrent neural network-based model can return estimations of precipitation kinetics and mechanical properties ~1000 times faster than the physics-based codes. This work demonstrates, as a proof of concept, that reactor material service lifetimes under variable service conditions can be predicted for a statistics-based model from a practicably obtainable dataset.

36 MATERIALS SCIENCE↗

AERO-MAP: a data compilation and modeling approach to understand spatial variability in fine- and coarse-mode aerosol composition

Abstract. Aerosol particles are an important part of the Earth climate system, and their concentrations are spatially and temporally heterogeneous, as well as being variable in size and composition. Particles can interact with incoming solar radiation and outgoing longwave radiation, change cloud properties, affect photochemistry, impact surface air quality, change the albedo of snow and ice, and modulate carbon dioxide uptake by the land and ocean. High particulate matter concentrations at the surface represent an important public health hazard. There are substantial data sets describing aerosol particles in the literature or in public health databases, but they have not been compiled for easy use by the climate and air quality modeling community. Here, we present a new compilation of PM2.5 and PM10 surface observations, including measurements of aerosol composition, focusing on the spatial variability across different observational stations. Climate modelers are constantly looking for multiple independent lines of evidence to verify their models, and in situ surface concentration measurements, taken at the level of human settlement, present a valuable source of information about aerosols and their human impacts complementarily to the column averages or integrals often retrieved from satellites. We demonstrate a method for comparing the data sets to outputs from global climate models that are the basis for projections of future climate and large-scale aerosol transport patterns that influence local air quality. Annual trends and seasonal cycles are discussed briefly and are included in the compilation. Overall, most of the planet or even the land fraction does not have sufficient observations of surface concentrations – and, especially, particle composition – to characterize and understand the current distribution of particles. Climate models without ammonium nitrate aerosols omit ∼ 10 % of the globally averaged surface concentration of aerosol particles in both PM2.5 and PM10 size fractions, with up to 50 % of the surface concentrations not being included in some regions. In these regions, climate model aerosol forcing projections are likely to be incorrect as they do not include important trends in short-lived climate forcers.

Mahowald, Natalie M. (ORCID:000000022873997X)↗

IM3 Projected US Data Center Locations

IM3 Projected US Data Center Locations This dataset contains model projections of new data center facilities in the contiguous United States (CONUS) through 2035 using the CERF – Data Centers model. Data center locations are modeled across four data center electricity demand growth scenarios (low, moderate, high, higher) and five market gravity scenarios (0%, 25%, 50%, 75%, 100%). Projected locations are intended to be regional representations of feasible siting locations in the future to assess potential grid and water stress impacts. The data center load growth scenarios correspond with the rates outlined in EPRI (2024) and include 3.71%, 5%, 10%, and 15% annual growth of electricity demand for data centers from 2023 values in 37 states across the CONUS. Market gravity scenarios correspond to the relative importance of proximity to data center markets or high population areas compared to locational cost in the siting algorithm. 0% market gravity means that siting decisions were entirely determined by the locational cost in each feasible location. 100% market gravity means that only market proximity was considered when siting. Other scenarios have weight placed on both components where total weight always equals 100%. Locational cost is dependent on facility cooling type and corresponding electricity cost, taxes, and other factors. Facility cooling type is spatially determined where high water stress and/or areas with high summer wet bulb temperatures are assumed to operate with mechanical cooling for a higher fraction of the year rather than evaporative cooling. Feasible data center siting areas are based on geospatial suitability raster data developed with open-source information. The following areas are excluded from siting: Areas within 300 m of a federal airport runway Waterbodies Areas with slope >16% Areas susceptible to sinkholes High coastal or inland flood risk areas Local, state, and federal parks, leisure areas, and cemeteries Areas >2 km away from electric substations Areas >5 km away from a municipal water supplier service area Areas >2 km away from high-speed fiber provider service territory Protected Areas Database of the United States (PAD-US) areas Railroads, major roadways, and minor roadways Military areas and training grounds NLCD developed lands Areas >0.8 km (0.5 miles) from NLCD developed lands Because we use open-source information, proprietary information that can influence siting decisions such as individual tax agreements with cities, detailed fiber line connectivity, electric grid power capacity agreements, and others, are not currently accounted for in the modeling process. Using specific building locations and footprints in the dataset for local planning purposes is not advised. Technical Information Geospatial data is provided in geojson format using the Albers Equal Area Conic (ESRI:102003) coordinate reference system. The datasets contain the following parameters: id - unique identification number within given scenario file growth_scenario – data center demand growth scenario market_gravity_weight – market gravity weight scenario (%) region – name of region (i.e., US State) total_cost_million_usd – locational siting cost ($million) campus_size_square_ft – total land acquired for data center facility (square ft) data_center_it_power_mw – IT power of data center facility (MW) mechanical_cooling_frac – fraction of year when data center uses mechanical cooling system water_cooling_frac– fraction of year when data center uses evaporative cooling system cooling_energy_demand_mwh – total annual facility energy demand for cooling (MWh) cooling_water_demand_mgy – total annual facility water demand for cooling (MG) cooling_water_consumption_mgy – total annual facility water consumed (MG) normalized_locational_cost – normalized total locational cost score for location normalized_gravity_score – normalized market gravity score for location weighted_siting_score – total weighted siting score of locational cost and gravity score geometry – polygon geometry of facility Acknowledgment IM3 is a multi-institutional effort led by Pacific Northwest National Laboratory and supported by the U.S. Department of Energy's Office of Science as part of research in MultiSector Dynamics, Earth and Environmental Systems Modeling Program. License This data is made available under a CCBY4.0 License Disclaimer This material was prepared as an account of work sponsored by an agency of the United States Government. Neither the United States Government nor the United States Department of Energy, nor the Contractor, nor any or their employees, nor any jurisdiction or organization that has cooperated in the development of these materials, makes any warranty, express or implied, or assumes any legal liability or responsibility for the accuracy, completeness, or usefulness or any information, apparatus, product, software, or process disclosed, or represents that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by trade name, trademark, manufacturer, or otherwise does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States Government or any agency thereof, or Battelle Memorial Institute. The views and opinions of authors expressed herein do not necessarily state or reflect those of the United States Government or any agency thereof. PACIFIC NORTHWEST NATIONAL LABORATORYoperated byBATTELLEfor theUNITED STATES DEPARTMENT OF ENERGYunder Contract DE-AC05-76RL01830

Mongird, Kendall (ORCID:0000000328077088)↗

IM3 Projected US Data Center Locations

IM3 Projected US Data Center Locations This dataset contains model projections of new data center facilities in the contiguous United States (CONUS) through 2035 using the CERF – Data Centers model. Data center locations are modeled across four data center electricity demand growth scenarios (low, moderate, high, higher) and five market gravity scenarios (0%, 25%, 50%, 75%, 100%). Projected locations are intended to be regional representations of feasible siting locations in the future to assess potential grid and water stress impacts. The data center load growth scenarios correspond with the rates outlined in EPRI (2024) and include 3.71%, 5%, 10%, and 15% annual growth of electricity demand for data centers from 2023 values in 37 states across the CONUS. Market gravity scenarios correspond to the relative importance of proximity to data center markets or high population areas compared to locational cost in the siting algorithm. 0% market gravity means that siting decisions were entirely determined by the locational cost in each feasible location. 100% market gravity means that only market proximity was considered when siting. Other scenarios have weight placed on both components where total weight always equals 100%. Locational cost is dependent on facility cooling type and corresponding electricity cost, taxes, and other factors. Facility cooling type is spatially determined where high water stress and/or areas with high summer wet bulb temperatures are assumed to operate with mechanical cooling for a higher fraction of the year rather than evaporative cooling. Feasible data center siting areas are based on geospatial suitability raster data developed with open-source information. The following areas are excluded from siting: Areas within 300 m of a federal airport runway or within an airport area boundary Waterbodies Areas with slope >16% Areas susceptible to sinkholes High coastal or inland flood risk areas Local, state, and federal parks, leisure areas, and cemeteries Areas >2 km away from electric substations Areas >5 km away from a municipal water supplier service area Areas >2 km away from high-speed fiber provider service territory USGS Protected Areas Database of the United States (PAD-US) GAP status 1, 2, or 3 areas US National Parks Wetlands USFWS critical habitats BIA land areas Railroads, major roadways, and minor roadways Military areas and training grounds NLCD developed lands Areas >0.8 km (0.5 miles) from NLCD developed lands Because we use open-source information, proprietary information that can influence siting decisions such as individual tax agreements with cities, detailed fiber line connectivity, electric grid power capacity agreements, and others, are not currently accounted for in the modeling process. Using specific building locations and footprints in the dataset for local planning purposes is not advised. Technical Information Geospatial data is provided in geojson format using the Albers Equal Area Conic (ESRI:102003) coordinate reference system. The datasets contain the following parameters: id - unique identification number within given scenario file growth_scenario – data center demand growth scenario market_gravity_weight – market gravity weight scenario (%) region – name of region (i.e., US State) total_cost_million_usd – locational siting cost ($million) campus_size_square_ft – total land acquired for data center facility (square ft) data_center_it_power_mw – IT power of data center facility (MW) mechanical_cooling_frac – fraction of year when data center uses mechanical cooling system water_cooling_frac– fraction of year when data center uses evaporative cooling system cooling_energy_demand_mwh – total annual facility energy demand for cooling (MWh) cooling_water_demand_mgy – total annual facility water demand for cooling (MG) cooling_water_consumption_mgy – total annual facility water consumed (MG) normalized_locational_cost – normalized total locational cost score for location normalized_gravity_score – normalized market gravity score for location weighted_siting_score – total weighted siting score of locational cost and gravity score geometry – polygon geometry of facility Acknowledgment IM3 is a multi-institutional effort led by Pacific Northwest National Laboratory and supported by the U.S. Department of Energy's Office of Science as part of research in MultiSector Dynamics, Earth and Environmental Systems Modeling Program. License This data is made available under a CCBY4.0 License Disclaimer This material was prepared as an account of work sponsored by an agency of the United States Government. Neither the United States Government nor the United States Department of Energy, nor the Contractor, nor any or their employees, nor any jurisdiction or organization that has cooperated in the development of these materials, makes any warranty, express or implied, or assumes any legal liability or responsibility for the accuracy, completeness, or usefulness or any information, apparatus, product, software, or process disclosed, or represents that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by trade name, trademark, manufacturer, or otherwise does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States Government or any agency thereof, or Battelle Memorial Institute. The views and opinions of authors expressed herein do not necessarily state or reflect those of the United States Government or any agency thereof. PACIFIC NORTHWEST NATIONAL LABORATORYoperated byBATTELLEfor theUNITED STATES DEPARTMENT OF ENERGYunder Contract DE-AC05-76RL01830

Mongird, Kendall (ORCID:0000000328077088)↗

Numerical Simulation of Biogenic Fluid Catalytic Cracking (BFCC) Regenerators at Different Scales with MFIX-Exa

Catalytic Fast Pyrolysis (CFP) is a process that converts biomass into liquid intermediates suitable for transportation fuels by rapidly heating it in the presence of a catalyst, aiming to produce stable oils with reduced oxygen content. During CFP, the catalyst can become deactivated by the accumulation of coke, a carbon-rich deposit formed from the decomposition of biomass components. Unlike in petroleum refining, regenerating coked catalysts from biomass pyrolysis requires specific approaches due to the different chemical nature of the coke formed. An experimental technique, Temperature Programmed Oxidation (TPO), was used to study the de-coking process by gradually increasing temperature while monitoring the production of CO and CO2, which provides data for kinetic modeling. Utilizing data from TPO experiments, coke combustion kinetic model was developed to describe the rate of coke removal at different temperatures, allowing for simulation of regeneration processes. Then kinetic model is integrated into MFIX-Exa for the simulation of Biogenic Fluid Catalytic Cracker (BFCC) regenerator at different scales, enabling analysis of catalyst flow, temperature distribution, and regeneration efficiency under various operating conditions.

biogenic fluid catalytic cracking↗

An Image-Plane Approach to Gravitational Lens Modeling of Interferometric Data

Strong gravitational lensing acts as a cosmic telescope, enabling the study of the high-redshift universe. Astronomical interferometers, such as the Atacama Large Millimeter/submillimeter Array (ALMA), have provided high-resolution images of strongly lensed sources at millimeter and submillimeter wavelengths. To model the mass and light distributions of lensing and source galaxies from strongly lensed images, strong lens modeling for interferometric observations is conventionally performed in the visibility space, which is computationally expensive. In this paper, we implement an image-plane lens modeling methodology for interferometric dirty images by accounting for noise correlations. We show that the image-plane likelihood function produces accurate model values when tested on simulated ALMA observations with an ensemble of noise realizations. We also apply our technique to ALMA observations of two sources selected from the South Pole Telescope survey, comparing our results with previous visibility-based models. Our model results are consistent with previous models for both parametric and pixelated source-plane reconstructions. We implement this methodology for interferometric lens modeling in the open-source software package lenstronomy.

Zhang, Nan [Illinois U., Urbana (main)] (ORCID:000↗

Report on the sensitivity kernel construction and updating the SPiRaL model with waveform data

The WAVEFORMS Initiative in the Ground-based Nuclear Detonation Detection (GNDD) program includes research leading towards the prediction of entire seismic and acoustic waveforms produced by natural and manmade events, including explosions. A key aspect of this research is the development of Earth (seismic) models at multiple scales including crustal, regional, and global. The work described here pertains to the global-scale seismic tomography effort led by LLNL.

58 GEOSCIENCES↗

Developing Data-Driven Synthetic Infrastructure Models for Resilience Analysis

Research on infrastructure resilience has produced promising methods to simulate and optimize complex networks to improve performance. However, restrictions on sharing infrastructure models and the steep cost of developing and maintaining infrastructure models presents a roadblock to adoption. To overcome this limitation, this research focuses on methods to create data-driven infrastructure models that will help improve infrastructure resilience and security. The analysis couples incomplete utility data, geospatial data, machine learning, and synthetic network generation methods to rapidly develop and update infrastructure models. The methods are validated using realistic utility models and site-specific data, with a focus on Puerto Rico due to its unique infrastructure challenges and available data. This research highlights promising opportunities for the use of synthetic network generation and machine learning to create infrastructure models when very little data is available. Results demonstrate that hybrid methods, which combine sparse utility data with synthetic models, can enhance model accuracy, and machine learning can predict model attributes using training data from other models. However, the complexity of infrastructure systems means that even minor changes in network connectivity can significantly impact simulation results. Resilience analysis using synthetic infrastructure models shows that while some system behaviors are preserved, the magnitude of disruptions may not be accurately represented, indicating the need for more research and validation before using synthetic models for critical infrastructure investment decisions. The framework outlined in this report represents a significant advance to infrastructure model development and could be applied to additional domains and sites. Future research will continue to streamline and validate methods to help reduce roadblocks to resilience analysis.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Data-driven nonlocal model for fragmentation in the crushing of solids

A technique is proposed for reproducing particle size distributions in three-dimensional simulations of the crushing and comminution of solid materials. The method is designed to produce realistic distributions over a wide range of loading conditions, especially for small fragments. In contrast to most existing methods, the new model does not explicitly treat the small-scale process of fracture. Instead, it uses measured fragment distributions from laboratory tests as the basic material property that is incorporated into the algorithm, providing a data-driven approach. The algorithm is implemented within a nonlocal peridynamic solver, which simulates the underlying continuum mechanics and contact interactions between fragments after they are formed. Finally, the technique is illustrated in reproducing fragmentation data from drop weight testing on sandstone samples.

58 GEOSCIENCES↗

Carbon-phosphorus cycle models overestimate CO 2 enrichment response in a mature Eucalyptus forest

The importance of phosphorus (P) in regulating ecosystem responses to climate change has fostered P-cycle implementation in land surface models, but their CO 2 effects predictions have not been evaluated against measurements. Here, we perform a data-driven model evaluation where simulations of eight widely used P-enabled models were confronted with observations from a long-term free-air CO 2 enrichment experiment in a mature, P-limited Eucalyptus forest. We show that most models predicted the correct sign and magnitude of the CO 2 effect on ecosystem carbon (C) sequestration, but they generally overestimated the effects on plant C uptake and growth. We identify leaf-to-canopy scaling of photosynthesis, plant tissue stoichiometry, plant belowground C allocation, and the subsequent consequences for plant-microbial interaction as key areas in which models of ecosystem C-P interaction can be improved. Together, this data-model intercomparison reveals data-driven insights into the performance and functionality of P-enabled models and adds to the existing evidence that the global CO 2 -driven carbon sink is overestimated by models.

54 ENVIRONMENTAL SCIENCES↗

Calibration of urban building energy model using smart meter data for district peak load prediction

Urban building energy modeling (UBEM) is a powerful approach to assessing baseline building energy performance and retrofits with new technologies across building stocks in cities. However, the accuracy of UBEM is often constrained by the limited availability of reliable data about building characteristics and operations, such as envelope efficiency levels, HVAC system performance, and end-use load patterns. Existing research has performed UBEM calibration using annual or monthly energy consumption data, which falls short when higher-resolution time series applications are needed, such as peak load prediction for utility operation planning. This study presents a new framework for calibrating building energy models at urban scale using smart meter data, targeting the accurate prediction of summer peak electricity loads to support robust grid planning. The framework first integrates various data sources to enhance baseline input assumptions for building models, and then calibrates the baseline models through a pattern-matching approach. A case study using CityBES and two years of AMI data from over 9000 residential customers in Portland, Oregon, demonstrated the workflow and its effectiveness. The calibrated models achieved a daily peak load mean absolute percentage error of 2.6 % during the heatwave in the calibration year, and 2.0 % in the validation year using another year of AMI data. Using the calibrated models, we analyzed the demand flexibility potential of the district building stock as an application of UBEM calibration. The findings affirm the appropriate use of UBEM for peak electric load forecasting and demand side management at the utility distribution system level.

AMI data↗

Prototype-Wise Sensitivity Analysis of Urban Building Energy Simulation Surrogate Modeling Accuracy

Urban Building Energy Modeling (UBEM) is an important reference for urban energy-related policymaking. Because of the significant impact of urban microclimates on the energy simulation, UBEM requires simulations of many microclimate-prototype pairs. Surrogate modeling is commonly used to reduce the cost of simulation computations. In UBEM surrogate modeling, it is important to determine the percentage of microclimates related to a prototype used for generating surrogate model training data. This study analyzes the prototype-wise variations and sensitivities of surrogate model estimation accuracy to the microclimate sampling ratios. The results of the study can help determine the number of simulations used for generating surrogate modeling data, avoid redundant simulations, and reduce the computational cost for UBEM surrogate modeling and its time.

Pan, Xiyu↗