Search NASA⌕ Search

SEARCH · Search NASA

Results for “data models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22

Challenges and Opportunities for Electric Utility Modeling and Asset Valuation Frameworks: Case Study on Valuing New Pumped Storage Hydropower

Asset valuation by electric utilities is becoming increasingly difficult in the rapidly changing electric sector. Rapid deployment of variable generation and inverter-based storage systems along with uncertain demand growth, climate, policies, and other factors create a challenging environment for understanding the value proposition of a new potential asset. This report describes an effort between the Tennessee Valley Authority (TVA) and three U.S. Department of Energy laboratories to perform a detailed review of utility modeling and analysis practices for asset valuation and identify challenges and opportunities for advancing its methods into the future. It focuses on a case study of new potential pumped storage hydropower (PSH) because of growing interest in new PSH capacity to provide energy balancing, firm capacity, and a range of ancillary services. Staff from the DOE labs conducted systematic interviews about current practices in capacity expansion modeling, production-cost modeling, hydrological modeling, and transmission stability modeling while also discussing how scenario analysis is conducted and how models and data are integrated. The effort resulted in a set of model, integration, and scenario recommendations that could be valuable to TVA, other utilities, system operators, and other stakeholders conducting integrated grid analysis. Individual model recommendations suggest exploring computational tradeoffs with detail and resolution across spatiotemporal structure, supply- and demand-side details, transmission overlays, market interactions, and ancillary services. Automated processes to pass data between models and conduct larger scenario suites could also enhance valuation practices by enabling a more consistent study of asset value across a broader range of uncertain future grid conditions where PSH could be particularly valuable. TVA and other industry stakeholders can learn from and adapt applied research-grade methods developed by DOE laboratories and other research institutions to improve decision making and accelerate progress towards a reliable, economic, sustainable energy system.

13 HYDRO ENERGY↗

Rare Lepton Decays and Differentiable Hadronization Models - From Signatures of New Physics to Data-driven Event Generation

This dissertation is partitioned into two parts: phenomenological studies focused on rare lepton decays as probes of heavy and light new physics, and the development of differentiable, data-driven hadronization models. Part I develops the phenomenology of new physics signatures stemming from rare charged lepton flavor violating decays probed by experiments at the intensity frontier. These include interactions mediated by both high-scale effective operators and light new physics, manifesting in multi-lepton final states ($\mu \to 5e$), elastic nuclear transitions ($\mu \to e$ conversion), baryon-number-violating muon capture, and time-dependent signals from ultralight dark matter ($\mu \to e \phi, \tau \to \ell \phi$). Part II develops two distinct strategies for advancing differentiable and data-driven hadronization models. One involves comprehensive reweighting frameworks for hadronization that enable efficient uncertainty estimation, facilitate parameter tuning, and interface naturally with differentiable programming paradigms. The other introduces machine-learning-based methods for extracting microscopic fragmentation dynamics directly from macroscopic observables through the deformation of existing models -- effectively providing solutions to the inverse problem of hadronization. Altogether, these studies advance the interpretability, flexibility, and precision of theoretical predictions for both high-intensity and high-energy experiments.

Menzo, Tony [Cincinnati U.] (ORCID:000000022013457↗

Towards provision of regularly updated climate data from the Coupled Model Intercomparison Project

The Coupled Model Intercomparison Project (CMIP) is a flagship of the World Climate Research Programme (WCRP). CMIP has become a recognised ‘brand’ in climate circles evolving over the last thirty years from a targeted research activity by a small number of climate modelling centres intercomparing their Earth System Model (ESM) simulations to a broad international coordinated research effort (Durack et al, 2025). CMIP is organized as a research activity leveraging funded and in-kind contributions from experts within modelling centres and the broader scientific community supported more recently by a fully-funded International Project Office. Within CMIP, Model Intercomparison Projects (MIPs) are community-designed to understand past, present and future climate. CMIP data provides a valuable resource for climate research and is routinely used to assess model representation of climate processes and test scientific hypotheses in the context of model uncertainty and (forced and internal) variability as evident from its prolific use in scientific publications1 . The impact relies on enabling infrastructure (most prominently via the Earth System Grid Federation (ESGF)), which allows sharing of simulation output, provision of the boundary conditions used in each simulation, and definition of the data standards that are essential to facilitating wide use of the data. The impact is supplemented by the wide-ranging scrutiny to which model simulations are subjected. Beyond its use in research, CMIP data is a key resource for communities producing derived climate information from downscaling and impact studies, such as the Coordinated Regional Downscaling Experiment (CORDEX; Gutowski et al., 2016) and the Intersectoral Impacts MIP (ISIMIP; Frieler et al., 2024). Government, academic and commercial entities also increasingly rely on CMIP and its downstream data for climate risk assessments and climate services (for example, Copernicus Climate Change Service and World Bank portal). This means that, although CMIP is a research activity, it increasingly serves a secondary and very relevant role as a provider of climate data – a long-recognised dichotomy (Stevens, 2024). Research and applications have distinct needs, with the former requiring flexibility and generality and the latter consistency. Here we explain how the design of the research activity has been adapted to reduce the burdens imposed by applications and how the research infrastructure might evolve to further enable scientific inquiry. We propose one possible approach to consistently providing model information and projections for applications in the future.

Environmental sciences↗

Use of Rig Parameter Data in Bit Constraint Models for Improved Drilling Performance at The Geysers

Surface parameter measurements are routinely used during deep well construction to monitor and guide drilling conditions for improved performance and reduced costs. However, these measurements are of reduced value without a standard to aid in evaluation and decision making. A method is demonstrated whereby drill bit constraint models are used to interpret drilling response parameters. Drill rig parameter data for well GDC-36 at the Geysers Geothermal Field Power were acquired by Geysers Power Company and drilling contractor Kenai Drilling using Pason US DataHub and evaluated. Drilling parameters are evaluated using laboratory-validated rock reduction models for predicting the phenomenological response of drag bits (Detournay and Defourny, 1992) along with other model constraints in computational algorithms. The method is used to evaluate overall bit performance, monitor bit integrity, and detect the presence of drillstring vibrations and other conditions contributing to bit failure; comparisons are made to observations of bit wear and damage. The method will be applied in real-time to improve decision-making on subsequent wells and has applicability to development of advanced analytics on future geothermal wells using real-time electronic drilling recorder (EDR) data for improved performance and reduced drilling costs.

15 GEOTHERMAL ENERGY↗

Leveraging Large Language Models for Real-World Data Evidence: A Framework for Automated Treatment Extraction and Data Harmonization

Background: The ability to comprehensively collect treatment information from cancer patient medical records would enable studies to evaluate real-world benefits and risks tied to specific treatments. Currently, it is difficult to system- atically collect high-quality treatment information because it is often stored in unstructured text. Manually extracting and standardizing drug and regimen data is time-intensive. Recent advances in large language models (LLMs) offer a potential solution for automated extraction of structured treatment information from clinical text. Objective: This study systematically evaluates the utility of four LLMs from the Llama family for automated extraction of oncology treatment information from clinical text. This information can guide researchers using cancer registry data to provide insights into cancer care and outcomes beyond clinical trials. Methods: Four instruction-tuned Llama models with varying parameter counts (1B, 3B, 8B, and 70B) were evaluated for their ability to extract treatment information from clinical documents. A unified oncology knowledge base integrating seven major public data sources was developed to standardize and normalize extracted entities—a critical step for harmonizing data from diverse sources. Extracted treatment data were compared against expert-annotated ground truth. Model performance was assessed using accuracy metrics (Precision, Recall, F1-Score) and opera- tional feasibility metrics, including processing speed and structural compliance of the output. Results: A strong positive correlation was observed between model size and extraction accuracy. F1-score improved from 0.609 for the 1B model to 0.710 (3B), 0.807 (8B), and 0.828 (70B). While larger models demonstrated superior accuracy and compliance, they incurred higher computational costs. The modest performance difference between 8B and 70B suggests diminishing returns with increasing model size. Conclusions: LLMs represent a viable technology for automating oncology treatment extraction. The 8B-parameter model emerged as a highly effective option, balancing high accuracy and computational efficiency. Selecting an appropriate LLM for deployment in cancer registries involves a trade-off between desired accuracy and available operational resources. Harmonizing extracted entities with the oncology knowledge base facilitates standardized integration into common data models, enhancing data quality for real-world evidence analyses.

artificial intelligence↗

Adsorption of terbium (III) on DGA and LN resins: Thermodynamics, isotherms, and kinetics

Two commercially available extraction chromatography (EXC) resins containing N,N,N’,N’-tetra-n-octyldiglycolamide (DGA Resin, Normal, 50 – 100 μm) and Bis(2-ethylhexyl) phosphate (LN Resin, 100 – 150 μm) were used as adsorbents to study fundamental adsorption properties such as thermodynamic values, equilibrium isotherms, and kinetic uptake models for terbium(III) adsorption. Weight distribution ratios (D w ) for terbium on DGA and LN resins were measured using a [ 160 Tb]Tb 3+ radiometric tracer in nitric acid as a function of acidity, temperature, initial analyte concentration, and equilibrium time. The D w values showed increasing binding affinity for DGA resin at high nitric acid concentrations and decreasing binding affinity for LN resins. Thermodynamic studies for DGA and LN resins revealed that the Gibbs free energy (ΔG) increased consistently with temperature. To model equilibrium data, increasingly higher parameter equilibrium isotherm models (Henry (1) < Langmuir, Freundlich (2) < Redlich-Peterson (3) < Fritz-Schluender (4)) were compared on their root mean squared errors (RMSE) and adjusted determination coefficients to determine the most applicable model. In all cases, the empirical four-parameter Fritz-Schluender isotherm demonstrated a superior fit. Similar comparisons for reaction-based kinetic models (Pseudo-first-order < Pseudo-second-order < Pseudo-n-order) revealed that the higher-order PNO model yielded a superior fit of kinetic data for both resins. Furthermore, in some cases, adsorption isotherms and kinetic models could also be modeled by a lower-order model with minimal change in error parameters. Weber-Morris plots revealed that two linear sections are observed for each resin, where the first linear segment is attributed to fast (film diffusion) adsorption of terbium, followed by slower intraparticle diffusion of terbium through the pores as the rate-limiting step. Based on the Weber-Morris plot, both film and intraparticle diffusion are involved in controlling the kinetic rate of adsorption for DGA and LN resins.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

Sensitivity Analysis of Drivers Water Shortage in the Los Angeles Region During Drought

The code and detailed step-by-step instructions for generating the model output data, processing results, and analysis and plotting are provided at https://github.com/IMMM-SFA/Ferencz_et_al_2026_ER_Water. The PyArtes model is a python adaptation of the Artes model. PyArtes uses many of the same input data and optimization model architecture as Artes. Documentation for the PyArtes model is provided in the Supplement to the paper. The primary data product are simulated monthly water shortages for indoor and outdoor demand under a large ensemble of drought scenarios (>13,000). The droughts are hypothetical and are not based on historical time series data of supply sources - though historical data did help inform ranges explored for supply parameters. Demands are informed by recent 2017-2021 water supply data. Demands used for the model can be accessed at https://github.com/IMMM-SFA/Ferencz_et_al_2026_ER_Water. Simulations resolve demand for over 90 water providers in the study region. The results report 36 months of water shortage data for each indoor and outdoor demand node. The study also developed a multilayer perceptron (MLP) neural network trained on a subset of the simulated shortage ensemble to emulate worst annual water shortage for a given set of parameter multipliers -- provided the parameter values fall within the ranges sampled in the ensemble. Emulated water shortages for synthetic ensembles are in the MLP-generated shortages folder. The MLP model was used to generate larger ensembles to support Sobol analysis that would have been extremely computationally expensive to simulate. Datasets provided in this repository*: Simulated shortages. These results are used for the analysis for Figures 5, 8, and 9 in the paper, and also to train the MLP emulator. .zip file containing outputs for the 13,312 scenario ensemble. Separate .csv files for indoor and outdoor shortage for each scenario. Rows = demand ids (~100), Columns = months (36) Units = acre-feet/month of shortage (shortage = monthly demand - supply). 1 acft = 1233.48 m^3 .csv files of aggregated shortages derived from the 13,312 ensemble Rows = scenarios (13,312), Columns = demand ids (~100) Units = acre-feet/year (either worst annual shortage or total shortage over the 3-year drought) .csv file of the parameter multipliers scenarios for the ensemble .csv file of the parameter ranges and baseline values the multipliers were applied to MLP-generated shortages. These results are used for Figures 4, 6, and 7 in the paper. mwd higher folder: scenario ensembles, emulated worst year total shortages (acft), and Sobol results Emulated shortages. Rows = scenarios, columns = demand ids, units acft Sobol results. Rows = demand ids, columns Sobol (S1, ST, or 95% confidence interval) value for each parameter mwd lower folder: scenario ensembles, emulated worst year total shortages (acft), and Sobol results same organization as mwd higher MLP performance: performance metrics (R^2, RMSE, BIAS, MAPE) for the testing subset (20% or 2,662 scenarios) and simulated vs emulated worst year shortage (acre-feet/year) for every demand node, MWD wholesale regions, and the entire study region (LAC). Supporting data for figures. Figure plotting scripts in the associated GitHub repo. These files support analysis and visualization. Geospatial Data used for plotting simulated water shortages and Sobol results. Dictionary of full names for demand nodes in the model and estimates of water supply by source type informed by Artes input files and California Urban Water Management Planning data: https://water.ca.gov/Programs/Water-Use-And-Efficiency/Urban-Water-Use-Efficiency/Urban-Water-Management-Plans *Readme files provided for each folder.

drought↗

National Energy Water Treatment & Speciation (NEWTS): A Water & Critical Mineral Database and Dashboard

The scarcity of water resources, the need for beneficial water reuse, and the challenges of wastewater treatment are becoming increasingly pressing in economic, social, and environmental domains. Addressing these concerns requires effective treatment strategies to manage wastewater streams and tackle environmental and economic issues. Furthermore, the recovery of critical minerals from the waste streams associated with energy production holds the promise of offsetting treatment costs and securing local sources of valuable minerals. However, relevant data on these waste streams are dispersed and challenging to locate. The process of ingesting such data into modeling software often involves multiple steps, requiring data restructuring to meet software-input requirements. The non-standardized reporting of water data makes data aggregation and reformatting a time-consuming process. Additionally, essential attributes necessary for modeling water treatment and mineral scale formation are frequently missing. Moreover, data gaps vary depending on the region of interest. Consequently, there is a pressing need for high-quality energy-water composition data that can be easily imported into water chemistry modeling software. To address this need, the National Energy Technology Laboratory has created the National Energy Water Treatment and Speciation (NEWTS) Database and Dashboard—a free online tool catering to community leaders and water researchers. NEWTS facilitates a comprehensive understanding of the composition of energy-related wastewater streams in the United States. The datasets provide detailed concentrations and speciation of major and minor aqueous compounds in energy-related wastewater streams, including power plant leachate, acid mine drainage, brackish water, and oil and gas produced water across the United States. Many of the aqueous species are critical minerals (Li, REEs) in high demand to modernize the world’s energy infrastructure. Many of the datasets also contain volumetric flow-rates needed to model the treatment and reuse scenarios in advanced aqueous chemistry software programs. The NEWTS Database and Dashboard offer public access to hitherto challenging-to-access datasets, presented in a standardized format that is tailored for easy input into aqueous chemistry modeling software. By performing the work needed to transform dispersed, disparate data sources into unified, model-ready datasets, NEWTS serves as an essential resource in advancing water treatment research and sustainable water resource management.

produced water management↗

Generalizing synthetic data-trained acoustic predictive models to real-world measurements

Acoustic Resonance Spectroscopy (ARS) is highly sensitive to structural properties such as material, geometry, and environmental conditions; as a consequence, it can noninvasively measure internal properties that are unobservable by most other methods. Because of its sensing capabilities and low implementation cost and complexity, ARS has potential as a paradigm shift in noninvasive sensing, characterization, and monitoring applications. However, extracting specific properties from ARS measurements, comprising the vibration spectrum of a test object, is challenging due to the sensitivity of the spectra to other structural changes not being measured, e.g. manufacturing tolerances, component coupling, environmental variation, etc. Neural Networks are promising tools for identifying trends in ARS measurements, but their training typically requires large datasets, which are often impractical to obtain for real-world systems. Synthetic data can be simulated efficiently, but discrepancies between synthetic and real-world data frequently lead to poor generalization when testing on the real-world data. We propose a novel ARS model training framework that enables networks trained exclusively on synthetic ARS data to generalize effectively to real-world measurements. Our approach leverages the Correlation Alignment (CORAL) technique to enforce the extraction of features common to both synthetic and real-world domains. As a case study, we demonstrate noninvasive ARS-based pressure measurements in sealed systems. Finite element method (FEM) simulations were used to generate synthetic training data across diverse vessel configurations and pressure conditions, and model performance was then tested on real-world measurements. We demonstrate that robust machine learning models for ARS can be developed without large real-world datasets, significantly broadening the applicability of ARS for noninvasive sensing. Moreover, the approach is extensible to other sensing modalities where synthetic data are abundant but real-world data are limited.

36 MATERIALS SCIENCE↗

Machine Learning Based Metamodel for Faster Life Cycle Assessment of Large Portfolio of Buildings

Managing a large portfolio of buildings involves decisions on reuse, retrofit, renovation, rehabilitation, and new construction, influenced by trade-offs between performance metrics such as cost, time, and operational flexibility over the building's life cycle. Traditional life cycle assessment tools for evaluating these metrics can be labor- and compute-intensive, requiring extensive data and modeling for each building. Metamodels (or surrogate models) using machine learning have been explored as faster alternatives, but training these models has been hindered by the limited availability of comprehensive data on key life cycle metrics. Recent advancements in machine learning, particularly deep learning techniques like zero-shot and few-shot learning, allow models to learn from sparse or limited data. We propose a machine learning-based metamodel that leverages these techniques for rapid estimation of key building life cycle metrics. This presentation will cover the model architecture, data collection, training, and validation processes, along with an ongoing case study applied to a large portfolio of buildings. We will discuss the model's performance in terms of accuracy, compute time, limitations, and its potential for expanding to additional life cycle metrics. This data-driven approach offers a promising direction for the rapid evaluation of large building portfolios.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Probabilistic Error Bounds for Low-Rank Tensor Decompositions Used in Large-Scale Data Analysis Applications (LDRD Final Report)

This report documents a research project on analyzing low-rank tensor models for data analysis that took place at Sandia National Laboratories from October 2023–September 2025. The focus of this work was to extend theoretical frameworks from statistics and probability theory for use with models for scalar, vector, and matrix data to models with tensor, or general multi-dimensional array, data. Through this work, we have provided a new set of tools for bounding errors on low-rank tensor models of both complete and sampled data. The remainder of this report is organized as follows. In Section 1, we describe the proposed work at the start of the project. Section 2 describes the research advances made as part of the project. Other research contributions in the form of conference presentations and software development is provided in Section 3. Workforce development at Sandia and Florida Atlantic University (via a subcontract on this project) is provided in Section 4.

97 MATHEMATICS AND COMPUTING↗

A portable application framework for energy management and information systems (EMIS) solutions using Brick semantic schema

This paper introduces a portable framework for developing, scaling and maintaining energy management and information systems (EMIS) applications using an ontology-based approach. Key contributions include an interoperable layer based on Brick schema, the formalization of application constraints pertaining metadata and data requirements, and a field demonstration. The framework allows for querying metadata models, fetching data, preprocessing, and analyzing data, thereby offering a modular and flexible workflow for application development. Its effectiveness is demonstrated through a case study involving the development and implementation of a data-driven anomaly detection tool for the photovoltaic systems installed at the Politecnico di Torino, Italy. During eight months of testing, the framework was used to tackle practical challenges including: (i) developing a machine learning-based anomaly detection pipeline, (ii) replacing data-driven models during operation, (iii) optimizing model deployment and retraining, (iv) handling critical changes in variable naming conventions and sensor availability (v) extending the pipeline from one system to additional ones.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Packages of Distributed Energy Technologies Demonstrating Demand Flexibility at Community Scale

The combination of increased electric load growth across all sectors, deferred electrical infrastructure investment, and other factors resulting in variable electric power supply, has created technical challenges to maintaining a resilient and reliable grid. Many federal, regional, and local efforts are in play to modernize the electric grid, including advancing building technologies and distributed energy resources (DERs) that are utilizing smarter controls to become responsive to both occupant and grid needs. This report reviews ten pilot projects demonstrating how groups of buildings combined with behind-the-meter (BTM) DERs such as electric vehicle (EV) charging, battery storage, flexible HVAC and domestic hot water systems, and photovoltaic systems can reliably and cost effectively provide grid services. Each of the ten pilot projects aim to deliver both energy efficiency and demand flexibility (DF) while supporting load growth. The ten demonstration teams are piloting flexible DER packages across diverse communities of residential and commercial buildings to address a variety of regional grid needs. The outcomes of these pilot projects will be used to inform future scaling through utility program development. This paper characterizes the ten teams, showcasing the decision-making process used by each group to develop their packages (Section 2), the grid services they plan to deliver (Section 3), the types of DER packages selected for deployment within building sectors (Section 4) and trends between building sector, DER types, and grid services In order to achieve community scale benefits, the pilot projects must utilize aggregated control mechanisms for coordinating buildings and DERs together. Several types of coordinated control architectures have evolved amongst the teams, influenced by use type, existing market conditions, and integration type. Three coordinated controls architectures have been characterized, highlighting their use cases, benefits, challenges, and tradeoffs in their design. These insights can aid utilities, control vendors, and developers in scaling community-level energy systems (Paul, 2024). Ultimately, the technology packages selected by the ten teams will be coordinated to provide power system services, also known as grid services. Insights from these demonstrations will be useful for grid operators, regulators, aggregators and other stakeholders as they look to deploy demand flexible resources as grid services in the future. The grid services that each team is targeting for demonstration are described in Section 3 and Section 4. Methods for evaluating the grid services have been described in the paper Metrics for Evaluating Grid Service Provision from Communities of Grid-interactive and Efficient Buildings and other DER (MacDonald, 2023). To identify technology packages for demonstration, Section 2 shows that project teams used a range of analysis approaches, including building energy modeling, AMI data analysis, cost-benefit frameworks, and utility pilot data. Some teams emphasized technical modeling to quantify grid impacts and demand reduction potential, while others prioritized economic evaluations, stakeholder input, or exploratory pilots to inform deployment decisions. This diversity reflects the need to tailor selection methods to project goals, available data, and organizational context. Section 5 discusses trends between the DER technologies deployed and the grid service provisions from each team. Residential buildings (multifamily and single family) lean towards technologies that enhance energy efficiency (e.g. weatherization upgrades, smart thermostats) and onsite power generation integration (e.g. solar PV). Commercial building demonstrations prioritize technologies that ensure operational reliability (e.g. battery storage) and centralized energy management systems and optimization solutions. Teams that are deploying controllable storage-based technologies are more likely to provide grid services that require a near real-time response. Teams incorporating load shifting technologies like smart thermostats with HEMs are likely to include energy markets participation and customer bill management offerings. Campus demonstrations are adopting diverse sets of DERs to emphasize renewable generation, paired with centralized control. This section also describes technologies that were considered during project planning but ultimately excluded from final deployment. These demonstrations reveal that effective DER package design should be tailored to building type, customer segment, and construction vintage. Multifamily buildings benefit from centralized HVAC upgrades and supervisory controls, while single-family homes are well-suited for individualized technologies like solar, storage, and smart home energy monitors. Commercial and campus settings prioritize EMIS integration and load optimization. New construction enables cost-effective integration of DER-ready infrastructure, whereas retrofits require deployments aligned with owner and tenant value streams. For utility program planners, early coordination with developers and building owners, paired with segmented and modular program offerings, can improve adoption, scalability, and grid impact.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Pavement condition and climatic data in southeast Texas: A dataset for evaluating flood impacts on pavement performance

Effective pavement maintenance is essential for economic stability, optimal network performance, and roadway safety. Achieving this requires thorough evaluation of pavement conditions, including structural integrity, surface roughness, and distress characteristics. Pavement performance indicators play a critical role in influencing vehicle safety and ride quality. Recent advances have emphasized the use of data-driven modeling to anticipate pavement behavior, with the goal of optimizing resource allocation and refining Maintenance and Rehabilitation (M&R) strategies through accurate condition assessment. A foundational requirement for these modeling efforts is the availability of standardized, high-quality datasets that can support robust and reproducible infrastructure analysis. This data article presents a comprehensive dataset assembled to facilitate pavement performance prediction, with a geographic focus on Southeast Texas, particularly the flood-vulnerable area of Beaumont. The dataset encompasses pavement and traffic attributes, meteorological records, flood simulation outputs, ground deformation measurements, and topographic indices, enabling detailed examination of both load-associated and non-load-associated degradation mechanisms. Data preprocessing was performed using ArcGIS Pro, Microsoft Excel, and Python to ensure consistency and usability in data-driven modeling applications, including machine learning workflows. Key contributions of this dataset include its utility in analyzing the climatic and environmental factors affecting pavement conditions, identifying critical predictive features, and enabling in-depth correlation analysis across diverse variables. By filling existing gaps in input variable selection resources, this dataset supports the development of predictive tools for estimating future maintenance demand and enhancing the resilience of pavement networks in flood-impacted areas. The resource highlights the importance of standardized datasets for advancing pavement management practices and provides a robust foundation for ongoing infrastructure performance modeling.

42 ENGINEERING↗

Applications of emulation and Bayesian methods in heavy-ion physics

Abstract Heavy-ion collisions provide a window into the properties of many-body systems of deconfined quarks and gluons. Understanding the collective properties of quarks and gluons is possible by comparing models of heavy-ion collisions to measurements of the distribution of particles produced at the end of the collisions. These model-to-data comparisons are extremely challenging, however, because of the complexity of the models, the large amount of experimental data, and their uncertainties. Bayesian inference provides a rigorous statistical framework to constrain the properties of nuclear matter by systematically comparing models and measurements. This review covers model emulation and Bayesian methods as applied to model-to-data comparisons in heavy-ion collisions. Replacing the model outputs (observables) with Gaussian process emulators is key to the Bayesian approach currently used in the field, and both current uses of emulators and related recent developments are reviewed. The general principles of Bayesian inference are then discussed along with other Bayesian methods, followed by a systematic comparison of seven recent Bayesian analyses that studied quark-gluon plasma properties, such as the shear and bulk viscosities. The latter comparison is used to illustrate sources of differences in analyses, and what it can teach us for future studies.

Paquet, Jean-François (ORCID:0000000187368171)↗

Microwave-Assisted Plastic Upcycling: Dynamic Data Reconciliation, Parameter Estimation, and Kinetic Modeling

Microwave (MW)-assisted catalytic pyrolysis offers a promising pathway for efficient plastic upcycling. This work develops an integrated modeling framework combining dynamic data reconciliation, a temperature-dependent rate model, and a yield model to represent the time-varying production rate of components in MW-assisted LDPE pyrolysis conducted in a batch reactor. An Arrhenius-type rate model with a temperature-dependent reaction order is developed. A biexponential correlation is proposed for the yield of gaseous products that enables to capture the evolving product formation behavior during conversion. In the yield correlation, one term is used to represent the initial increase in yield, reflecting the rapid formation of intermediate or primary products at the early stages of the reaction when a larger fraction of the reactant remains available. As conversion progresses, the influence of this term gradually diminishes. The other term accounts for the subsequent decrease in the predicted yield, representing secondary reactions such as further cracking or coke formation that reduce the concentration of certain products at higher conversion. The model is found to accurately represent reconciled experimental flow rate profiles from an in-house MW-assisted catalytic batch reactor for major products, including ethylene, ethane, 1-butene, and benzene, across 250−350 °C. Ethylene remains the dominant product but decreases from about 41.95% at 250 °C to 30.14% at 350 °C, while heavier products increase significantly, with 1-butene rising to nearly 8.37% and benzene reaching 2.17% at intermediate temperatures. The model shows that the ethylene production rate can be maximized at around 270 °C. The models developed in this work can be utilized for process optimization, reactor design and scale-up of microwave-assisted plastic conversion technologies, and economic analysis.

Damahe, Harish [West Virginia Univ., Morgantown, W↗