Search NASA⌕ Search

SEARCH · Search NASA

Results for “High performance computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 667 records · Page 37

Synthetic-domain computing and neural networks using lithium niobate integrated nonlinear phononics

Analogue computing uses the physical behaviours of devices to provide energy-efficient arithmetic operations. However, scaling up analogue computing platforms by simply increasing the number of devices leads to challenges such as device-to-device variation. Here, in this study, we report scalable analogue computing and neural networks in the synthetic frequency domain using an integrated nonlinear phononic platform on lithium niobate. This synthetic-domain computing is robust to device variations, as vectors and matrices are concurrently encoded at different frequencies within a single device, achieving a high throughput per area. Leveraging inherent nonlinearities, our device-aware neural network can perform a four-class classification task with an accuracy of 98.2%. The nonlinear phononic computing hardware also maintains consistent performance over a wide operational temperature range (characterized up to 192 °C). Our synthetic-domain computing combines single-device parallelism, inherent nonlinearity and environmental stability, and could be of use in edge computing applications in which power efficiency and environmental resilience are crucial.

Ji, Jun [Virginia Polytechnic Inst. and State Univ↗

Bayesian Calibration of Stochastic Agent Based Model via Random Forest

Agent-based models (ABM) provide an excellent framework for modeling outbreaks and interventions in epidemiology by explicitly accounting for diverse individual interactions and environments. However, these models are usually stochastic and highly parametrized, requiring precise calibration for predictive performance. When considering realistic numbers of agents and properly accounting for stochasticity, this high-dimensional calibration can be computationally prohibitive. This paper presents a random forest-based surrogate modeling technique to accelerate the evaluation of ABMs and demonstrates its use to calibrate an epidemiological ABM named CityCOVID via Markov chain Monte Carlo (MCMC). The technique is first outlined in the context of CityCOVID's quantities of interest, namely hospitalizations and deaths, by exploring dimensionality reduction via temporal decomposition with principal component analysis (PCA) and via sensitivity analysis. The calibration problem is then presented, and samples are generated to best match COVID-19 hospitalization and death numbers in Chicago from March to June in 2020. Further, these results are compared with previous approximate Bayesian calibration (IMABC) results, and their predictive performance is analyzed, showing improved performance with a reduction in computation.

60 APPLIED LIFE SCIENCES↗

In situ multi-tier auto-ignition detection applied to dual-fuel combustion simulations

Here we use an anomaly detection methodology that is centered on analyzing fourth-order joint moments (co-kurtosis), particularly focusing on its application in auto-ignition of combustion problems with large numbers of species. Unsupervised anomaly detection is challenging to generalize across problem types and domains. A recent technique, centered on analyzing information in the fourth-order joint moment co-kurtosis, has shown promise, especially for high-dimensional scientific data. In this work we present developments to the co-kurtosis based anomaly detection method needed to make it effective and scalable for large-scale distributed scientific data, such as those generated by massively parallel simulations. An in situ co-kurtosis algorithm is employed as the anomaly detection method for identifying ignition kernels in simulations of turbulent combustion. Here, we extend an existing methodology which identifies regions of the domain where anomalies are present, and add another tier of anomaly detection where the individual samples contributing to the anomaly are identified. We apply this algorithm on-the-fly to a variety of turbulent reacting flow problems and compare it to the widely used (but significantly more expensive) chemical explosive mode analysis (CEMA). We demonstrate the ability of the method to detect and identify the onset of low and high temperature ignition which can be used for computational steering, as chemical and combustion anomalies occur intermittently at spatio-temporal locations unknown a priori. Finally, we apply our lightweight in situ algorithm to an exascale high-fidelity simulation with a total of 2.4 Trillion degrees of freedom, performed using an adaptive mesh refinement solver. Furthermore, through a scalability analysis, we show that the relative computational cost of this in-situ anomaly detection algorithm compared to an iteration of the reacting flow solver is negligible.

97 MATHEMATICS AND COMPUTING↗

Efficient mapping between void shapes and stress fields using Deep Convolutional Neural Networks with sparse data

Establishing fast and accurate structure-to-property relationships is an important component in the design and discovery of advanced materials. Physics-based simulation models like the finite element method (FEM) are often used to predict deformation, stress, and strain fields as a function of material microstructure in material and structural systems. Such models may be computationally expensive and time intensive if the underlying physics of the system is complex. This limits their application to solve inverse design problems and identify structures that maximize performance. In such scenarios, surrogate models are employed to make the forward mapping computationally efficient to evaluate. However, the high dimensionality of the input microstructure and the output field of interest often renders such surrogate models inefficient, especially when dealing with sparse data. Deep convolutional neural network (CNN) based surrogate models have shown great promise in handling such high-dimensional problems. In this paper, a single ellipsoidal void structure under a uniaxial tensile load represented by a linear elastic, high-dimensional and expensive-to-query, FEM model. We consider two deep CNN architectures, a modified convolutional autoencoder framework with a fully connected bottleneck and a UNet CNN, and compare their accuracy in predicting the von Mises stress field for any given input void shape in the FEM model. Additionally, a sensitivity analysis study is performed using the two approaches, where the variation in the prediction accuracy on unseen test data is studied through numerical experiments by varying the number of training samples from 20 to 100.

surrogate modeling; convolutional neural networks;↗

Ultra-low thermal resistance and pressure drop copper and copper-tungsten diamond-shaped pin fin cold plates for liquid cooling of electronics

Modern and future data centers face increasing cooling challenges due to increasing chip thermal design power and die size, along with the need to reduce energy consumption used for cooling. High performance cooling solutions that maintain a low chip junction temperature are needed to ensure electronics reliability. This work develops an ultra-low thermal resistance and low pressure drop 75 mm × 75 mm cold plate, intended for next-generation electronics cooling. The cold plate features an array of diamond-shaped pin fins and integrated copper tungsten heat spreader, selected for its low coefficient of thermal expansion which reduces thermomechanical deformation and allows for closer integration of the cold plate with silicon dies. Starting with 300 candidate designs, three-dimensional computational fluid dynamics simulations predict the thermal-hydraulic performance of cold plate subsections. The highest performing geometries are evaluated with high fidelity simulations. Four cold plates are manufactured for experiments: three with diamond-shaped pin fins and one with straights fins for comparison purposes. The cold plates are fabricated from copper-tungsten (CuW), copper (Cu), or aluminum-silicon-magnesium alloy (AlSi10Mg). The diamond-shaped pin fins achieve a roughly 15 % lower thermal resistance compared to the conventional straight fin microchannel. The highest performing design achieves a chip-to-coolant (including thermal interface material) thermal resistance of 9.0 K/kW in CuW and 6.9 K/kW in Cu under a 1 kW heat load with an inlet-to-outlet pressure drop of 9.0 kPa and water as the working fluid. This work demonstrates ultra-low thermal resistance and pressure drop cold plates for large die, high heat load applications, and shows that CuW is an attractive cold plate material for improved reliability in next generation data center cooling.

Coefficient of thermal expansion↗

An Instrumented Capsule Design to Measure Thermal Conductivity in Miniature UO2 Specimens

Numerous separate effects irradiations of miniature nuclear fuel specimens have been conducted in the High Flux Isotope Reactor (HFIR) under the experimental platform designated as MiniFuel. MiniFuel is a static irradiation capability in which microstructural evolution and fuel performance phenomena are observed during postirradiation examination thereby offering a snapshot of the terminal fuel characteristics. This approach inherently requires fielding an irradiation where the experimental conditions are determined using predictive models and the pertinent outcomes are measured at the end of the test. Static irradiations can provide useful insights to the relationships between fuel performance and the pivotal irradiation conditions, namely temperature and burnup, but the ability to monitor fuel performance in situ would further support fuel development and qualification. To this end, an instrumented experiment design is being developed at Oak Ridge National Laboratory to capture thermal conductivity degradation and fission gas release during HFIR irradiation. These phenomena will be monitored using unique capsule designs that each target a different phenomenon. This paper details the thermal conductivity capsule (TCC) design and its expected performance envelope as determined using computer models. Each TCC will contain a miniature UO2 disc specimen (~0.5 mm thick × 5 mm diameter) sandwiched between metallic slugs with embedded thermocouples. The coupling of in situ temperature measurements, known thermal conductivity of the metallic components, and heat generation rates computed using high-fidelity neutronics models make the thermal conductivity measurement possible. This paper describes the reactor physics and heat transfer models used to predict the capsule’s performance and the methodology for calculating the fuel specimen’s thermal conductivity from the thermocouple measurements.

Gorton, Jacob [ORNL] (ORCID:0000000269806083)↗

Computational flow modeling of triply periodic minimal surfaces as feed channel spacers in ultra-high pressure reverse osmosis applications

Triply periodic minimal surfaces (TPMS) are a special class of mathematical surfaces characterized by a high surface area-to-volume ratio. They have generated considerable interest in fields such as acoustics, heat transfer, and membrane-based filtration processes. This study evaluates the performance of four different TPMS designs—Schoen Gyroid, Schoen Crossed Layers of Parallels (CLP), Schoen Transverse Crossed Layers of Parallels (tCLP), and Schwarz-Primitive—when used as feed channel spacers under ultra-high pressure reverse osmosis (UHPRO) conditions, at approximately 200 bar. Our experimentally validated computational fluid dynamics model reveal different flow patterns within the feed channels for each of the four TPMS designs, leading to varying hydrodynamic and permeation properties. Under the simulated UHPRO conditions, the Gyroid and tCLP designs yield up to a 23% increase in average permeate velocity and a 14% reduction in average membrane-surface concentration relative to a non-woven spacer of the same porosity. Furthermore, the enhanced performance comes with an increased feed channel pressure drop, although it only constitutes less than 4% of the operating pressure when extrapolated for a meter-long membrane module. Additionally, the study analyzes the effects of varying inlet velocity and spacer porosity on membrane performance. Overall, this research provides valuable insights into the potential use of TPMS spacers in UHPRO applications.

36 MATERIALS SCIENCE↗

Control And Optimization Modular Modeling Application For Nuclear Deployment

The purpose of the COMMAND code is to provide a flexible, scalable tool for use in developing, integrating, and testing the technologies necessary for achieving autonomous operations of advanced nuclear reactors. The code enables users to efficiently implement custom simulations and experiments by combining key methods from different software modules. These modules are focused on: modeling and simulation tools, such as nuclear simulation tools used for high-fidelity modeling (e.g., Reactor Excursion and Leak Analysis Program [RELAP5-3D] and Monte Carlo N-Particle [MCNP]); machine learning and optimization tools (e.g., anomaly detection and data-driven modeling techniques); advanced control in its digital, high-performance, and supervisory control forms (e.g., proportional integral derivative (PID) control and model predictive control (MPC); and integration with hardware through industrial communication protocols. To ensure flexibility and scalability, COMMAND was designed to be both modular—the software “pieces” all inherit from generic building blocks and can be combined and connected to create complicated simulations—and high performing—designed for parallel processing, enabling simulations and experiments to take advantage of multi-core computers, servers, and nodes. The code is written in the Python programming language due to the language's popularity, active community, and open-source and cross-platform nature. Maintaining consistency with other simulation tools used within the nuclear energy community, users implement simulations and experiments through text input files, which define components, parameters, connections, etc., through lines of text. Given that COMMAND is written in Python, these input files are native Python scripts, and so use the standard Python structure and formatting. This also enables users to take advantage of Python's extensive package library to develop custom capabilities for their specific use cases.

Faber, Jacob [Idaho National Laboratory (INL), Ida↗

A theoretical/computational framework to measure SiO2 and MgO viscosity at high pressure

The convection of the mantle of Earth and super-Earths is important for many terrestrial phenomena, from plate tectonics to outgassing. Rheological properties, such as viscosity, regulate the transport of thermal energy and mass. However, the viscosity of mantle-relevant materials, such as MgO, at relevant pressures (>120 GPa) are not well constrained. The objective of this work was to develop a computational platform to simulate novel experiments aiming to measure the viscosity of MgO at high pressures. Experiments performed by our collaborators at Johns Hopkins University and Lawrence Livermore National Laboratory use the OMEGA EP laser facility to shock a corrugated MgO interface to 170 GPa, with resulting velocity evolution governed by the viscous Richtmyer-Meshkov instability. We used an in-house hydrocode to simulate this process and thus provide a bound on viscosity by comparing our simulations results to these experiments, as well as to examine the physical processes at play. We simulated the experiments with different values of MgO viscosity (from inviscid to 10,000 Pa∙s) while taking into account the unsteady laser pulse, the material's equation of state, and the rate-dependent constitutive relation for MgO, and the materials' equations of state. Our results suggest that MgO at these conditions has a viscosity within an order of magnitude of 5000 Pa∙s.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Solar Spectrum Conversion for an Algae Bioreactor (CRADA Final Report)

This project focused on developing advanced optical coatings to improve solar energy utilization. The research aimed to create lanthanide-doped upconversion nanoparticles (UCNPs) capable of capturing unused near-infrared (NIR) light from the sun and converting it into visible light (blue and red photons) that can be used for photosynthesis. The primary goal was to identify, synthesize, and integrate highly efficient UCNPs into a transparent thin-film device. Through a comprehensive workflow involving computer simulations, high-throughput robotic synthesis, and detailed optical characterization, the project successfully developed a high-performance material. The key technical achievement was the creation of a core-shell UCNP (NaYF₄:20%Yb³⁺, 2%Er³⁺ coated with a 10 nm NaYF₄ shell) that demonstrated a quantum yield of 3.2% for converting 980 nm NIR light into visible light. Transparent thin films fabricated from these nanoparticles showed excellent optical properties, confirming their potential for practical applications. This research adds to the scientific understanding of energy transfer in lanthanide materials and demonstrates a technically effective method for creating efficient light-converting coatings. The primary benefit to the public lies in the potential for these coatings to enhance the efficiency of solar-driven processes, such as boosting the growth of algae in photobioreactors for biofuel production.

14 SOLAR ENERGY↗

Towards Robust Calibration of the AWSD Reactive Burn Model

Calibration of a reactive burn model for detonation of high explosive is an important step towards predictive hy drodynamic simulations of detonation. A typical calibration consists of varying model parameters (e.g., rate constants, activation energies) until results of hydrodynamic simulations match the experimental data for a certain set of ex periments. Hydrodynamic simulations of the dependence of steady detonation velocity on the radius of a cylindrical high-explosive charge - often used in such calibrations - can be computationally expensive. In this work, we propose a method where such expensive simulations are performed infrequently, and only to parameterize and refine a surrogate model for the dependence of the detonation velocity on calibrated parameters. The method is developed, implemented and applied to an example problem - calibration of the AWSD reactive burn model for important high explosive PBX 9502. Two different flavors of the surrogate model are investigated, and the calibration is performed successfully.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Effect of high scandium doping in barium zirconate on nickel diffusion and performance of proton-conducting solid oxide electrolyzer cells

Proton-conducting solid oxide electrolyzer cells (p-SOECs) are emerging but promising technologies for hydrogen production. However, due to the lack of a robust electrolyte, p-SOECs struggle simultaneously to display high performance, Faradaic efficiency, and durability. Motivated by its high proton concentrations and stability as a barium zirconate, we have investigated BaZr 0.6 Sc 0.4 O 3-δ (BZSc40) as a potential next-generation p-SOEC electrolyte. Here, we found elevated levels of NiO diffusion through BZSc40 electrolytes during high-temperature sintering, attributed to the large oxygen vacancy concentrations present in BZSc40, as revealed by first-principle computational results. Controlling NiO diffusion is critical, as it can facilitate densification and grain size growth, but it may also detrimentally impact performance by causing electronic leakage. By optimizing sintering temperature when fabricating BZSc40 cells, we successfully controlled NiO diffusion, achieving sufficient electrolyte densification along with high performance and Faradaic efficiency. BZSc40 cells reached −0.99 A/cm 2 at 1.3 V and 600 °C and exhibited enhanced durability with a 3.37 mV/kh degradation rate at −0.8 A/cm 2 over a 200-h testing period. BZSc40 electrolytes demonstrated superior performance over BaZr 0.8 Y 0.2 O 3-δ (BZY20). In addition to elevated current densities and grain sizes, BZSc40 cells achieved Faradaic efficiencies of 76 % compared to 54 % for BZY20 at −0.2 A/cm 2 and 600 °C. This work lays the foundation for BZSc40 as a potential electrolyte due to its advantages over BZY20 while demonstrating the significance of controlling NiO diffusion when fabricating p-SOECs.

Electrolyzer↗

Protein-ligand binding affinity prediction using multi-instance learning with docking structures

Recent advances in 3D structure-based deep learning approaches demonstrate improved accuracy in predicting protein-ligand binding affinity in drug discovery. These methods complement physics-based computational modeling such as molecular docking for virtual high-throughput screening. Despite recent advances and improved predictive performance, most methods in this category primarily rely on utilizing co-crystal complex structures and experimentally measured binding affinities as both input and output data for model training. Nevertheless, co-crystal complex structures are not readily available and the inaccurate predicted structures from molecular docking can degrade the accuracy of the machine learning methods. We introduce a novel structure-based inference method utilizing multiple molecular docking poses for each complex entity. Our proposed method employs multi-instance learning with an attention network to predict binding affinity from a collection of docking poses. We validate our method using multiple datasets, including PDBbind and compounds targeting the main protease of SARS-CoV-2. The results demonstrate that our method leveraging docking poses is competitive with other state-of-the-art inference models that depend on co-crystal structures. This method offers binding affinity prediction without requiring co-crystal structures, thereby increasing its applicability to protein targets lacking such data.

97 MATHEMATICS AND COMPUTING↗

Vision Foundation Models in Remote Sensing: A survey

Artificial intelligence (AI) technologies have profoundly transformed the field of remote sensing (RS), revolutionizing data collection, processing, and analysis. Traditionally reliant on manual interpretation and task-specific models, RS research has been significantly enhanced by the advent of foundation models (FMs)—large-scale pretrained AI models capable of performing a wide array of tasks with unprecedented accuracy and efficiency. This article provides a comprehensive survey of FMs in the RS domain. We categorize these models based on their architectures, pretraining datasets, and methodologies. Through detailed performance comparisons, we highlight emerging trends and the significant advancements achieved by those FMs. Additionally, we discuss technical challenges, practical implications, and future research directions, addressing the need for high-quality data, computational resources, and improved model generalization. Our research also finds that pretraining methods, particularly self-supervised learning (SSL) techniques like contrastive learning (CL) and masked autoencoders (MAEs), remarkably enhance the performance and robustness of FMs. This survey aims to serve as a resource for researchers and practitioners by providing a panorama of advances and promising pathways for the continued development and application of FMs in RS.

data models↗

Low-Cost Heliostat for High-Flux Small-Area Receivers (Final Technical Report)

This project analyzed a two-stage heliostat concept consisting of a tracking stage and a concentrating stage. The tracking stage uses mirrors mounted on a common drive that move to track the sun. The concentrating stage consists of stationary mirrors that each have a unique angle to direct rays towards a small-area, high-flux, point-focused receiver. By splitting the collection and concentrating process into two stages, multiple small, inexpensive mirrors can share a structure and be controlled by a single drive in the tracking stage. The project effort developed modeling techniques that were specifically relevant to this two-stage heliostat concept. Both field-level and unit-level models were developed. The field-level model does not explicitly consider unit-level losses which are predicted by the unit-level model and then integrated into the field-level model through a correlation referred to as an efficiency modifier. This approach is referred to as the two-model approach; the development and demonstration of this two-model approach for a multi-stage heliostat technology is a key outcome of this work. The field-level model is used to design a field that hits a specific design day power given a set of heliostat design parameters. An oversized field is simulated and then heliostat units are removed based on their annual energy production in order to generate the highest performing field. The field reduction procedure fits a smooth curve fit to annual energy production as a function of position in the field which has the effect of reducing the noise that is otherwise caused by the Monte Carlo ray tracing technique. This approach is referred to as the annual energy fit method and substantially reduces computational run time for a given field level modeling accuracy. The annual energy fit approach enables the selection of a properly sized, high-performing field using orders of magnitude fewer rays than would otherwise be possible and the development of this approach is a second key outcome of this work. These models are used within a genetic optimization algorithm in order to optimize the geometric parameters associated with a heliostat in order to achieve the lowest cost per unit of collected design day power. The cost modeling that underlies the optimization is a simple, scaling type analysis backed up by a much more detailed Design for Manufacture and Assembly (DFMA) analysis. Although the figure of merit used for optimization was not cost per mirror area, this metric is reasonable to use as a means of comparison. The optimally designed 500 kW design has a tracking mirror specific cost of $181.85/m 2 , which is significantly larger than the target value and also larger than the current state of the art. The cost of the torque-tube type linkages contributed substantially to the overall cost. Based on this observation, potentially attractive alternative design configuration utilizing a capstan type actuation system should be investigated. Finally, NREL compared the performance of the two-stage heliostat to the performance of a focused and different sized flat conventional heliostats and showed that, as expected, additional losses versus the convention heliostat caused by a worse cosine efficiency, two stages of reflection, and interstage interactions. The two-stage heliostat requires around 75% more reflective area than a flat 1x1 meter conventional heliostat (similar to a focused heliostat) and 40% more than a flat 2x2 meter conventional heliostat.

14 SOLAR ENERGY↗

A three-dimensional laser ray-tracing methodology for radiation-hydrodynamics simulations

We report on a methodology for performing laser ray-tracing in three spatial dimensions for radiation-hydrodynamics simulation codes. Our method, which is an extension of that developed in Haines et al., Comput. Fluids 201, 104478 (2020), utilizes an automatically generated separate mesh for the laser ray-tracing from the radiation-hydrodynamics mesh. This enables the laser mesh to be tailored to minimize ray noise with significantly fewer rays than would be required when the ray-tracing is performed on the radiation-hydrodynamics mesh, primarily by allowing the use of high-aspect-ratio cells that are not suitable for hydrodynamics solvers. For a planar target, we show that our method provides a ≈ 100× reduction in computational expense to achieve a fixed level of ray noise relative to ray-tracing directly on the radiation-hydrodynamics mesh. The relatively low ray requirement also enables efficient computation of cross-beam energy transfer. Each cell in the logically cubic laser mesh is a non-convex dodecahedron with triangular sides, and numerical integration of the ray trajectories and inverse bremsstrahlung is performed by mapping each cell to the unit cube. We will describe our methodology in detail as well as its implementation in the xRAGE radiation-hydrodynamics code, discuss performance, and present the results from applying the methodology to test problems with analytic solutions for laser ray-tracing through a quadratic density gradient with an analytic solution as well as for a laser-driven heat front. In 3D radiation-hydrodynamics simulations of laser-driven experiments performed on the National Ignition Facility, laser ray-tracing with our methodology uses less than 1% of total computational time while introducing acceptably low levels of ray noise.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

A Comparative Study of Physics‐Informed and Data‐Driven Neural Networks for Compound Flood Simulation at River‐Ocean Interfaces: A Case Study of Hurricane Irene

Simulating compound flooding (CF) at the river-ocean interface within large-scale Earth System Models (ESMs) presents significant challenges due to complex interactions between river discharge, storm surge, and tides. This study assesses the comparative advantages of physics-informed and data-driven machine learning (ML) approaches for enhancing local ESM performance. We systematically compare data-driven neural network models (i.e., CNNs, U-Net, Long Short-Term Memory (LSTM), Gated Recurrent Unit), and physics-informed neural network (PINN) models, including vanilla PINN and a finite-difference-based PINN (FD-PINN). Specifically, FD-PINN is introduced to enhance computational efficiency, accelerating vanilla PINNs by ∼6.5 times while improving accuracy. To enhance data-driven model training, a new data-generation approach is developed to sample historical fluvial and coastal flood events, which ensures a robust data set for extreme event prediction. The models are evaluated using a realistic one-dimensional river domain extracted from an ESM's river mesh and the Hurricane Irene event as an independent test case. Results show that FD-PINN achieves accurate predictions with significantly reduced computational costs relative to vanilla PINNs. Among data-driven models, the best overall performance is achieved by a CNN-LSTM hybrid, which balances accuracy and efficiency. While a fully connected CNN (CNN-FC) provides the best accuracy, it incurs high computational cost. Architectures lacking strong temporal modeling tend to underperform on unseen events. These findings highlight the importance of sequence-aware designs for robust generalization. This study reveals the trade-offs between physics-informed and data-driven models and proposes an adaptive hybrid framework for integrating ML into ESMs to enhance local flood simulations.

Earth Systems Modeling↗

Leveraging Large Language Models for Real-World Data Evidence: A Framework for Automated Treatment Extraction and Data Harmonization

Background: The ability to comprehensively collect treatment information from cancer patient medical records would enable studies to evaluate real-world benefits and risks tied to specific treatments. Currently, it is difficult to system- atically collect high-quality treatment information because it is often stored in unstructured text. Manually extracting and standardizing drug and regimen data is time-intensive. Recent advances in large language models (LLMs) offer a potential solution for automated extraction of structured treatment information from clinical text. Objective: This study systematically evaluates the utility of four LLMs from the Llama family for automated extraction of oncology treatment information from clinical text. This information can guide researchers using cancer registry data to provide insights into cancer care and outcomes beyond clinical trials. Methods: Four instruction-tuned Llama models with varying parameter counts (1B, 3B, 8B, and 70B) were evaluated for their ability to extract treatment information from clinical documents. A unified oncology knowledge base integrating seven major public data sources was developed to standardize and normalize extracted entities—a critical step for harmonizing data from diverse sources. Extracted treatment data were compared against expert-annotated ground truth. Model performance was assessed using accuracy metrics (Precision, Recall, F1-Score) and opera- tional feasibility metrics, including processing speed and structural compliance of the output. Results: A strong positive correlation was observed between model size and extraction accuracy. F1-score improved from 0.609 for the 1B model to 0.710 (3B), 0.807 (8B), and 0.828 (70B). While larger models demonstrated superior accuracy and compliance, they incurred higher computational costs. The modest performance difference between 8B and 70B suggests diminishing returns with increasing model size. Conclusions: LLMs represent a viable technology for automating oncology treatment extraction. The 8B-parameter model emerged as a highly effective option, balancing high accuracy and computational efficiency. Selecting an appropriate LLM for deployment in cancer registries involves a trade-off between desired accuracy and available operational resources. Harmonizing extracted entities with the oncology knowledge base facilitates standardized integration into common data models, enhancing data quality for real-world evidence analyses.

artificial intelligence↗