Search NASA⌕ Search

SEARCH · Search NASA

Results for “performance modeling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

A parcel-level evaluation of distributed wind opportunity in the contiguous United States

This study examines the potential for distributed wind (DW) energy across the contiguous United States, leveraging advancements in the National Renewable Energy Laboratory's distributed wind model, dWind. The novel modeling approach described here utilizes a high-resolution dataset and analyzes over 150 million parcels, a significant improvement from prior methods that extrapolated results from a smaller random sample. This achievement is enabled through key model performance improvements, such as transitioning to multiprocessing, which reduces runtime by 97 %. This optimized, high-resolution approach allows the inspection of technology deployment potential and impact on a variety of scales tailored to individual properties and regions. The results here align with prior work showing substantial opportunity for energy generation using DW technologies. Key findings reveal a substantial increase from prior results in estimated technical and economic potential for DW. Metrics tuned to highlight economic potential also show increased incentives supporting rural adoption. Results are spatially aggregated for usability and published via the U.S. Department of Energy Wind Data Portal and a custom scenario visualization platform, aiding policymakers, industry, and property owners in assessing DW viability across various scenarios and spatial scales.

17 WIND ENERGY↗

Idealized simulations of wind farm interactions with intermittent turbulence in stable boundary layer conditions

Stable atmospheric boundary layer conditions typically correspond to weak turbulence levels, but intermittent periods of elevated turbulence can occur during otherwise quiescent conditions. The interaction between intermittent turbulence and wind turbines is not well understood because of sparse observations, as well as the difficulty in realistically resolving small-scale turbulence during strongly stable conditions with numerical simulations. In this study, an explicit filtering and reconstruction approach for large-eddy simulation (LES) is used to simulate weakly and strongly stable conditions, with surface cooling rates of −0.2 and −2.0 K h −1 , respectively. This approach can sustain resolved background turbulence at relatively coarse grid spacing and stronger stratification compared to conventional closures, permitting more realistic intermittent stable boundary layer (SBL) turbulence. The idealized LES capability of the Weather Research and Forecasting model is employed with turbine rotors parameterized using generalized actuator disks to examine (1) how the presence of turbine wakes affects SBL evolution and (2) the effect of intermittent turbulence on power production and wake recovery. Wakes increase mixing and deepen the SBL, with a stronger effect under strongly stable conditions, primarily because the SBL is shallower and closer to the top of the wind turbine rotor layer. Intermittent turbulence does not have a significant impact on mean power generation and wake recovery because the relevant intermittent turbulent structures in this study only affect the bottom half of the rotor disk. Power production is, however, more variable during periods of elevated turbulence, demonstrating the impact of SBL intermittency. This study uses an idealized configuration, focusing on LES model performance and physical understanding, with the goal of informing future simulations of the conditions observed during the American Wake Experiment.

Energy - Wind↗

Guideline for Characterizing and Evaluating a Candidate Project Site for Solar Thermal Applications

This document presents a structured procedure for characterizing and evaluating candidate project sites for concentrating solar power (CSP) and solar heat for industrial processes (SHIP) applications. The objective is to provide project developers, researchers, and other stakeholders with a consistent, technology-agnostic framework for early-stage site assessment, enabling informed decision-making prior to significant investment in project development. Site selection is a critical factor in project success or failure for both CSP and SHIP projects. Key factors such as solar resource availability, land characteristics, environmental and regulatory constraints, infrastructure availability, and community context are determined by the choice of project site and can materially impact project performance, cost, schedule, and overall viability. This procedure is designed to systematically evaluate these factors, identify potential fatal flaws, and prioritize the most favorable candidate sites for further development. The process begins with rapid screening-level evaluation, using publicly available data to assess solar resource, land availability and suitability, zoning and land-use compatibility, and exclusion zones such as protected lands or sensitive habitats. Sites that meet the minimum screening criteria advance to a more detailed characterization. Subsequent sections of this report provide guidance for a next-level assessment of the most important technical and environmental parameters, including: 1) Solar resource quality, variability, and uncertainty using multiyear datasets and, where appropriate, on-site measurement campaigns; 2) Meteorological conditions such as wind, temperature, extreme weather events, and soiling impacts; 3) Land characteristics including slope, shading, and geotechnical conditions; and 4) Environmental and regulatory considerations, including permitting processes, endangered species, cultural resources, and visual impacts. The procedure also addresses infrastructure and integration considerations, including: 1) Grid interconnection requirements for CSP power generation projects; 2) Electrical and operational integration for SHIP facilities; 3) Water availability, quality, and permitting constraints, which are particularly critical for CSP in arid regions; and 4) Site access, construction logistics, and availability of workforce and supporting services. Recognizing the importance of social and economic context, the procedure includes evaluation of community engagement factors, such as stakeholder sentiment, proximity to sensitive visual receptors, workforce development opportunities, and local economic incentives. The outputs of these assessments are synthesized in a cost and risk evaluation, translating site characteristics into expected impacts on capital cost, operating cost, schedule, and technical risk. This is complemented by screening-level performance modeling, including 8760 simulations and long-term projections, to quantify expected energy or thermal output, assess variability thereof, and support comparison between candidate sites. Finally, the procedure provides high-level guidance on a structured go/no-go decision framework, categorizing sites based on identified risks and constraints, and outlining a clear path forward to feasibility studies and front-end engineering design for viable projects. By standardizing the site characterization process across both CSP and SHIP applications, this guideline aims to: 1) Improve consistency and transparency in early-stage project evaluation; 2) Reduce development risk and avoid investment in nonviable project sites; 3) Support collaboration between developers, researchers, and public agencies; and 4) Accelerate successful deployment of concentrating solar technologies for both power generation and industrial process heat.

14 SOLAR ENERGY↗

Multiobjective Constrained Symbolic Regression for Predictive Modeling of Material Creep Behavior

When creep testing is repeated on samples of the same alloy under the same parametric conditions (i.e., stress and temperature), the resulting strain/time curves can vary from each other considerably as shown in Figure 1 [1]. The time required to creep test a material to rupture can extend to the order of years. Because of this, a numerical model that can quickly analyze the incomplete results of an ongoing experiment to predict 1) the incomplete portion of the strain/time curve leading up to the rupture point and 2) the rupture point itself would be of great utility to the materials community. Such a model has the potential to save 1) the time required to finish running the experiment to rupture 2) the associated monetary cost of finishing said experiment. Furthermore, it would be advantageous if the predictive model could give a parametric function modeling strain/time curves for material scientists to investigate the impact of the temperature and stress parameters on the resulting creep behavior. This work introduces a piecewise symbolic regression algorithm to predict the remainder of the strain/time curve. Preliminary results show good model performance.

36 MATERIALS SCIENCE↗

Pre-Transient Characterization of Historic EBR-II Pins for Transient Testing

Current interest in sodium-cooled fast reactor (SFR) designs, such as TerraPower’s Natrium Reactor, has highlighted the need for advanced reactor fuel technology development. Modern U-Zr and U- Pu-Zr pin designs are primary candidates to fuel SFRs and boast high fuel utilization capacity, increased fuel-cladding compatibility, and improved safety through inherent feedback mechanisms. Despite over 60 years of metallic fuel irradiation, uncertainties exist in the performance of the fuel system, particularly under transient overpower (TOP) and loss of flow (LOF) scenarios. Throughout historical testing within the Experimental Breeder Reactor II (EBR-II) and the Fast Flux Test Facility (FFTF), fuel behavior has demonstrated benign response to transient reactor conditions; however, accurate predictions of failure thresholds to inform operational limitations rely heavily on fuel composition, burnup, and irradiation history. In expanding TOP and LOF testing, the Transient Heat sink Overpower Response (THOR) Capsule will be used to test modern fuel technologies in a static sodium environment in the Transient Reactor Test (TREAT) Facility. The THOR capsule is highly instrumented and will provide time-dependent thermal behavior of SFR fuel pins subjected to accident conditions within TREAT. The THOR-Metallic (THOR- M) campaign aims to validate and expand historical TOP and LOF testing on high burnup U-Zr and U-Pu- Zr fuel alloys previously irradiated in EBR-II by running the rods to failure. This contribution focuses primarily on the pre-transient engineering-scale destructive and non- destructive characterization that has been conducted on both the test and sibling pins used for the TOP and LOF tests. All pins underwent visual examination, neutron radiography, element contact profilometry, and precise gamma scan. The sibling pins used for each test were further analyzed using gas assay, sampling, and recharge analysis (GASR), and optical microscopy. The results from each technique confirmed that the fuel pins were intact and devoid of any atypical developments when compared to historical data. Additionally, the analyzed measurements establish a baseline for comparison to post-transient analysis. Key fuel behaviors quanitifed include axial elongation of the fuel column, diametral strain of the pin, patterns in fluff structure geometry, changes in axial isotope distribution, evolution of constituent redistribution, porosity, and fission gas release. The pre-transient measurements and changes attributed to transient behavior from post-transient measurement will be compared to historical data to capture the behavioral dependence on composition, burnup, and irradiation history. Results from this work advance the initiatives of the THOR-M campaign, which aid in informing fuel performance models and establishing safety criteria for SFR operational limits. The novel combination of test environment, in-situ instrumentation, and comprehensive suite of characterization methods provides greater understanding of transient fuel behavior. Overall, information on the time and condition of pin failure for high burnup U-Pu-Zr will greatly expand the limited existing TOP and LOF test data.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

BISON Simulated and Experimental Fission Product Release Comparisons from Reradiated AGR-3/4 Compacts During High Temperature Heating Tests

The fuel performance modeling code BISON was used to predict the release of fission products iodine-131 (131I), xenon-133 (133Xe), and krypton-85 (85Kr) from four re-irradiated AGR-3/4 fuel compacts containing tristructural isotropic (TRISO) coated particles during high-temperature isothermal heating tests. The AGR-3/4 fuel compacts were irradiated in the Advanced Test Reactor (ATR) as part of the third and fourth series of planned experiments to support the Advanced Gas Reactor (AGR) Program. They were subsequently stored and re-irradiated in the Neutron Radiography (NRAD) reactor for approximately five days and then stored for another five to eight days before being subjected to isothermal heating tests in the Fuel Accident Condition Simulation (FACS furnace) for 200 to 300 hours at temperatures between 1000°C and 1600°C to evaluate fission product release at elevated temperatures. New nuclide-specific fission product source term models for the three nuclides of interest were developed using the reactor multiphysics code Griffin and implemented into BISON to support this work. The new source term models were incorporated into coupled compact- and particle-scale BISON simulations, which predict spatially- and temporally-resolved radionuclide generation, radioactive decay, transport, and release throughout the entire irradiation history, including the initial ATR irradiation, NRAD re-irradiations, FACS heating tests, and intermediate periods spent in storage. The experimentally measured fission product release from the heating tests were compared to modeling release predictions calculated by BISON to evaluate how well the code compares to experimental results. Overall, the experimental measured and BISON predicted comparative release results varied but generally agreed to within 5 particle equivalents. Comparative release results identified general observations to take into consideration to help refine future models and reduce uncertainties associated with both the measurement results and predictive results. This includes developing new uranium oxycarbide (UCO) specific kernel diffusivities for the three isotopes examined to more accurately reflect the material properties of the fuel form. Deriving new diffusivities will aid in producing a more informed BISON model

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Pre-Transient Characterization of Historic EBR-II Pins for Transient Testing

Current interest in sodium-cooled fast reactor (SFR) designs, such as TerraPower’s Natrium Reactor, has highlighted the need for advanced reactor fuel technology development. Modern U-Zr and U- Pu-Zr pin designs are primary candidates to fuel SFRs and boast high fuel utilization capacity, increased fuel-cladding compatibility, and improved safety through inherent feedback mechanisms. Despite over 60 years of metallic fuel irradiation, uncertainties exist in the performance of the fuel system, particularly under transient overpower (TOP) and loss of flow (LOF) scenarios. Throughout historical testing within the Experimental Breeder Reactor II (EBR-II) and the Fast Flux Test Facility (FFTF), fuel behavior has demonstrated benign response to transient reactor conditions; however, accurate predictions of failure thresholds to inform operational limitations rely heavily on fuel composition, burnup, and irradiation history. In expanding TOP and LOF testing, the Transient Heat sink Overpower Response (THOR) Capsule will be used to test modern fuel technologies in a static sodium environment in the Transient Reactor Test (TREAT) Facility. The THOR capsule is highly instrumented and will provide time-dependent thermal behavior of SFR fuel pins subjected to accident conditions within TREAT. The THOR-Metallic (THOR- M) campaign aims to validate and expand historical TOP and LOF testing on high burnup U-Zr and U-Pu- Zr fuel alloys previously irradiated in EBR-II by running the rods to failure. This contribution focuses primarily on the pre-transient engineering-scale destructive and non- destructive characterization that has been conducted on both the test and sibling pins used for the TOP and LOF tests. All pins underwent visual examination, neutron radiography, element contact profilometry, and precise gamma scan. The sibling pins used for each test were further analyzed using gas assay, sampling, and recharge analysis (GASR), and optical microscopy. The results from each technique confirmed that the fuel pins were intact and devoid of any atypical developments when compared to historical data. Additionally, the analyzed measurements establish a baseline for comparison to post-transient analysis. Key fuel behaviors quanitifed include axial elongation of the fuel column, diametral strain of the pin, patterns in fluff structure geometry, changes in axial isotope distribution, evolution of constituent redistribution, porosity, and fission gas release. The pre-transient measurements and changes attributed to transient behavior from post-transient measurement will be compared to historical data to capture the behavioral dependence on composition, burnup, and irradiation history. Results from this work advance the initiatives of the THOR-M campaign, which aid in informing fuel performance models and establishing safety criteria for SFR operational limits. The novel combination of test environment, in-situ instrumentation, and comprehensive suite of characterization methods provides greater understanding of transient fuel behavior. Overall, information on the time and condition of pin failure for high burnup U-Pu-Zr will greatly expand the limited existing TOP and LOF test data.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Pre-Transient Characterization of MLOF-1 Test Pin

This study focuses on the pre-transient characterization of U-10Zr test and sibling fuel pins for the THOR-M-LOF test series. Using neutron radiography, element contact profilometry (ECP), precise gamma scan (PGS), and gas assay, sampling, and recharge (GASR) analysis, it was confirmed that the fuel pins were intact and suitable for testing. Key fuel behaviors quantified include axial elongation, diametral strain, fluff structure geometry, axial isotope distribution, and fission gas release. Any deviations from historically expected behaviors were investigated and attributed to factors other than the irradiation behavior of the fuel pin. These pre-transient measurements establish a baseline for future post-transient analysis, which will be used to inform fuel performance models and safety criteria for sodium-cooled fast reactors (SFRs). The results will enhance understanding of transient fuel behavior and expand limited data on LOF scenarios.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

AGR-5/6/7 Ceramography Report

The Advanced Gas Reactor (AGR) Fuel Development and Qualification Program was established in 2002 to conduct research and development on tri-structural isotropic (TRISO)-coated particle fuel for High Temperature Gas-cooled Reactors (HTGRs). All AGR irradiations have been completed: AGR-1 (Collin 2015), AGR-2 (Collin 2014), AGR-3/4 (Collin 2016) and AGR-5/6/7 (Pham 2021). The objectives of each program are presented in Figure 1.1. The purpose of the AGR-5/6/7 program was to provide a baseline fuel qualification data set to support licensing, deployment, and operation of HTGRs in the United States. To achieve these goals, the program includes fuel fabrication, irradiations of TRISO fuels and high-temperature materials, safety testing and post-irradiation examination (PIE), fuel performance modeling, and fission product (FP) transport (Sharp 2020). The AGR-5/6/7 work continues as a final irradiation program of the Advanced Reactor Technologies (ART) program, that began in February 2018 and ended in July 2020. The AGR-5/6/7 fuel compacts were irradiated for a total of approximately 360.9 effective full power days (EFPDs) (Stempien 2023), resulting in final burn-up values, on a per-compact basis, ranging from 5.66% to 15.26% fissions per initial heavy metal atom (FIMA), and fast fluence values ranging from 1.62×1025 n/m2 to 5.55×1025 n/m2 (E >0.18 MeV) (Stempien 2022)

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Mechanistic Fission Gas Release Uncertainty Induced by Microstructure Data

Fission Gas Release (FGR) is an important engineering safety parameter for nuclear fuel. While fuel performance modeling with BISON currently relies on mechanistic models to predict it, comparison with experimental data shows both under or over prediction depending on operation mode (steady or transient).Predicting microstructure data is essential to accurately predicts the engineering scale parameters. An important source of uncertainty in mechanistic models arises from the missing captured physics. Continuous validation and refinement of these models against experimental data are also necessary to ensure their reliability and accuracy in predicting engineering parameters.

36 - MATERIALS SCIENCE↗

High Spatial Resolution Mapping of Retained Fission Gas

Fission gas isotopic analysis provides quantitative high precision determination of irradiated nuclear fuel burnup, offers diagnostic value, and informs fuel performance models. A measurement capability has been developed at Idaho National Laboratory (INL) for the release of retained fission gas using a focused laser and static noble gas mass spectrometry (MS) analysis. This high resolution (10s microns spot size) capability was demonstrated using Xe implanted metal foils.

07 - ISOTOPES AND RADIATION SOURCES↗

Archi: Agentic Operations at the CMS Experiment

We present Archi, an open-source, end-to-end framework for scientific collaborations that combines the systematic ingestion and organization of heterogeneous data sources with the deployment of configurable, private, and extensible agents that retrieve and reason over them. An instance of Archi has been deployed for the Computing Operations team of the CMS experiment at CERN's LHC since February 2026 as a support agent for technical operators, offering retrieval and analysis capabilities by combining documentation, historical data, and live monitoring systems. We evaluate the system on operator feedback and a question set collected from production usage, graded by human and automated panels. The system proves effective at operational tasks, resolving real-world queries posed by CMS operators. We also observe that locally-hosted, open-weight models perform competitively, enabling fully private management of sensitive data.

Lugato, Pietro [MIT; CERN]↗

Near-Optimal Performance of Stochastic Model Predictive Control

Here, this article presents a regret analysis for stochastic model predictive control (SMPC) in linear systems with quadratic performance index and additive and multiplicative uncertainties. Under a finite support assumption, the problem can be cast as a finite-dimensional quadratic program, but the problem becomes quickly intractable as the problem size grows exponentially in the horizon length. SMPC aims to compute approximate solutions by solving a sequence of problems with truncated prediction horizons and committing the solution in a receding-horizon fashion. Although this approach is widely used in practice, its performance relative to the optimal solution is not well understood. This article reports for the first time a rigorous near-optimal performance guarantee of SMPC: under stabilizability and detectability conditions, the regret of SMPC is exponentially small in the prediction horizon length, allowing SMPC to achieve near-optimal performance at a substantially reduced computational expense.

93E20, 93B45↗

Cross-Domain Reasoning for Neuromorphic Model Design

Designing performant neuromorphic models requires reasoning across neuroscience, neuromorphic computing, and machine learning, making it a natural target for cross-domain hypothesis generation. Our primary contribution is a multi-corpus knowledge graph spanning all three domains, which we show substantially increases cross-domain retrieval novelty over single-corpus baselines. We additionally introduce NeuKReAct, an agentic reasoning framework that iteratively retrieves from this graph and synthesizes design hypotheses via a step-by-step blackboard architecture, enabling structured compartmentalization of design decisions. Lastly, we introduce an execution head that translates hypotheses into structured design documents and runnable code. We evaluate novelty using a combinatorial creativity metric that measures cross-domain retrieval distance across the citation graph. Our results confirm that corpus breadth is the dominant driver of novelty. Moreover, we highlight a concrete instance of the novelty-utility tradeoff within NeuKReAct, underscoring a need for joint creativity evaluation, balancing both novelty and utility.

Ramavarapu, Vikram [ORNL] (ORCID:0009000188757213)↗

Bridging interfacial properties and cell performance: A multiscale model for proton-exchange-membrane fuel cells

Here, to elucidate the impact of local interfaces on mass-transport resistance and overall cell performance of low-loaded proton-exchange-membrane fuel cells (PEMFCs), we present a multiscale modeling framework incorporating a novel modified agglomerate model. The model considers three distinct Pt-electrolyte interfaces: Pt on the carbon surface covered by either ionomer or water film and Pt inside carbon nanopores. Detailed mass-transport voltage-loss breakdowns reveal that coupled agglomerate-interface-scale mass transport dominates the mass-transport loss. The ionomer poisons the exterior-Pt surface through suppressing O 2 adsorption and intrinsic ORR activity, leading to low current-density performance. Conversely, interior-Pt interface enhances the kinetic performance but limits high current-density performance due to its low interfacial permeability. The exterior-Pt/water interface demonstrates superior kinetic performance and mass transport, though its practical implementation requires ensuring proton transport. By coupling the multiscale CL properties with ink parameters, the model identifies an optimal I to C ratio of approximately 0.5, a moderate value where the ionomer content is sufficient to guarantee proton transport without fully covering the Pt surface and forming large agglomeration, thus allowing the utilization of the Pt-water interface and avoiding high mass-transport loss. Overall, the model helps unravel limiting phenomena across different operating regimes and provides routes for optimizing performance.

Cell diagnostic↗

Comparative Performance of Gaussian Plume and Backward Lagrangian Stochastic Models for Near-Field Methane Emission Estimation Using a Single Controlled Release Experiment

Methane (CH 4 ) is a major component of natural gas and a potent greenhouse gas. Increasing atmospheric methane concentrations are attributed to emissive anthropogenic activities by an average of 13 ppb per yr since 2020 and are linked to a changing global climate. Mitigating CH 4 emissions from oil and gas production sites has recently become a target to reduce overall greenhouse gas emissions; however, monitoring the efficacy of mitigation strategies depends on accurate quantification of CH 4 emissions at the facility-level. Near-field quantification of methane (CH 4 ) emissions from oil and gas (O&G) facilities remains challenging due to the effects of atmospheric variability and sensor configuration on atmospheric dispersion models. This study evaluates the performance of two atmospheric dispersion models, the Gaussian plume (GP) and backward Lagrangian stochastic (bLS), by comparing calculated CH 4 emissions to controlled single-point emissions between 0.4 and 5.2 kg CH 4 h −1 . Emissions were calculated by both models using 121 individual sets of measurements comprising five-minute averaged downwind methane mixing ratios and matching meteorological data. The comparison shows that the bLS approach achieved a higher proportion of emission estimates within a factor of two (FAC2) of the known emission rates compared to the GP approach. The emissions calculated by the bLS model also had a lower multiplicative error and reduced bias relative to GP. Other error-based metrics further confirmed the bLS model performed better, as it yielded lower RMSE and MAE than GP. Statistical analysis of the emission data shows that the lateral and vertical alignment of the source and the sensor plays a critical role in emission estimations, as measurements made closer to the plume centerline and at a distance between 40 and 80 m downwind yielded the best FAC2 agreement. High wind meander degraded the ability of both approaches to generate representative emissions, particularly with the GP approach, as it violates the modeling approach’s assumption of steady-state emissions. Data suggest emissions calculated by the bLS model are comprehensively in better agreement, but the computational demands of the modeling approach and integration into fenceline systems limit real-time applicability. While these results provide insight into model performance under controlled near-field conditions, their applicability to more complex or heterogeneous oil and gas production environments (e.g., the regions Marcellus or Unita Basins) remains limited and uncertain.

gaussian plume↗

Porting ATLAS Fast Calorimeter Simulation to GPUs with Performance Portable Programming Models

FastCaloSim is a parameterized simulation of the particle energy response and of the energy distribution in the ATLAS calorimeter. It is a relatively small and self-contained package with massive inherent parallelism and captures the essence of GPU offloading via important operations like data transfer, memory initialization, floating point operations, and reduction. It was identified by the High Energy Physics Center for Computational Excellence project as a good testbed for evaluating the performance and ease of portability of programming models. In this paper, we will discuss the results of our evaluation of the porting process to Kokkos, SYCL, Alpaka, OpenMP and std::par (nvc++), and compare performance on NVIDIA, AMD and Intel GPUs, as well as multicore CPUs.

97 MATHEMATICS AND COMPUTING↗

Predictive Indicators of the Performance of Large Language Models

In several mission contexts, it is desirable to estimate the performance of large language models (LLMs) on tasks that we cannot run directly. In light of published “scaling laws” our hypothesis is that some tasks should be consistently more challenging than others based on characteristics of the task. The goal of this project was to begin quantifying how much information about LLM performance can be gained from the features of a model and a task. Two of our statistical models struggled to converge. Pass/fail test results may provide limited information for inference beyond model quality and task difficulty, but we see no evidence at this time for significant feature interaction effect sizes, arguing for simple models. Future work extending the models to capitalize on perplexity of ground truth answers is suggested. This project also introduces “Depth of Knowledge Variant Testing” as a strategy for more finely assessing language models on open domain question and answer tasks. We developed sets of questions that ask a language model to produce similar information while demonstrating increasing depth of knowledge, and also relabeled existing Q&A test questions with their depth of knowledge. Our results suggest further consideration of Bloom’s taxonomy and further refinement of prompts to properly elicit information at varying depths. In the course of this work, we set up a basic infrastructure for standardizing tasks and testing many language models on these tasks. In addition to testing the predictive quality of model features and performance across test suites, with this project we have introduced two new task features to contextualize each test question: the Dewey Classification main category of information covered, and the Bloom’s taxonomy level that corresponds to the depth of knowledge probed by the question. Splits across these and other features produced over five hundred task subtypes with distinct feature vectors, which we tested on half a dozen models.

97 MATHEMATICS AND COMPUTING↗