Design Task Group Update
Slides detailing Primary Subtask 1 and 2 with the Graphite Review, accomplishments, studies, and on-going subtasks for the timeline of January 2023 through May 2023.
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Slides detailing Primary Subtask 1 and 2 with the Graphite Review, accomplishments, studies, and on-going subtasks for the timeline of January 2023 through May 2023.
The American Society of Mechanical Engineers Boiler Pressure and Vessel Code (ASME BPVC) Section III, Division 5, Article HHA-3000 outlines graphite core component and graphite core assembly design guidelines. Graphite core components are defined as ?components manufactured from graphite that are installed to form a graphite core assembly within the reactor pressure vessel of a high temperature, graphite moderated fission reactor.? (p. 413) Graphites? inherent defect distributions do not allow for deterministic material reliability. Rather, graphite has variable strength distributions which change by grade. Article HHA-3000 outlines two semi-probabilistic methods, the full and simplified assessments, which set design load limit targets for each of three component structural reliability classes. The Design Task Group was officially recognized as a specialized task group within ASME November of 2023, though we?ve been collaborating since 2022. The purpose of the Design Task Group is to correct, clarify, and make HHA-3000 function as intended. The Design Task Group will sunset once we?ve achieved our objectives. The Design Task Group was specifically told to not write new Code. While there may be more precise and more accurate methods to determine reliability targets, the current methods are conservative, relatively simple to implement, and have thus far been considered satisfactory for setting design reliability targets. Much of the ground-work to write proposal files and background documents for records to make the changes needed to achieve our objective have been completed. The Design Task Group has documented much of their work through papers, presentations, and memorandums. Three memorandums in which INL team members had substantial contributions are found in the Appendices: FEA Modeling for the Baseline Program, Evaluating the Effects on Margin of Updating the Threshold and Shape Parameters in the Full Assessment, and Interpretations of the Full and Simplified Assessments in ASME BPVC. Most of the on-going work to achieve the Design Task Group?s objective will be addressing comments on existing records and moving records through the balloting process. The Design Task Group met bi-weekly mostly through the end of FY2023. Since February 2024, the Design Task Group has mostly been completed with solving and documenting the technical issues associated with the assessments. Unless new tasks are identified, the remaining work of the Design Task Group will be political and editorial.
This report provides an update on the status of the composite code rule development for the American Society of Mechanical Engineers (ASME) Boiler Pressure Vessel (BPV) code Section III Division 5 Subsection HA Subpart B (HAB) and Subsection HH Subpart B (HHB) during FY 2025 (from October 2024 to September 2025). In the context of this document, reference to “the code” is specific to these subsections unless otherwise specified. A composites task group under the ASME Nonmetallic Design and Materials Working Group (WG-NDM), with the support of external experts, was established and is now working toward addressing previously identified areas.
Graphite is an important material being used for core components in next-generation high-temperature gas-cooled nuclear reactors. The selection of graphite grade for a specific Designer is a complex task, dependent on reactor conditions, component functionality, and required reliability. The American Society of Mechanical Engineers (ASME) provides two semi-probabilistic design-by-analysis assessments to evaluate graphite core components against design reliability targets. The simplified assessment uses a 2-parameter Weibull distribution to describe the graphite grade’s tensile-strength distribution to establish component stress limits. The full assessment uses the 3-parameter Weibull distribution and a modified Weakest-Link Theory approach to calculate a component design probability of failure. The paper defines recommended assessment rules, which are the as-written simplified assessment and the full assessment with parameter lower bounds, the modulus update with threshold reduction, and the 2027 grouping rules. Code rules are applied to three grades: 2114, IG-110, and NBG-18. The baseline margin calculation is developed using the experimental tensile dogbone specimen. Percent margin is defined as the percent reduction in the median experimental load to obtain the allowable load per ASME assessments. Under the recommended rules, the SRC–1 margin in the simplified assessment ranged from 40.2 % to 52.7 % among the grades in this study and from 36.1 % to 49.8 % in the full assessment. The full assessment only decreases the margins by 2.5–4.5 % for the SRC-1 components and 0–1.5 % for the SRC-2 components for this baseline case. Margin is inversely related to material median strength (i.e., the strongest grade, 2114, has the lowest margin).
The Artificial Intelligence Enhanced Co-Design for Next Generation Microelectronics virtual workshop was held April 4-5, 2023, and attended by subject matter experts from universities, industry, and national laboratories. This was the third in a series of workshops to motivate the research community to identify and address major challenges facing microelectronics research and production. The 2023 workshop focused on a set of topics from materials to computing algorithms, and included discussions on relevant federal legislation and such as the Creating Helpful Incentives to Produce Semiconductors and Science Act (CHIPS Act) which was signed into law in the summer of 2022. Talks at the workshop included edge computing in radiation environments, new materials for neuromorphic computing, advanced packaging for microelectronics, and new AI techniques. We also received project updates from several of the Department of Energy (DOE) microelectronics co-design projects funded in the fall of 2021, and from three of the Energy Frontier Research Centers (EFRCs) that had been funded in the fall of 2022. The workshop also conducted a set of breakout discussions around the five principal research directions (PRDs) from the 2018 Department of Energy workshop report: 1) define innovative material, device, and architecture requirements driven by applications, algorithms, and software; 2) revolutionize memory and data storage; 3) re-imagine information flow unconstrained by interconnects; 4) redefine computing by leveraging unexploited physical phenomena; 5) reinvent the electricity grid through new materials, devices, and architectures. We tasked each breakout group to consider one primary PRD (and other PRDs as relevant topics arose during discussions) and to address questions such as whether the research community has embraced co-design as a methodology and whether new developments at any level of innovation from materials to programming models requires the research community to reevaluate the PRDs developed back in 2018.
Large language models (LLMs) are increasingly adapted to downstream tasks via reinforcement learning (RL) methods like Group Relative Policy Optimization (GRPO), which often require thousands of rollouts to learn new tasks. We argue that the interpretable nature of language often provides a much richer learning medium for LLMs, compared to policy gradients derived from sparse, scalar rewards. To test this, we introduce GEPA (Genetic-Pareto), a prompt optimizer that thoroughly incorporates natural language reflection to learn high-level rules from trial and error. Given any AI system containing one or more LLM prompts, GEPA samples trajectories (e.g., reasoning, tool calls, and tool outputs) and reflects on them in natural language to diagnose problems, propose and test prompt updates, and combine complementary lessons from the Pareto frontier of its own attempts. As a result of GEPA's design, it can often turn even just a few rollouts into a large quality gain. Across six tasks, GEPA outperforms GRPO by 6% on average and by up to 20%, while using up to 35x fewer rollouts. GEPA also outperforms the leading prompt optimizer, MIPROv2, by over 10% (e.g., +12% accuracy on AIME-2025), and demonstrates promising results as an inference-time search strategy for code optimization. We release our code at https://github.com/gepa-ai/gepa.
Information obtained from Fukushima Daiichi Nuclear Power Station (Daiichi) is required to inform future Decontamination and Decommissioning (D&D) activities, improving the ability of the Tokyo Electric Power Company Holdings, Incorporated (TEPCO Holdings) to characterize potential hazards and to ensure the safety of workers involved with cleanup activities. This information also has important implications for the safety and operation of U.S. commercial nuclear power plants. This document summarizes results from the Fiscal Year 2023 (FY2023) U.S. effort to review Daiichi information and extract insights to enhance the safety of existing and future nuclear power plant designs. This U.S. effort, which was initiated in 2014 by the Department of Energy Office of Nuclear Energy, is completed by a group of experts in reactor safety and plant operations that identify examination needs and evaluate recent Daiichi examination data to address these needs. Fukushima-related information and associated discussions during these meetings benefit operating, new, and advanced reactors. Significant safety insights have been and are continuing to be obtained in several areas: system and component performance, radionuclide surveys and sampling, debris end-state location, combustible gas effects, and plant operations and maintenance. In addition to reducing uncertainties related to severe accident modeling progression, these insights have and continue to be used to update guidance for severe accident prevention, mitigation, and emergency planning. Furthermore, Daiichi-related activities, such as code modeling improvements and analysis, testing, and new technology deployment efforts, have the potential to offer additional benefits to the operating fleet and new LWR and non-LWR designs. U.S. evaluations of obtained examination information and input regarding future Daiichi examinations are of interest to several organizations within Japan. Since its inception, the U.S. has provided consensus input for high priority time-sequenced examination tasks and supporting research activities. In their Mid-to-Long-term Examination Plan for 1F investigations, TEPCO included all remaining U.S. consensus information requests and additional information requests they identified. TEPCO periodically provides reports on the status of these requests (reflecting D&D priorities, new insights from investigations, and new technologies that become available). Hence, U.S. experts agreed that it was appropriate for TEPCO to track and prioritize these information requests as D&D progresses. U.S. experts will continue to review and comment on the information obtained from examinations and, as needed, provide additional details and relevant background material to support future examinations. As documented in this report, several other items, such as additional details on information requests pertaining to ex-vessel examinations, relevant references from prior research, additional documents to provide insights regarding recent investigation findings, and reviews of recently released documents, were agreed to during the FY2023 meeting.
Steel Thread is a NA-22 venture that seeks to build trustworthy, reliable AI models that can be used in a wide variety of nonproliferation tasks. A key aspect of building these models is developing appropriate benchmarks and evaluation methods, which will enable the venture to identify and adapt models to provide the most value in the nonproliferation domain. Benchmarks must be relevant to key tasks in this domain, such as question answering, information retrieval, document summarization and classification, consensus analysis, and image and data analysis. This report 1) provides an overview of benchmark design, evaluation, and challenges; 2) reviews a variety of open benchmarks, with a focus on language models and tasks; and 3) identifies benchmarks that are most relevant to Steel Thread. This report is intended to serve as a basis for further efforts to classify and evaluate benchmarks and their correlation with success on nonproliferation-specific tasks. The Steel Thread venture has defined benchmarks to be a particular combination of a dataset (or datasets) and a metric (or metrics) conceptualized as representing one or more specific tasks or sets of abilities for a specific modality. It is adopted by a research community as a shared framework for comparing methods.1 It includes 1) Data: Labeled (a designated subset not used for training, which could be all the data), 2) Metric: A way to quantify performance, 3) Task/Ability: What the benchmark is testing, 4) Protocol: A structured and repeatable evaluation process, 5) Baseline/Reference Model: For comparison; could be statistical, rule-based, SME-derived, or another model, and 6) Maintenance Plan: to update with new information over time; important for long-term utility. For further clarity, the definition includes what a benchmark, in this context, is not. It is not a corpus of training data, specific to a model (it is intended to apply to a range of models), a universal evaluation of performance, a guarantee that the ‘top’ model on the leaderboard will be the best fit for every specific use case, an all-encompassing proof of a model’s universal quality, nor is it a one-size-fits-all measure of success. It does not cover every real-world constraint (like operational, ethical, or cost considerations), a systems integration test, or a unit test. This definition was inspired by and resulted from discussions within the Steel Thread Benchmarking Task Force. This group was formed to define what we would mean as a benchmark within Steel Thread but persisted as the need to develop a thorough understanding of the large and expanding existing benchmarking space. This technical report is a result of the group’s divide and conquer approach to exploring this space. The release of benchmarks might not be progressing as quickly as model development, but it is moving very fast, as many benchmarks quickly become saturated, when state-of-the-art models score so close to the benchmark’s ceiling that their results are virtually indistinguishable. At that point, the test no longer differentiates between new systems, so researchers usually stop reporting scores as the benchmark no longer informs about improvements from the next generation of models. In the OpenAI announcement of GPT-5, they reported results on six flagship public benchmarks (AIME 2025, SWE-bench Verified, Aider Polyglot, MMMU, HealthBench Hard, GPQA) but the full system-card covers roughly thirty-five separate evaluations, comprising hundreds of test task items in total. There have been some efforts to summarize benchmarks in specific fields, like for text-to-image generation, but these surveys have had a narrow methodology scope. Therefore, a comprehensive survey of all benchmarks or even all benchmarks that could be relevant to Steel Thread is outside of the scope of this report. We chose some specific benchmarks to investigate in detail.
According to US Census Bureau, the elderly population (individuals 65 years and older) has grown by over a third during the past decade (2010 to 2019), and by 3.2% from 2018 to 2019. It is essential for policymakers and planners to understand transportation issues associated with the elderly to meet their increasing travel demands. These issues include transportation and mobility of the elderly population, factors impacting their travel behavior, and transportation safety. In this study, Oak Ridge National Laboratory was tasked by the New York State Department of Transportation (NYSDOT) to conduct a detailed examination of travel behaviors and identify patterns and trends of its elderly residents. The National Household Travel Survey (NHTS) was used as the primary data source to analyze subjects and address questions such as: Are there differences in traveler demographics between the elderly population and those of younger age groups who live in various New York State (NYS) regions, e.g., New York City (NYC), other urban areas of NYS, or other parts of the country? How do they compare with the population at large? Are there any regional differences (e.g., urban versus rural)? Do any unique travel characteristics or patterns exist within the elderly group? How did these patterns change over time? In addition to the analysis of NHTS data, roadway travel safety concerns associated with elderly travelers were also investigated. Specifically, data on crashes involving the elderly (including drivers, passengers, and pedestrians) as captured in the Fatal Analysis Reporting System database was analyzed to examine elderly drivers and elderly pedestrian travel safety issues in NYS. This study report provides a summary of travel behavior and social-demographic characteristics of NYS elderly residents. These statistics could be used to examine equity issue concerning elderly New Yorkers, as well as to evaluate how well their mobility needs are being met. With a deeper understanding of issues and needs that this special population group is facing, policymakers and transportation planners would be able to make informed decisions on transportation investments and design services that could better address them.