Search NASA⌕ Search

SEARCH · Search NASA

Results for “design task group update”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Design Task Group Update

Slides detailing Primary Subtask 1 and 2 with the Graphite Review, accomplishments, studies, and on-going subtasks for the timeline of January 2023 through May 2023.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

ASME Design Code Rule Changes for Nuclear Graphite

The American Society of Mechanical Engineers Boiler Pressure and Vessel Code (ASME BPVC) Section III, Division 5, Article HHA-3000 outlines graphite core component and graphite core assembly design guidelines. Graphite core components are defined as ?components manufactured from graphite that are installed to form a graphite core assembly within the reactor pressure vessel of a high temperature, graphite moderated fission reactor.? (p. 413) Graphites? inherent defect distributions do not allow for deterministic material reliability. Rather, graphite has variable strength distributions which change by grade. Article HHA-3000 outlines two semi-probabilistic methods, the full and simplified assessments, which set design load limit targets for each of three component structural reliability classes. The Design Task Group was officially recognized as a specialized task group within ASME November of 2023, though we?ve been collaborating since 2022. The purpose of the Design Task Group is to correct, clarify, and make HHA-3000 function as intended. The Design Task Group will sunset once we?ve achieved our objectives. The Design Task Group was specifically told to not write new Code. While there may be more precise and more accurate methods to determine reliability targets, the current methods are conservative, relatively simple to implement, and have thus far been considered satisfactory for setting design reliability targets. Much of the ground-work to write proposal files and background documents for records to make the changes needed to achieve our objective have been completed. The Design Task Group has documented much of their work through papers, presentations, and memorandums. Three memorandums in which INL team members had substantial contributions are found in the Appendices: FEA Modeling for the Baseline Program, Evaluating the Effects on Margin of Updating the Threshold and Shape Parameters in the Full Assessment, and Interpretations of the Full and Simplified Assessments in ASME BPVC. Most of the on-going work to achieve the Design Task Group?s objective will be addressing comments on existing records and moving records through the balloting process. The Design Task Group met bi-weekly mostly through the end of FY2023. Since February 2024, the Design Task Group has mostly been completed with solving and documenting the technical issues associated with the assessments. Unless new tasks are identified, the remaining work of the Design Task Group will be political and editorial.

97 MATHEMATICS AND COMPUTING↗

Interface Consistency: Phase I Results & Phase II Status

Future exploration missions will rely on designing and developing vehicles and complex systems from within NASA and through multiple external commercial partners to meet mission goals. Despite existing consistency-related agency requirements, NASA’s approach to commercial spaceflight development encourages providers’ flexibility and innovation. This strategy is resulting in significant design diversity across Artemis vehicles. Design best practices and guidelines champion interface consistency to promote mental model development and knowledge transfer. However, research investigating the benefits of consistency is mixed, and little is known about its role in complex systems. Determining the level of risk that system diversity presents is difficult, as there is no established method for quantifying the degree of consistency within and across interfaces, nor is there information about the differential impacts of different types of inconsistency. Phase I of this project (Characterization and Measurement) served as a starting point to better understand the construct of consistency, its application, and the range of studies and methods for measuring it. The project team created a taxonomy of consistency to apply to interfaces as a framework to guide the development of tools to assess intersystem consistency. Checklist and cognitive walkthrough methods were developed for use by human factors (HF) and human-computer interaction (HCI) experts. The Intersystem Consistency Scale (ICS) was developed for interface evaluations with crew. A pilot study evaluated the methods’ ability to distinguish differences between Artemis-like prototype pairs exhibiting either high or low design consistency. In addition, click errors and time on task were collected within the ICS (crew-like) group. Results from our exploratory analysis and lessons learned from the pilot study will be discussed. The project team will also present the status of Phase II (Risk Assessment, Standards and Guidelines). This includes incorporating feedback to redesign the assessment tools, and inputs from displays and training Subject Matter Experts to update tasks and prototype designs. The team will present the risk assessment study design to identify the types and levels of inconsistency that pose the greatest risk to performance. Plans to apply these results toward agency standards and guideline recommendations will also be discussed.

Human-Computer Interaction↗

Optimal design of structures with multiple design variables per group and multiple loading conditions on the personal computer

A finite element based programming system for minimum weight design of a truss-type structure subjected to displacement, stress, and lower and upper bounds on design variables is presented. The programming system consists of a number of independent processors, each performing a specific task. These processors, however, are interfaced through a well-organized data base, thus making the tasks of modifying, updating, or expanding the programming system much easier in a friendly environment provided by many inexpensive personal computers. The proposed software can be viewed as an important step in achieving a 'dummy' finite element for optimization. The programming system has been implemented on both large and small computers (such as VAX, CYBER, IBM-PC, and APPLE) although the focus is on the latter. Examples are presented to demonstrate the capabilities of the code. The present programming system can be used stand-alone or as part of the multilevel decomposition procedure to obtain optimum design for very large scale structural systems. Furthermore, other related research areas such as developing optimization algorithms (or in the larger level: a structural synthesis program) for future trends in using parallel computers may also benefit from this study.

Nguyen, D. T.↗

ASME Composite Code Rule Development Status

This report provides an update on the status of the composite code rule development for the American Society of Mechanical Engineers (ASME) Boiler Pressure Vessel (BPV) code Section III Division 5 Subsection HA Subpart B (HAB) and Subsection HH Subpart B (HHB) during FY 2025 (from October 2024 to September 2025). In the context of this document, reference to “the code” is specific to these subsections unless otherwise specified. A composites task group under the ASME Nonmetallic Design and Materials Working Group (WG-NDM), with the support of external experts, was established and is now working toward addressing previously identified areas.

36 MATERIALS SCIENCE↗

MACH 3: Past and future approaches to intelligent tutoring

In 1986, the U.S. Army Research Institute created an intelligent tutoring system as a proof-of-concept for artificial intelligence applications in Army training. The Maintenance Aid Computer HAWK Intelligent Institutional Instructor (MACH 3) taught student mechanics to maintain and troubleshoot the AN/MPQ-57 High Power Illuminator Radar (HPIR) of the HAWK Air Defense Missile System. In 1989, TRADOC Analysis Command compared the effectiveness of MACH 3 to traditional paper-based troubleshooting drills. For the study, all students received lecture and hands-on training as usual. However, during troubleshooting drills, students traced faults using either MACH 3 or the traditional paper-based method. Class records showed that the MACH 3 group completed significantly more troubleshooting tasks and progressed through tasks of greater difficulty than the paper-based group. Upon completion of training, students took written, practical, and oral essay tests. Mean test scores showed that students performed similarly regardless of the drill method used. However, significantly different standard deviations showed that the MACH 3 group performed more consistently than the paper-based group. Furthermore, significantly different time measures showed that the MACH 3 group reached faster troubleshooting solutions on the actual radar transmitter than the paper-based group. We will present the study results and discuss how updating the design of the MACH 3 can include desktop computing in a virtual environment.

Acchione-Noel, Sylvia↗

Evaluating design safety margins in the American Society of Mechanical Engineers graphite core components design-by-analysis assessments

Graphite is an important material being used for core components in next-generation high-temperature gas-cooled nuclear reactors. The selection of graphite grade for a specific Designer is a complex task, dependent on reactor conditions, component functionality, and required reliability. The American Society of Mechanical Engineers (ASME) provides two semi-probabilistic design-by-analysis assessments to evaluate graphite core components against design reliability targets. The simplified assessment uses a 2-parameter Weibull distribution to describe the graphite grade’s tensile-strength distribution to establish component stress limits. The full assessment uses the 3-parameter Weibull distribution and a modified Weakest-Link Theory approach to calculate a component design probability of failure. The paper defines recommended assessment rules, which are the as-written simplified assessment and the full assessment with parameter lower bounds, the modulus update with threshold reduction, and the 2027 grouping rules. Code rules are applied to three grades: 2114, IG-110, and NBG-18. The baseline margin calculation is developed using the experimental tensile dogbone specimen. Percent margin is defined as the percent reduction in the median experimental load to obtain the allowable load per ASME assessments. Under the recommended rules, the SRC–1 margin in the simplified assessment ranged from 40.2 % to 52.7 % among the grades in this study and from 36.1 % to 49.8 % in the full assessment. The full assessment only decreases the margins by 2.5–4.5 % for the SRC-1 components and 0–1.5 % for the SRC-2 components for this baseline case. Margin is inversely related to material median strength (i.e., the strongest grade, 2114, has the lowest margin).

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

AI-Enhanced Co-Design for Next-Generation Microelectronics: Innovating Innovation (Workshop Report)

The Artificial Intelligence Enhanced Co-Design for Next Generation Microelectronics virtual workshop was held April 4-5, 2023, and attended by subject matter experts from universities, industry, and national laboratories. This was the third in a series of workshops to motivate the research community to identify and address major challenges facing microelectronics research and production. The 2023 workshop focused on a set of topics from materials to computing algorithms, and included discussions on relevant federal legislation and such as the Creating Helpful Incentives to Produce Semiconductors and Science Act (CHIPS Act) which was signed into law in the summer of 2022. Talks at the workshop included edge computing in radiation environments, new materials for neuromorphic computing, advanced packaging for microelectronics, and new AI techniques. We also received project updates from several of the Department of Energy (DOE) microelectronics co-design projects funded in the fall of 2021, and from three of the Energy Frontier Research Centers (EFRCs) that had been funded in the fall of 2022. The workshop also conducted a set of breakout discussions around the five principal research directions (PRDs) from the 2018 Department of Energy workshop report: 1) define innovative material, device, and architecture requirements driven by applications, algorithms, and software; 2) revolutionize memory and data storage; 3) re-imagine information flow unconstrained by interconnects; 4) redefine computing by leveraging unexploited physical phenomena; 5) reinvent the electricity grid through new materials, devices, and architectures. We tasked each breakout group to consider one primary PRD (and other PRDs as relevant topics arose during discussions) and to address questions such as whether the research community has embraced co-design as a methodology and whether new developments at any level of innovation from materials to programming models requires the research community to reevaluate the PRDs developed back in 2018.

97 MATHEMATICS AND COMPUTING↗

An On-line Technology Information System (OTIS) for Advanced Life Support

OTIS is an on-line communication platform designed for smooth flow of technology information between advanced life support (ALS) technology developers, researchers, system analysts, and managers. With pathways for efficient transfer of information, several improvements in the ALS Program will result. With OTIS, it will be possible to provide programmatic information for technology developers and researchers, technical information for analysts, and managerial decision support. OTIS is a platform that enables the effective research, development, and delivery of complex systems for life support. An electronic data collection form has been developed for the solid waste element, drafted by the Solid Waste Working Group. Forms for other elements (air revitalization, water recovery, food processing, biomass production and thermal control) will also be developed, based on lessons learned from the development of the solid waste form. All forms will be developed by consultation with other working groups, comprised of experts in the area of interest. Forms will be converted to an on-line data collection interface that technology developers will use to transfer information into OTIS. Funded technology developers will log in to OTIS annually to complete the element- specific forms for their technology. The type and amount of information requested expands as the technology readiness level (TRL) increases. The completed forms will feed into a regularly updated and maintained database that will store technology information and allow for database searching. To ensure confidentiality of proprietary information, security permissions will be customized for each user. Principal investigators of a project will be able to designate certain data as proprietary and only technical monitors of a task, ALS Management, and the principal investigator will have the ability to view this information. The typical OTIS user will be able to read all non-proprietary information about all projects.Interaction with the database will occur over encrypted connections, and data will be stored on the server in an encrypted form. Implementation of OTIS will initiate a community-accessible repository of technology development information. With OTIS, ALS element leads and managers will be able to carry out informed technology selection for programmatic decisions. OTIS will also allow analysts to make accurate evaluations of technology options. Additionally, the range and specificity of information solicited will help educate technology developers of program needs. With augmentation, OTIS reporting is capable of replacing the current fiscal year-end reporting process. Overall, the system will enable more informed R&TD decisions and more rapid attainment of ALS Program goals.

Levri, Julie A.↗

GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning

Large language models (LLMs) are increasingly adapted to downstream tasks via reinforcement learning (RL) methods like Group Relative Policy Optimization (GRPO), which often require thousands of rollouts to learn new tasks. We argue that the interpretable nature of language often provides a much richer learning medium for LLMs, compared to policy gradients derived from sparse, scalar rewards. To test this, we introduce GEPA (Genetic-Pareto), a prompt optimizer that thoroughly incorporates natural language reflection to learn high-level rules from trial and error. Given any AI system containing one or more LLM prompts, GEPA samples trajectories (e.g., reasoning, tool calls, and tool outputs) and reflects on them in natural language to diagnose problems, propose and test prompt updates, and combine complementary lessons from the Pareto frontier of its own attempts. As a result of GEPA's design, it can often turn even just a few rollouts into a large quality gain. Across six tasks, GEPA outperforms GRPO by 6% on average and by up to 20%, while using up to 35x fewer rollouts. GEPA also outperforms the leading prompt optimizer, MIPROv2, by over 10% (e.g., +12% accuracy on AIME-2025), and demonstrates promising results as an inference-time search strategy for code optimization. We release our code at https://github.com/gepa-ai/gepa.

97 MATHEMATICS AND COMPUTING↗

Response Timing: Low Airspeed (Energy) Alerting

In response to NTSB Safety Recommendation A-14-043, the FAA was asked to task a panel of human factors, aviation operations, and aircraft design specialists, such as the Avionics Systems Harmonization Working Group (ASHWG), to develop design requirements for context-dependent low energy alerting systems for airplanes engaged in commercial operations. This recommendation is in association to the July 6, 2013, accident involving a Boeing 777-200ER, Korean registration HL7742, operating as Asiana Airlines flight 214, which was on approach to runway 28L when it struck a seawall at San Francisco International Airport (SFO), San Francisco, California. The FAA asked for human factor support from NASA in developing design requirements for context-dependent low energy alerting systems for airplanes engaged in commercial operations. Recently the ASHWG chair asked for continued support in examining flightcrew alert response timing and requested that NASA update the previous working paper on this subject. This product/paper will be provided to the ASHWG as part of the ongoing assignments from each of the ASHWG members in support of developing requirements and guidance for context-dependent low energy alerting systems for airplanes. The ASHWG is made up of interested parties from government and industry.

response timing↗

Generic POCC architecture: Revised recommended refinements and object-oriented interfaces

This document is a sequel to the report entitled Generic POCC Architectures, dated April 5, 1989, prepared by CTA under Contract NAS5-31500, Task 28-11600. That document presented a generic architecture based upon current technology, and a series of three refinements based on object-oriented analysis principles and expectations for POCC evolution. The current document revisits the object-oriented analysis of POCC's. We have reassessed the functional groupings that best adhere to object-oriented principles and have revised the recommended architecture accordingly. We present an updated view of the recommended generic POCC architecture using the same graphical models as the previous document: entity-relationship diagrams, dataflow diagrams, and composition graphs. In addition, we present another view in the form of entity-interface diagrams (EID's). EID's may be viewed as a precursor to object diagrams which are the basic construct of the general object-oriented design (GOOD) methodology. The entity-interface diagrams, together with their textual annotations, constitute our specification of object-oriented interfaces in the generic architecture.

Source record↗

Data Interconnection Exploration

This project involves innovating the way the Quality Assurance organization at the Jet Propulsion Laboratory (JPL) tracks the receiving inspection of JPL Critical Items (JCI) procured from suppliers. The Quality Assurance organization uses a Microsoft Access based front-end to enter items into an inspection queue. The current design has grown to be much more than was expected when the inspection queue system was created in 2007. The goal is to migrate to a web based solution that will allow the handling of more data, multiple input interfaces, and customizable display options. This has been achieved by implementing a divided tables scheme, the use of ColdFusion programming language, and the usage of Web 2.0 technologies such as AJAX and jQuery. These updates will allow for future expansion as well as the standardization of inspection queues for the Quality Assurance organization. When engineers and scientists order flight hardware, they cannot use it straight away. After a spacecraft or a satellite has been launched, it will be impossible to repair; therefore, the parts used must be inspected to make sure they are operating correctly and within the required specifications. The task of inspecting these JPL Critical Items (JCI) falls on the Quality Assurance organization. One of the tools the receiving inspection group uses is a database to maintain a queue of inspections that need to be done and of inspections that have been completed. The current Access interface has worked well but increased amounts of data, multiple input interfaces, and changes to how data is handled have made it necessary to seek a new way of organizing items for inspection. In addition, the different inspection groups have recently been merged and each has different ways of keeping track of inspection information. The ultimate goal is for the new group to have one shared inspection queue and website. To achieve this goal, I have been working with Ian Luczon of Procurement Quality Assurance and 3 Myers Hawkins, another intern, to not only create the new queue for receiving inspection but also a central web portal for the newly combined inspection group.

flight hardware↗

Countermeasures for Mitigation of Sensorimotor Decrements Following Head-Down Bed Rest

BACKGROUND Astronauts experience postflight disturbances in postural and locomotor control due to sensorimotor adaptations during spaceflight. These alterations may have adverse consequences if a rapid egress is required after landing. Although exercise is partially effective for mitigating cardiovascular and muscular deconditioning, additional countermeasures are needed to further preserve sensorimotor function for exploration missions. We have identified proprioceptive training and electrical muscle stimulation (EMS) as promising in-flight countermeasures. Since prolonged head down bed rest (HDBR) is a spaceflight analog for body unloading and causes postural and locomotor control decrements that parallel those observed after spaceflight, it can be used to accelerate the development of these countermeasures. METHODS This study will determine the effects of proprioceptive training and EMS on functional task performance and sensorimotor function following 60 days of 6° HDBR. Subjects will be randomly assigned to one of four groups: 1) an EMS arm, 2) a proprioceptive training arm, 3) an exercise plus proprioceptive training arm, and 4) a control arm. The EMS countermeasure will include daily bilateral stimulation of the quadriceps femoris muscle (30 minutes per session). Proprioceptive training will be performed three days per week (20 minutes per session) consisting of body-loaded postural tasks in the horizontal position on an air bearing sled. Exercise training will mimic current protocols used on the International Space Station, but treadmill aerobic exercise will be replaced with additional cycling aerobic exercise. Primary outcome measures will include pre and post HDBR functional tests that are representative of high priority exploration mission tasks and require high demand for dynamic control of postural stability. Additional measures will be used to identify the key physiological factors contributing to countermeasure benefits. Given the constrained samples size, a Bayesian modelling approach will be used to quantify the probability that there is an effect of a given magnitude. COUNTERMEASURE UPDATES Proprioceptive countermeasure design enhancements and human in the loop pilot testing continued through the Crew Health Countermeasures (CHC) Systems Capability Leadership Team (SCLT). The primary goals of this work were to enhance the visual feedback system’s capabilities and develop a proprioceptive training program for 60 days of HDBR. Six healthy non-astronaut volunteers participated in four pilot training sessions to systematically examine how each training variable (e.g., axial load, foot placement, and software profile) affects the overall proprioceptive challenge. The resulting training program will maintain an appropriate challenge during 60 days of HDBR by progressively decreasing the subject’s base of support, increasing tilt board target distances, and increasing axial loads using both subjective verbal feedback and objective performance data. RELEVANCE The deliverable from this project will be proof-of-concept sensorimotor countermeasure designs for functional task performance with full assessment of efficacy in a spaceflight analog. If the countermeasures are effective, they will be translated for validation with the suite of operationally implemented in-flight countermeasures.

T R Macaulay↗

U.S. Efforts in Support of Examinations at Fukushima Daiichi - November 2022 Meeting Notes and Information Request Status

Information obtained from Fukushima Daiichi Nuclear Power Station (Daiichi) is required to inform future Decontamination and Decommissioning (D&D) activities, improving the ability of the Tokyo Electric Power Company Holdings, Incorporated (TEPCO Holdings) to characterize potential hazards and to ensure the safety of workers involved with cleanup activities. This information also has important implications for the safety and operation of U.S. commercial nuclear power plants. This document summarizes results from the Fiscal Year 2023 (FY2023) U.S. effort to review Daiichi information and extract insights to enhance the safety of existing and future nuclear power plant designs. This U.S. effort, which was initiated in 2014 by the Department of Energy Office of Nuclear Energy, is completed by a group of experts in reactor safety and plant operations that identify examination needs and evaluate recent Daiichi examination data to address these needs. Fukushima-related information and associated discussions during these meetings benefit operating, new, and advanced reactors. Significant safety insights have been and are continuing to be obtained in several areas: system and component performance, radionuclide surveys and sampling, debris end-state location, combustible gas effects, and plant operations and maintenance. In addition to reducing uncertainties related to severe accident modeling progression, these insights have and continue to be used to update guidance for severe accident prevention, mitigation, and emergency planning. Furthermore, Daiichi-related activities, such as code modeling improvements and analysis, testing, and new technology deployment efforts, have the potential to offer additional benefits to the operating fleet and new LWR and non-LWR designs. U.S. evaluations of obtained examination information and input regarding future Daiichi examinations are of interest to several organizations within Japan. Since its inception, the U.S. has provided consensus input for high priority time-sequenced examination tasks and supporting research activities. In their Mid-to-Long-term Examination Plan for 1F investigations, TEPCO included all remaining U.S. consensus information requests and additional information requests they identified. TEPCO periodically provides reports on the status of these requests (reflecting D&D priorities, new insights from investigations, and new technologies that become available). Hence, U.S. experts agreed that it was appropriate for TEPCO to track and prioritize these information requests as D&D progresses. U.S. experts will continue to review and comment on the information obtained from examinations and, as needed, provide additional details and relevant background material to support future examinations. As documented in this report, several other items, such as additional details on information requests pertaining to ex-vessel examinations, relevant references from prior research, additional documents to provide insights regarding recent investigation findings, and reviews of recently released documents, were agreed to during the FY2023 meeting.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

AI Model Benchmarking for Nonproliferation Applications: Steel Thread Benchmarking Task Force Technical Report (Rev. 2)

Steel Thread is a NA-22 venture that seeks to build trustworthy, reliable AI models that can be used in a wide variety of nonproliferation tasks. A key aspect of building these models is developing appropriate benchmarks and evaluation methods, which will enable the venture to identify and adapt models to provide the most value in the nonproliferation domain. Benchmarks must be relevant to key tasks in this domain, such as question answering, information retrieval, document summarization and classification, consensus analysis, and image and data analysis. This report 1) provides an overview of benchmark design, evaluation, and challenges; 2) reviews a variety of open benchmarks, with a focus on language models and tasks; and 3) identifies benchmarks that are most relevant to Steel Thread. This report is intended to serve as a basis for further efforts to classify and evaluate benchmarks and their correlation with success on nonproliferation-specific tasks. The Steel Thread venture has defined benchmarks to be a particular combination of a dataset (or datasets) and a metric (or metrics) conceptualized as representing one or more specific tasks or sets of abilities for a specific modality. It is adopted by a research community as a shared framework for comparing methods.1 It includes 1) Data: Labeled (a designated subset not used for training, which could be all the data), 2) Metric: A way to quantify performance, 3) Task/Ability: What the benchmark is testing, 4) Protocol: A structured and repeatable evaluation process, 5) Baseline/Reference Model: For comparison; could be statistical, rule-based, SME-derived, or another model, and 6) Maintenance Plan: to update with new information over time; important for long-term utility. For further clarity, the definition includes what a benchmark, in this context, is not. It is not a corpus of training data, specific to a model (it is intended to apply to a range of models), a universal evaluation of performance, a guarantee that the ‘top’ model on the leaderboard will be the best fit for every specific use case, an all-encompassing proof of a model’s universal quality, nor is it a one-size-fits-all measure of success. It does not cover every real-world constraint (like operational, ethical, or cost considerations), a systems integration test, or a unit test. This definition was inspired by and resulted from discussions within the Steel Thread Benchmarking Task Force. This group was formed to define what we would mean as a benchmark within Steel Thread but persisted as the need to develop a thorough understanding of the large and expanding existing benchmarking space. This technical report is a result of the group’s divide and conquer approach to exploring this space. The release of benchmarks might not be progressing as quickly as model development, but it is moving very fast, as many benchmarks quickly become saturated, when state-of-the-art models score so close to the benchmark’s ceiling that their results are virtually indistinguishable. At that point, the test no longer differentiates between new systems, so researchers usually stop reporting scores as the benchmark no longer informs about improvements from the next generation of models. In the OpenAI announcement of GPT-5, they reported results on six flagship public benchmarks (AIME 2025, SWE-bench Verified, Aider Polyglot, MMMU, HealthBench Hard, GPQA) but the full system-card covers roughly thirty-five separate evaluations, comprising hundreds of test task items in total. There have been some efforts to summarize benchmarks in specific fields, like for text-to-image generation, but these surveys have had a narrow methodology scope. Therefore, a comprehensive survey of all benchmarks or even all benchmarks that could be relevant to Steel Thread is outside of the scope of this report. We chose some specific benchmarks to investigate in detail.

97 MATHEMATICS AND COMPUTING↗

ISEP: A Joint SRAG/CCMC Collaboration to Improve Mitigation of Space Weather Effects on Crew Health in the Exo-LEO Era

The Space Radiation Analysis Group (SRAG) at Johnson Space Center (JSC) is tasked with monitoring changes to space weather and mitigating any resultant impacts to crew health and safety. As human spaceflight goals extend from Low-Earth Orbit (LEO) missions like the International Space Station (ISS) to the moon, Mars and beyond, SRAG will need to update their current approach for crew monitoring of and protection from radiation exposure due to energetic Solar Particle Events (ESPEs). Challenges faced in planning exo-LEO missions include the lack of protection from the Earth’s geomagnetic field employed by the ISS in addition to limited communication capability between the crew and the ground. In the event of an ESPE, the current ISS trajectory ensures that the vehicle is only traveling through fields of higher radiation exposure for a brief period of time; the Earth’s geomagnetic field prevents the penetration of the high-energy particles of concern throughout the majority of the orbit. Exo-LEO missions, on the other hand, require that the vehicle travel through free space, exposing vehicle and crew to the full impact of the ESPE. NASA has combined multiple approaches to resolve this radiation exposure issue. New vehicles are designed to take advantage of advances in particle transport modeling capabilities and shielding technology, allowing redistribution of mass throughout the vehicle to areas of thinner shielding when the energetic particle flux has increased to levels of concern. Although vehicle shielding is an important aspect of radiation exposure protection, there is a continued requirement to monitor and predict the space weather environment. To this end, SRAG maintains a console position in Mission Control with 24/7 mission support capability. In the event of increased solar activity, SRAG collaborates with the Flight Control Team (FCT) to determine if crew action (i.e., shelter) is required. During any increase in solar activity, the FCT needs three pieces of information to effectively decide the crew response in light of other required mission tasks: if an event (ESPE) will occur, how ‘intense’ an observed event will be, and how long will an observed event will last. An ideal alert system limits false alarms, therefore causing the crew to take action unnecessarily, without ignoring events that pose a hazard to the crew. SRAG’s current operational concept for ISS missions focuses on short-term forecasts, best described as ‘now-casting’. Console operators are in daily communication with the Space Weather Prediction Center (SWPC) for situational awareness purposes. When conditions exist that may lead to increased solar activity, operators receive notifications from SWPC. In the case of a well-connected ESPE, the console operator may only have on the order of minutes to several hours to notify the FCT of the event and provide a recommendation for crew action. As NASA shifts to exo-LEO missions, the increased time in free space as well as the reduced ability to communicate with the crew will force a transition in crew protection strategy that emphasizes improvments to both the accuracy and the lead time in forecasting capabilities.

Barzilla, Janet E.↗

From LDEF to a national Space Environment and Effects (SEE) program: A natural progression

As the LDEF program draws to a close, it leaves in place the fundamental building blocks for a Space Environment and Effects (SEE) program. Results from LDEF data analyses and investigations now form a substantial core of knowledge on the long term effects of the space environment on materials, system and structures. In addition, these investigations form the basic structure of a critically-needed SEE archive and database system. An agency-wide effort is required to capture all elements of a SEE program to provide a more comprehensive and focused approach to understanding the space environment, determining the best techniques for both flight and ground-based experimentation, updating the models which predict both the environments and those effects on subsystems and spacecraft, and, finally, ensuring that this multitudinous information is properly maintained, and inserted into spacecraft design programs. Many parts and pieces of a SEE program already exist at various locations to fulfill specific needs. The primary purpose of this program, under the direction of the Office of Advanced Concepts and Technology (OACT) in NASA Headquarters, is to take advantage of these parts; apply synergisms where possible; identify and when possible fill-in gaps; coordinate and advocate a comprehensive SEE program. The SEE program must coordinate and support the efforts of well-established technical communities wherein the bulk of the work will continue to be done. The SEE program will consist of a NASA-led SEE Steering Committee, consisting of government and industry users, with the responsibility for coordination between technology developers and NASA customers; and Technical Working Groups with primary responsibility for program technical content in response to user needs. The Technical Working Groups are as follows: Materials and Processes; Plasma and Fields; Ionizing Radiation; Meteoroids and Orbital Debris; Neutral External Contamination; Thermosphere, Thermal, and Solar Conditions; Electromagnetic Effects; Integrated Assessments and Databases. Specific technology development tasks will be solicited through a NASA Research Announcement to be released in May of 1994. The areas in which tasks are solicited include: (1) engineering environment definitions, (2) environments and effects design guidelines, (3) environments and effects assessment models and databases, and (4) flight/ground simulation/technology assessment data.

Bowles, David E.↗