Search NASASearch

SEARCH · Search NASA

Results for “Modal Test and Analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

27 records · Page 2

AI Model Benchmarking for Nonproliferation Applications: Steel Thread Benchmarking Task Force Technical Report (Rev. 2)

Steel Thread is a NA-22 venture that seeks to build trustworthy, reliable AI models that can be used in a wide variety of nonproliferation tasks. A key aspect of building these models is developing appropriate benchmarks and evaluation methods, which will enable the venture to identify and adapt models to provide the most value in the nonproliferation domain. Benchmarks must be relevant to key tasks in this domain, such as question answering, information retrieval, document summarization and classification, consensus analysis, and image and data analysis. This report 1) provides an overview of benchmark design, evaluation, and challenges; 2) reviews a variety of open benchmarks, with a focus on language models and tasks; and 3) identifies benchmarks that are most relevant to Steel Thread. This report is intended to serve as a basis for further efforts to classify and evaluate benchmarks and their correlation with success on nonproliferation-specific tasks. The Steel Thread venture has defined benchmarks to be a particular combination of a dataset (or datasets) and a metric (or metrics) conceptualized as representing one or more specific tasks or sets of abilities for a specific modality. It is adopted by a research community as a shared framework for comparing methods.1 It includes 1) Data: Labeled (a designated subset not used for training, which could be all the data), 2) Metric: A way to quantify performance, 3) Task/Ability: What the benchmark is testing, 4) Protocol: A structured and repeatable evaluation process, 5) Baseline/Reference Model: For comparison; could be statistical, rule-based, SME-derived, or another model, and 6) Maintenance Plan: to update with new information over time; important for long-term utility. For further clarity, the definition includes what a benchmark, in this context, is not. It is not a corpus of training data, specific to a model (it is intended to apply to a range of models), a universal evaluation of performance, a guarantee that the ‘top’ model on the leaderboard will be the best fit for every specific use case, an all-encompassing proof of a model’s universal quality, nor is it a one-size-fits-all measure of success. It does not cover every real-world constraint (like operational, ethical, or cost considerations), a systems integration test, or a unit test. This definition was inspired by and resulted from discussions within the Steel Thread Benchmarking Task Force. This group was formed to define what we would mean as a benchmark within Steel Thread but persisted as the need to develop a thorough understanding of the large and expanding existing benchmarking space. This technical report is a result of the group’s divide and conquer approach to exploring this space. The release of benchmarks might not be progressing as quickly as model development, but it is moving very fast, as many benchmarks quickly become saturated, when state-of-the-art models score so close to the benchmark’s ceiling that their results are virtually indistinguishable. At that point, the test no longer differentiates between new systems, so researchers usually stop reporting scores as the benchmark no longer informs about improvements from the next generation of models. In the OpenAI announcement of GPT-5, they reported results on six flagship public benchmarks (AIME 2025, SWE-bench Verified, Aider Polyglot, MMMU, HealthBench Hard, GPQA) but the full system-card covers roughly thirty-five separate evaluations, comprising hundreds of test task items in total. There have been some efforts to summarize benchmarks in specific fields, like for text-to-image generation, but these surveys have had a narrow methodology scope. Therefore, a comprehensive survey of all benchmarks or even all benchmarks that could be relevant to Steel Thread is outside of the scope of this report. We chose some specific benchmarks to investigate in detail.

97 MATHEMATICS AND COMPUTING

Hydride and Seek: Comparing Crystallographic Hydride Placement Techniques with an Open-Shell Cobalt Complex

Locating hydrides is crucial in organometallic chemistry but difficult to do accurately using X-ray diffraction. Electron diffraction has been proposed as a way to overcome this problem but has not been systematically compared to neutron diffraction and to quantum crystallography (Hirshfeld atom refinement, HAR) to test this hypothesis. Here, we present a comparative analysis of methods for a terminal cobalt hydride complex by comparing a single-crystal neutron diffraction reference structure to results from single-crystal X-ray diffraction with and without Hirshfeld atom refinement (HAR, NoSpherA2), density functional theory (DFT), and electron diffraction (3D-ED/MicroED) refined under kinematical and dynamical formalisms. Conventional X-ray diffraction gives lower precision than neutron diffraction as expected. Despite expected improvements, HAR gives systematic deviation from the neutron benchmark. Interestingly, optimized DFT equilibrium geometries are closer to the neutron value than the value from HAR. On the other hand, electron diffraction with a high-quality data set coupled with dynamical refinement localizes the hydride in difference maps and gives excellent agreement with the neutron data. Dynamical refinement is crucial, as kinematical refinement does not allow assignment of a hydride peak. This cross-modal comparison defines the conditions under which 3D-ED/MicroED delivers high-precision metal–hydride distances for this open-shell cobalt hydride.

anions

Anomaly Detection for Online Monitoring of Thermocouple Sensors in the Advanced Test Reactor

This study explores data-driven anomaly detection methods to analyze sensor fail- ures in the Advanced Gas Reactor (AGR) nuclear fuel irradiation experiments. Specifically, we examine failures of thermocouples (TCs), which are critical for mon- itoring and controlling in-reactor temperatures during operation. Failures were pri- marily observed during abrupt power transitions and manifested as sensor drop-outs, drifts, or unexplained behavior. We applied three time-series analysis techniques— rolling mean smoothing, matrix profile, and vector auto-regression (VAR)—to de- tect anomalies in TC data prior to failure events. The rolling mean method effec- tively highlighted deviations aligned with reported failures, while the matrix profile provided partial early warning but sometimes flagged normal fluctuations during power-down periods. VAR shows potential in capturing multivariate dependencies but requires further calibration. A rare case of TC drift was also documented, which did not result in failure, underscoring the challenge of building predictive models with sparse positive examples. Our findings demonstrate that traditional statistical tools can aid anomaly detection but have limited predictive power without richer training data. We propose future directions including synthetic data generation, real- time surrogate modeling, and multi-modal feature integration. This work provides a foundation for applying robust anomaly detection frameworks to mission-critical sensor systems in experimental settings.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS

SpaceNet 9—Cross-Sensor Alignment of Optical and SAR Imagery

Precise registration of high-resolution synthetic aperture radar (SAR) and optical imagery is necessary for realizing the full potential and benefits of multimodal image analysis. However, two significant challenges presently exist. First, there is a lack of annotated datasets and benchmarks available for high-resolution SAR–optical image registration. Second, an assessment of efficient and reliable image registration methods that can precisely align these modalities is lacking. Here, we present a holistic description of the SpaceNet 9 Challenge and its results. We present a description of the dataset and baseline algorithm along with the results of the challenge, including a description of the winning algorithms. We release the SpaceNet 9 dataset along with open-sourcing the winning algorithms and baseline. The objective of SpaceNet 9 was to compute a dense displacement map that indicates the shift needed to align pixels in an optical image to the pixels in a SAR image. The challenge launched in April 2025 and was active for approximately two months. The top five solutions reduced image alignment error from approximately 34 m to under 13 m for public and private test data, with the best results obtaining a registration error of only 8.5 and 6.7 m on the public testing and private testing dataset, respectively. Usage of pretrained image matching models, robust outlier rejection with RANSAC, and estimating local displacement were common among the top solutions. The results of this challenge provide insight into high-resolution SAR–optical image registration and offer opportunities for future benchmarking in this domain. The baseline algorithm, winning solutions, and datasets are available at https://spacenet.ai/sn9-challenge/.

benchmark datasets

FY 2026 Midyear Report: Seismic Monitoring of Underground Vibration Sources Using Distributed Acoustic Sensing and Seismometers

Safeguards-relevant temporal changes in underground facilities can be observed using geophysical monitoring techniques. Seismic waves, in particular, provide valuable insights into subsurface activities and can serve as an important tool for detecting anomalous events that may indicate containment breaches at geological repositories. This midyear report summarizes ongoing efforts to automatically and rapidly detect and locate anomalous vibration signals that could be indicative of potential containment breaches. Previous work during FY25 focused on compiling continuous seismic datasets from two underground sites and developing a database of continuous waveforms and ground-truth event data derived from multiple sensing modalities. Building on this foundation, we are adapting anomaly detection and geolocation algorithms to explore methods for monitoring underground activities using two relatively low-maintenance sensing technologies: a dense surface geophone array deployed at the Pleasant Gap mine in Pennsylvania, and a three-dimensional fiber-optic cable array for distributed acoustic sensing (DAS) installed in the subsurface at the Sanford Underground Research Facility (SURF) in South Dakota. This report summarizes work conducted during the first two quarters of FY26, during which we refined a dynamic power spectral density (PSD)-based detector, applied it independently to each geophone station, and then combined the per‑station detections with density-based spatial clustering of applications with noise (DBSCAN) to cluster events and produce spatial maps over a nine‑day interval. In addition, we outline plans for a field trial at the Waste Isolation Pilot Plant (WIPP) in New Mexico to compare traditional seismic monitoring approaches with DAS techniques and to evaluate the benefits of combined data analysis. Activities during the past two quarters have included the preparation and submission of a Field Test Plan to WIPP for approval, as well as submission to headquarters for review and feedback.

58 GEOSCIENCES

Cascading economic losses from port disruptions under capacity constrained multimodal freight networks

This study quantifies how throughput disruptions at major seaports cascade through capacity-constrained multimodal freight networks and interregional production systems. We couple an agent-based model (ABM) multimodal freight simulation that resolves rerouting, terminal queueing, and inventory drawdown under binding modal and facility capacities with a multiregional output loss input-output (MRIIM) model that propagates realized delivery shortfalls across regions and sectors. The framework is demonstrated for the Port of Los Angeles using Freight Analysis Framework flows and Bureau of Economic Analysis input-output accounts and is evaluated over a 52-week horizon under deterministic sector targeted shocks and stochastic disruption realizations with uncertain severity and duration. Results indicate nonlinear amplification: realized national losses concentrate in manufacturing and transportation/warehousing even when exogenous port shocks are dispersed, suggesting that congestion spillback and limited short-run substitution can dominate the initial shock allocation. We further evaluate a tabular reinforcement-learning (Q-learning) intervention layer that selects among a small set of implementable system level levers (truck-to-rail and truck-to-barge shift settings) without overriding shipper routing, finding that such interventions reduce total losses for moderate disruptions but yield diminishing returns once substitute modes approach capacity. By linking operational freight behavior to system wide impacts under uncertainty, the proposed ABM-MRIIM pipeline provides a reusable workflow for port disruption stress testing, identification of structurally critical sectors/corridors, and evaluation of resilience interventions under realistic capacity limits.

42 ENGINEERING

Comparative Uptake Patterns of Radioactive Iodine and [18F]-Fluorodeoxyglucose (FDG) in Metastatic Differentiated Thyroid Cancers

Background: Metastatic differentiated thyroid cancer (DTC) represents a molecularly heterogeneous group of cancers with varying radioactive iodine (RAI) and [ 18 F]-fluorodeoxyglucose (FDG) uptake patterns potentially correlated with the degree of de-differentiation through the so-called “flip-flop” phenomenon. However, it is unknown if RAI and FDG uptake patterns correlate with molecular status or metastatic site. Materials and Methods: A retrospective analysis of metastatic DTC patients (n = 46) with radioactive 131-iodine whole body scan (WBS) and FDG-PET imaging between 2008 and 2022 was performed. The inclusion criteria included accessible FDG-PET and WBS studies within 1 year of each other. Studies were interpreted by two blinded radiologists for iodine or FDG uptake in extrathyroidal sites including lungs, lymph nodes, and bone. Cases were stratified by BRAF V600E mutation status, histology, and a combination of tumor genotype and histology. The data were analyzed by McNemar’s Chi-square test. Results: Lung metastasis FDG uptake was significantly more common than iodine uptake (WBS: 52%, FDG: 84%, p = 0.04), but no significant differences were found for lymph or bone metastases. Lung metastasis FDG uptake was significantly more prevalent in the papillary pattern sub-cohort (WBS: 37%, FDG: 89%, p = 0.02) than the follicular pattern sub-cohort (WBS: 75%, FDG: 75%, p = 1.00). Similarly, BRAF V600E+ tumors with lung metastases also demonstrated a preponderance of FDG uptake (WBS: 29%, FDG: 93%, p = 0.02) than BRAF V600E- tumors (WBS: 83%, FDG: 83%, p = 1.00) with lung metastases. Papillary histology featured higher FDG uptake in lung metastasis (WBS: 39%, FDG: 89%, p = 0.03) compared with follicular histology (WBS: 69%, FDG: 77%, p = 1.00). Patients with papillary pattern disease, BRAF V600E+ mutation, or papillary histology had reduced agreement between both modalities in uptake at all metastatic sites compared with those with follicular pattern disease, BRAF V600E- mutation, or follicular histology. Low agreement in lymph node uptake was observed in all patients irrespective of molecular status or histology. Conclusions: The pattern of FDG-PET and radioiodine uptake is dependent on molecular status and metastatic site, with those with papillary histology or BRAF V600E+ mutation featuring increased FDG uptake in distant metastasis. Further study with an expanded cohort may identify which patients may benefit from specific imaging modalities to recognize and surveil metastases.

60 APPLIED LIFE SCIENCES

Analysis and optimization of seismic monitoring networks with Bayesian optimal experimental design

SUMMARY Monitoring networks increasingly aim to assimilate data from a large number of diverse sensors covering many sensing modalities. Bayesian optimal experimental design (OED) seeks to identify data, sensor configurations or experiments which can optimally reduce uncertainty and hence increase the performance of a monitoring network. Information theory guides OED by formulating the choice of experiment or sensor placement as an optimization problem that maximizes the expected information gain (EIG) about quantities of interest given prior knowledge and models of expected observation data. Therefore, within the context of seismo-acoustic monitoring, we can use Bayesian OED to configure sensor networks by choosing sensor locations, types and fidelity in order to improve our ability to identify and locate seismic sources. In this work, we develop the framework necessary to use Bayesian OED to optimize a sensor network’s ability to locate seismic events from arrival time data of detected seismic phases at the regional-scale. This framework requires five elements: (i) A likelihood function that describes the distribution of detection and traveltime data from the sensor network, (ii) A prior distribution that describes a priori belief about seismic events, (iii) A Bayesian solver that uses a prior and likelihood to identify the posterior distribution of seismic events given the data, (iv) An algorithm to compute EIG about seismic events over a data set of hypothetical prior events, (v) An optimizer that finds a sensor network which maximizes EIG. Once we have developed this framework, we explore many relevant questions to monitoring such as: how to trade off sensor fidelity and earth model uncertainty; how sensor types, number and locations influence uncertainty; and how prior models and constraints influence sensor placement.

58 GEOSCIENCES

Optimal Stopping Ages for Colorectal Cancer Screening

Importance Prior studies have shown that the benefits, harms, and costs of colorectal cancer (CRC) screening at older ages are associated with a patient’s sex, health, and screening history. However, these studies were hypothetical exercises and not directly informed by data on CRC risk. Objective To identify the optimal stopping ages for CRC screening by sex, comorbidity, and screening history from a cost-effectiveness perspective. Design, Setting, and Participants This economic evaluation first validated the MISCAN-Colon (Microsimulation Screening Analysis–Colon) model against community-based CRC incidence and mortality rates for 2 subcohorts of the PRECISE (Optimizing Colorectal Cancer Screening Precision and Outcomes in Community-Based Populations) cohort. Subsequently, different CRC screening scenarios were simulated in older individuals. Cohorts of US adults aged 76 to 90 years varied by sex and comorbidity status (none, low, moderate, or severe). Statistical and sensitivity analyses were performed from March 2023 to May 2024. Exposures CRC screening histories including fecal immunochemical test (FIT) or colonoscopy, such as a negative colonoscopy result from 10, 15, 20, 25, or 30 years before the index age; 1 to 5 negative FIT results within 5 years of the index age, with different patterns of recency; or a combination of negative colonoscopy and negative FIT results. Main Outcomes and Measures The main outcomes included estimated lifetime clinical outcomes, incremental costs, and quality-adjusted life-years gained (QALYG) associated with 1 additional FIT or colonoscopy. Optimal stopping age for screening, defined as the oldest age for which the incremental cost-effectiveness ratio was still below the willingness-to-pay threshold of $\$$100 000 per QALYG, was evaluated. Results The first of the 2 PRECISE subcohorts used in validating the simulation model included 25 974 adults (15 060 females [58.0%]; 54.7% aged 76 to 80 years) with a negative colonoscopy result 10 years before the index date. The second subcohort consisted of 118 269 adults (67 058 females [56.7%]; 90.5% aged 76 to 80 years) with a negative FIT result 1 year before the index date. Older age, male sex, higher comorbidity levels, and recent CRC screenings were associated with reduced incremental benefit and cost-effectiveness of additional screening. For the reference cohort of 76-year-old females without comorbidities and a negative colonoscopy result 10 years before the index age, 1 additional colonoscopy cost $\$$38 226 per QALYG. For cohorts with otherwise equivalent characteristics, associated costs increased to $\$$1 689 945 per QALYG for females at age 90 years without comorbidities and a negative colonoscopy results 10 years before the index age, $\$$51 604 per QALYG for males at age 76 years without comorbidities and a negative colonoscopy result 10 years before the index age, and $\$$108 480 per QALYG for females at age 76 years with severe comorbidities and a negative colonoscopy result 10 years before the index age and decreased to $\$$16 870 per QALYG for females without comorbidities and a negative colonoscopy result 30 years before the index age. The optimal stopping ages across different cohorts ranged from younger than 76 to 86 years for colonoscopy and younger than 76 to 88 years for FIT. Conclusions and Relevance In this economic evaluation, age, sex, screening history, comorbidity, and future screening modality were associated with the clinical outcomes, cost-effectiveness, and optimal stopping age for CRC screening. These results can inform guideline development and patient-directed informed decision-making.

Harlass, Matthias [Erasmus Erasmus University Medi