Search NASASearch

DOE OSTI · code-177019

MTRE

Abstract

Multi-Token Reliability Estimation (MTRE) is a lightweight, white-box hallucination detector for vision-language models. Instead of using only the first output token, MTRE aggregates logits from the first ~10 tokens and feeds them to a small attention-based reliability head; per-token scores are combined via a sequential log-likelihood-ratio test with early-stopping, and an MTRE-t variant calibrates thresholds via cross-fitting. MTRE reports average gains of +9.4% Accuracy and +14.8% AUROC over common baselines across MAD-Bench, MM-SafetyBench, MathVista, and arithmetic/counting tasks, while adding ~4.3M params and ~1% inference overhead (~26 MB VRAM, ~0.94 ms per detection). Key limitation: requires access to early token logits and is evaluated on a handful of open-source 7B VLMs.

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Bhattarai, Manish [Los Alamos National Labs], Vu, Minh Nhat, Zolllicoffer, Geigh. 2026-01-12. MTRE. https://doi.org/10.11578/dc.20260306.3

Cite the original work for its findings. Save a collection to share your selection of sources.