DOE OSTI · 3157834
QCalEval: Benchmarking Vision-Language Models for Quantum Calibration Plot Understanding
Abstract
Quantum computing calibration depends on interpreting experimental data, and calibration plots provide the most universal human-readable representation for this task, yet no systematic evaluation exists of how well vision-language models (VLMs) interpret them. We introduce QCalEval, the first VLM benchmark for quantum calibration plots: 243 samples across 87 scenario types from 22 experiment families, spanning superconducting qubits and neutral atoms, evaluated on six question types in both zero-shot and in-context learning settings. The best general-purpose zero-shot model reaches a mean score of 72.3, and many open-weight models degrade under multi-image in-context learning, whereas frontier closed models improve substantially. A supervised fine-tuning ablation at the 9-billion-parameter scale shows that SFT improves zero-shot performance but cannot close the multimodal in-context learning gap. As a reference case study, we release NVIDIA Ising Calibration 1, an open-weight model based on Qwen3.5-35B-A3B that reaches 74.7 zero-shot average score.
Keep this discovery
Explore connections, maps & timelines
Cao, Shuxiang, Zhang, Zijian [U. Toronto (main)], Agarwal, Abhishek [Unlisted, US], Bratrud, Grace [Fermilab; Northwestern U.], Beysengulov, Niyaz R. [Unlisted, US], Cole, Daniel C., Frieiro, Alejandro Gómez [Unlisted, DE], Glen, Elena O. [Unlisted, US], Hsu, Hao [Unlisted, DE], Huang, Gang [LBL, Berkeley], Jow, Raymond, Shaji, Greshma [Unlisted, DE], Lubowe, Tom, Zhu, Ligeng, Calderón, Luis Mantilla [U. Toronto (main)], Pancotti, Nicola, Pendleton, Joel, Severin, Brandon, Staub, Charles Etienne [Harvard U.], Sussman, Sara [Fermilab], Vepsäläinen, Antti [Unlisted, DE], Vora, Neel Rajeshbhai [LBL, Berkeley], Xu, Yilun [LBL, Berkeley], Bernales, Varinia [U. Toronto (main)], Bowring, Daniel [Fermilab], Kyoseva, Elica, Rungger, Ivan [Unlisted, US; Royal Holloway, U. of London], Semeghini, Giulia [Harvard U.], Stanwyck, Sam, Costa, Timothy, Aspuru-Guzik, Alán [U. Toronto (main)], Svore, Krysta. 2026-04-28. QCalEval: Benchmarking Vision-Language Models for Quantum Calibration Plot Understanding. https://www.osti.gov/biblio/3157834
Cite the original work for its findings. Save a collection to share your selection of sources.