Abstract

Vision-language models (VLMs) are often reported to outperform task-specific vision backbones for unmanned aerial vehicle (UAV) power-line defect assessment. We test that claim on ElecVQA-Bench, a 56,972-item benchmark derived from the public InsPLAD dataset, across six evaluation choices: partition, evaluated item set, label space, replication, input resolution, and side information. On a matched partition, a Swin Transformer and the strongest adapted VLM differ by only 0.03 points at binary screening. At seven-way defect typing, increasing the vision backbones from 224 px to the measured pixel budget of the VLM preprocessor narrows the gap against InternVL3.5-8B from +20.53 to -0.57 points for ResNet-50 and from +23.67 to +4.70 points for Swin-T. A pixel-budget audit shifts Qwen3-VL-8B macro recall by 10.78 points, yet a source-pixel-matched InternVL control still leaves Qwen ahead by 7.43 to 13.61 points while using 56% fewer visual tokens, so neither source pixels nor token budget explains the difference between the two VLMs. A two-seed global replication changes Qwen binary accuracy and seven-way macro recall by 0.86 and 1.02 points. After split-specific retraining, Qwen does not lead at crop or image level, and a 14-tower, three-seed replication reverses the sign across seeds, giving mean common-six macro recall of 0.9085 for Qwen against 0.9509 for ResNet-50. No split regime yields a family-level advantage that survives multiple-comparison correction. The study supports a benchmark-audit contribution rather than a general claim of VLM superiority.

Keywords

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Zhang, L., Xiang, S., Kuang, J., & Yi, P. (2026). Beyond Balanced Accuracy: A Resolution and Parity-Controlled Benchmark for Vision-Language and Vision-Only Defect Assessment in UAV Power-Line Inspection. https://omanscience.com/en/articles/beyond-balanced-accuracy-a-resolution-and-parity-controlled-benchmark-for-vision-language-and-vision-only-defect-assessment-in-uav-power-line-inspecti

MLA 9

Zhang, Linghao, et al. "Beyond Balanced Accuracy: A Resolution and Parity-Controlled Benchmark for Vision-Language and Vision-Only Defect Assessment in UAV Power-Line Inspection." https://omanscience.com/en/articles/beyond-balanced-accuracy-a-resolution-and-parity-controlled-benchmark-for-vision-language-and-vision-only-defect-assessment-in-uav-power-line-inspecti.

Chicago (author–date)

Zhang, Linghao, Siyu Xiang, Junwei Kuang, and Peiyu Yi. 2026. "Beyond Balanced Accuracy: A Resolution and Parity-Controlled Benchmark for Vision-Language and Vision-Only Defect Assessment in UAV Power-Line Inspection." https://omanscience.com/en/articles/beyond-balanced-accuracy-a-resolution-and-parity-controlled-benchmark-for-vision-language-and-vision-only-defect-assessment-in-uav-power-line-inspecti.

Harvard

Zhang, L., Xiang, S., Kuang, J. and Yi, P. (2026) 'Beyond Balanced Accuracy: A Resolution and Parity-Controlled Benchmark for Vision-Language and Vision-Only Defect Assessment in UAV Power-Line Inspection', Available at: https://omanscience.com/en/articles/beyond-balanced-accuracy-a-resolution-and-parity-controlled-benchmark-for-vision-language-and-vision-only-defect-assessment-in-uav-power-line-inspecti.

Vancouver

Zhang L, Xiang S, Kuang J, Yi P. Beyond Balanced Accuracy: A Resolution and Parity-Controlled Benchmark for Vision-Language and Vision-Only Defect Assessment in UAV Power-Line Inspection. https://omanscience.com/en/articles/beyond-balanced-accuracy-a-resolution-and-parity-controlled-benchmark-for-vision-language-and-vision-only-defect-assessment-in-uav-power-line-inspecti

IEEE

L. Zhang, S. Xiang, J. Kuang, and P. Yi, "Beyond Balanced Accuracy: A Resolution and Parity-Controlled Benchmark for Vision-Language and Vision-Only Defect Assessment in UAV Power-Line Inspection," https://omanscience.com/en/articles/beyond-balanced-accuracy-a-resolution-and-parity-controlled-benchmark-for-vision-language-and-vision-only-defect-assessment-in-uav-power-line-inspecti.