نسخة أولية وصول مفتوح
DIBench: Benchmarking Decision Integrity of GUI-based Mobile Agents Under Deceptive Injections
As GUI-based mobile agents rapidly progress, rigorous safety evaluation of their autonomous decision-making in realistic app interfaces becomes increasingly critical. Existing benchmarks mainly focus on execution-level anomalies using task success or hijack rates, but fail to capture the in-task goal deviation risk in …