الباحثون

Qingqing Ye

المنشورات 6

نسخة أولية وصول مفتوح

DIBench: Benchmarking Decision Integrity of GUI-based Mobile Agents Under Deceptive Injections

Li Hu, Kanghua Mo, Yingbin Jin وآخرون · 2026

As GUI-based mobile agents rapidly progress, rigorous safety evaluation of their autonomous decision-making in realistic app interfaces becomes increasingly critical. Existing benchmarks mainly focus on execution-level anomalies using task success or hijack rates, but fail to capture the in-task goal deviation risk in …

نسخة أولية وصول مفتوح

SINGED: Correct Outputs Do Not Certify Safe Execution in LLM Agents

Xiaoyu Xu, Zi Liang, Minxin Du وآخرون · 2026

Tool-using language-model agents select and execute third-party artifacts. Different implementations can return the requested output while producing hidden execution effects that task-, attack-, or choice-based evaluations may miss. We study functional counterfeits: implementations that match benign alternatives on the …

نسخة أولية وصول مفتوح

Machine Unlearning for Large Language Models: Foundations, Advances, and Agentic Extensions

Xiaoyu Xu, Minxin Du, Li Bai وآخرون · 2026

Machine unlearning aims to remove target influence while preserving other capabilities. This survey compares methods, benchmarks, and evidence across large language models and systems using retrieval, memory, tools, and interacting agents. A five-layer framework connects removal requests, system boundaries, target loca …

نسخة أولية وصول مفتوح

FeatMark: Feature-level Watermark Protection against Mimicry Attacks with Diffusion Models

Haoyang Li, Ruoxi Sun, Qingqing Ye وآخرون · 2026

Text-to-image diffusion models enable data-efficient "mimicry" attacks, wherein adversaries fine-tune the model on a handful of public photos to synthesize convincing forgeries of a target individual. A common countermeasure is to embed imperceptible, low-energy watermarks, yet recent studies show these signatures are …

نسخة أولية وصول مفتوح

TraceGuard: Adaptive Multimodal Poison Filtering through Cross-Feature Rank Agreement

Haoyang Li, Yaxin Xiao, Linyan Dai وآخرون · 2026

Multimodal training relies on image-text corpora collected from external sources, creating opportunities for attackers to poison the data. Stealthy attacks can preserve plausible image-text pairs while concealing the differences used by detectors, so apparently clean data can still redirect the trained model. We theref …

المؤلفون المشاركون