Authors

Perouz Taslakian

Publications 1

Preprint Open access

Visual Memory Attacks Can Persist Through The KV Cache

Modern language model systems operate autonomously over increasingly long contexts containing untrusted text and images. Can an adversarial input continue to steer a model even after that input is removed from its context? We show that attacks can be trained to persist through the key/value (KV) cache of subsequent tok …

Co-authors