Authors

Charith Peris

Publications 1

Preprint Open access

Controlled Decoding Attacks on Black-Box LLMs

Jesson Wang, Shawn Li, Wei Yang et al. · 2026

Manipulating next-token probabilities during generation can bypass the safety alignment of large language models. Existing approaches, however, rely on access to model weights or numerical token probabilities and therefore do not apply to interfaces that return only sampled text. Reconstructing probabilities from sampl …

Co-authors