Preprint Open access
From Expected Harmfulness to Likelihood: A Probabilistic Reformulation of Jailbreaking LLM Agents
When the harmfulness of an LLM agent's output can be quantified, a natural jailbreaking objective is to maximize expected harmfulness over admissible input modifications. An alternative approach constructs or selects harmful target outputs and modifies the input to increase their likelihood. We establish a precise conn …