الباحثون

Vaggos Chatziafratis

المنشورات 3

نسخة أولية وصول مفتوح

Jailbreaking Open-Weight LLMs via Random Embedding Perturbations

While open-weight models have enjoyed steady progress in capabilities and wide adoption across multiple domains, their safety remains an important concern. One key feature is the ability to refuse or deflect harmful, malicious, or insensitive prompts. In this paper, we expose safety vulnerabilities across six common op …

المؤلفون المشاركون