الباحثون

Manit Baser

المنشورات 2

نسخة أولية وصول مفتوح

The Fragility of Trigger-Tag Mechanisms for Misuse Detection in Open-Weight LLMs

Toluwani Aremu, Manit Baser, Mohan Gurusamy وآخرون · 2026

Open-weight language models can be downloaded, modified, and deployed beyond their developers' control, limiting the effectiveness of centrally enforced safeguards. Recent work has therefore proposed \emph{trigger-tag} mechanisms that produce a detectable signal when a model is used under a target condition, such as ge …

نسخة أولية وصول مفتوح

The Tokens Remember: When Tokenization Bypasses Knowledge Editing and Unlearning

Open-weight LLMs give downstream users control over the inference stack, but this flexibility can undermine post-release guarantees that sensitive knowledge has been modified or removed. Model editing and machine unlearning are used to modify or remove targeted knowledge without retraining models from scratch. However, …

المؤلفون المشاركون