Preprint Open access
SAFESHIELD: A Decision-Organization Framework for Deployment-Time Safety of Small Language Models
Deployment-time safety of language models is commonly implemented through runtime guardrails such as input moderation, routing, retrieval verification, and output filtering. Existing deployment frameworks provide increasingly capable mechanisms for these functions, but offer limited guidance on how the safety decisions …