Preprint Open access
Safety risks in conversations with adolescents are not always explicit. A request may appear harmless unless a model considers the user's age, circumstances, and earlier turns. Existing Chinese safety benchmarks mainly target general users and give limited attention to adolescent safety. Single-turn tests also miss ris …
Preprint Open access
As adolescents increasingly use LLMs in everyday life, ensuring safe and developmentally appropriate responses has become essential. However, existing LLM guardrails primarily target explicit harmful content in isolated prompts or responses and are less effective at identifying implicit, context-dependent developmental …
Preprint Open access
Large language model agents can correctly judge that an action should be blocked while still preferring to take it. We ask why this judgment-action disconnect arises, and whether explicit safety judgment causally governs subsequent action preference. Across three open-weight language models, safety-predictive informati …
Preprint Open access
Safety evaluations often ask whether a model recognizes that an action is unsafe, whereas agent evaluations ask what the model chooses to do. Using safety judgments as evidence about action selection therefore raises a measurement question: does the influence of the same safety-relevant fact persist across response int …