Abstract

Output-only safety monitoring sees only the end of a model's computation, yet the model computes its answer before emitting it: what it is internally poised to say is safety-critical. Jacobian-space (J-space) readouts, linear maps from hidden states to the output vocabulary via the model's input-output Jacobian, have been proposed as a window on behaviorally accessible internal content, and a first safety protocol (JADR) showed that danger recognition is readable there. What remains unknown is how this accessible content relates to the latent-space safety geometry studied by representation engineering, and how safety training shapes it. We introduce two quantitative bridges: (i) a transport-amplification profile $A_\ell$, measuring how strongly a latent safety direction is carried toward output space by each layer's Jacobian, and (ii) a paired base-vs-tuned protocol that attributes accessibility to training provenance. Across four small model pairs (135M to 1.5B, three lens-fit seeds), tuned checkpoints concentrate recognition in upper-middle layers, and at 0.5B DPO installs refusal while J-space recognition drops to chance: training can widen the accessibility gap exactly where behavioral safety looks best. A deployment audit closes the loop: monitor-aware GCG-style suffixes suppress the prompt-side monitor at zero behavioral cost at every scale; only a learned monitor rung resists its own adaptive re-attack; the training-time defense fails its re-attack at every penalty weight; steering shows the amplification profile is descriptive, not causal; and continuous-prefix optimization fails where discrete search succeeds. Safety monitoring, training, and evaluation must operate on accessibility itself, not on outputs alone.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Mosafer, M. (2026). From Latent Space to Jacobian Space: Measuring, Evading, and Training Against Safety-Content Accessibility. https://omanscience.com/en/articles/from-latent-space-to-jacobian-space-measuring-evading-and-training-against-safety-content-accessibility

MLA 9

Mosafer, Mohammad. "From Latent Space to Jacobian Space: Measuring, Evading, and Training Against Safety-Content Accessibility." https://omanscience.com/en/articles/from-latent-space-to-jacobian-space-measuring-evading-and-training-against-safety-content-accessibility.

Chicago (author–date)

Mosafer, Mohammad. 2026. "From Latent Space to Jacobian Space: Measuring, Evading, and Training Against Safety-Content Accessibility." https://omanscience.com/en/articles/from-latent-space-to-jacobian-space-measuring-evading-and-training-against-safety-content-accessibility.

Harvard

Mosafer, M. (2026) 'From Latent Space to Jacobian Space: Measuring, Evading, and Training Against Safety-Content Accessibility', Available at: https://omanscience.com/en/articles/from-latent-space-to-jacobian-space-measuring-evading-and-training-against-safety-content-accessibility.

Vancouver

Mosafer M. From Latent Space to Jacobian Space: Measuring, Evading, and Training Against Safety-Content Accessibility. https://omanscience.com/en/articles/from-latent-space-to-jacobian-space-measuring-evading-and-training-against-safety-content-accessibility

IEEE

M. Mosafer, "From Latent Space to Jacobian Space: Measuring, Evading, and Training Against Safety-Content Accessibility," https://omanscience.com/en/articles/from-latent-space-to-jacobian-space-measuring-evading-and-training-against-safety-content-accessibility.