Abstract
Screening a phone call for fraud needs a trustworthy probability after every caller turn, in milliseconds. Jev-style typed decisions promise exactly that: declared options go in, one calibrated probability per option comes out of a single forward pass, with no generated text. We test an open implementation of this readout, JevLite, on scam-call screening: Qwen3-4B is LoRA-tuned so that the temperature-scaled softmax over two answer-label logits is P(scam). On 41 held-out CallScreenBench scenarios (577 per-turn decisions) a three-seed ensemble reaches AUROC .974 with calibration error .052, non-inferior to an LLM judge (MiniMax-M3) at a pre-registered .02 margin, with no false alarms on legitimate calls, decisions 1.14 turns earlier under the same hang-up rule, and 64.5 ms per decision on one consumer GPU, 4.9x lower than the same backbone fine-tuned to generate its answer. The gain is in the readout and calibration, not accuracy: a fine-tuned ModernBERT encoder is not significantly worse, the recipe was selected with test-set exposure, and all callers are synthetic. We claim no architectural novelty; the contribution is the application and an evaluation reporting calibration, false alarms and decision timing alongside AUROC.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Ren, S., Zewde, K., Shen, X., Zhou, Y., Ng, D., Raj, A., Duong, T., Zhang, Y., & Tiangratanakul, N. (2026). Open-Jev Judgments on CallScreenBench: Calibrated One-Pass Scam Screening with a Small Language Model. https://omanscience.com/en/articles/open-jev-judgments-on-callscreenbench-calibrated-one-pass-scam-screening-with-a-small-language-model
MLA 9
Ren, Simiao, et al. "Open-Jev Judgments on CallScreenBench: Calibrated One-Pass Scam Screening with a Small Language Model." https://omanscience.com/en/articles/open-jev-judgments-on-callscreenbench-calibrated-one-pass-scam-screening-with-a-small-language-model.
Chicago (author–date)
Ren, Simiao, Kidus Zewde, Xingyu Shen, Yuchen Zhou, Dennis Ng, Ankit Raj, Tommy Duong, Yuxin Zhang, and Neo Tiangratanakul. 2026. "Open-Jev Judgments on CallScreenBench: Calibrated One-Pass Scam Screening with a Small Language Model." https://omanscience.com/en/articles/open-jev-judgments-on-callscreenbench-calibrated-one-pass-scam-screening-with-a-small-language-model.
Harvard
Ren, S., Zewde, K., Shen, X., Zhou, Y., Ng, D., Raj, A., Duong, T., Zhang, Y. and Tiangratanakul, N. (2026) 'Open-Jev Judgments on CallScreenBench: Calibrated One-Pass Scam Screening with a Small Language Model', Available at: https://omanscience.com/en/articles/open-jev-judgments-on-callscreenbench-calibrated-one-pass-scam-screening-with-a-small-language-model.
Vancouver
Ren S, Zewde K, Shen X, Zhou Y, Ng D, Raj A, et al. Open-Jev Judgments on CallScreenBench: Calibrated One-Pass Scam Screening with a Small Language Model. https://omanscience.com/en/articles/open-jev-judgments-on-callscreenbench-calibrated-one-pass-scam-screening-with-a-small-language-model
IEEE
S. Ren, K. Zewde, X. Shen, Y. Zhou, D. Ng, A. Raj, T. Duong, Y. Zhang, and N. Tiangratanakul, "Open-Jev Judgments on CallScreenBench: Calibrated One-Pass Scam Screening with a Small Language Model," https://omanscience.com/en/articles/open-jev-judgments-on-callscreenbench-calibrated-one-pass-scam-screening-with-a-small-language-model.