Abstract

Screening a phone call for fraud needs a trustworthy probability after every caller turn, in milliseconds. Jev-style typed decisions promise exactly that: declared options go in, one calibrated probability per option comes out of a single forward pass, with no generated text. We test an open implementation of this readout, JevLite, on scam-call screening: Qwen3-4B is LoRA-tuned so that the temperature-scaled softmax over two answer-label logits is P(scam). On 41 held-out CallScreenBench scenarios (577 per-turn decisions) a three-seed ensemble reaches AUROC .974 with calibration error .052, non-inferior to an LLM judge (MiniMax-M3) at a pre-registered .02 margin, with no false alarms on legitimate calls, decisions 1.14 turns earlier under the same hang-up rule, and 64.5 ms per decision on one consumer GPU, 4.9x lower than the same backbone fine-tuned to generate its answer. The gain is in the readout and calibration, not accuracy: a fine-tuned ModernBERT encoder is not significantly worse, the recipe was selected with test-set exposure, and all callers are synthetic. We claim no architectural novelty; the contribution is the application and an evaluation reporting calibration, false alarms and decision timing alongside AUROC.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Ren, S., Zewde, K., Shen, X., Zhou, Y., Ng, D., Raj, A., Duong, T., Zhang, Y., & Tiangratanakul, N. (2026). Open-Jev Judgments on CallScreenBench: Calibrated One-Pass Scam Screening with a Small Language Model. https://omanscience.com/en/articles/open-jev-judgments-on-callscreenbench-calibrated-one-pass-scam-screening-with-a-small-language-model

MLA 9

Ren, Simiao, et al. "Open-Jev Judgments on CallScreenBench: Calibrated One-Pass Scam Screening with a Small Language Model." https://omanscience.com/en/articles/open-jev-judgments-on-callscreenbench-calibrated-one-pass-scam-screening-with-a-small-language-model.

Chicago (author–date)

Ren, Simiao, Kidus Zewde, Xingyu Shen, Yuchen Zhou, Dennis Ng, Ankit Raj, Tommy Duong, Yuxin Zhang, and Neo Tiangratanakul. 2026. "Open-Jev Judgments on CallScreenBench: Calibrated One-Pass Scam Screening with a Small Language Model." https://omanscience.com/en/articles/open-jev-judgments-on-callscreenbench-calibrated-one-pass-scam-screening-with-a-small-language-model.

Harvard

Ren, S., Zewde, K., Shen, X., Zhou, Y., Ng, D., Raj, A., Duong, T., Zhang, Y. and Tiangratanakul, N. (2026) 'Open-Jev Judgments on CallScreenBench: Calibrated One-Pass Scam Screening with a Small Language Model', Available at: https://omanscience.com/en/articles/open-jev-judgments-on-callscreenbench-calibrated-one-pass-scam-screening-with-a-small-language-model.

Vancouver

Ren S, Zewde K, Shen X, Zhou Y, Ng D, Raj A, et al. Open-Jev Judgments on CallScreenBench: Calibrated One-Pass Scam Screening with a Small Language Model. https://omanscience.com/en/articles/open-jev-judgments-on-callscreenbench-calibrated-one-pass-scam-screening-with-a-small-language-model

IEEE

S. Ren, K. Zewde, X. Shen, Y. Zhou, D. Ng, A. Raj, T. Duong, Y. Zhang, and N. Tiangratanakul, "Open-Jev Judgments on CallScreenBench: Calibrated One-Pass Scam Screening with a Small Language Model," https://omanscience.com/en/articles/open-jev-judgments-on-callscreenbench-calibrated-one-pass-scam-screening-with-a-small-language-model.