Abstract

Large language models are increasingly used where small syntactic errors matter, yet character-level reasoning is still evaluated mostly through isolated probes and aggregate accuracy. We introduce SyntaxBench, a diagnostic benchmark and statistical evaluation framework for character-level reasoning. It contains five core tasks, character counting, letter containment, palindrome detection, edit distance, and longest-string selection, plus index_to_span, a harder substring-extraction stress test. The five core tasks use paired English and character-length-matched random-string inputs. index_to_span documents share a 200-500 word band and are not character-length matched. All six tasks use zero-, one-, and four-shot prompts. We evaluate eight open-weight models from 2B to 32B parameters across 11 reasoning-mode configurations. The framework reports exact-match and relaxed accuracy, Cohen's kappa, paired McNemar tests with odds ratios, bootstrap confidence intervals, Kendall's tau, class-conditional metrics, tokenization analysis, and multiple-comparison-corrected tests. Three findings stand out. First, tokenization shapes accuracy: random strings are more character-visible than English strings (1.892 vs. 3.169 characters per token), and character-counting accuracy falls as English words occupy more tokens. Second, reasoning mode is not uniformly helpful: Gemma4-31B is nearly unchanged across modes on the near-saturated tasks, while Qwen3.6-27B is worse with thinking on palindrome detection (0.952 non-thinking vs. 0.886 thinking at four-shot). Third, index_to_span remains largely unsolved; the best four-shot exact-match accuracy is 6.75%. Character-level evaluation needs controlled inputs, paired tests, and analyses of tokenization and reasoning mode rather than aggregate accuracy alone.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Larni, M., Azar, S. E., Nahed, P., & Taghva, K. (2026). SyntaxBench: A Statistical Diagnostic Framework for Character-Level Reasoning in Large Language Models. https://omanscience.com/en/articles/syntaxbench-a-statistical-diagnostic-framework-for-character-level-reasoning-in-large-language-models

MLA 9

Larni, Mohsen, et al. "SyntaxBench: A Statistical Diagnostic Framework for Character-Level Reasoning in Large Language Models." https://omanscience.com/en/articles/syntaxbench-a-statistical-diagnostic-framework-for-character-level-reasoning-in-large-language-models.

Chicago (author–date)

Larni, Mohsen, Sobhan Ebrahimi Azar, Pouyan Nahed, and Kazem Taghva. 2026. "SyntaxBench: A Statistical Diagnostic Framework for Character-Level Reasoning in Large Language Models." https://omanscience.com/en/articles/syntaxbench-a-statistical-diagnostic-framework-for-character-level-reasoning-in-large-language-models.

Harvard

Larni, M., Azar, S. E., Nahed, P. and Taghva, K. (2026) 'SyntaxBench: A Statistical Diagnostic Framework for Character-Level Reasoning in Large Language Models', Available at: https://omanscience.com/en/articles/syntaxbench-a-statistical-diagnostic-framework-for-character-level-reasoning-in-large-language-models.

Vancouver

Larni M, Azar SE, Nahed P, Taghva K. SyntaxBench: A Statistical Diagnostic Framework for Character-Level Reasoning in Large Language Models. https://omanscience.com/en/articles/syntaxbench-a-statistical-diagnostic-framework-for-character-level-reasoning-in-large-language-models

IEEE

M. Larni, S. E. Azar, P. Nahed, and K. Taghva, "SyntaxBench: A Statistical Diagnostic Framework for Character-Level Reasoning in Large Language Models," https://omanscience.com/en/articles/syntaxbench-a-statistical-diagnostic-framework-for-character-level-reasoning-in-large-language-models.