Abstract
Large language models are increasingly used where small syntactic errors matter, yet character-level reasoning is still evaluated mostly through isolated probes and aggregate accuracy. We introduce SyntaxBench, a diagnostic benchmark and statistical evaluation framework for character-level reasoning. It contains five core tasks, character counting, letter containment, palindrome detection, edit distance, and longest-string selection, plus index_to_span, a harder substring-extraction stress test. The five core tasks use paired English and character-length-matched random-string inputs. index_to_span documents share a 200-500 word band and are not character-length matched. All six tasks use zero-, one-, and four-shot prompts. We evaluate eight open-weight models from 2B to 32B parameters across 11 reasoning-mode configurations. The framework reports exact-match and relaxed accuracy, Cohen's kappa, paired McNemar tests with odds ratios, bootstrap confidence intervals, Kendall's tau, class-conditional metrics, tokenization analysis, and multiple-comparison-corrected tests. Three findings stand out. First, tokenization shapes accuracy: random strings are more character-visible than English strings (1.892 vs. 3.169 characters per token), and character-counting accuracy falls as English words occupy more tokens. Second, reasoning mode is not uniformly helpful: Gemma4-31B is nearly unchanged across modes on the near-saturated tasks, while Qwen3.6-27B is worse with thinking on palindrome detection (0.952 non-thinking vs. 0.886 thinking at four-shot). Third, index_to_span remains largely unsolved; the best four-shot exact-match accuracy is 6.75%. Character-level evaluation needs controlled inputs, paired tests, and analyses of tokenization and reasoning mode rather than aggregate accuracy alone.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Larni, M., Azar, S. E., Nahed, P., & Taghva, K. (2026). SyntaxBench: A Statistical Diagnostic Framework for Character-Level Reasoning in Large Language Models. https://omanscience.com/en/articles/syntaxbench-a-statistical-diagnostic-framework-for-character-level-reasoning-in-large-language-models
MLA 9
Larni, Mohsen, et al. "SyntaxBench: A Statistical Diagnostic Framework for Character-Level Reasoning in Large Language Models." https://omanscience.com/en/articles/syntaxbench-a-statistical-diagnostic-framework-for-character-level-reasoning-in-large-language-models.
Chicago (author–date)
Larni, Mohsen, Sobhan Ebrahimi Azar, Pouyan Nahed, and Kazem Taghva. 2026. "SyntaxBench: A Statistical Diagnostic Framework for Character-Level Reasoning in Large Language Models." https://omanscience.com/en/articles/syntaxbench-a-statistical-diagnostic-framework-for-character-level-reasoning-in-large-language-models.
Harvard
Larni, M., Azar, S. E., Nahed, P. and Taghva, K. (2026) 'SyntaxBench: A Statistical Diagnostic Framework for Character-Level Reasoning in Large Language Models', Available at: https://omanscience.com/en/articles/syntaxbench-a-statistical-diagnostic-framework-for-character-level-reasoning-in-large-language-models.
Vancouver
Larni M, Azar SE, Nahed P, Taghva K. SyntaxBench: A Statistical Diagnostic Framework for Character-Level Reasoning in Large Language Models. https://omanscience.com/en/articles/syntaxbench-a-statistical-diagnostic-framework-for-character-level-reasoning-in-large-language-models
IEEE
M. Larni, S. E. Azar, P. Nahed, and K. Taghva, "SyntaxBench: A Statistical Diagnostic Framework for Character-Level Reasoning in Large Language Models," https://omanscience.com/en/articles/syntaxbench-a-statistical-diagnostic-framework-for-character-level-reasoning-in-large-language-models.