الباحثون

Simon Ostermann

المنشورات 2

نسخة أولية وصول مفتوح

FineWeb-CLaR: Culture, Language, and Region Annotations for Benchmark-Aligned Corpus Auditing

Cultural evaluation coverage and robustness in language models are difficult to diagnose because pretraining corpora and cultural benchmarks are rarely indexed with comparable metadata. Benchmarks increasingly target culturally situated phenomena at the level of languages, regions, and locale-specific practices, while …

نسخة أولية وصول مفتوح

From Concept Alignment to Causal Grounding: An Intervention Test of Chain-of-Thought Faithfulness

Qianli Wang, Yilong Wang, Dennis Wei وآخرون · 2026

Chain-of-thought (CoT) can sound plausible yet be unfaithful to the model's underlying reasoning. Most prior work probes CoT faithfulness through input--output behavior or input attributions, leaving internal computation largely underexplored. We instead cast faithfulness as internal concept grounding: Does a large lan …

المؤلفون المشاركون