الباحثون

Yutao Wu

المنشورات 2

نسخة أولية وصول مفتوح

VEX-Bench: Benchmarking Verification Complexity of LLM-Generated Misinformation

Hanxun Huang, Yutao Wu, Qizhou Wang وآخرون · 2026

Large language models (LLMs) have made misinformation inexpensive to produce but not to verify, creating a growing asymmetry in the information ecosystem. Under tight time, labor, and budget constraints, media organizations, platforms, and fact-checkers rely on screening to prioritize which content to verify. We introd …

نسخة أولية وصول مفتوح

Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures

Ruoqi Guo, Yi Liu, Gelei Deng وآخرون · 2026

Detectors of alignment failures screen deployed language models and score alignment benchmarks. Most are generative judges that spend a decoding pass on every criterion, and classifiers that read token probabilities, such as Llama Guard, still score one fixed label per call. Jev, a model trained with reinforcement lear …

المؤلفون المشاركون