الباحثون

Chengcheng Wan

المنشورات 3

نسخة أولية وصول مفتوح

CheckerBench: Can Long-Horizon Agents Synthesize Static-Analysis Checkers?

Hang He, Li Wang, Hao Chen وآخرون · 2026

Static-analysis checker synthesis requires agents to interpret a defect specification, inspect a repository, implement analyzer-specific logic, and refine the checker through repeated compilation and analysis feedback. Existing coding-agent benchmarks focus on tasks such as patch generation or vulnerability detection a …

نسخة أولية وصول مفتوح

Harmful SFT Leaves a Continuous Trace in LLM Checkpoint Updates

Ziqun Bao, Xinyu Zhang, Yuchen Shao وآخرون · 2026

Safety auditing of post-trained large language models typically relies on model behavior, requiring model execution and depending on the coverage of available evaluations. This work asks a different question: Do the target behaviors optimized during supervised fine-tuning (SFT) leave readable evidence directly in check …

نسخة أولية وصول مفتوح

WEFT: Scaling Tool-Use Post-Training for General-Purpose Agents

Bo Mao, Hang He, Linting Wang وآخرون · 2026

Recent efforts to scale tool-use post-training have largely centered on the synthesis of executable environments, which constitute only one component of a broader agentic interaction system comprising the environment, task, agent harness, and evaluator. Scaling environments in isolation, however, does not guarantee com …

المؤلفون المشاركون