الباحثون

يوشين شاو

المنشورات 3

نسخة أولية وصول مفتوح

CheckerBench: Can Long-Horizon Agents Synthesize Static-Analysis Checkers?

Static-analysis checker synthesis requires agents to interpret a defect specification, inspect a repository, implement analyzer-specific logic, and refine the checker through repeated compilation and analysis feedback. Existing coding-agent benchmarks focus on tasks such as patch generation or vulnerability detection a …

نسخة أولية وصول مفتوح

Harmful SFT Leaves a Continuous Trace in LLM Checkpoint Updates

Safety auditing of post-trained large language models typically relies on model behavior, requiring model execution and depending on the coverage of available evaluations. This work asks a different question: Do the target behaviors optimized during supervised fine-tuning (SFT) leave readable evidence directly in check …

نسخة أولية وصول مفتوح

OneStreamer: Unifying Perception, Memory, and Proactive Response in Streaming Video Interaction

Xiangyu Zeng, Yuandong Yang, Zhiqiu Zhang وآخرون · 2026

Streaming video LLMs must retain evidence before its relevance to future tasks is known and respond when sufficient evidence becomes available. The challenge is to form reusable factual memory without compromising real-time perception. We introduce OneStreamer, which jointly learns query-independent evidence recording …

المؤلفون المشاركون