الباحثون

Yue Huang

المنشورات 4

نسخة أولية وصول مفتوح

MADBench: Benchmarking the Security of Multi-Agent Debate

Yuwan Liu, Jiaming Zhang, Yue Huang وآخرون · 2026

Multi-agent debate (MAD) can improve large language model (LLM) reasoning by allowing multiple agents to exchange and critique their answers to the same task. However, the interactions that enable agents to correct mistakes can also spread adversarial errors and steer the agents toward an incorrect answer. Although som …

نسخة أولية وصول مفتوح

MiniRep: Robust Reputation-Based Aggregation for Multi-Agent Debate

Jiaming Zhang, Yuwan Liu, Yue Huang وآخرون · 2026

Autonomous agents powered by large language models (LLMs) are rapidly evolving into an open agentic ecosystem. To support trustworthy collaboration, industry initiatives increasingly assess agent reputation from past behavior and provide performance leaderboards. However, reputation derived from past performance may no …

نسخة أولية وصول مفتوح

SecProbe: Adaptive Evaluation of Coding Agents on Cybersecurity Vulnerabilities

Xiaonan Luo, Yue Huang, Kehan Guo وآخرون · 2026

Assessing cybersecurity vulnerability awareness in coding agents requires evaluations that reveal capability gaps and remain informative as models evolve. Static benchmarks offer fixed coverage and difficulty, while scarce vulnerable repositories and costly expert authoring limit their renewal at scale. We introduce Se …

نسخة أولية وصول مفتوح

Reward Hacking Challenges Oversight of Autonomous Research Agents

Yue Huang, Zhangchen Xu, Yuchen Ma وآخرون · 2026

Autonomous research agents can design experiments, evaluate results, and write reports, giving them control over both a scientific result and the evidence used to support it. This creates a risk of reward hacking: meeting the reward criteria without achieving the intended goal. We study (1) how often models reward-hack …

المؤلفون المشاركون