الباحثون

Maxm Pan

المنشورات 3

نسخة أولية وصول مفتوح

ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds

Ming Zhang, Zhenghao Xiang, Peizhong Gao وآخرون · 2026

Scientific discovery begins where known problems end. There, AI systems must engage in exploration: framing hypotheses, designing experiments, and iterating on the results. However, evaluating this ability is difficult: (1) how to verify whether a genuinely new hypothesis holds, and (2) how to determine whether a syste …

نسخة أولية وصول مفتوح

Beyond Outcomes: Dual-View Relational Learning for Efficient Agent Benchmarking

Xinshuai Guo, Junjie Wu, Dolly Deng وآخرون · 2026

Agent benchmarks are substantially more costly to evaluate than conventional LLM benchmarks. Benchmark compression is therefore a natural solution, yet existing methods primarily model redundancy in task--model final-score distributions, which is important in agentic evaluation. To address this limitation, we analyze l …

المؤلفون المشاركون