نسخة أولية وصول مفتوح
SkillScriptBench: Benchmarking Self-Evolution of Executable Agent Skill Packages Beyond Markdown
Executable Agent Skills combine natural-language instructions and scripts into reusable packages for LLM agents, and revising them requires fixing errors without breaking correct behavior. Existing benchmarks do not systematically distinguish documentation repair, script repair, and preservation when evaluating skill s …