الباحثون

Cheng Yang

المنشورات 7

نسخة أولية وصول مفتوح

SWE-Journey: Towards More Realistic Evaluation of Coding Assistants through Long-Horizon, Multi-Turn Interaction

Hexuan Deng, Yue Wang, Wenyu Jiang وآخرون · 2026

Coding assistants such as Claude Code and Codex have become a major application of LLM agents, yet existing benchmarks remain far from real-world use, particularly in task horizon and interaction length. Code assistants require completing long chains of development work in continuously evolving repositories, while repe …

نسخة أولية وصول مفتوح

Visual Abstention in Unified Multimodal Models

Chufan Shi, Cheng Yang, Tiannuo Yang وآخرون · 2026

Unified multimodal models (UMMs) integrate understanding and generation, yet their generative behavior is rarely governed by what they understand about the task. We formalize visual abstention: when a requested visual transformation is impossible under the task's rules, the model should recognize that no valid solution …

نسخة أولية وصول مفتوح

MASBench: Benchmarking LLM-based Multi-Agent Collaboration under Partial Observability

Qizhi Chu, Zekai Yu, Sijie Wen وآخرون · 2026

Large language models (LLMs) have progressively evolved into the core of autonomous agents. Building on this progress, LLM-based multi-agent systems (MAS) coordinate multiple agents into a synergistic team to accomplish complex tasks that exceed the capabilities of individual agents. The effectiveness of such systems d …

نسخة أولية وصول مفتوح

Recursive Self-Improvement in Unified Multimodal Models

Huijuan Wang, Chufan Shi, Cheng Yang وآخرون · 2026

Unified multimodal models (UMMs) understand and generate both text and images, which lets a model produce its own training data. Existing self-improvement in UMMs keeps supervision on the visual side, where image understanding judges image generation. We propose recursive cross-capability self-improvement (RSI), a trai …

نسخة أولية وصول مفتوح

AutoGUIWorld: Image Generators as Visual World Models for GUI Agent

Cheng Yang, Yifan Wu, Yutao Huang وآخرون · 2026

GUI agents require high-quality interaction trajectories to learn how software environments respond to actions, maintain state, and support multi-step workflows. However, the diversity of available trajectories is constrained by the applications, interface states, and workflows accessible in the underlying environments …

نسخة أولية وصول مفتوح

InfoEdit: Probing Global Layout Reasoning in Infographic Editing

Cheng Yang, Chufan Shi, Huijuan Wang وآخرون · 2026

Multimodal foundation models edit natural photographs at production quality, yet the same models struggle with structured visual content such as infographics. Unlike photographs, infographics encode information through logical relations; editing one element often requires surrounding elements to be adapted. We refer to …

المؤلفون المشاركون