الباحثون

Wentao Zhang

المنشورات 9

نسخة أولية وصول مفتوح

AgentEvolver: System-Wide Self-Evolution Through Task Execution

Wentao Zhang, Fuchao Yang, Yilei Zhao وآخرون · 2026

An agent can complete a task without improving how it works. Turning task experience into reusable capability requires connecting the changed component to its evaluation and subsequent use. We present AgentEvolver, a system for developing capabilities during task execution while keeping the foundation model fixed. Eigh …

نسخة أولية وصول مفتوح

SemOPT: Fixing Semantic Errors in LLM-based Optimization Modeling via Reward-Guided Search

Zetong Zhou, Wentao Zhang, Jingyuan Wang وآخرون · 2026

Operations research supports decision-making in domains such as energy, economics, and healthcare. Solving operations research problems typically begins with optimization modeling, which translates a natural-language problem description into executable solver code. LLMs offer a promising way to automate this process, b …

نسخة أولية وصول مفتوح

A Benchmark & Dataset for Detecting AI-Manipulated Visual Evidence in the Court System

Kelly McConvey, Sajad Ebrahimi, Nima Jamali وآخرون · 2026 · 10.1145/3799682.3840194

Photographic evidence is becoming increasingly vulnerable to forms of alteration and fabrication that existing legal and technical workflows are not well equipped to evaluate. Surveillance frames, dashcam stills, and phone photographs may be used to establish presence, sequence, causation, damage, or identity, yet cont …

نسخة أولية وصول مفتوح

Do Coding Agents Reuse Existing Code or Reinvent the Wheel?

Dongsheng Ma, Sizhe Wang, Xinyi Huang وآخرون · 2026

Coding agents are increasingly deployed for iterative development on real repositories, yet existing evaluation barely answers a basic question: \emph{do coding agents reuse existing code or reinvent the wheel?} The question matters: every duplicated implementation is a fix applied twice and agents produce code far fas …

نسخة أولية وصول مفتوح

GenMem: Generative Symbolic Memory for Self-Evolving Harness

Xinke Jiang, Tao Feng, Weixuan Xu وآخرون · 2026

Long-term memory supports the self-evolution of LLM agents by retaining experience and skills across tasks and enabling their retrieval, reuse, and revision in subsequent long-horizon decision-making. Yet existing memory management approaches remain limited to discriminative retrieval and to address the sparse, hierarc …

نسخة أولية وصول مفتوح

MaPP: A Unified Marginalized Posterior-Predictive Framework for Data-Efficient RLVR

Yangyang Ren, Haodong Zhu, Sheng Xu وآخرون · 2026

Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models but incurs substantial costs from rollouts and policy updates. Online prompt selection improves efficiency by using per-prompt Bayesian posteriors to predict difficulty and prioritize informative prompts. …

نسخة أولية وصول مفتوح

OmniEdu: Open Foundation Models for Learning and Teaching

Hao Liang, Qihan Lin, Meiyi Qiang وآخرون · 2026

Educational foundation models must solve problems, understand curriculum structure, diagnose learner difficulties, and provide appropriate instructional support. Existing educational language models often focus on either problem solving or tutoring, with training mixtures organized by source or task rather than capabil …

نسخة أولية وصول مفتوح

DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression

DeepSeek-AI, Anyi Xu, B. Li وآخرون · 2026

The widespread adoption of long-horizon agents has made model workloads increasingly input-heavy. Although prior work has substantially reduced the cost of long-context computation, prefill remains computationally expensive, and large KV caches continue to strain HBM and SSD capacity and data-transfer bandwidth. Togeth …

المؤلفون المشاركون