الباحثون

Yang Liu

المنشورات 37

نسخة أولية وصول مفتوح

PyCache Trap: The Inspection-Execution Gap in Agent Skill Scanners

Jie Liao, Simeng Qin, Wenqi Ren وآخرون · 2026

Agent skills combine instructions with executable resources, giving third-party packages access to an agent's runtime. Existing skill scanners inspect documentation and visible source, but Python may execute a bundled bytecode cache with different behavior. We study this gap between inspection and execution through PyC …

نسخة أولية وصول مفتوح

Robotic Boomerang Throwing via Model-Based Release Design

Throwing objects that generate aerodynamic lift can greatly extend robot throwing beyond ballistic flight. A returning boomerang is a challenging example because its flight depends strongly on the release velocity, attitude, and spin, while robotic manipulators cannot readily reproduce the rapid motions used in human t …

نسخة أولية وصول مفتوح

Fast Planning for Multi-object Multi-target Throwing

Zhengming Zhu, Yang Liu, Xiao Gao وآخرون · 2026

Robot throwing has emerged as a promising technique for improving efficiency in logistics and warehouse automation, by enlarging the workspace and speeding up the process. To significantly increase the throwing system's throughput, we develop strategies for throwing multiple objects in one swipe. Such multi-object mult …

نسخة أولية وصول مفتوح

Visual-Invariance-Augmented Feature Optimal Alignment for Transferable Adversarial Attacks against Closed-Source MLLMs

Xiaojun Jia, Simeng Qin, Yiming Li وآخرون · 2026

Multimodal large language models (MLLMs) remain vulnerable to transferable adversarial examples, especially in black-box settings where only open-source surrogate models are accessible. Existing targeted transfer attacks mainly align adversarial and target samples using global image-level features, such as encoder [CLS …

نسخة أولية وصول مفتوح

Sparse-View 4D Gaussian Splatting via Spatiotemporal Priors and Generative Assistance

Shengqi Wang, Zhengxian Yang, Kaiwen Tian وآخرون · 2026

We present a 4D Gaussian Splatting framework for the Sparse-View Track of the SIGGRAPH Asia 2026 Volumetric Video Challenge, which requires dynamic scene reconstruction from only six cameras with wide baselines. To achieve robust dynamic reconstruction under such sparse views, our framework integrates three components. …

نسخة أولية وصول مفتوح

MASBench: Benchmarking LLM-based Multi-Agent Collaboration under Partial Observability

Qizhi Chu, Zekai Yu, Sijie Wen وآخرون · 2026

Large language models (LLMs) have progressively evolved into the core of autonomous agents. Building on this progress, LLM-based multi-agent systems (MAS) coordinate multiple agents into a synergistic team to accomplish complex tasks that exceed the capabilities of individual agents. The effectiveness of such systems d …

نسخة أولية وصول مفتوح

Federated Agent Optimization

Qiang Yang, Zhiqiang Kou, Xueyi Zhang وآخرون · 2026

Large language model (LLM) agents increasingly operate in private environments and accumulate valuable experience from task execution, tool use, feedback, and local knowledge. Yet such experience is distributed across organizations and cannot be directly shared because of privacy and proprietary constraints. Convention …

نسخة أولية وصول مفتوح

Recovering the View: Benchmarking Physical Active Vision for Occlusion Recovery in Robotic Manipulation

Kaijun Luo, Yudi Huang, Qijun Zhong وآخرون · 2026

Physical active vision allows robots to change their viewpoint when task-relevant observations become unreliable, yet existing manipulation benchmarks provide limited support for studying how policies recover from occlusion during execution. We introduce BAVO-Bench (Bimanual Active Vision under Occlusion), a bimanual a …

نسخة أولية وصول مفتوح

Which Models Work Well Together? Measuring Heterogeneity for LLM Team Selection

Liangyu Teng, Hengsong Liu, Juncen Guo وآخرون · 2026

The performance ceiling of an LLM team is constrained not only by individual model capabilities, but also by inter-member error resonance and predictive differences. Although heterogeneous teaming is often observed to be effective in practice, existing approaches lack complementarity metrics that are computable, interp …

نسخة أولية وصول مفتوح

Long-Term Memory-Guided Enhancement for Target Perception in Audio-Language Models

Zhenhong Zhou, Xuanyue Zhao, Youji Liu وآخرون · 2026

Audio large language models (ALLMs) can reason about the content of audio recordings to perform complex tasks. However, these capabilities usually collapse in real-world environments when background noise and competing sources mix the target sound. Inspired by long-term memory in human listening, we propose Long-Term M …

نسخة أولية وصول مفتوح

When Upstream Messages Override Correct Answers: A Controlled Study of Multi-Agent LLM Collaboration

Yaxin Gong, Gangyi Zhang, Chongming Gao وآخرون · 2026

Multi-agent LLM systems rely on message passing among specialized agents to accomplish complex tasks. However, an upstream agent may provide useful information or an incorrect answer that causes a downstream agent to override a correct answer supported by its own evidence. Prior work has not clearly separated the benef …

نسخة أولية وصول مفتوح

Breaking Babel: A Self-Evolving Multi-Agent System for Long-Form Subtitle Translation

Haibo Jin, Xinjie Li, Najmeh Sadoughi وآخرون · 2026

Long-form subtitle translation requires reasoning over discourse and cultural context spanning episodes or entire series, while maintaining consistent terminology and style. Existing single-LLM methods are largely sentence-level, and multi-agent systems often use static workflows that do not adapt to scene complexity o …

نسخة أولية وصول مفتوح

SmartMemory: Detecting On-chain-off-chain Communication Inconsistency for Smart Contract via Memory-based Agent

Zeqin Liao, Yuhong Nan, Henglong Liang وآخرون · 2026

Smart contracts underpin decentralized finance, where growing demand for on-chain/off-chain communication(OFC) has driven diverse applications such as cross-chain bridges, real-world asset tokenization, and fiat-backed stablecoins. TheOFC-related security incidents in these applications are increasingly frequent, but p …

نسخة أولية وصول مفتوح

When Words Fall Short: Iterative Synergy Between Verbalized Reasoning and Hidden Features for LLM Confidence Estimation

Yekun Xu, Ante Wang, Jingyi Ren وآخرون · 2026

Confidence estimation is crucial for developing trustworthy large language models (LLMs), with most methods following estimator-based or verbalization-based paradigms. While recent research increasingly focuses on improving verbalized self-reports of confidence, we challenge the prevailing view that this approach surpa …

نسخة أولية وصول مفتوح

Beyond Verbalized Confidence: Calibrating Reasoners with Differentiable Readouts

Chenxiao Fan, Chongming Gao, Gangyi Zhang وآخرون · 2026

Reinforcement learning with verifiable rewards (RLVR) trains reasoning models to produce correct answers, but does not ensure that their stated confidence is calibrated. The resulting models are systematically overconfident. Recent methods train calibration inside the RLVR loop by having the model state a numerical con …

نسخة أولية وصول مفتوح

GAUGE: Group-Wise View-Inconsistency Rectification for Feed-Forward 4D Tracking

Zhuoqian Feng, Weixing Chen, Ziliang Chen وآخرون · 2026

Feed-forward models regress dense 3D point trajectories directly from monocular video, yet the residual after global alignment is substantial and lacks a structural explanation. Measured on dynamic query points across models and datasets, the error concentrates along the view direction, while the scale correction each …

المؤلفون المشاركون