الباحثون

Hao Wang

المنشورات 32

نسخة أولية وصول مفتوح

PageWeaver: KV-Guided Query Unions for Sparse Attention

Zhiyuan Li, Zihan Li, Zefang Yuan وآخرون · 2026

Dynamic sparse attention limits the KV pages selected by each query, but a small support does not necessarily yield efficient GPU work. Query unions share page loads and populate Tensor Core tiles; their cost depends on which queries are grouped together. We present PageWeaver, an execution design that uses selected-pa …

نسخة أولية وصول مفتوح

Harness Evolution Hits a Ceiling: When Weight Training Should Begin

Yuan Tian, Bing Hu, Hao Wang وآخرون · 2026

Improving a long-horizon LLM agent means evolving the harness around a frozen model or training its weights. We let a self-evolving harness make the system stronger first, then cross seed and evolved harnesses with base and trained weights to learn which gains the trained model keeps and which still need the runtime. W …

نسخة أولية وصول مفتوح

Geometry-Supervised Visual Representation Learning for Multi-Phenotype Lesion Interpretation in Medical VLMs

Hao Wang, Qiwei Zeng, Shuchang Ye وآخرون · 2026

Medical vision-language models (VLMs) have shown increasing potential for clinical image interpretation. However, these models still struggle to interpret multi-phenotype lesions whose diagnosis requires the joint assessment of multiple pathological phenotypes. Existing vision-language alignment methods produce visual …

نسخة أولية وصول مفتوح

$Δ$Representation: Geometry Supervised Representation Learning of Phenotypes via Counterfactual Reasoning for Medical VLMs

Hao Wang, Qiwei Zeng, Jinghao Lin وآخرون · 2026

Medical vision-language models (VLMs) have shown increasing potential for radiological image interpretation. Medical VLMs encode radiological images into visual representations that capture both anatomical and phenotypic information for diagnosis. Existing approaches improve pathological phenotype representations throu …

نسخة أولية وصول مفتوح

EvoSignal: LLM-Guided Evolutionary Design of Modular Traffic Signal Control Programs

Leizhen Wang, Peibo Duan, Zhenlin Qin وآخرون · 2026

Effective traffic signal control (TSC) requires policies that respond to changing traffic demand and network conditions while meeting different control objectives. However, adapting existing strategies often involves repeated manual design and adjustment, making it difficult to systematically explore better control rul …

نسخة أولية وصول مفتوح

TestJack: Should you trust the results in coding benchmarks? Agentic Coding Benchmarks Auditing via Evaluator Evolution

Shuangjie Yao, Hao Wang, Koushik Sen وآخرون · 2026

Large language model (LLM) agents are rapidly reshaping software engineering, accompanied by an explosion of new code benchmarks. Yet nearly all existing benchmarks still rely on the same decades-old criterion: a solution is correct if it passes a fixed set of unit tests. Such tests are often insufficient: they check o …

نسخة أولية وصول مفتوح

Sequential Probabilistic Uncertainty Estimation for Parallel Multi-Agent Reasoning Systems

Tunyu Zhang, Zihao Zhao, Yusong Zhao وآخرون · 2026

LLM-based multi-agent systems (MAS) have attracted growing attention for improving reasoning through interaction among multiple agents. In this work, we focus on parallel multi-agent reasoning systems, where several agents solve the same problem over multiple rounds and aggregate their outputs into a final answer. Desp …

نسخة أولية وصول مفتوح

Work While They Sleep: Exploiting Evaluation Latency for Fully Bayesian Optimization

Black-box optimization problems are ubiquitous across science and engineering, often dealing with expensive objective functions. This objective latency has two consequences during optimization: (i) the objective evaluation dominates execution time, and (ii) sample-efficient algorithms are crucial to accelerate developm …

نسخة أولية وصول مفتوح

ReDiffNet: Differential RGB-Infrared Learning for Low-Light UAV Oriented Vehicle Detection

Qifan Zhang, Ziran Zhou, Ruijie Li وآخرون · 2026

Low-light UAV-based RGB-infrared oriented small-vehicle detection is important for nighttime traffic monitoring, emergency response, and urban inspection. Illumination variations, headlight glare, local shadows, and thermal-response degradation cause spatially varying modality reliability, while the small visual extent …

نسخة أولية وصول مفتوح

When Debate Helps: Proposal Supply and Verification-Aware Readout in Multi-Agent Reasoning

Zihao Zhao, Tunyu Zhang, Haizhou Shi وآخرون · 2026

Multi-agent debate can improve reasoning, yet often fails to beat simple majority voting. We argue that successful debate requires two distinct mechanisms: proposal supply must surface a correct answer, and readout must identify that answer when voting misses it. We formalize the first requirement through recoverable h …

نسخة أولية وصول مفتوح

AvoKV-E: Payload-Aware KV Cache Eviction for Long Reasoning

Han Yu, Wenhui Zhu, Xiwen Chen وآخرون · 2026

Long-output reasoning shifts the KV-cache bottleneck from the fixed prompt to the generated trace. Existing reasoning-cache eviction methods largely treat cached entries as routing objects, estimating whether an old key will still be read, will recur, or can be replaced. This routing-only view overlooks two effects: lo …

نسخة أولية وصول مفتوح

LatentQuant: Preserving the Policy-Facing Latent Contract under NVFP4 VAE Quantization

Ziye Deng, Lufang Chen, Shuyu Feng وآخرون · 2026

Recent world action models (WAMs) reuse pretrained video VAEs whose encoder latents directly condition downstream action policies. Quantization must therefore preserve not only reconstruction fidelity but also the policy-facing latent contract expected by the frozen policy. Direct NVFP4 leaves W4A4 quantization error u …

نسخة أولية وصول مفتوح

GIFTBench: Diagnosing Generalization in Image Forgery Localization and Informing Model Design

Baoke Dou, Ziye Wang, Hao Wang وآخرون · 2026

Reliable evaluation of image forgery localization (IFL) requires assessing models under diverse distribution changes, yet existing benchmarks often cover limited manipulation conditions or entangle multiple factors in cross-dataset evaluation. Consequently, aggregate performance provides an incomplete view of localizat …

نسخة أولية وصول مفتوح

EWAM: Emergent Depth-Wise Specialization in a Unified Embodied Model -- From Semantic Understanding through Visual Foresight to Action

Hao Wang, Jiajun Wen, Jingzhi Liu وآخرون · 2026

Vision-language-action (VLA) policies emphasize semantic understanding, whereas world-action models (WAMs) learn predictive representations of environment dynamics. Systems that expose a policy to both sources often still concentrate action computation on a single expert. We present EWAM, an action-centric unified embo …

نسخة أولية وصول مفتوح

Denoising Surface: Modeling and Predicting Inference Cost for Diffusion LLM Serving

Haoyu Zheng, Fangcheng Fu, Binhang Yuan وآخرون · 2026

As diffusion large language models (dLLMs) become more capable, they are moving from research settings to real-world \textit{serving}, where request management (such as scheduling and resource allocation) relies on accurate estimation of per-request inference cost. However, common cost proxies fall short for dLLMs: out …

نسخة أولية وصول مفتوح

Explicit Trajectory Diversity for RL-Based Post-Training of LLM Agents

Huaiyu Fu, Heng Cao, Hao Wang وآخرون · 2026

LLM agents often admit multiple high-quality solutions to the same task, differing in reasoning structure, tool-use pattern, or interaction trajectory. Yet existing notions of diversity in LLM post-training are mostly implicit, arising from general stochasticity and regularization mechanisms rather than explicitly targ …

نسخة أولية وصول مفتوح

D-Scope: Decomposing and Steering Diffusion Transformers with Sparse Autoencoders

Xinyue Xu, Jiahao Zhang, Lijie Hu وآخرون · 2026

Sparse autoencoders (SAEs) reveal visual structure in diffusion transformers (DiTs), but interpreting a feature does not establish whether it can be used to control generation. We introduce D-Scope (Diffusion Scope), a framework that connects feature interpretation to generation control through shared visual evidence. …

المؤلفون المشاركون