الباحثون

Yang Yang

المنشورات 22

نسخة أولية وصول مفتوح

SpikingVLA: Asynchronous Spiking Vision-Language-Action Models

Jingya Wang, Dehao Zhang, Shuai Wang وآخرون · 2026

ANN-to-SNN conversion offers a practical route toward energy-efficient spiking Vision-Language-Action (VLA) models by bypassing the substantial cost of training large-scale SNNs from scratch. However, existing methods often require many timesteps to maintain competitive performance, resulting in substantial inference l …

نسخة أولية وصول مفتوح

MIRA: A Musical Intent Refinement Agent for Aligning Text-to-Music Generation with User Intent

Zekai Liu, Zhilin Wang, Xuzheng He وآخرون · 2026

Text-to-music systems produce increasingly convincing audio, yet evaluation reveals little about whether the result matches user intent. A global text-audio relevance score can overlook the implicit intent in underspecified prompts and mask failures in specific requirements, such as instrumentation, structure, rhythm, …

نسخة أولية وصول مفتوح

Patient, Place, Prior (P$^3$): What Counts as Personalization in Medical World Models?

Xingrui Gu, Hanxue Gu, Yuxiang Zhang وآخرون · 2026

Longitudinal models forecast how a patient's imaging state evolves, but accuracy does not show whether the patient's observed trajectory drives the prediction. A population-average forecast may be useful but cannot establish a patient-specific world-model claim. We introduce Patient, Place, Prior (P$^3$), an audit aski …

نسخة أولية وصول مفتوح

Internalizing Agent Experience into Diffusion Model Weights via On-Policy Context Distillation

Wenxuan Wang, Zekai Liu, Weinan Zhang وآخرون · 2026

Wrapping an image generation model in an agentic harness can effectively boost Text-to-Image task performance: the harness can leverage memory, skills, workflow orchestration, result verification, and iterative refinement to continually construct and revise prompts, thereby eliciting better images. These gains, however …

نسخة أولية وصول مفتوح

Grounded in Time: A Multi-Source Dataset and Benchmark for Temporal Grounding in Robotic Manipulation

Yi Wang, Yang Yang, Guangqi Xu وآخرون · 2026

Robotic manipulation often requires inferring task-relevant states from past interactions when the current observation alone is insufficient to determine the appropriate action. Despite progress in benchmarking memory-augmented vision-language-action (VLA) models, application-oriented tasks requiring history-dependent …

نسخة أولية وصول مفتوح

Balancing Multimodal Learning via Functional Progress

Zhongjing Gu, Fengqiang Wan, Yiming Cui وآخرون · 2026

Multimodal learning often suffers from modality imbalance, where the joint optimization process is dominated by a single modality. Existing methods typically estimate modality imbalance from score disparities derived from prediction uncertainty or optimization statistics. However, due to distinct prediction uncertainty …

نسخة أولية وصول مفتوح

Consensus-Aware Multi-Source Fusion for Reference-Guided Camouflaged Object Detection

Junyang Xia, Luocheng Zhang, Wenwen Pan وآخرون · 2026

Reference-guided camouflaged object detection aims to segment a target whose visual appearance closely resembles its surroundings by exploiting auxiliary reference samples. The task remains difficult because reference samples contain inconsistent target cues, while generic visual representations are not inherently alig …

نسخة أولية وصول مفتوح

Spike-driven Vision-Language-Action Model

Shuai Wang, Malu Zhang, Mingquan Liu وآخرون · 2026

Vision-language-action (VLA) models bridge multimodal understanding and robotic control, advancing the dominant paradigm for embodied intelligence. However, most existing models rely on large Transformers, whose latency and energy costs hinder deployment on resource-constrained platforms. Through sparse event-driven co …

نسخة أولية وصول مفتوح

Semantic-Aware Joint Source-Channel Optimization for Encoder-Agnostic Digital Video Communication

Xiangben Zhu, Caili Guo, Yang Yang وآخرون · 2026

Video semantic communication has attracted increasing attention as a promising approach to improving video transmission efficiency. However, most existing approaches rely on computationally intensive deep learning-based video encoders and decoders, which hinders their deployment in resource-constrained scenarios. To ad …

نسخة أولية وصول مفتوح

Merlin Plus: A Large-Scale, Multi-Cancer, Image-Mask-Report Dataset

Multi-cancer segmentation in computed tomography (CT) is fundamentally limited by the scarcity of tumor masks across different organs. We present Merlin Plus, the first large-scale CT dataset with radiologist-created tumor masks across 9 organs. Merlin Plus extends the Merlin dataset by adding 1,153 per-voxel tumor mas …

نسخة أولية وصول مفتوح

KlinikeBench: Evaluating Language Models Beyond Diagnostic Accuracy

Xueting Fang, Zehui Li, Yang Yang وآخرون · 2026

Most clinical benchmarks evaluate language models (LMs) on diagnosis using complete case descriptions. In clinical practice, however, patients present information in different ways, and clinicians must obtain relevant history and determine which examinations are needed before reaching a diagnosis. Diagnostic accuracy a …

نسخة أولية وصول مفتوح

Beyond Geometry: Benchmarking and Consistency Reasoning for 3D Logical Anomaly Detection

Zhiqiang Qin, He Xie, Junfei Yi وآخرون · 2026

Existing 3D industrial anomaly detection mainly targets local geometric deviations. In contrast, many industrial anomalies violate object-level design or assembly rules, which we define as 3D logical anomalies. To address these challenges, we introduce the Industrial Logical Anomaly Detection Dataset (ILGAD), the first …

نسخة أولية وصول مفتوح

RT-Super: Learning Tumor Segmentation from Longitudinal Images and Reports

Pedro R. A. S. Bassi, Wenxuan Li, Hanxue Gu وآخرون · 2026

Multi-tumor segmentation is important for early cancer detection and allows radiologists to visualize, verify, and understand AI predictions. However, tumor segmentation masks are expensive, time-consuming, and unavailable for many tumor types in public data. Instead, hospitals have vast, readily available data that ca …

نسخة أولية وصول مفتوح

SAGE: A Statistical Acceptance Gate for Self-Evolving Agents

Yihao Wang, Linhan Xia, Rui Liu وآخرون · 2026

Large Language Model (LLM)-based agents increasingly self-evolve by editing a persistent skill document that encodes their workflow, tool-use rules, and decision logic. This loop has two steps, an optimizer that proposes a candidate edit and a gate that accepts or rejects it. Prior work has concentrated on the optimize …

نسخة أولية وصول مفتوح

MoR-MLLM: Mixture of Recursions for Efficient Multimodal Large Language Models

Pengcheng Zheng, Chaoning Zhang, Jiaxin Yan وآخرون · 2026

Multimodal Large Language Models (MLLMs) have demonstrated remarkable reasoning capabilities across vision and language tasks. However, their massive computational and memory demands hinder real-world deployment. While recent efforts reduce costs by employing lightweight language backbones, existing paradigms remain co …

نسخة أولية وصول مفتوح

Learning Through Game: Skewed Transfer of Tabular Knowledge to Strengthen Image Model

Multimodal tabular-image learning is gaining growing attention, yet it faces challenges due to tabular data unavailable at test time. A practical solution involves transferring tabular knowledge to images during training to enhance the performance of image models at inference. However, the overlooked yet important chal …

نسخة أولية وصول مفتوح

SelfCue: Making a 3D CT Report Generator Say What It Already Knows

Renjie Liang, Yang Yang, Jinqian Pan وآخرون · 2026

Progress in 3D CT report generation is usually sought in increasingly sophisticated architectures and larger pools of training data. We find instead that a 3D CT report generator already holds what its report leaves out, and loses it when the hidden state becomes tokens. Over the 18 CT-RATE abnormalities, this hidden-t …

نسخة أولية وصول مفتوح

CommitFlow: Semantic Commitment Verification and Local Correction for Long-Horizon Robot Manipulation VLA Execution

Zixiang Zhao, Yansong Feng, Yang Yang وآخرون · 2026

Although vision-language-action (VLA) policies have advanced rapidly, long-horizon execution may still progress to the next task stage before the required physical effect has been established. We call this a mismatch between semantic commitments, physical conditions that a stage must establish or maintain, and the actu …

المؤلفون المشاركون