الباحثون

Sheng Li

المنشورات 4

نسخة أولية وصول مفتوح

Beyond Speech Captions: Speech-Rewarded Style Planning for Conversational Text-to-Speech

Shiao Zhu, Lianbo Liu, Sizhen Lyu وآخرون · 2026

Natural-language style descriptions provide an interpretable interface between large language models (LLMs) and controllable text-to-speech (TTS). However, using descriptions as pseudo-labels compresses target acoustics into text, and descriptive fidelity need not imply effective control of a particular synthesizer. We …

نسخة أولية وصول مفتوح

PixelDense: Dense Prediction as Representation Alignment for Pixel Diffusion

Lehan Yang, Daiqing Qi, Wenhao Zhang وآخرون · 2026

Representation alignment (REPA) accelerates diffusion transformer training, but its alignment targets are almost exclusively semantic encoders such as DINOv2 and CLIP. Recent analysis points to spatial structure, not global semantics, as the carrier of the alignment effect, yet dense-prediction foundation models traine …

نسخة أولية وصول مفتوح

Contamination, Prior, or Evidence? Decomposing and Training Evidence Use in Whole-Slide Vision-Language Models

Wenhao Zhang, Zhongliang Zhou, Shiyuan Zhang وآخرون · 2026

Pathology vision-language models (VLMs) are conventionally evaluated by accuracy, but accuracy alone does not measure evidence use: it may conflate dataset contamination, prior knowledge, and image evidence. In a motivating study of lymph-node metastasis prediction, we found that most public pathology VLMs showed minim …

المؤلفون المشاركون