الباحثون

Shun Zhang

المنشورات 8

نسخة أولية وصول مفتوح

Conversational Voice Aesthetic Model with Reinforcement Learning from Human Listeners

Xilin Jiang, Shun Zhang, Tejas Jayashankar وآخرون · 2026

We introduce Conversational Voice Aesthetic Model, a speech large language model for describing the voice aesthetics of real or synthetic speech responses in natural conversational contexts. Given a context and a response speech, CVAM describes salient moments that characterize the voice and predicts nine categorical a …

نسخة أولية وصول مفتوح

FOSLS-deRhaNN: native de Rham neural classes for H(div) and H(curl) with applications to first-order system least-squares neural network methods for partial differential equations

Shun Zhang · 2026

We construct neural approximation classes native to the graph spaces H(div) and H(curl), in two and three dimensions and, for H(div), in any dimension. Every realization lies in the space for all parameter values, and with kinked potentials, such as ReLU networks, the admissible jumps appear at finite width. The classe …

نسخة أولية وصول مفتوح

When Updating Stops Being Learning: Rethinking LLM Self-Evolution via learnable information gain

Chenxu Wang, Chaozhuo Li, Xinze Shi وآخرون · 2026

Self-evolution lets large language models (LLMs) improve iteratively using their own generated data, but often suffers from self-evolution degeneration: performance improves, plateaus, then declines. Existing methods address this issue at the component level, targeting either the Questioner or the Solver, and overlook …

نسخة أولية وصول مفتوح

Spoken Language Models that Think Aloud

Junyi Ao, Kainan Peng, Mingbo Ma وآخرون · 2026

While Chain-of-Thought (CoT) reasoning has improved the capability of language models, directly applying it to Spoken Language Models (SLMs) may introduce long silent intervals under the serial "think-then-speak" paradigm, disrupting real-time spoken interaction. To address this issue, we propose an asynchronous think- …

نسخة أولية وصول مفتوح

CIBuzzBench: A Benchmark for Cross-Lingual Understanding of Chinese Internet Buzzwords

Yifan Wang, Junyu Lu, Qifan Wang وآخرون · 2026

Chinese social media has generated a vast and continually evolving lexicon of internet buzzwords whose meanings are often non-literal and deeply rooted in local cultural and pragmatic contexts. Existing research has primarily focused on interpreting these buzzwords within Chinese, leaving largely unexplored whether LLM …

نسخة أولية وصول مفتوح

TeleAntiFraud 2.0: A Refreshable, Profile-Grounded, and Audio-Based Benchmark for Telecom Fraud Detection

Huiyuan Liu, Zhiming Ma, Yanxing Liu وآخرون · 2026

Telecom fraud scripts evolve rapidly and are often designed to resemble routine service conversations, creating two key requirements for audio-based telecom-fraud evaluation. First, benchmarks must incorporate newly observed scam patterns without overwriting previously established test sets. Second, they must distingui …

نسخة أولية وصول مفتوح

FRAUDSkill: Structured Frozen-Weight Skill Optimization for Audio Anti-Fraud Detection

Chengxian Hu, Zhiming Ma, Mingjun Pan وآخرون · 2026

Large audio-language models have shown promise for anti-fraud detection by directly processing speech and reasoning over fraud-related evidence. Their deployment, however, requires predictions to follow a predefined label space and a structured decision protocol consisting of service-scenario identification, fraud dete …

نسخة أولية وصول مفتوح

RiskChainBench: A Benchmark for Obfuscated Platform Message Restoration and Evidence-Grounded Web Investigation

ZhuoXin Liu, Zhiming Ma, Ying Zhang وآخرون · 2026

Platform abuse campaigns conceal redirection instructions with emojis, homophones, character decomposition, and redundant symbols, then route users through disguised links to services associated with pornography, fraud, gambling, or illicit transactions. Existing benchmarks evaluate obfuscated text and risky webpages s …

المؤلفون المشاركون