Authors

Ying Zhang

Publications 7

Preprint Open access

OceanMind: A multi-agent AI system for ocean diagnosis

Time-dependent, three-dimensional (3D) oceanic multi-variables define coherent states of the evolving ocean to facilitate ocean diagnosis and advance ocean science to better inform environmental and hazard management. However, extracting quantitative evidence from these variables requires substantial and complex analyt …

Preprint Open access

ScopeIF: Improving Scope-Aware Precise Instruction-Following in Large Language Models via Graded Reward Modeling

Bosi Wen, Yilin Niu, Xiaoying Ning et al. · 2026

Precise instruction-following is a fundamental ability of large language models (LLMs), requiring their outputs to strictly satisfy objective constraints in input instructions. In complex application scenarios, these constraints often possess diverse scopes that govern specific response segments rather than the entire …

Preprint Open access

Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures

Ruoqi Guo, Yi Liu, Gelei Deng et al. · 2026

Detectors of alignment failures screen deployed language models and score alignment benchmarks. Most are generative judges that spend a decoding pass on every criterion, and classifiers that read token probabilities, such as Llama Guard, still score one fixed label per call. Jev, a model trained with reinforcement lear …

Preprint Open access

HYDRA: Proactive Android Malware Drift Adaptation via Hierarchical Graph Contrastive Learning

Han Chen, Hanchen Wang, Hongmei Chen et al. · 2026 · 10.1145/3830454.3832684

Concept drift, driven by the rapid evolution of Android malware, severely degrades the performance of machine learning detectors. Current adaptation strategies are often reactive, responding only after performance has dropped and imposing a significant manual annotation burden, or they are proactive but rely on unstabl …

Preprint Open access

Preserving What Matters: Semantic Scaffolds Beyond Saturation in Summarization Evaluation

Summarization ships in countless production systems, making model selection a routine decision that depends on measuring summary quality. Existing metrics struggle to support this: ROUGE captures only surface overlap, while LLM-as-judge scores saturate to near-identical values that fail to rank models effectively. We o …

Preprint Open access

RiskChainBench: A Benchmark for Obfuscated Platform Message Restoration and Evidence-Grounded Web Investigation

ZhuoXin Liu, Zhiming Ma, Ying Zhang et al. · 2026

Platform abuse campaigns conceal redirection instructions with emojis, homophones, character decomposition, and redundant symbols, then route users through disguised links to services associated with pornography, fraud, gambling, or illicit transactions. Existing benchmarks evaluate obfuscated text and risky webpages s …

Co-authors