Authors

Yue Zhao

Publications 7

Preprint Open access

AI4Fire: Evaluating Large Language Models on Wildfire Tasks

Yue Zhao, Xiyang Hu, Zuobin Xiong et al. · 2026

Large language models (LLMs) are entering wildfire management, where overstated evaluations can cost property and lives. How do they perform on wildfire tasks, with and without grounding? Bare means a model receives the task input alone. Grounded means it also receives one task-specific addition: for smoke detection, a …

Preprint Open access

Auditable Claims about AI Agents

Yue Zhao, Jiate Li, Li Li et al. · 2026

Organizations make claims about their AI agents: a person approves every external email, every action is logged, an evaluation shows the agent is safe to deploy. Article 12 of the EU AI Act requires high-risk systems to allow the automatic recording of events but does not say which records settle a given claim. The pos …

Preprint Open access

Controlled Decoding Attacks on Black-Box LLMs

Jesson Wang, Shawn Li, Wei Yang et al. · 2026

Manipulating next-token probabilities during generation can bypass the safety alignment of large language models. Existing approaches, however, rely on access to model weights or numerical token probabilities and therefore do not apply to interfaces that return only sampled text. Reconstructing probabilities from sampl …

Preprint Open access

PhoenixSR: Generative Heterogeneous Distillation Unleashes Efficient Models for Real-World Super-Resolution

Xin Di, Mingyu Shi, Yuanfei Bao et al. · 2026

Real-world image super-resolution (SR) requires recovering perceptually realistic high-resolution images from complex low-resolution observations while preserving faithful content. Diffusion-based SR benefits from strong generative priors but incurs substantial computational overhead, whereas feed-forward CNN and Trans …

Preprint Open access

JevOut: Natural Context Can Flip Decision Models

Zixiang Xu, Zirui Song, Chiyu Zhang et al. · 2026

An ordinary-looking background detail can turn a correct model decision into a confident mistake. We demonstrate this fragility in four decision systems, including Jev, across seven datasets covering knowledge, reasoning, and tool routing. Within 64 accepted target evaluations per decision, we uncover short context add …

Co-authors