Authors

Xiyang Hu

Publications 6

Preprint Open access

AI4Fire: Evaluating Large Language Models on Wildfire Tasks

Yue Zhao, Xiyang Hu, Zuobin Xiong et al. · 2026

Large language models (LLMs) are entering wildfire management, where overstated evaluations can cost property and lives. How do they perform on wildfire tasks, with and without grounding? Bare means a model receives the task input alone. Grounded means it also receives one task-specific addition: for smoke detection, a …

Preprint Open access

Auditable Claims about AI Agents

Yue Zhao, Jiate Li, Li Li et al. · 2026

Organizations make claims about their AI agents: a person approves every external email, every action is logged, an evaluation shows the agent is safe to deploy. Article 12 of the EU AI Act requires high-risk systems to allow the automatic recording of events but does not say which records settle a given claim. The pos …

Preprint Open access

SyncRA: Learning Temporal Correspondence in Omni-Modal Models

Zelong Xu, Yan Li, Wenhe Hu et al. · 2026

Recent omni-modal models demonstrate strong perception of audio and visual inputs, yet often struggle to connect what they hear with what they see at the same moment. This weakness in temporal correspondence can cause models to associate spoken cues with the wrong visual scenes, producing plausible answers grounded in …

Preprint Open access

Joint and Cross-Modal Video-Audio Generation and Editing: A Unified Formulation and Design Taxonomy

Video and audio are perceived together, yet most generative models treat them in isolation. We examine methods that model the two modalities jointly, generate one from the other, or edit them in a coupled manner, organized around a single question: how is the output kept coherent across modalities in time and semantics …

Preprint Open access

JevOut: Natural Context Can Flip Decision Models

Zixiang Xu, Zirui Song, Chiyu Zhang et al. · 2026

An ordinary-looking background detail can turn a correct model decision into a confident mistake. We demonstrate this fragility in four decision systems, including Jev, across seven datasets covering knowledge, reasoning, and tool routing. Within 64 accepted target evaluations per decision, we uncover short context add …

Co-authors