Authors

Wei Zhang

Publications 27

Preprint Open access

When AI Finds Hidden Messages, Does It Report?

When an assistant encounters a message for another AI, does it tell its user? Four fixed model-provider deployments perform simulated source tasks in 1,280 ordinary-note and 128 enhanced-note sessions. Harmless and harmful messages have matched plaintext and ROT13 versions, with no-message controls. Observers receive n …

Preprint Open access

Loopy: Low-Bit Quantization Framework for Looped Language Models

Zeyu LI, Yipu ZHANG, Jintao Chen et al. · 2026

Looped language models provide a parameter-efficient way to scale iterative test-time computation by repeatedly executing a shared recurrent core. Post-training quantization (PTQ) can reduce the memory footprint and inference cost of looped language models, but errors introduced by a quantized shared core affect subseq …

Preprint Open access

S$^3$N: A Spherical Spiral Scanning Network for Weather Forecasting

Fan Yan, Chen Hui, Weisi Lin et al. · 2026

Machine learning-based weather prediction (MLWP) has achieved strong performance in global weather forecasting. Recent Hierarchical Equal Area isoLatitude Pixelation (HEALPix)-based methods use the HEALPix (HP) grid to avoid area distortion near the poles of conventional latitude-longitude (LL) grids. However, existing …

Preprint Open access

Behavior Pack Optimization for Video MLLM Post-Training

Zhaolu Kang, Shiyu Liu, Tailong Luo et al. · 2026

Video multimodal large language models (MLLMs) keep climbing video question answering benchmarks, yet shuffling the frames, masking the segment that supports the answer, or occluding the target object barely changes their predictions. The accuracy rests on appearance and language priors, not on the temporal evidence th …

Preprint Open access

Octrees as an Explicit 3D Language

Ran Dan, Si-Tong Wei, Pengfei Xiong et al. · 2026

Existing 3D large language models (LLMs) compromise on two fronts: they compress shapes into latent codebook indices or coordinate text, which removes spatial structure from what the model observes, and they acquire the 3D modality by fine-tuning the backbone, which overwrites its general language ability. We present O …

Preprint Open access

Every Batch Is Its Own Validation Set: Leave-One-Out Gradient Matching for Online Data Selection in LLM Fine-Tuning

Hongyu Chen, Xinyi Luo, Ming Zhao et al. · 2026

Online batch selection fine-tunes a language model on the most useful part of each candidate batch. Selectors that match the gradient of the candidate batch are attractive because they need no held-out data, yet they rarely beat training on the whole batch. We show why. In-sample gradient matching uses every example as …

Co-authors