الباحثون

Ziming Liu

المنشورات 4

نسخة أولية وصول مفتوح

Emergent Inverse-Depth Scaling From Nonlinearity In Attention

Zirui Peng, Yizhou Liu, Ziming Liu وآخرون · 2026

Scaling laws describe power-law improvements in model performance with dataset size and parameter count, yet their underlying mechanisms are not fully understood. To explain the parameter count scaling, existing theory posits power-law scaling with model depth. In linear-attention models, this scaling is tied to a powe …

نسخة أولية وصول مفتوح

Forking: Sudden Overfitting Under Replay

Shanbin Yu, Shaoyang Guo, Haoran Zhao وآخرون · 2026

This paper studies forking, a generalization failure discovered in NanoGPT autoresearch. Under data replay, models with an over-encoding n-gram memory branch show a sharp separation of training and validation loss at epoch boundaries, resembling the shape of forks. We study this phenomenon in a controlled vanilla NanoG …

نسخة أولية وصول مفتوح

Why Do Conventional World Models Fail to Learn Cellular Automata?

Although conventional world models - auto-regressive or diffusion models based on transformers or convolutional networks - may learn surface statistics of world dynamics, can they learn the exact world dynamics from its observed history? Leveraging cellular automata as a simple testbed, we find the answer to be no in m …

نسخة أولية وصول مفتوح

ArchitectureIQ: On the Measure of Training Intuition

Zirui Ren, Shaoyang Guo, Chencheng Tang وآخرون · 2026

Top researchers have good intuition, but do language models have as good intuition about model training as top AI researchers? To measure model intuition of LLMs and humans, we introduce the ArchitectureIQ benchmark. Each question presents a synthetic dataset and several training recipes, and the test-taker is asked to …

المؤلفون المشاركون