Authors

Ziming Liu

Publications 4

Preprint Open access

Emergent Inverse-Depth Scaling From Nonlinearity In Attention

Zirui Peng, Yizhou Liu, Ziming Liu et al. · 2026

Scaling laws describe power-law improvements in model performance with dataset size and parameter count, yet their underlying mechanisms are not fully understood. To explain the parameter count scaling, existing theory posits power-law scaling with model depth. In linear-attention models, this scaling is tied to a powe …

Preprint Open access

Forking: Sudden Overfitting Under Replay

This paper studies forking, a generalization failure discovered in NanoGPT autoresearch. Under data replay, models with an over-encoding n-gram memory branch show a sharp separation of training and validation loss at epoch boundaries, resembling the shape of forks. We study this phenomenon in a controlled vanilla NanoG …

Preprint Open access

Why Do Conventional World Models Fail to Learn Cellular Automata?

Although conventional world models - auto-regressive or diffusion models based on transformers or convolutional networks - may learn surface statistics of world dynamics, can they learn the exact world dynamics from its observed history? Leveraging cellular automata as a simple testbed, we find the answer to be no in m …

Preprint Open access

ArchitectureIQ: On the Measure of Training Intuition

Top researchers have good intuition, but do language models have as good intuition about model training as top AI researchers? To measure model intuition of LLMs and humans, we introduce the ArchitectureIQ benchmark. Each question presents a synthetic dataset and several training recipes, and the test-taker is asked to …

Co-authors