الباحثون

Shaoyang Guo

المنشورات 3

نسخة أولية وصول مفتوح

Forking: Sudden Overfitting Under Replay

Shanbin Yu, Shaoyang Guo, Haoran Zhao وآخرون · 2026

This paper studies forking, a generalization failure discovered in NanoGPT autoresearch. Under data replay, models with an over-encoding n-gram memory branch show a sharp separation of training and validation loss at epoch boundaries, resembling the shape of forks. We study this phenomenon in a controlled vanilla NanoG …

نسخة أولية وصول مفتوح

Why Do Conventional World Models Fail to Learn Cellular Automata?

Although conventional world models - auto-regressive or diffusion models based on transformers or convolutional networks - may learn surface statistics of world dynamics, can they learn the exact world dynamics from its observed history? Leveraging cellular automata as a simple testbed, we find the answer to be no in m …

نسخة أولية وصول مفتوح

ArchitectureIQ: On the Measure of Training Intuition

Zirui Ren, Shaoyang Guo, Chencheng Tang وآخرون · 2026

Top researchers have good intuition, but do language models have as good intuition about model training as top AI researchers? To measure model intuition of LLMs and humans, we introduce the ArchitectureIQ benchmark. Each question presents a synthetic dataset and several training recipes, and the test-taker is asked to …

المؤلفون المشاركون