Abstract
In this work, we show that a single Transformer block, applied recurrently, can match the accuracy of a full-depth vision encoder at comparable inference FLOPs without intermediate feature distillation. reViT restores depth-specific transformations by representing the FFN at each recurrent depth as a convex combination of a small shared expert bank. A continuous normalized-depth coordinate programs this mixture, defining a resampleable trajectory through FFN parameter space. We evaluate this design in two regimes: supervised ImageNet-1k training and distillation from a DINOv2 teacher. Across both regimes, controlled adaptations identify weight-space merging as the strongest tested MoE family at a matching one-FFN budget, ahead of the token-dispatch and output-mixture alternatives. Trained from scratch, reViT-B/16 attains DeiT III accuracy with about 70\% fewer stored parameters. An 8-experts model distilled using only the teacher's output features retains nearly all of its DINOv2 teacher's linear-probe accuracy and transfers across classification, segmentation, and depth prediction. Elastic-depth training allows one checkpoint (trained model) to operate at multiple tested depths by resampling the same normalized coordinate interval. For fixed-depth deployment, the recurrent block can be materialized as a conventional dense graph, removing online routing and merging without changing the one-FFN-per-depth compute but expanding deployment storage.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Bulat, A., Ouali, Y., & Tzimiropoulos, G. (2026). One Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed Experts. https://omanscience.com/en/articles/one-block-multiple-depths-recurrent-vision-transformers-with-depth-programmed-experts
MLA 9
Bulat, Adrian, et al. "One Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed Experts." https://omanscience.com/en/articles/one-block-multiple-depths-recurrent-vision-transformers-with-depth-programmed-experts.
Chicago (author–date)
Bulat, Adrian, Yassine Ouali, and Georgios Tzimiropoulos. 2026. "One Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed Experts." https://omanscience.com/en/articles/one-block-multiple-depths-recurrent-vision-transformers-with-depth-programmed-experts.
Harvard
Bulat, A., Ouali, Y. and Tzimiropoulos, G. (2026) 'One Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed Experts', Available at: https://omanscience.com/en/articles/one-block-multiple-depths-recurrent-vision-transformers-with-depth-programmed-experts.
Vancouver
Bulat A, Ouali Y, Tzimiropoulos G. One Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed Experts. https://omanscience.com/en/articles/one-block-multiple-depths-recurrent-vision-transformers-with-depth-programmed-experts
IEEE
A. Bulat, Y. Ouali, and G. Tzimiropoulos, "One Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed Experts," https://omanscience.com/en/articles/one-block-multiple-depths-recurrent-vision-transformers-with-depth-programmed-experts.