Abstract

End-to-end backpropagation has been the dominant mode of training in deep learning, allowing for the coordination of parameter updates across layers of a neural network. Prior studies have explored alternative -- and, in some cases, simpler -- training mechanisms, showing that they can sometimes achieve performance similar to backpropagation. However, the architectural conditions under which locally optimized networks, which avoid end-to-end backpropagation of error, can learn representations comparable to those learned through end-to-end training remain unclear. We aim to answer this question in the context of self-supervised learning, an important framework for large-scale pretraining in artificial intelligence. Here, we investigate how network width and depth affect the efficacy of greedy layer-wise and end-to-end self-supervised training in convolutional networks. We find that in wider networks, the benefits of end-to-end backpropagation over greedy layer-wise training shrink: in relatively shallow and very wide networks, we even observed higher performance in models trained with greedy layer-wise training. Subsequent analysis of the representations formed by these networks shows that very wide greedy-trained networks exhibit more favorable representational geometry than do networks trained end-to-end with backpropagation. This work shows that width can compensate for restricted credit assignment and identifies differences in representational geometry as a potential mechanism for their improved performance.

Keywords

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Mansur, S., & Zylberberg, J. (2026). Increasing Width Allows Greedy Layer-wise Training to Rival End-to-End Backpropagation in Self-Supervised Learning. https://omanscience.com/en/articles/increasing-width-allows-greedy-layer-wise-training-to-rival-end-to-end-backpropagation-in-self-supervised-learning

MLA 9

Mansur, Syon, and Joel Zylberberg. "Increasing Width Allows Greedy Layer-wise Training to Rival End-to-End Backpropagation in Self-Supervised Learning." https://omanscience.com/en/articles/increasing-width-allows-greedy-layer-wise-training-to-rival-end-to-end-backpropagation-in-self-supervised-learning.

Chicago (author–date)

Mansur, Syon, and Joel Zylberberg. 2026. "Increasing Width Allows Greedy Layer-wise Training to Rival End-to-End Backpropagation in Self-Supervised Learning." https://omanscience.com/en/articles/increasing-width-allows-greedy-layer-wise-training-to-rival-end-to-end-backpropagation-in-self-supervised-learning.

Harvard

Mansur, S. and Zylberberg, J. (2026) 'Increasing Width Allows Greedy Layer-wise Training to Rival End-to-End Backpropagation in Self-Supervised Learning', Available at: https://omanscience.com/en/articles/increasing-width-allows-greedy-layer-wise-training-to-rival-end-to-end-backpropagation-in-self-supervised-learning.

Vancouver

Mansur S, Zylberberg J. Increasing Width Allows Greedy Layer-wise Training to Rival End-to-End Backpropagation in Self-Supervised Learning. https://omanscience.com/en/articles/increasing-width-allows-greedy-layer-wise-training-to-rival-end-to-end-backpropagation-in-self-supervised-learning

IEEE

S. Mansur, and J. Zylberberg, "Increasing Width Allows Greedy Layer-wise Training to Rival End-to-End Backpropagation in Self-Supervised Learning," https://omanscience.com/en/articles/increasing-width-allows-greedy-layer-wise-training-to-rival-end-to-end-backpropagation-in-self-supervised-learning.