Preprint Open access
Linear attention is a tractable model for understanding the mechanisms governing in-context learning in transformers. For linear regression tasks, recent asymptotic analyses have characterised its learning and generalisation behaviour. We extend this theory to nonlinear single-index targets, $y=f(x^\top w)+\varepsilon …
Preprint Open access
Recursive feature machines (RFMs) learn representations of data by alternating between fitting a predictor to a dataset and updating features of that predictor using the average gradient outer product (AGOP). Connections between AGOPs and feature learning in neural networks motivate linear RFMs as a simple setting for …
Preprint Open access
Approximate Message Passing (AMP) algorithms are attractive as they are computationally efficient and simultaneously admit a precise characterization in terms of the low-dimensional "state-evolution" recursion. In this work, we establish quantitative universality for AMP with centered rank-one sensing matrices $Z_i=(x_ …