Abstract

Pixel-space diffusion models avoid the lossy VAE of latent models, which suggests an advantage on downstream tasks where fine-grained detail matters. We test this claim along both routes to a pixel-space backbone. We pretrain Iris-3B, a 3B-parameter pixel-space text-to-image transformer, from scratch through a $256\to512\to1024$ curriculum, after first ablating the prediction target and representation alignment at $256^2$ to decide what to scale. We also convert a pretrained latent model, FLUX.2 Klein base 4B, to pixel space. We fine-tune both families for monocular depth estimation and for image restoration/super-resolution. We find no significant improvement from using a pixel-space generative prior. Fine-tuned for depth with one matched direct-regression recipe, Iris-3B is level with the latent FLUX.2 Klein and the converted pixel FLUX.2 Klein falls behind it, and on $4\times$ DIV2K restoration neither pixel model beats a latent FLUX.2 Klein fine-tune, the converted one trailing it slightly. We document the recipes, the failure modes and the remaining confounds behind this negative result. Nevertheless, Iris-3B shows that pixel-space pretraining with the pixel-transformer (PiT) head of PixelDiT scales to 3B parameters and to text-to-image quality competitive with latent models, matching Qwen-Image on OneIG under the official evaluators at $1024^2$. We release its weights and training code in the hope that they help pave the way for further work on pixel-space generation.

Keywords

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Cai, H. L., & Garabito, C. (2026). Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning. https://omanscience.com/en/articles/iris-3b-going-beyond-the-latent-with-pixel-space-diffusion-training-conversion-and-fine-tuning

MLA 9

Cai, Hanqiu Li, and Chema Garabito. "Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning." https://omanscience.com/en/articles/iris-3b-going-beyond-the-latent-with-pixel-space-diffusion-training-conversion-and-fine-tuning.

Chicago (author–date)

Cai, Hanqiu Li, and Chema Garabito. 2026. "Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning." https://omanscience.com/en/articles/iris-3b-going-beyond-the-latent-with-pixel-space-diffusion-training-conversion-and-fine-tuning.

Harvard

Cai, H. L. and Garabito, C. (2026) 'Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning', Available at: https://omanscience.com/en/articles/iris-3b-going-beyond-the-latent-with-pixel-space-diffusion-training-conversion-and-fine-tuning.

Vancouver

Cai HL, Garabito C. Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning. https://omanscience.com/en/articles/iris-3b-going-beyond-the-latent-with-pixel-space-diffusion-training-conversion-and-fine-tuning

IEEE

H. L. Cai, and C. Garabito, "Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning," https://omanscience.com/en/articles/iris-3b-going-beyond-the-latent-with-pixel-space-diffusion-training-conversion-and-fine-tuning.