نسخة أولية وصول مفتوح
Juno: Taming Predictive Latents for Vision-Language-Action Models
Joint-embedding predictive architectures (JEPAs) predict masked or future observations in representation space, offering a natural source of predictive latents for vision-language-action (VLA) models. Yet making these latents useful across pretraining, policy learning, and deployment requires addressing three failures: …