نسخة أولية وصول مفتوح
Vela: Scaling Vision-Language-Action Models with Adaptive Action Curve Parametrization
Most vision-language-action models represent future motion as fixed-rate action chunks, tying temporal resolution and prediction horizon to a fixed output budget. This pointwise representation wastes capacity on highly correlated neighboring actions, leaves temporal continuity and smoothness to be learned implicitly, a …