Preprint Open access
Continuous Conditioning of VLAs with Augmenting EMG and Visual Task Descriptors
Vision-Language-Action (VLA) models rely strongly on language for describing task information, despite having multimodal inputs. We hypothesize that other modalities in the state space may present opportunities for supplemental task conditioning, which may be particularly relevant in cluttered or otherwise ambiguous sc …