Preprint Open access
Quantum processing of classical data fundamentally relies on efficiently mapping classical datasets onto the amplitudes of quantum states. Preparing large amplitude-encoded states, however, remains a major bottleneck in quantum computing. In this paper, we introduce a novel state-preparation framework we call Q-MERGE, …
Preprint Open access
Vision-language-action models benefit from the understanding and reasoning capabilities of pretrained vision-language models, but action-only supervision provides limited grounding in world dynamics. Conversely, world-action models inherit spatiotemporal priors from video generation models, yet remain limited in semant …
Preprint Open access
Efficient action generation in vision-language-action (VLA) models requires capturing both coarse action structure and fine-grained details. Discrete action tokens provide compact structural representations but sacrifice precision, while continuous action tokens offer high precision but often require multiple denoising …
Preprint Open access
Human motion is shaped by both external acoustic events and behavioral intent: spatial audio conveys environmental cues that elicit or guide a response, while text specifies the desired action and how it should be performed. In this paper, we study the novel task of human motion synthesis jointly conditioned on spatial …
Preprint Open access
Humanoid robots can step over, squeeze past, and duck under obstacles, but learning to select and coordinate these behaviors from onboard perception remains challenging. Many existing approaches rely on task-specific reinforcement-learning objectives or curated motion libraries, making broad behavioral coverage costly. …