Preprint Open access
DVD: Dynamic Vector Decoding for Efficient MLLM-based Perception
Multimodal large language models have made remarkable progress in bridging vision and language, facilitating various perception tasks essential for human-machine interaction, robotics, and autonomous driving. However, existing MLLM-based perception methods predominantly rely on text-based coordinate representation, whi …