الباحثون

Saugat Adhikari

المنشورات 2

نسخة أولية وصول مفتوح

Foresight: planning future perception in streaming VLMs without retraining

Existing streaming vision-language models (VLMs) continuously perceive and reason over visual streams, but their computational pathways remain fixed throughout inference. Consequently, they cannot adapt computation to evolving scene dynamics, where different future events demand different levels and forms of perception …

نسخة أولية وصول مفتوح

iADD: Improving Alignment and Diversity in Diffusion Policy Optimization

Reinforcement learning based post training of diffusion models, such as Denoising Diffusion Policy Optimization (DDPO), optimizes a reverse diffusion process under a reward function. However, current approaches to reward optimizations do so at the cost of diversity and quality. In this paper, we provide better tradeoff …

المؤلفون المشاركون