الباحثون

هان ليو

المنشورات 4

نسخة أولية وصول مفتوح

DiMOS: Doob-Guided Inference-Time Multi-Objective Search for Scientific Design

Scientific design often requires jointly satisfying multiple objectives and constraints. Pretrained masked diffusion models provide a generative foundation for this task, but fine-tuning them to meet these objectives and constraints incurs additional training costs, motivating inference-time guidance with frozen models …

نسخة أولية وصول مفتوح

D-DOIT: Training-free Adaptation of Discrete Diffusion via Doob's h-Transform

جيكي وو, قجي زو, ويمين وو وآخرون · 2026

We propose D-DOIT (Discrete Doob-Oriented Inference-time Transformation), a training-free and efficient adaptation method for discrete diffusion models with generic rewards. D-DOIT formulates adaptation as sampling from a reward-tilted target distribution and realizes this transport through Doob's h-transform of the di …

نسخة أولية وصول مفتوح

ProgressCompass: Embodied Progress Reward Models Are Lost Without the Right Context

Embodied agents now take on ever longer tasks. For long tasks, knowing only whether a task finally succeeds or fails says little; the steps along the way matter. Progress Reward Models (PRMs) score how far a task has come at every step, and serve as dense rewards, verifiers and monitors. Yet in long tasks the current f …

نسخة أولية وصول مفتوح

Video2Skill: From Streaming Experience to Reusable Embodied Skills

Manipulation behaviors vary widely across objects and scenes, but they share a small set of reusable skills, and planning with these skills helps embodied agents generalize to new tasks. Yet an agent can only plan with skills it knows. Recovering skills from observed experience, the inverse of planning, builds this kno …

المؤلفون المشاركون