الباحثون

Jakob Foerster

المنشورات 2

نسخة أولية وصول مفتوح

Q-Shaped Options for Hierarchical Reinforcement Learning

Learning to tackle long-horizon, goal-conditioned tasks requires an agent to reason over extended timescales and act across a broad range of states. In principle, Hierarchical Reinforcement Learning (HRL) addresses both challenges through the interaction between action (temporal) and state (spatial) abstraction. First, …

نسخة أولية وصول مفتوح

GitSwarm: Decentralized Compounding Inference

Vedant Shah, Ankur Samanta, Paras Dahal وآخرون · 2026

Long-horizon problem solving and scientific research require computation to accumulate across successive attempts. Partial solutions, experimental findings, and unsuccessful approaches can inform later work, yet most inference-time computation is organized around individual trajectories or candidates rather than a pers …

المؤلفون المشاركون