الباحثون

Rasul Tutunov

المنشورات 4

نسخة أولية وصول مفتوح

Memento 3: Model-Based Recursive Self-Improvement through Reflective Rulebooks

Haoyu Zhao, Zhengxu Yu, Zhiyuan He وآخرون · 2026

Learning to act in unfamiliar environments requires agents to infer how the world works and revise that understanding as new evidence arrives. Yet limited observations can support multiple world models that explain past interactions but predict different outcomes in unseen states. We introduce Memento 3, building on th …

نسخة أولية وصول مفتوح

Tail-Influence Sampling for CVaR Policy Evaluation

Policies with similar mean returns can differ sharply in rare failures, yet estimating lower-tail conditional value-at-risk (CVaR) accurately can require many costly rollouts. When different conditional components of a stochastic workflow can be queried separately, we ask how to allocate a fixed evaluation budget to es …

نسخة أولية وصول مفتوح

Composable Decoding on the Probability Simplex: Theory and Implementation

Decoding for large language models is typically treated as a collection of isolated sampling strategies, with limited theoretical understanding of the behaviours they induce and how their underlying objectives relate. We formulate decoding as an optimisation problem over next-token distributions on the probability simp …

المؤلفون المشاركون