الباحثون

Philip Torr

المنشورات 7

نسخة أولية وصول مفتوح

Thinking Inertia: LLMs Keep Thinking When Told Not To

Dianqiao Lei, Kevin Qinghong Lin, Pan Lu وآخرون · 2026

Large Language Models (LLMs) increasingly ship with explicit "thinking modes", yet their counterpart, "no-thinking", has received far less attention. We study LLMs' no-thinking behavior along two axes. a. How to measure no-thinking? Prior work typically defines no-thinking through proxies such as a disabled thinking mo …

نسخة أولية وصول مفتوح

BazaarBench: Delegation Safety in Decentralized C2C Marketplaces Run by LLM Agents

Ziyan Wang, Shuqing Shi, James Oldfield وآخرون · 2026

In decentralized consumer-to-consumer (C2C) marketplaces, people list goods, negotiate with strangers, and rate one another, so trust rests on reputation. Large language model (LLM) agents now act for users, raising risks to their money, privacy, and reputation. We introduce BazaarBench, a simulated C2C marketplace and …

نسخة أولية وصول مفتوح

Towards Better Exploration in Sequential Test-Time Scaling

Joseph Rance, Fabio Pizzati, Juil Sock وآخرون · 2026

Test-time scaling improves language model reasoning by spending additional compute at inference. However, both classes of existing methods often fail to continue improving over long timescales. Parallel methods repeatedly sample independent answers from the model, scaling poorly on problems the model is unlikely to sol …

نسخة أولية وصول مفتوح

Training Object Permanence in World Models

Haotian Zhang, Fengyuan Yu, Dezhi Luo وآخرون · 2026

Object permanence and solidity are hallmarks of human cognitive priors. Recent studies show that video generation models, a paradigmatic class of current world models, have begun to show emerged reasoning abilities, making them ideal candidates for building human-like physical intelligence. Do video models have emerged …

نسخة أولية وصول مفتوح

Beyond Visual Quality: A Study of Test-Time Planning with World Action Models

Jianhao Yuan, Yu Yuan, Benjamin Ramtoula وآخرون · 2026

World action models generate actions together with visual predictions of their consequences. These paired outputs create the potential for planning by sampling multiple actions from one state, comparing their imagined outcomes, and choosing the action with the most promising predicted outcome. However, how to use imagi …

نسخة أولية وصول مفتوح

PaperDoctor: Evidence-Grounded and Actionable Feedback for Scientific Papers in Progress

Kevin Qinghong Lin, Siyuan Hu, Pan Lu وآخرون · 2026

Autoresearch agents are reshaping the research ecosystem, but they can also let flawed claims enter the literature at scale. Human advisors catch such issues in drafts through careful, traceable feedback, yet advisor-style assessment requires extensive manual effort and does not scale. To shift automated paper assessme …

المؤلفون المشاركون