نسخة أولية وصول مفتوح
Large Language Models (LLMs) increasingly ship with explicit "thinking modes", yet their counterpart, "no-thinking", has received far less attention. We study LLMs' no-thinking behavior along two axes. a. How to measure no-thinking? Prior work typically defines no-thinking through proxies such as a disabled thinking mo …
نسخة أولية وصول مفتوح
Assistance from AI tools has supported and improved human performance across domains. However, recent research suggests that these immediate benefits may entail future costs, including diminished performance when AI assistance is no longer available. We study how human-AI interaction behaviors correlate with immediate …
نسخة أولية وصول مفتوح
Scientific figures often encode quantitative results that are not readily available in machine-readable form, making accurate plot digitization important for verifying and reusing published findings. Yet it remains unclear how accurately current models recover plotted values from real scientific figures, as existing be …
نسخة أولية وصول مفتوح
Autoresearch agents are reshaping the research ecosystem, but they can also let flawed claims enter the literature at scale. Human advisors catch such issues in drafts through careful, traceable feedback, yet advisor-style assessment requires extensive manual effort and does not scale. To shift automated paper assessme …