الباحثون

Tomas Pfister

المنشورات 7

نسخة أولية وصول مفتوح

ASPIRE: Agentic Safety & Prompt Injection Red-teaming Engine

Pengfei He, Deep Mitra, Vishesh Sharma وآخرون · 2026

LLM agents retrieve untrusted content and act through tools, creating indirect prompt-injection risks that can cause unauthorized actions or persistent state changes. Existing automated red-teaming largely optimizes payloads for pre-specified scenarios, leaving latent vulnerabilities across the agent's behavior space u …

نسخة أولية وصول مفتوح

SEER: Self-Evolving Event Reasoning and Retrieval for Time Series Forecasting

Mingtian Tan, Palash Goyal, Mihir Parmar وآخرون · 2026

Real-world time series are frequently driven by exogenous events and structural shifts, rendering conventional forecasting based solely on historical numerical observations insufficient. While language models can retrieve external news, standard retrieval-augmented approaches struggle with high noise, missing signals, …

نسخة أولية وصول مفتوح

VeriHarness: Scaling Agentic Verification for Long-Horizon Tasks

Caiqi Zhang, Rujun Han, Zifeng Wang وآخرون · 2026

As LLM agents undertake increasingly complex, long-horizon tasks, verifying their outputs becomes increasingly challenging. We study how verification capability can be strengthened with a fixed base model, without access to reference answers or grading rubrics at test time. Repeated sampling yields multiple rollouts th …

نسخة أولية وصول مفتوح

AIM: Agentic Idea Management for Automated Research

Frontier LLMs are increasingly used to automate scientific research through iterative search. We distinguish idea-driven search from solution-driven search and identify three core challenges: organizing evolving research ideas, selecting promising directions, and maintaining alignment between ideas and their implementa …

نسخة أولية وصول مفتوح

PreviewDiff: Multimodal Critic-Guided Search over Diffusion Latents

Diffusion models can produce striking images and videos, but they still struggle with the compositional details that make a generation faithful to a prompt, such as object counts, attribute binding, spatial relations, and temporally grounded actions. A common way to improve prompt satisfaction is to spend more compute …

نسخة أولية وصول مفتوح

RRSI: Regularized Recursive Self-Improvement of Agent Harnesses

Peng Xia, Rujun Han, Zifeng Wang وآخرون · 2026

An LLM agent's capability is largely magnified by its harness, namely the prompts, control flow, tooling, memory, and context management surrounding the frozen backbone model. Recent methods increasingly automate this process by iteratively proposing and selecting component-wise edits of an agent harness, practically e …

المؤلفون المشاركون