الباحثون

Stefan Feuerriegel

المنشورات 4

نسخة أولية وصول مفتوح

OrthoGen: A Generative Orthogonal Learner for Time-Varying Treatments

Estimating conditional distributional potential outcomes (CDPOs) over time is important in medicine (e.g., to estimate patient-specific risks under different treatment sequences). However, this task is challenging because of time-varying confounding, yet existing adjustment strategies for this task are limited. In this …

نسخة أولية وصول مفتوح

Efficient Best-of-N policy evaluation for inference-time alignment

Best-of-N (BoN) is a common inference-time alignment method that selects the highest-scoring response among N samples from a reference model. Evaluating BoN policies from logged data is challenging under sample-only access because standard off-policy estimators require density ratios that depend on unavailable response …

نسخة أولية وصول مفتوح

Prediction-powered Neural Architecture Search

Pascal Janetzky, Yuxin Wang, Michael Klar وآخرون · 2026

Evaluating candidate architectures in neural architecture search (NAS) faces an inherent trade-off: on the one hand, reliable performance labels are limited because training and evaluating architectures is costly; on the other hand, zero-cost proxies (ZCPs) are cheap to compute at large scale but can be noisy. Yet, how …

نسخة أولية وصول مفتوح

Reward Hacking Challenges Oversight of Autonomous Research Agents

Yue Huang, Zhangchen Xu, Yuchen Ma وآخرون · 2026

Autonomous research agents can design experiments, evaluate results, and write reports, giving them control over both a scientific result and the evidence used to support it. This creates a risk of reward hacking: meeting the reward criteria without achieving the intended goal. We study (1) how often models reward-hack …

المؤلفون المشاركون