الباحثون

Prithwish Jana

المنشورات 5

نسخة أولية وصول مفتوح

FreeEvolve: Learning to Evolve Beyond Fixed Loops

Lecheng Kong, Like Hui, Nikos Kanakaris وآخرون · 2026

Agent evolvers automate the design of the prompts, skills and workflows around language model agents, yet the optimization process they follow is still designed by hand: a fixed search loop decides how candidates are evaluated, which are kept and when the search stops. We propose FREEEVOLVE, which automates this proces …

نسخة أولية وصول مفتوح

AIProver: Agentic Auto-Formalization of Mathematical Research via Certificate-Driven Evolving Harness

Prithwish Jana, Viet Bach Hoang, Logan Luna وآخرون · 2026

Proof auto-formalization translates natural-language (NL) theorems and proofs into a formal language (FL) such as Lean, enabling mechanical verification. Despite rapid progress, research-level proofs often depend on concepts missing from leading proof assistant libraries (e.g., Lean's Mathlib), and successful compilati …

نسخة أولية وصول مفتوح

Hermes: Learning Contextual Reasoning Unlocks Test-Time Scaling

Xinyu Li, Mononito Goswami, Hao Liu وآخرون · 2026

Test-time scaling improves model performance by allocating additional compute during inference. Using this compute effectively across multiple context windows requires deciding how to allocate fresh contexts and what information to carry between them. We call a model's ability to make these decisions contextual reasoni …

نسخة أولية وصول مفتوح

MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution

Prithwish Jana, Mononito Goswami, Hao Liu وآخرون · 2026

Modern agentic systems combine an AI model with a harness that controls execution and environmental interactions. Harness design strongly affects long-horizon performance, yet its combinatorial search space demands substantial human effort that must be repeated as models change. Existing automated methods explore this …

نسخة أولية وصول مفتوح

On the Off-Policy Teacher in On-Policy Distillation

Langlin Huang, Hao Liu, Mononito Goswami وآخرون · 2026

On-policy distillation (OPD) has recently emerged as a promising post-training paradigm in which the student learns from trajectories generated by its own policy under dense teacher supervision. However, OPD introduces a fundamental asymmetry: although the sampled trajectories are on-policy for the student, they are of …

المؤلفون المشاركون