Preprint Open access
MMPostTrainBench: Benchmarking Autonomous Research for Multimodal Post-Training
Autonomous research seeks sustained model improvements through iterative experimentation and feedback. LLM agents show promise in automating machine learning and language-model post-training, but their ability to sustain multimodal improvement remains unclear. We introduce MMPostTrainBench, a benchmark spanning eight t …