الباحثون

Chenyan Xiong

المنشورات 4

نسخة أولية وصول مفتوح

Effective Synthetic Data Curation Requires Group-Level Signals

Synthetic data now is essential to LLM training, used to strengthen advanced capabilities such as autonomous and long-horizon task execution. Yet recent work shows that training on it at scale can degrade model generation, making it important to decide what synthetic data is worth training on. While current data curati …

نسخة أولية وصول مفتوح

Delta-Matching: Closing the Final Gap of Native 8-bit Training for LLMs

Haozhan Tang, Hao Kang, Han Cai وآخرون · 2026

Reliable FP8 attention remains a barrier to fully native 8-bit large language model training. We derive how forward-backward inconsistencies produce stale delta and empirically show how it distorts training dynamics. Our stale-delta hybrid runs show a modest loss gap at 569M parameters but substantial loss increases an …

نسخة أولية وصول مفتوح

ReScraper: Unified Scraping and Cleaning of Web Data for Effective LLM Pretraining

Zichun Yu, Jiarui Yan, Shlok Sanghvi وآخرون · 2026

LLM pretraining corpora are normally cleaned by a stack of hand-written heuristics. A heuristic scraper extracts the main content from HTML, and dozens of rule-based filters then clean it, so corpus quality is capped by the coarseness and accuracy of the rules. In this work, we propose ReScraper, a unified language mod …

المؤلفون المشاركون