الباحثون

Cristian McGee

المنشورات 2

نسخة أولية وصول مفتوح

Trust the Direction, Search the Step: Zero-and-First-Order Methods for LLM Fine-Tuning

Step-size selection remains a central challenge in large-scale neural network optimization; conservative steps slow convergence, while aggressive steps can destabilize it. We combine \textbf{Z}ero-and-\textbf{F}irst-\textbf{O}rder optimization~(ZFO) and propose a lightweight framework that decouples direction selection …

نسخة أولية وصول مفتوح

TACO: Ternary Absolute-max Column-wise One-sparse Optimizer for LLM Fine-Tuning

Full-parameter fine-tuning of large language models (LLMs) incurs substantial optimizer state memory overhead, limiting the model sizes that fit on modern GPUs. Existing approaches either compress optimizer state, abandon first-order gradients, or change the update geometry while retaining dense state. The recently int …

المؤلفون المشاركون