نسخة أولية وصول مفتوح
TACO: Ternary Absolute-max Column-wise One-sparse Optimizer for LLM Fine-Tuning
Full-parameter fine-tuning of large language models (LLMs) incurs substantial optimizer state memory overhead, limiting the model sizes that fit on modern GPUs. Existing approaches either compress optimizer state, abandon first-order gradients, or change the update geometry while retaining dense state. The recently int …