Authors

Fei Shen

Publications 8

Preprint Open access

BARQ: Balanced Codebook Refinement for Low-Bit LLM Quantization

Chenhang Cui, Xu Xie, Linrui Xu et al. · 2026

As large language models (LLMs) grow in parameter count, model storage and parameter memory traffic have become major bottlenecks to efficient deployment. Codebook-based weight quantization reduces these costs, but imbalanced nearest-codeword assignments during fitting can leave some codewords insufficiently updated, l …

Preprint Open access

Focusing Condition: Inference-Time Self-Contrastive Steering Elicits Better Conditional Text Embeddings in LLMs

Extracting conditional text embeddings from large language models (LLMs) is a promising paradigm, as it requires neither additional data nor fine-tuning. Existing methods incorporate conditions into prompts to guide LLMs to focus on specific aspects and elicit conditional text embeddings. However, relying solely on pro …

Co-authors