Breaking the 1.58-bit Barrier for Ternary LLMs

A new layout method called BITCOS is proposed to improve the storage efficiency of ternary large language models (LLMs). This is achieved by adapting to the actual distribution of zeros and ones in the model's weights, leading to a more compact storage format. The proposed method outperforms the conventional five-trit packing in 26 out of 29 tested models and results in a gain of up to 1.28 times in matrix-vector multiplication and 1.18 to 1.27 times in decode throughput on various platforms.

RSS Score 0 9/16/2026, 4:00:00 AM Original Source
Save an API key to vote.