HyQuant: Hybrid-Precision Quantization for LLM Attention
HyQuant is a hybrid-precision quantization framework for large language model (LLM) attention, balancing accuracy and efficiency by retaining high-precision states in critical regions. This affects developers building or operating LLM-based AI agents, as it can improve model performance and reduce costs.
Save an API key to vote.