☆ Yσɠƚԋσʂ ☆@lemmy.ml to Technology@lemmy.mlEnglish · 5 months agoA SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformersgithub.comexternal-linkmessage-square0linkfedilinkarrow-up16arrow-down11cross-posted to: Aii@programming.dev
arrow-up15arrow-down1external-linkA SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformersgithub.com☆ Yσɠƚԋσʂ ☆@lemmy.ml to Technology@lemmy.mlEnglish · 5 months agomessage-square0linkfedilinkcross-posted to: Aii@programming.dev