cm0002@literature.cafe to AI - Artificial intelligence@programming.devEnglish · 5 months agoA SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformersgithub.comexternal-linkmessage-square0linkfedilinkarrow-up11arrow-down10cross-posted to: technology@lemmy.ml
arrow-up11arrow-down1external-linkA SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformersgithub.comcm0002@literature.cafe to AI - Artificial intelligence@programming.devEnglish · 5 months agomessage-square0linkfedilinkcross-posted to: technology@lemmy.ml