AI Potluck
Infrastructure / Compilers & Model Optimization

GPT-QModel

ModelCloud.ai

GPT-QModel is a quantization toolkit for large language models with hardware-accelerated kernels for NVIDIA CUDA, AMD ROCm, Huawei Ascend NPUs, Intel XPUs and x86 and Apple CPUs. Quantized checkpoints load through Hugging Face Transformers, vLLM and SGLang. It continues the GPTQ line of post-training quantization work.

Verified 2026-08-18 via the GitHub API and the LICENSE body.

Openness

5 high confidence
5.0
license
Apache-2.0(OSI)
source
public
core-gated
ungated

Apache-2.0 license body confirmed. The repository is public and unarchived and builds the whole product, and the README describes no paid tier, enterprise edition or license-gated build beside it, so source is public and the core ungated.

Adoption

2 high confidence
2.0

27,639 PyPI downloads of `gptqmodel` in the trailing 30 days, which lands in the 10K-100K band of the software usage scale, level 2.

Capability

3 medium confidence
3.0

Banded on the category feature matrix as one optimization family, applied broadly. Placed two bands below the apache-tvm anchor, on a matrix that bands on how much of the model-to-hardware transformation pipeline a product performs, over how many inputs and targets.

  • https://github.com/ModelCloud/GPTQModel/blob/main/README.md recorded 2026-08-18

    README still documents an LLM quantization toolkit with hardware-accelerated kernels for NVIDIA CUDA, AMD ROCm, Huawei Ascend NPU, Intel XPU and Intel, AMD and Apple CPUs, loadable through Transformers, vLLM and SGLang.

Verified 2026-08-18