GPT-QModel
ModelCloud.aiGPT-QModel is a quantization toolkit for large language models with hardware-accelerated kernels for NVIDIA CUDA, AMD ROCm, Huawei Ascend NPUs, Intel XPUs and x86 and Apple CPUs. Quantized checkpoints load through Hugging Face Transformers, vLLM and SGLang. It continues the GPTQ line of post-training quantization work.
Verified 2026-08-18 via the GitHub API and the LICENSE body.
Openness
5 high confidence- license
- Apache-2.0(OSI)
- source
- public
- core-gated
- ungated
Apache-2.0 license body confirmed. The repository is public and unarchived and builds the whole product, and the README describes no paid tier, enterprise edition or license-gated build beside it, so source is public and the core ungated.
- https://github.com/ModelCloud/GPTQModel/blob/main/LICENSE recorded 2026-08-18
LICENSE body opens with ModelCloud.ai SPDX copyright headers declaring SPDX-License-Identifier Apache-2.0 and then runs the verbatim Apache License 2.0 text.
- https://api.github.com/repos/ModelCloud/GPTQModel recorded 2026-08-18
Repo metadata - license spdx_id NOASSERTION, private false, archived false, default branch main - for ModelCloud/GPTQModel.
- https://github.com/ModelCloud/GPTQModel/blob/main/README.md recorded 2026-08-18
README describes an LLM quantization toolkit with hardware-accelerated kernels for NVIDIA CUDA, AMD ROCm, Huawei Ascend NPU, Intel XPU and Intel, AMD and Apple CPUs, loadable through Transformers, vLLM and SGLang, with no paid tier, enterprise edition or license-key-gated build beside it.
Adoption
2 high confidence27,639 PyPI downloads of `gptqmodel` in the trailing 30 days, which lands in the 10K-100K band of the software usage scale, level 2.
- https://pypistats.org/api/packages/gptqmodel/recent recorded 2026-08-18
last_month downloads = 27,639 for gptqmodel
Capability
3 medium confidenceBanded on the category feature matrix as one optimization family, applied broadly. Placed two bands below the apache-tvm anchor, on a matrix that bands on how much of the model-to-hardware transformation pipeline a product performs, over how many inputs and targets.
- https://github.com/ModelCloud/GPTQModel/blob/main/README.md recorded 2026-08-18
README still documents an LLM quantization toolkit with hardware-accelerated kernels for NVIDIA CUDA, AMD ROCm, Huawei Ascend NPU, Intel XPU and Intel, AMD and Apple CPUs, loadable through Transformers, vLLM and SGLang.
Verified 2026-08-18