AI Potluck
Infrastructure / Compilers & Model Optimization

Humming

inclusionAI (Ant Group)

Humming is a lightweight, JIT-compiled GEMM kernel library for quantized inference, supporting dense and MoE GEMM across FP16/BF16/FP8/FP4/INT8/INT4 activation and weight-type combinations on NVIDIA GPUs from Turing onward.

Verified 2026-09-02 via the GitHub API and the LICENSE body.

Openness

5 high confidence
5.0
license
Apache-2.0(OSI)
source
public
core-gated
ungated

Apache-2.0 license body confirmed. The repository is public and unarchived and builds the whole product, and the README describes no paid tier, enterprise edition or license-gated build beside it, so source is public and the core ungated.

Adoption

1 low confidence
1.0

219 GitHub stars, which lands in the <1K stars band of the stars scale, level 1. Humming has no registry package and is built from source, so no download count exists and the stars scale applies.

Capability

3 medium confidence
3.0

Banded on the category feature matrix as one optimization family, applied broadly - the breadth of dtype and hardware coverage puts it beside cutlass and onednn rather than a single-purpose kernel. Placed two bands below the apache-tvm anchor.

  • https://github.com/inclusionAI/humming/blob/main/README.md recorded 2026-09-02

    README describes Humming as a high-performance, lightweight, JIT-compiled GEMM kernel library for quantized inference, supporting dense and MoE GEMM across a wide activation/weight type matrix, with no paid tier, enterprise edition or license-gated build beside the published source.

Verified 2026-09-02