cuLA
inclusionAI (Ant Group)cuLA (CUDA Linear Attention) provides hand-tuned CUDA kernels, written in CuTe DSL and CUTLASS C++, for linear attention variants including GLA, KDA, GDN and Lightning Attention, targeting NVIDIA Hopper and Blackwell GPUs. It shares its interface with the flash-linear-attention library as an early-stage standalone precursor to a planned integration.
Verified 2026-09-02 via the GitHub API and the LICENSE body.
Openness
5 high confidence- license
- Apache-2.0(OSI)
- source
- public
- core-gated
- ungated
Apache-2.0 license body confirmed. The repository is public and unarchived and builds the whole product, and the README describes no paid tier, enterprise edition or license-gated build beside it, so source is public and the core ungated.
- https://github.com/inclusionAI/cuLA/blob/main/LICENSE recorded 2026-09-02
LICENSE body carries the verbatim Apache License 2.0 text under an Ant Group / inclusionAI copyright header.
- https://api.github.com/repos/inclusionAI/cuLA recorded 2026-09-02
Repo metadata for cula's canonical repository - private false, archived false, fork false.
- https://github.com/inclusionAI/cuLA/blob/main/README.md recorded 2026-09-02
README describes cuLA as hand-tuned CUDA kernels for linear attention variants targeting NVIDIA Blackwell and Hopper, explicitly early-stage with room for further kernel optimization, and names no paid tier, enterprise edition or license-gated build beside the published source.
Adoption
1 low confidence544 GitHub stars, which lands in the <1K stars band of the stars scale, level 1. cuLA has no registry package and is built from source, so no download count exists and the stars scale applies.
- https://api.github.com/repos/inclusionAI/cuLA recorded 2026-09-02
Repo metadata - stargazers_count = 544 - for inclusionAI/cuLA.
Capability
2 medium confidenceBanded on the category feature matrix as a narrow kernel set: linear-attention kernels for one vendor's GPUs, self-described as early-stage. Placed two bands below the tensorrt rung.
- https://github.com/inclusionAI/cuLA/blob/main/README.md recorded 2026-09-02
README describes cuLA as hand-tuned CUDA kernels for linear attention variants targeting NVIDIA Blackwell and Hopper, explicitly early-stage with room for further kernel optimization, and names no paid tier, enterprise edition or license-gated build beside the published source.
Verified 2026-09-02