Unsloth
Unsloth AIPerformance-optimized fine-tuning library that rewrites kernels in Triton/CUDA to deliver 2-5x faster training with 60-80% less memory on a single GPU, makes 7B/13B LoRA fine-tuning feasible on a consumer RTX 4090. Picked when hardware is constrained and the workload fits one GPU; Unsloth's kernels also slot under TRL trainers so it complements rather than replaces the broader ecosystem. 65K GitHub stars and ~1500 commits in 90 days makes it the fastest-moving repo in the category; widely promoted on r/LocalLLaMA as the default for local fine-tuning.
Unsloth (unslothai/unsloth), confirmed live June 2026 (~65.8K stars). Memory/speed-optimized fine-tuning via custom Triton kernels; SFT, RL (GRPO), pretraining, quantization; claims up to 2x faster / 70% less VRAM. Core Apache-2.0; optional Unsloth Studio UI under AGPL-3.0.
Openness
5 high confidence- license
- Apache-2.0(OSI, core unsloth pkg)+AGPL-3.0(OSI, optional Unsloth Studio UI)
- source
- public(unslothai/unsloth)
OSI-licensed throughout (Apache-2.0 core, AGPL-3.0 for the optional Studio UI); both are OSI licenses so class is open_source, not open_core.
- https://github.com/unslothai/unsloth recorded 2026-06-04
Apache-2.0 core + AGPL-3.0 Studio UI; fine-tuning framework
Adoption
4 high confidence~2.43M PyPI downloads last month; widely used for low-VRAM/consumer-GPU fine-tuning of open models; ~65.8K stars corroborate. Level 4 on download volume.
- https://pypistats.org/packages/unsloth recorded 2026-06-04
Downloads last month: 2,432,077
- https://github.com/unslothai/unsloth recorded 2026-06-04
~65.8K GitHub stars (corroborating)
Capability
4 high confidenceBest-in-class single-/consumer-GPU efficiency (kernel-level), broad method coverage; C4, not the frontier multi-node-scale anchor.
- https://github.com/unslothai/unsloth recorded 2026-06-04
Triton-kernel 2x speed / 70% VRAM claims, GRPO/FP8 RL, 500+ models
Unchanged since 2026-06-24 (last edited, not re-checked)