AI Potluck
Back to Gap Map Infrastructure / Classic ML & computer vision

XGBoost

DMLC
open source / Overall score: 3.8

Optimized gradient-boosted decision tree library for classification, regression, ranking and survival analysis, running on a single machine, on GPUs, and distributed across Kubernetes, Spark, Dask and Ray. Bindings cover Python, R, the JVM and several other languages. It began as a University of Washington research project and is governed by a project management committee.

Openness

5 high confidence
5.0
license
Apache-2.0(OSI)
source
public(github.com/dmlc/xgboost)
core features withheld
no — committee-governed

Apache-2.0, built from the public repository, and run by a project management committee that takes sponsorship for its CI costs. No company sells an edition of it.

Adoption

5 high confidence
5.0

Measured on monthly PyPI downloads of the xgboost package; the R, JVM and other bindings are not counted.

Capability

2 high confidence
2.0

XGBoost does one kind of model, boosted trees, and does it very well, at scale and on GPUs. Anything else - another learner, the comparison between them, the pipeline around them - comes from a toolkit like scikit-learn, whose estimator interface it implements.

  • https://xgboost.readthedocs.io/en/stable/ recorded 2026-09-26

    Docs landing tutorials: Introduction to Boosted Trees, Learning to Rank, DART, Monotonic Constraints, Feature Interaction Constraints, Survival Analysis with Accelerated Failure Time, Categorical Data, Multiple Outputs, Random Forests in XGBoost; distributed on Kubernetes, Spark, Dask, PySpark and Ray; GPU support; scikit-learn estimator interface.

Verified 2026-09-26