AI Potluck
Back to Gap Map Infrastructure / Classic ML & computer vision

NLTK

NLTK Project
open source / Overall score: 4.2(strong)

The Natural Language Toolkit, a suite of Python modules for tokenization, stemming, tagging, chunking, parsing, classification, language modeling and semantic reasoning, with interfaces to more than 50 corpora and lexical resources such as WordNet. It supports research and development in natural language processing, and is a community project led by Steven Bird since 2001.

Openness

5 high confidence
5.0
license
Apache-2.0(OSI
source
public(github.com/nltk/nltk)
core features withheld
no — community-driven

The code is Apache-2.0 and builds whole from the public repository, which has no enterprise, commercial or pro directory. Its documentation is under a non-commercial Creative Commons license and each corpus carries its own terms, but those cover the manuals and the data, not the library.

Adoption

5 high confidence
5.0

Measured on monthly PyPI downloads of the nltk package.

Capability

3 medium confidence
3.0

NLTK covers the classical NLP pipeline end to end with many learners - taggers, parsers, classifiers - and the corpora to train them on. It is a toolkit to build with, and leaves the choice among its methods to the user.

  • https://www.nltk.org/ recorded 2026-09-26

    Homepage: "easy-to-use interfaces to over 50 corpora and lexical resources such as WordNet, along with a suite of text processing libraries for classification, tokenization, stemming, tagging, parsing, and semantic reasoning".

  • https://www.nltk.org/py-modindex.html recorded 2026-09-26

    Module index: ccg, chat, chunk, classify, cluster, collocations, corpus, grammar, inference, lm, metrics, parse, probability, sem, sentiment, stem, tag, tbl, tokenize, translate, tree, wsd.

Verified 2026-09-26