AI Potluck
Back to Gap Map Infrastructure / Classic ML & computer vision

Gensim

RaRe Technologies
open source / Overall score: 3.6

Python library for topic modeling, document indexing and similarity retrieval over large text corpora, streaming data so memory use does not grow with corpus size. It implements LDA, LSI, HDP and NMF topic models, word2vec, doc2vec and fastText embeddings, and similarity search, with a downloader for pretrained word vectors such as GloVe and fastText. It is in stable maintenance mode.

Openness

5 high confidence
5.0
license
LGPL-2.1(GNU Lesser General Public License v2.1, OSI-approved)
source
public(github.com/piskvorky/gensim)
core features withheld
no — funded by sponsorship

Gensim is under the LGPL 2.1, an OSI-approved copyleft license that binds only redistribution, and builds from the public repository, now under its author's account. Commercial support is offered through sponsorship, and no edition is sold. New features are no longer accepted. LGPL-2.1 is OSI-approved, so the license tier is osi.

Adoption

4 high confidence
4.0

Measured on monthly PyPI downloads of the gensim package.

Capability

3 high confidence
3.0

Gensim is a toolkit of several text-modeling methods and the tools to train and compare them on the user's own corpus, in the way NLTK is for classical NLP. Its pretrained vectors are third-party downloads, and choosing and tuning a model stays with the user.

Verified 2026-09-27