Licenses
The code is MIT. The embedded models, rarity tables, and Snowball word lists keep their own terms.
- Part-of-speech models are CC BY-SA. The Italian model also cites CC BY-SA 3.0 and CC BY 4.0. They were trained on Universal Dependencies treebanks and contain feature weights, not the treebank sentences.
- Rarity tables are CC BY, from the Leipzig Corpora Collection news and Wikipedia samples. Cite Goldhahn, Eckart, and Quasthoff, LREC 2012.
- Snowball stopword lists are BSD-3-Clause.
The sources, treebank names, and corpus filenames are in lang/THIRD_PARTY.md. NOTICE points at that file.