Per-language word-triple tables (<lang>_trigrams_100k.txt) for next-word prediction — issue #334.
Format: w1 w2 w3<TAB>count, lowercased, top 100 000 by count, written in key order so the app can binary-search the file without sorting it on load.
Source: Leipzig Corpora Collection (https://wortschatz-leipzig.de), CC BY — © Universität Leipzig / Sächsische Akademie der Wissenschaften / InfAI. Please cite D. Goldhahn, T. Eckart & U. Quasthoff (2012), "Building Large Monolingual Dictionaries at the Leipzig Corpora Collection", LREC 2012. Counted from each package's *-sentences.txt by tools/glide-dict/generate_ngrams.py.
Packages used: en = eng_news_2024_1M · de = deu_news_2022_1M