A frequency list of the 10,000 most common lemmas in European Portuguese film and TV dialogue, built from the Portugal-tagged portion of the OpenSubtitles 2018 corpus — 636 million tokens of subtitle text.
Ready-to-import Anki decks are included.
| Entries | 10,000 rows (two files of 5,000): 9,390 distinct words, 311 of them listed once per part of speech, plus 295 multi-word expressions |
| Corpus | OPUS OpenSubtitles v2018, pt (Portugal-tagged) |
| Corpus size | 118,469,705 lines · 636,214,452 tokens · 838,358 unique surface types |
| Lemmatization accuracy | 98.5% on a 332-word human-reviewed sample |
| Part of speech | From a neural tagger (Stanza), by majority vote over up to 50 sentences per word |
| Licence | Data and docs (out/, eval/, reports/, archive/, README, CHANGELOG, DATA_SOURCES): CC BY-SA 4.0 — see LICENSE-DATA. Code (scripts/, tests/, config.yaml, requirements.txt): MIT — see LICENSE |
An earlier version of this list (April 2026) was never published; the
pipeline that produced it was lost. It is kept for comparison only in
archive/out_april_2026/ (tag
v0-april-2026), and CHANGELOG.md records how this list
differs from it.
Learners of European Portuguese who want to study vocabulary in the
order they are likely to actually hear it. If you have worked through a
beginner course and want a principled answer to "what should I learn
next?", a frequency list built from dialogue is a reasonable guide — and
the Anki decks in out/ let you start tonight.
It is particularly aimed at learners who keep finding that the Portuguese in their textbook is Brazilian, or that their word list is drawn from newspapers rather than conversation.
Researchers and tool-builders may find the lists useful as a rough register baseline, provided you read the Limitations below first and treat the numbers accordingly.
It is not a course, a dictionary, or a substitute for either. The
glosses in out/glosses.tsv are a one-to-three-word
English hint per entry and one real corpus sentence, written by a language
model in a single pass — a snapshot of what that model said on one day,
not a lexicographer's definition. They have not been checked entry by
entry. Treat a gloss as a starting point and a dictionary as the
authority. (The "enrichable" Anki files leave the gloss columns empty if
you would rather write your own.)
Please read this section before you rely on anything here. This is a hobbyist project, and being honest about what it is not is more useful than overselling what it is.
Every number here comes from one source: subtitles for films and TV shows. That has real consequences.
- Subtitles are written approximations of speech, not speech. They are condensed to fit reading speed, cleaned of disfluencies, hesitations, repairs, and overlaps, and they lose everything about how something was said. Real conversation looks different.
- The register is dramatic dialogue. Film and TV over-represent crime, conflict, romance, and emergencies. Matar ("to kill") lands at rank 121, morrer at 153, arma at 248 and polícia at 267 — higher than any corpus of everyday life would put them. The vocabulary of work, admin, school and errands correspondingly sits lower.
- Much of it is translated. A large share of these subtitles are translations of English-language film and TV, so the phrasing can carry translation artifacts and English-shaped idiom rather than natively produced Portuguese.
- The
pttag is corpus metadata, not verified provenance. It is a strong signal of European Portuguese, but Brazilian material can leak through. An explicit BP exclusion list is applied, and it is a blunt instrument — it removes obvious markers, not every Brazilianism. - It is single-genre and one snapshot in time. No news, no academic prose, no fiction, no non-fiction, no actual recorded conversation, and nothing newer than the 2018 corpus release.
- Neighbouring lines are not a scene. The corpus file is a stream of subtitle lines that has been deduplicated, reordered and interleaved across different subtitle versions of the same films, so two adjacent lines usually do not belong to one exchange. Counting and tagging are unaffected, since both work a line at a time, but every "sample sentence" in this project is an isolated line: do not read the lines around one as its context, and do not build an annotation task that relies on them.
The standard published reference is Mark Davies and Ana Maria Raposo Preto-Bay's A Frequency Dictionary of Portuguese (Routledge, 2008). If you are choosing between them, choose Davies — it is peer-reviewed, professionally edited, and comes with the things this list lacks.
| This list | Davies | |
|---|---|---|
| Corpus | ~636M tokens, subtitles only | ~20M tokens, balanced across registers |
| Registers | Film/TV dialogue | Spoken, fiction, newspaper, academic |
| Variety | Portugal-tagged, BP markers excluded post hoc; headwords in post-1990 spelling | Both varieties, with BP/EP marked per entry |
| Glosses & examples | One to three English senses per entry and one example sentence taken verbatim from the corpus — written by a language model in one pass, unchecked entry by entry | English glosses and example sentences throughout, written and edited by lexicographers |
| Lemmatization | Automatic: a neural tagger (Stanza) reading each word in up to 50 real sentences, checked against a dictionary. 98.5% correct on a 332-word hand-checked sample | Curated |
| POS tags | Automatic, from the same tagger: majority vote over each word's sample sentences, and a word used as two parts of speech is listed once for each. Not hand-checked | Professionally tagged |
| Contractions (do, ao, pela) | Kept as their own entries | Split into preposition + article |
| Review | Automatic output. Spot-checked against a hand-reviewed sample, with about 830 individual decisions made by hand (gold set, gender pairs, proper-noun drops); not edited entry by entry | Edited and peer-reviewed |
| Cost | Free | A book you buy |
What this list offers that Davies does not is scale on a single register and a strictly European-Portuguese slant: 30× more tokens, all of it dialogue, with Brazilian markers filtered out. That makes it a decent complement to Davies for a learner targeting spoken EP. It is not a replacement, and it has not been edited the way a published dictionary is.
These are specific and worth knowing before you study from the list.
- Part of speech is tagger output, not hand-checked. The
poscolumn is Stanza's tag, by majority vote over each word's sample sentences (50 for the 4,347 most frequent word forms, 5 for the rest). The tagger is not always consistent with itself: 3% of its tags contradicted the lemma it gave in the same sentence (a VERB tag on the noun pergunta) and were discarded. A word the tagger regularly uses two ways gets one row per part of speech, with its count divided between them:a(article and preposition),que(conjunction and pronoun),este(determiner and pronoun),morto(adjective and noun). The tags follow Universal Dependencies conventions, with one exception made for learners: a preposition is never counted as a conjunction (UD tags para in para fazer as one). Entries the tagger never saw intact keep a rule-based guess, except the contractions, which are all published asdet(see below). - The glosses are a model
snapshot, not a reference work. Every gloss and translation in
out/glosses.tsvwas written by Claude (claude-opus-5) in a single pass, one request per entry, from the entry, its part of speech, its frequency and up to twenty real corpus sentences containing it. The prompt is in the repository (scripts/gloss_prompt.md) because the glosses cannot be read as evidence without it. Nobody checked all 10,000 by hand. Two consequences worth knowing: the model chose which sense to put first, so the order reflects its judgement of spoken frequency rather than a count; and re-running the same prompt on a newer model would not give the same answers, so unlike the lists themselves the glosses are not reproducible byte for byte. The example sentences are the one part that is verifiable — each is a corpus line, copied unaltered but for a leading dialogue dash and surrounding quotation marks, and the build rejects any the model did not copy exactly. Where the model edited a sentence rather than quoting it (24 entries), the example was dropped and the entry keeps its gloss alone; 200 entries have no example because none of the sampled lines showed the entry in the sense being glossed. - Contractions are labelled by what they contract, not by the tagger.
Stanza expands do, nas, pelo and the rest into two words before
tagging, so the surface token collects almost no usable evidence: before
this release nas, no and 26 others read
unk, and the ones with a stray tag read worse — pelo as a noun, deste as a verb, contigo split across three parts of speech. Each form now gets one decided answer: preposition + article or demonstrative isdet(do, nas, pelo, num, deste, naquela — 58 forms), preposition + pronoun ispron(dele, comigo, disso, nisto, daquilo — 19), and preposition + adverb of place isadv(daqui, daí, dali). This is a decision, not tagger output, and it is the one part of theposcolumn that does not come from Stanza. serandirare separated by a rule, not by the tagger.foi,fui,fomos,foramandforabelong to ser ("was") or ir ("went") depending on the sentence, and the tagger assigns them to ser almost every time (fui ao Paquistão included). So the tagger only decides whether an occurrence is a verb at all — the adverb fora, "outside", is not — and a rule reading the neighbouring words picks the verb: ir before a, ao, para, embora or an infinitive, or after lá; ser before a participle, an adjective, a noun phrase or nothing. On the 100 lines used to develop it, it is right on 95.9% of the verb uses it decides; on a held-out 100 lines labeled afterwards by a native-speaker Portuguese teacher, 93.2% (68 of 73), abstaining on 8%. An abstention takes the form's own ser/ir ratio. Every one of those five held-out errors calls an ir use ser (já fomos, já foram todos?), so ir is still undercounted a little. The split is estimated from 1,000 random lines per form, so each form's share is good to about ±3 percentage points. It misses idioms (ele não foi nessa) and elliptical questions (sempre foram?). Ser counts 21.3M tokens and ir 9.0M.- Lemmatization is not perfect. 98.5% on the hand-checked sample means
something like 150 of the 10,000 entries may carry a lemma error. The
known pattern: Stanza sometimes reads a verb form as a noun (
procura,mentes,sacaskept as nouns) or picks the wrong verb (vejamas vir, not ver). - Split counts are estimates. Participles (
educados: the verb educar or the adjective educado) and words used as two parts of speech have their counts divided according to how the tagger reads a sample of sentences: 50 for word forms with 10,000+ occurrences, which gives 2% steps, and 5 for the rest (20% steps). - Six known duplicates remain. Missing-accent forms are folded into
their accented twin only when the accented form is at least 10× as
frequent. Six pairs fall below that and appear twice:
camera/câmera,frigorifico/frigorífico,amen/ámen,bla/blá,mafia/máfia,karate/karaté. They are accepted as known typo pairs rather than hidden. Treat the unaccented one as a typo of the other. - Folding by frequency can over-merge. The same 10× rule folds a rare
unaccented form into a common accented one even when the rare form is a
real word. Every fold is listed in
reports/QUALITY.md, with the borderline ones (10–20×) in their own section. - Proper nouns are removed by capitalization, and some will slip
through. A word is treated as a name if it is capitalized away from the
start of a line almost every time. Dictionary words that the filter
removed at 5,000+ occurrences were reviewed one by one —
deus,sr,nataland 26 others were restored — but names below that threshold were not reviewed. Everything removed with 100+ occurrences is listed inout/dropped_proper_nouns.txt. - English words are only partly filtered. English plurals and many
English names are removed; other English words that turn up in
Portuguese dialogue (
ok,zombie) are kept, deliberately in the case ofok. - MWE counts are an overlay, not a replacement. An MWE's
raw_freqis the bigram count, and it is not subtracted from its constituent words' lemma counts. This is the usual frequency-list convention, but it means the column does not sum to the corpus.
Two rank bands — 1–5000 and 5001–10000 — each published in four formats. Files without a number in the name are the first band.
| File | What it is |
|---|---|
ep_spoken_top5000.tsv / .csv |
Ranks 1–5000. The main list. |
ep_spoken_5001_10000.tsv / .csv |
Ranks 5001–10000. |
TSV and CSV hold identical data; pick whichever your tools prefer. Columns:
| Column | Meaning |
|---|---|
rank |
1-based rank by raw_freq. |
lemma |
Lemma form, or the surface multi-word expression for MWE entries. |
is_mwe |
1 for a multi-word expression, 0 otherwise. |
pos |
Part of speech from the tagger. One of det, prep, conj, pron, adv, intj, num, noun, adj, verb, mwe, unk. A word used as two parts of speech has one row for each. See caveats. |
raw_freq |
Lemma occurrence count across the full corpus. |
freq_per_million |
Occurrences per million tokens (of 636,214,452). |
| File | What it is |
|---|---|
EP_Spoken_Frequency.apkg |
Ready to import: twenty decks of 500 words, the note type and the card templates in one package. See Importing the Anki decks. |
ep_spoken_anki_minimal.tsv |
Rank, Lemma_PT, POS, Is_MWE, Raw_Freq, Freq_per_million, Tags |
ep_spoken_anki_enrichable.tsv |
The same, plus empty Gloss_EN, Example_PT, Example_EN for you to fill in |
ep_spoken_5001_10000_anki_minimal.tsv |
Ranks 5001–10000, minimal fields |
ep_spoken_5001_10000_anki_enrichable.tsv |
Ranks 5001–10000, enrichable fields |
| File | What it is |
|---|---|
glosses.tsv |
One row per entry: gloss (one to three English senses, commonest first), example_pt (a corpus line containing the entry, verbatim), example_en (its translation), flags, source. |
flags is zero or more of vulgar, bp-leaning (the form or sense is
mainly Brazilian), archaic, name-like (in this corpus the token is
probably mostly a name) and uncertain (the model was not confident).
Slang and vulgarities are glossed plainly: this is a record of how people
speak, not a teaching aid. See the caveat below.
source says where the row came from, and the deck publishes it as
Gloss_Source. It is model for 9,985 rows. The other 15 are in
eval/gloss_overrides.tsv, where a reviewer
overruled the model: override-edited (6) for a corpus line a native speaker
edited, override-chosen (7) for a different corpus line she chose, and
override (2) for a gloss written by hand — caseiro, which the model
declines to gloss at all, and largo.
| File | What it is |
|---|---|
qc_sample.csv |
A spot-check sample across the first band — a quick way to eyeball output quality without opening a 5,000-row file. |
dropped_proper_nouns.txt |
Words removed as proper nouns (with 100+ occurrences), plus the BP exclusions, English plurals and fragments removed by the other filters. |
The build reports are in reports/: COMPARISON.md (this
list against the unpublished April 2026 version), QUALITY.md (the quality gate and every
accent fold), and backend_scores.md (how the lemmatizers compared).
From pt.txt.gz to out/, in one command (python -m scripts.build):
-
Tokenize and count. A Portuguese-aware tokenizer strips subtitle artifacts and lowercases. Enclitic clusters are split and the verb is restored:
fazê-locounts as fazer + lo, anddar-lhe-iaas daria + lhe. -
Second pass. Bigram counts for multi-word expressions, sample sentences for every frequent word (50 for the 4,347 word forms with 10,000+ occurrences, 5 for the rest), and capitalization statistics for the proper-noun filter.
-
Tag and lemmatize. Stanza reads the 63,253 word forms frequent enough to reach the list inside their sample sentences — about 500,000 sentences, covering 62,494 of those forms — and gives each occurrence a lemma and a part of speech. The majority lemma wins, after discarding votes where the lemma and the tag contradict each other. When Stanza's answer is not a dictionary word (it occasionally invents forms like agradeçar), simplemma is used instead; simplemma also handles the rare tail. A small table of hand-verified corrections covers errors Stanza makes at high frequency (
dóiis doer, not dizer). -
Resolve context-dependent forms per occurrence. Participles are split between the verb and the adjective according to how Stanza tags them in their sample sentences.
foi,fui,fomos,foramandforaare split between ser and ir by a separate next-word rule, because the tagger cannot tell them apart (see caveats). -
Apply the lemmatization conventions in eval/conventions.md:
- contractions (
do,ao,pela) are their own entries; - adjectives fold to the masculine, but nouns for people keep the
feminine (
senhora,rapariga), and a feminine with its own meaning always stays (música,sexta) — decided pair by pair in eval/gender_pairs.tsv; - diminutives stay separate (
coisinha); - comparatives and irregular superlatives are their own entries
(
maior,melhor,ótimo); - spelling-reform variants merge under the post-1990 spelling
(
acção→ação,óptimo→ótimo), while words that keep their consonant in European spelling (facto,contacto) are left alone.
Pronouns are their own entries too (
me, not eu). - contractions (
-
Fold remaining duplicates. Each lemma is checked against the lemmatizer again, and regular plurals are folded onto their singular (guarded:
óculos,fériasandcaisare not plurals of anything). An unaccented form folds into its accented twin when the accented form is at least 10× as frequent, and wrong or Brazilian accents (näo,prêmio) fold into the European spelling. -
Filter. BP-leaning words are removed from the exclusion list —
aeromoça,bacana,banheiro,celular,galera,geladeira,legal,mano,massa,né,oi,trem,você,vocês,xícara,ônibus, with their unaccented variants. Proper nouns are removed by capitalization, subject to the reviewed decisions in eval/proper_noun_drops_review.tsv. English plurals are removed, and so are one- and two-letter fragments that are neither words nor interjections. -
Assign part of speech. Each word's count is divided across the parts of speech the tagger gave it, the same way participle counts are divided across lemmas. A second part of speech becomes its own row only if it holds at least 25% of the word's count and at least 3 tagged sentences; otherwise it joins the majority.
-
Rank and emit the top 10,000 rows by count, with multi-word expressions — kept by log-likelihood (Dunning G² ≥ 2500) and a collocation share of at least 5% — ranked alongside them. 295 of them make the top 10,000: 191 in the first band and 104 in the second.
- A human-reviewed gold set. eval/lemma_gold.tsv is 400 words sampled from the corpus across the frequency range, each checked by hand. 332 are scored; 68 were marked as names, foreign words or junk. The published lemmas are correct for 98.5% of them.
- A quality gate. Every build looks for entries that are probably the
same word counted twice — a form that lemmatizes to another entry, an
accent variant of another entry, a regular plural of another entry — and
fails if it finds any that are not explicitly accepted. This release
passes. Two kinds of pair are accepted in
config.yaml, each listed: words that are genuinely different (avô/avó,pôr/por,cirurgiã/cirurgia), and the six known typo pairs from the caveats, which sit below the 10× folding threshold. Any new suspect fails the build. - Two labeled ser/ir samples, 100 corpus lines each, 20 for each of
foi,fui,fomos,foramandfora. eval/ser_ir_sample.tsv was used to develop the rule: its labels were proposed by an AI model acting as a Portuguese informant and checked by hand, and the rule scores 95.9% on the verb uses it decides (71 of 74). eval/ser_ir_sample2.tsv is held out: labeled afterwards by a native-speaker Portuguese teacher, never used to shape the rule, and therefore the honest estimate — 93.2% (68 of 73), abstaining on 8%. Stanza's verb test keeps the adverb fora away from the rule in 19 of 20 held-out cases; the exception is lá fora, which it mistagged as a verb. The build uses the rule only while a score of at least 95% on the development set exists for its current version; both tables are in reports/serir_scores.md. - Automatic gates on the glosses. Every gloss run checks that each published row got a gloss, that each example sentence is one of the corpus lines the model was shown — character for character — and that the sentence really contains the entry; a run that fails any of those is not published. Counts, per-flag totals and what the run cost are in reports/gloss_gates.md. The glosses' meaning is not gated: a sample is set aside for a reader in eval/gloss_review.tsv.
- Determinism. The same corpus, configuration and library versions produce byte-identical output. The glosses are the exception — see the caveat.
You need Python 3.13 and several GB of free disk (the corpus, Stanza's models and PyTorch). The first build takes about an hour on an 8-core machine; later builds reuse cached passes and take a minute.
git clone https://github.com/professorcdp/EPspokenfrequency.git
cd EPspokenfrequency
python3 -m venv .venv
./.venv/bin/pip install -r requirements.txt
./.venv/bin/python -c "import stanza; stanza.download('pt')"Download the corpus (1.1 GB) into data/ as described in
DATA_SOURCES.md, then:
./.venv/bin/python -m scripts.buildThis writes out/ and reports/ and exits with status 0. If you change
the configuration and the build exits with status 1, the quality gate has
found a new suspected duplicate: it is listed in reports/QUALITY.md, and
every output has still been written.
requirements.txt pins the exact versions used for this release. With
them the build reproduces out/ byte for byte; a different Stanza or
simplemma release can change individual lemmas. To rebuild the April-style
list for comparison, run python -m scripts.build --stage 1 (output in
build/stage1/). scripts/README.md covers the other
options.
Use out/EP_Spoken_Frequency.apkg. In Anki, choose File → Import,
select it, and you are done: the note type, the card templates and both
decks come with it.
- Twenty decks of 500 words each, under one parent and numbered so they
sort in rank order:
Portuguese::European Portuguese - Spoken::01 · 1–500,::02 · 501–1000, and so on to::20 · 9501–10000. Study them in order and you meet the commonest words first. - One note type,
Portuguese (EP) – Spoken Frequency, with the fieldsRank,Lemma_PT,POS,Gloss_EN,Example_PT,Example_EN,Is_MWE,Raw_FreqandFreq_per_million. - Card 1 (Portuguese → English) shows the word large, with its part of speech beneath it. The back adds your gloss and example sentences if you have filled them in, with the rank and frequency in small print.
- Card 2 (English → Portuguese) exists only for notes whose
Gloss_ENis filled in. The list ships without glosses, so at first you only have Card 1. Type a gloss into a note and its reverse card appears. - Every note carries the tags below, the same as in the TSV files.
Importing a later release. Notes are identified by their word and part of speech, so a new release matches the notes you already have instead of adding duplicates, and your review history is kept. Anki's import screen asks whether to Update notes. If newer or Always brings in the new ranks and counts, but it replaces every field, including glosses and examples you have typed in. To keep your own glosses, choose Never; your ranks then stay as they were.
If you want your own fields or card design, import the *_anki_*.tsv
files instead. They carry Anki's header directives, so Anki configures
most of the import itself: separator, note type, target deck, field
mapping and tags column are all declared in the file.
-
Create the note type first; Anki will not invent it for you. Go to Tools → Manage Note Types → Add, name it
Portuguese (EP) – Spoken Frequency, and give it fields matching the file you plan to import:- minimal:
Rank,Lemma_PT,POS,Is_MWE,Raw_Freq,Freq_per_million - enrichable:
Rank,Lemma_PT,POS,Gloss_EN,Example_PT,Example_EN,Is_MWE,Raw_Freq,Freq_per_million
(The
Tagscolumn is not a field; Anki maps it to note tags automatically.) Then design your own card templates. - minimal:
-
Choose File → Import and select one of the
out/*_anki_*.tsvfiles. -
The import screen pre-fills from the file's headers. Check that the Field separator is Tab, the Notetype is
Portuguese (EP) – Spoken Frequency, the Deck isPortuguese::European Portuguese - Spoken - First 5000(or... - 5001 to 10000), and Allow HTML in fields is off. -
Click Import.
Do not mix the two paths in one Anki profile: the TSV import creates its
own notes, separate from the .apkg's, so you would study every word
twice.
The .apkg, unless you have a reason not to. It has room for your own
translations and example sentences, which is the approach that actually
builds retention, and adding a gloss gives you the reverse card for free.
On the advanced path, take ep_spoken_anki_enrichable.tsv if you will
add glosses and examples, or the minimal file if you only want the ranked
word list and will look things up elsewhere. The TSV files hold 5,000 words
each, so you will want to slice them by tag (below) as you go.
The twenty decks are the order. Work through 01 · 1–500 before 02 · 501–1000, and you learn the words roughly in the order you will hear
them. The numbers keep the decks in order in Anki's deck list, and each
deck has its own daily limits, so a deck you have not started sends you no
new cards.
Every note also carries tags, for slicing more finely than a deck — all the verbs in the first thousand, say — from the Anki browser:
| Tag | Meaning |
|---|---|
freq::0001-0500, freq::0501-1000, … |
500-word band (matches the decks) |
freq500::… |
500-word band (alias) |
freq1000::0001-1000, … |
1000-word band |
pos::det, pos::verb, … |
Part of speech |
source::opensubtitles_ep_v2018 |
Provenance |
variety::EP, register::spoken |
Variety and register |
The tags are on the TSV notes too, and are the only way to study those in order, since each TSV file imports into a single 5,000-note deck.
out/ Published frequency lists and Anki decks
reports/ Build reports: comparison, quality gate, backend scores
archive/out_april_2026/ The unpublished April 2026 list, kept for comparison
eval/ Gold set, conventions, and the hand-reviewed decision files
scripts/ The pipeline (python -m scripts.build)
tests/ Unit tests (pytest; no corpus needed)
config.yaml Every threshold, list and switch the pipeline uses
data/ Corpus download target — gitignored, never committed
DATA_SOURCES.md Corpus citation and download instructions
CHANGELOG.md What changed between versions
data/ is excluded from version control because the corpus is a 1.1 GB
file that has no business in git history. Download it yourself following
DATA_SOURCES.md.
The data and the code are licensed separately:
- The data — the lists and Anki files in
out/, andeval/,reports/andarchive/— is CC BY-SA 4.0 (LICENSE-DATA), and so is the documentation: this README, CHANGELOG.md and DATA_SOURCES.md. Use it, change it, redistribute it: keep the attribution and share derivatives alike. - The code —
scripts/(with its data tables),tests/,config.yamlandrequirements.txt— is MIT (LICENSE), so you can reuse the pipeline in any project, keeping the copyright notice.
To attribute the data: European Portuguese Spoken Frequency List, by Jim Ranalli (Professor Caloiro), licensed under CC BY-SA 4.0. Derived from OPUS OpenSubtitles v2018 (Lison & Tiedemann 2016).
If you publish anything based on this, please cite the underlying corpus:
Lison, P., & Tiedemann, J. (2016). OpenSubtitles2016: Extracting Large Parallel Corpora from Movie and TV Subtitles. LREC 2016.
Full citations, BibTeX and upstream terms are in DATA_SOURCES.md.