== top-50 EN by corpus count == '▁the' '▁of' '▁and' '▁in' '▁for' '▁is' '▁many' '▁How' '▁as' '▁he' '▁on' '▁has' '▁are' '▁that' '▁each' '▁If' '▁she' '▁how' '▁The' '▁his' '▁her' '▁much' '▁with' '▁have' '▁does' '▁at' '▁than' '▁did' '▁more' '▁total' '▁He' '▁will' '▁can' '▁per' '▁day' '▁it' '▁from' '▁two' '▁number' '▁be' '▁times' '▁was' '▁one' '▁they' '▁She' 'ing' '▁cost' '▁In' '▁this' '▁three' == random-30 EN == 'itting' '▁everything' '▁have' 'ulated' '▁articles' '▁theater' '▁paintings' '▁wallet' '▁Keith' '▁melt' '▁below' 'ative' '▁shelf' '▁many' '▁corner' '▁colony' '▁meat' '▁Robin' 'ations' '▁loose' '▁received' '▁offer' '▁family' 'sg' '▁via' 'gio' '▁coal' '▁Ire' '▁Sand' '▁Jen' == top-50 PL by corpus count == '▁w' '▁z' 'ł' 'ę' 'ą' 'ż' 'z' 'u' 'w' '▁i' 'zy' 'k' 'ie' 'j' 'cz' 'c' '▁nie' 'i' 'ć' 'ow' '▁na' '▁się' 'nie' '▁jest' '▁k' 'ów' '▁n' 'od' 'ś' '▁od' '▁o' 'ch' 'ne' 'ym' 'ia' 'em' '▁u' '▁wy' '▁prz' '▁po' 'ki' 'ści' 'ze' 'ad' 'ania' 'aw' 'na' 'ach' 'ce' 'ak' == random-30 PL == '▁cy' 'bach' '▁ni' 'ril' 'tu' '▁pref' '▁pod' '▁raz' 'ZE' 'wn' 'god' 'zw' 'plement' 'ww' '▁az' 'icy' 'lane' 'bi' 'oca' '▁Ele' '▁kar' 'pul' '▁pob' '▁Poz' 'gor' 'ła' 'oka' '▁kle' '▁logo' 'ze' == sample OTHER with high counts in both corpora (ambiguous) == '▁t' 'in' 'er' '▁a' 'he' 'on' 're' '▁s' 'en' 'or' 'es' 'an' '▁c' 'is' 'it' 'ou' '▁d' 'al' 'ar' '▁p' '▁f' 'ed' '▁b' '▁m' 'le' 'as' 'ic' '▁h' 'ion' '▁to' 'et' 'el' '▁l' 'ent' 'il' 'ro' '▁re' 'id' '▁I' '▁e'