A curated set of ~4,540 simplified Chinese characters for advanced language learners.
The MteH corpus is designed as an "endgame corpus" for advanced students. Basically, if you learn these characters, you're practically "done for life" studying simplified Chinese characters (congratulations!). Obviously, there are more simplified Chinese characters than this (in proper nouns, scientific terms, chengyu, Chinese history, online usernames, etc.), but at a certain point you've got to draw the line and say "this is my endgame".
Currently, MteH focuses entirely on simplified Chinese characters, especially those you’ll encounter in mainland China and in HSK exams.
- MteH corpus (v0.1.3) (plain text)
- Handwriting practice (PDFs to print out)
- Extra characters (good to know, but not part of MteH)
There is also an Anki Deck (here) which should already work, but should be thought of as a work-in-progress. (On a computer, AnkiDraw allows you to handwrite. On AnkiDroid, the in-built whiteboard feature enables handwriting.)
The MteH corpus is built to minimize "missing" characters; any characters not included are extremely rare or niche. The current update merges the following diverse corpora. Not all of the characters are included (there are way too many); omitted characters are documented in the missing chars report).
| # | Corpus | #chars | Source / Reference |
|---|---|---|---|
| 1 | HSK 1.0 | 2,866 | pre-2010, 11 levels |
| 2 | HSK 2.0 | 2,663 | post-2010, 6 levels |
| 3 | HSK 3.0 | 3,000 | 2021 version 3.0 standards, 9 levels |
| 4 | HSK 3.1 | 3,088 | 2025 version 3.0 standards, 9 levels |
| 5 | TOCFL | 3,027* | Taiwan's TOCFL 3100 + 33 traditional chars |
| 6 | K-5 | 1,817 | K-5 word frequency |
| 7 | 通用规范汉字表 | 3,500 | Ministry of Education (2013) |
| 8 | 现代汉语常用字表 | 3,500 | Ministry of Education (1988) |
| 9 | 普通话水平测试 | 3,788 | Putonghua Proficiency Test for native Mandarin fluency |
| 10 | Taiwan MoE | 4,661* | Taiwan Ministry of Education 常用字 |
| 11 | primary school | 2,999 | Zhang et al. (2024) |
| 12 | 语文 | 3,500 | China compulsory education 语文 (2022) |
| 13 | Singapore | 1,655 | Singapore primary schools (2015) |
| 14 | age of acquisition | 4,237 | Cai et al. (2022) |
| 15 | psycholinguistic | 3,253 | Chang et al. (2016) |
| 16 | Heisig | 3,018 | Heisig & Richardson, Remembering Simplified Hanzi I–II |
| 17 | Hoenig | 2,177 | Learn & Remember 2,178 Characters and Their Meanings |
| 18 | Jun Da | 4,486* | modern Chinese corpus |
| 19 | SUBTLEX | 4,464* | film and TV subtitle corpus |
| 20 | Tsai | 4,329* | Usenet newsgroups (1993-1994) |
| 21 | CKIP | 3,392* | CKIP (Chinese Knowledge and Information Processing) research group |
| 22 | Wikipedia | 3,476* | Chinese Wikipedia |
| 23 | Chinese.SE | 4,973* | Chinese Stack Exchange (Jan 2026) |
| 24 | classical | 1,968* | prior to the end of the Han dynasty |
| 25 | THUOCL | 3,469* | mostly Sogou webpages |
| 26 | Leeds | 4,230* | Internet corpus |
| 27 | BLCU | 5,000* | "balanced", written Chinese |
| 28 | LWC | 4,130* | Sina Weibo |
| 29 | food | 1,182 | food-related terms |
| 30 | species | 4,086 | species names |
| 31 | Chinese surnames | 1,745 | 1,807 Chinese surnames |
| 32 | Chinese names | 2,269 | 1,200,000 Chinese names |
| 33 | city-geo | 1,277 | mainland China city terms |
| 34 | company | 5,403* | company proper nouns |
| 35 | med-orgs | 4,826 | medical organizations |
| 36 | MCT | 1,180 | Medical Chinese Test |
| 37 | BCT | 1,752 | Business Chinese Test |
| 38 | chengyu convention | 2,226 | characters in "chengyu convention" chengyu |
| 39 | Xinhua | 5,357 | Xinhua chengyu and xiehouyu |
Those marked * have extraction steps (documented in their respective readmes), typically this involves the selection of top-N words/characters, conversion from traditional to simplified, or sporadic bugs.
Characters are ordered in Unicode order, grouping visually or structurally related forms as much as possible.
MteH also incorporates:
- Character structure data and character drawings from Make Me a Hanzi and cjkvi-ids
- Frequency data from Jun Da’s modern corpus
- Images from Pexels, Wikimedia, etc.
Statistics and debug reports: missing chars; corpus histogram; debug; modifications; syllables.
© 2025 Rebecca J. Stones
Licensed under CC BY-SA 4.0