Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

301 Commits
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation








talk

A Phonetic Encoding




Overview

Talk is a phonetic front end and tokenization layer for speech and language models. Its machine encoding turns text or IPA into a single code point per sound, so a model gets a finite, normalized phone vocabulary instead of raw IPA or byte-pair fragments that split a pronunciation across several tokens. If you build TTS, ASR, G2P, or any multilingual pipeline that has to reason about how words sound, this gives you one clean token per phone that converts losslessly to and from IPA, plus syllabification for free.

It is also a readable, keyboard-typable romanization. Talk uses the Latin script with diacritics to encode most of Earth's natural language features, enough that you can write every language in the same Latin-oriented system and land close to a realistic pronunciation, including nasalized vowels, tense consonants, clicks, and tones. It reads like the pronunciation guides people already know, so a non-linguist can get the gist without special training.

There are two forms:

  • ASCII: For writing in a text editor without fancy symbols.
  • Machine: A single Hangul code point per sound, for compact tokenization (useful for AI models and lookups).

Talk is not trying to compete with or replace IPA. It is for writing pronunciations in plain ASCII and tokenizing them in a standard way for tools like NLP, AI, and databases. It aims for about 95% accuracy, which is plenty for a tokenizer or a learner, and it speaks IPA fluently for the cases where you need the last 5%.

Libraries

Implementations of Talk, one per language. Each converts between IPA, Talk, and the machine (token) encoding. The TypeScript library also splits words into syllables.

Language Library Status
TypeScript @cluesurf/talk
Python cluesurf-talk

Examples

The machine column is the machine encoding: talk.machine(...) packs each sound into a single Hangul code point (base plus all its modifiers), so a whole word becomes a compact, one-code-point-per-sound string.

ascii machine
txando^ 켶콣뱿켔켤붇
surdjyo^ 콉삟콞켤콃콩붇
HEth~Ah 켾뀟턕눯콀
siqk 콉롟켕켼
txya@+a-a++u 켶콣콩뱌뱶뱷삟
hwpo$kUi^mUno$s 콀콧켬뺄켼뙏띗켒뙏켔뺄콉
sinho^rEsi 콉롟켔콀붇콞뀟콉롟
batO_'aH 켧뱿켶뎏켚뱿켾
aiyuQaK 뱿롟콩삟켛뱿켻
s'oQya&te 콉켚뺏켛콩뱩켶멯
t!arEba 켱뱿콞뀟켧뱿
txhaK!EnEba 켶콣콀뱿켺뀟켔뀟켧뱿
txh~im 켶콣켑롟켒
txy~h~im 켶턹켑롟켒
mh!im 킡롟켒

Encoding

Here are the modifiers on consonants, vowels, and symbols (like punctuation), and how they look in ASCII.

Modifiers

There are a few affixes on consonants and vowels. There are 5 tone affixes (extra low, low, neutral, high, and extra high), which can be combined in standard ways.

category symbol meaning
consonant h~ Aspiration (when added after a consonant)
consonant w~ Labialization (when added after a consonant)
consonant y~ Palatalization (when added after a consonant)
consonant G~ Velarization (when added after a consonant)
consonant Q~ Pharyngealization (when added after a consonant)
consonant ! Makes consonant ejective
consonant ? Makes consonant implosive
consonant @ Makes consonant tense (korean)
consonant . Makes consonant a stop consonant (korean, so when you end on t., it is making the mouth shape of t but not really pronouncing it)
consonant * Click consonant: p* (ʘ), t* (ǀ), k* (ǃ), l* (ǁ), d* (ǂ), c* (𝼊), K* (ʞ)
consonant capital Consonant variant
vowel ^ Stressed vowel
vowel _ Long vowel
vowel ! Short vowel
vowel & Nasal vowel
vowel @ Non-syllablic vowel
vowel capital Vowel variant
vowel $ Vowel variant
vowel + High tone (mandarin high)
vowel ++ Extra high tone
vowel - Low tone (mandarin low)
vowel -- Extra low tone
vowel / Rising tone (vietnamese sắc, mandarin rising)
vowel // Rising tone 2 (vietnamese ngã)
vowel \ Falling tone (vietnamese huyền, mandarin falling)
vowel \\ Falling tone 2 (vietnamese nặng)
vowel /\ Rising falling tone
vowel \/ Falling rising tone (vietnamese hỏi)
symbol = When preceding, can write a literal symbol, like =. is a period, =+ is a plus, etc..

Consonants

This is the full set of IPA consonants we map, in ASCII, and machine (Hangul) form. Every consonant, however modified, is a single Hangul code point.

Note: GitHub markdown doesn't really render the diacritics that nicely, some are misaligned. We will have a font to remedy this for websites.

IPA ascii machine
mh!
m m
ɱ̊ m~h!
ɱ m~
n
n̪̊ n~h!
n~
nh!
n n
n̠̊ nh!
n
ɳ̊ Nh!
ɳ N
ɲ̊ ny~h!
ɲ ny~
ŋ̊ qh!
ŋ q
ɴ̥ qh!
ɴ q
p p
b b
p~
b~
t
d
t~
d~
t t
d d
ʈ T
ɖ D
c ky~
ɟ gy~
k k
ɡ g
q K
ɢ g
ʔ '
s~
z~
s s
z z
ʃ x
ʒ j
ʂ X
ʐ J
ɕ xy~
ʑ jy~
ɸ F
β V
f f
v v
θ̼ c
ð̼ C
θ c
ð C
θ̠ c
ð̠ C
ɹ̠˔ u$
ɻ˔ u$
ç hy~
ʝ y
x H
ɣ G
χ H
ʁ G
ħ Hh~
ʕ Q
h h
ɦ hh~
β̞ V
ʋ V
ð̞ C
ɹ u$
ɹ̠ u$
ɻ u$
j y
ɰ W
ⱱ̟ V
V
ɾ̥ rh!
ɾ r
ɽ̊ Rh!
ɽ R
ɢ̆ g
ʙ̥ bbh!
ʙ bb
rh!
r r
r
ɽ̊r̥ Rh!rh!
ɽr Rr
ʀ̥ GGh!
ʀ GG
ɬ̪ S~
ɬ S
ɮ Z
ʎ̝ ly~
l~
lh!
l l
l
ɭ̊ Lh!
ɭ L
ʎ̥ ly~h!
ʎ ly~
ɺ̥ lh!
ɺ l
ʍ wh!
w w
ɥ yw~
ɧ H
ɫ lG~
ɓ b?
ɗ d?
ʄ g?y~
ɠ g?
ʛ g?
ɓ̥ b?h!
ɗ̥ t?
ʄ̥ g?y~h!
ɠ̊ g?h!
ʛ̥ g?h!
p!
t!
ʈʼ T!
ky~!
k!
K!
f!
s!
ʂʼ X!
ɕʼ xy~!
H!
χʼ H!
ɸʼ F!
θʼ c!
ʃʼ x!
ʄ̊ g?y~h!
ɬʼ S!
ʘ p*
ǀ t*
ǃ k*
𝼊 c*
ǂ d*
ʞ K*
ǁ l*

Vowels

This is the full set of IPA vowels we map. Every vowel, with any tone, length, stress, or nasalization, is a single Hangul code point.

IPA ascii machine
i i
y i$
ɨ i$
ʉ u
ɯ O
u u
ɪ I
ʏ i$
ʊ O
e e
ø a$
ɘ I
ɵ O
ɤ O
o o
e
ø̞ a$
ə U
ɤ̞ O
o
ɛ E
œ e$
ɜ O
ɞ U
ʌ U
ɔ o$
æ A
ɐ a 뱿
a a 뱿
ɶ e$
ä a 뱿
ɑ a 뱿
ɒ a 뱿

Every vowel takes the same modifier set (stress, length, nasal, tone, etc.). Here are all the combos for the letter a, and the same pattern applies to all vowels.

IPA ascii machine
æ A
œ a$
ˈa a^
a_
a!
a&
a@
a+
a++
a-
a--
a˧˥ a/
a˩˥ a//
a˥˧ a\
a˥˩ a\\
a˩˥˩ a/\
a˥˩˥ a\/

Note: Exact tone sequences can be represented with sequences like a+a++a--, where one vowel is spread across multiple tones. But common tones, across languages, can take advantage of the shortened syntax/encoding.

IPA Diacritics

The IPA defines a set of diacritics that modify a base symbol, grouped by type below. The state column says whether Talk maps the diacritic to a phonetic feature, or ignores it, and talk is the resulting Talk affix when it is mapped.

symbol type meaning state talk
◌ʰ secondary articulation aspirated mapped h~
◌ʷ secondary articulation labialized mapped w~
◌ʲ secondary articulation palatalized mapped y~
◌ˠ secondary articulation velarized mapped G~
◌ˤ secondary articulation pharyngealized mapped Q~
◌ⁿ secondary articulation nasal release mapped n
◌̥ phonation voiceless mapped h!
◌̬ phonation voiced mapped
◌̤ phonation breathy voice ignored
◌̰ phonation creaky voice ignored
◌̹ phonation more rounded ignored
◌̜ phonation less rounded ignored
◌̟ relative articulation advanced ignored
◌̠ relative articulation retracted ignored
◌̈ relative articulation centralized ignored
◌̽ relative articulation mid-centralized ignored
◌̝ relative articulation raised ignored
◌̞ relative articulation lowered ignored
◌̘ tongue root advanced tongue root ignored
◌̙ tongue root retracted tongue root ignored
◌̪ articulation dental mapped ~
◌̺ articulation apical ignored
◌̻ articulation laminal ignored
◌̃ nasality nasalized mapped &
◌̩ syllabicity syllabic ignored
◌̯ syllabicity non-syllabic ignored
◌ː length long mapped _
◌ˑ length half-long ignored
◌̆ length extra-short ignored
◌̚ release no audible release mapped .
◌ʼ airstream ejective mapped !
◌ˈ stress primary stress mapped ^
◌ˌ stress secondary stress ignored
◌˥ tone extra high mapped ++
◌˦ tone high mapped +
◌˧ tone mid mapped
◌˨ tone low mapped -
◌˩ tone extra low mapped --
◌‖ prosody major phrase break ignored
◌↗ prosody global rise ignored
◌↘ prosody global fall ignored

On the ignored ones. IPA itself is not perfectly exact. Real speech has far more subtle variation than any discrete character set can capture, and the ignored diacritics above (relative articulation, breathy/creaky voice, tongue-root position, half-long, and so on) mark distinctions so fine-grained that they are hard to even hear reliably, hard to transcribe consistently, and hard to reproduce the same way across recordings. For a discrete model of sound aimed at AI and coding, they are mostly unnecessary, so Talk lets them fall through rather than inventing symbols for detail it cannot use.

Syllables

Talk also splits a word into syllables, each a sequence of onset, nucleus, and coda clusters. IPA and X-SAMPA give you symbols and stop there, so this comes built in, which the TTS, speech-recognition, and language-learning use cases all need.

Some rarer sounds are not yet handled by the syllable system. The current gaps are listed in base/syllable/unsupported.csv (dentals, retroflex plosives, implosives, ejectives, and clicks).

Token encoding

Which Chinese characters actually render in the base fonts that ship on ordinary computers, rather than showing the "missing-glyph" box basically (so you can actually see the glyphs not missing font glyphs, most of the time). Coverage is read from each font's cmap (the ground truth for whether a glyph exists), then intersected per script: Simplified with Simplified, Traditional with Traditional.

The base list is 84,107 Han ideographs used in Chinese, taken from the Unihan database by Chinese source or reading.

Font System Script Characters
PingFang macOS SC + TC 30,979
Microsoft YaHei Windows SC 27,761
SimSun Windows SC 27,565
Microsoft JhengHei Windows TC 27,534
MingLiU Windows TC 27,531

Characters that render on every common desktop of a script (the intersection):

Set Fonts Characters
Simplified common PingFang, YaHei, SimSun 27,565
Traditional common PingFang, JhengHei, MingLiU 27,528

Per-font and intersection lists live in base/symbol/chinese/.

License

MIT

ClueSurf

Made by ClueSurf, meditating on the universe ¤. Follow the work on YouTube, X, Instagram, Substack, Facebook, and LinkedIn, and browse more of our open-source work here on GitHub.