Wednesday, September 16, 2026

Announcing the Unicode® Standard, Version 18.0

Version 18.0 of the Unicode Standard is now available. The Unicode Standard is the foundation for all digital communications, and this new version supports a wider range of text encoding needs. This major update includes new characters and code charts, updated data files, and updated specifications that define many fundamental aspects of text processing.

This version adds 13,007 new characters, including nine new emoji characters as well as many other characters and symbols, bringing the total number of encoded characters to 172,808.

Among the most anticipated new characters are three new currency symbols:

  • U+20C2 RUFIYAA SIGN
  • U+20C3 UAE DIRHAM SIGN
  • U+20C4 OMANI RIAL SIGN

Each of these symbols was authorized for public use by the respective monetary authority over a year ago, but usage has been hindered by lack of a standardized encoding. With the release of Unicode 18.0, vendors are now able to implement support for these symbols.

The largest set of new characters is for the historical Seal (or “Small Seal”) script. This set of 11,328 ideographic characters has important cultural significance in China, dating back to the Qin Dynasty (around 200 BCE). Another new cultural heritage script from China is Jurchen, used in northeastern China during the Jin Dynasty.

See the delta code charts for details on all the new scripts and characters. For additional details regarding new emoji, see Emoji Recently Added, v18.0

No new algorithms have been introduced in this release, but new data files have been added for Seal and other East Asian scripts, along with a new Unicode Standard Annex documenting these new data files: UAX #60, Data for East Asian Scripts.

Conformance language related to variation selectors has been updated, making clearer which uses of variation selectors are or are not conformant to the Unicode Standard. 

Recommendations for implementations were also added to make non-conformant uses of variation selectors visible in text. This is important as research has shown that sequences of invisible variation selector characters can be used to attack modern AI applications.

For complete details on Unicode Version 18.0, see https://www.unicode.org/versions/Unicode18.0.0/




Friday, September 4, 2026

Unicode CLDR 49 Alpha available for testing

Building construction emoji imageThe Unicode CLDR 49 Alpha is now available for integration testing.

CLDR provides key building blocks for software to support the world's languages (dates, times, numbers, sort-order, etc.) For example, all major browsers and all modern mobile phones use CLDR for language support. (See Who uses CLDR?)

The alpha has already been integrated into the development versions of ICU and ICU4X. We would especially appreciate feedback from non-ICU consumers of CLDR data on Migration issues. Feedback can be filed using CLDR Tickets.

Some of the most significant changes in this release include the following (for more detail, see the CLDR 49 release note page):
  • Updated for Unicode 18 including annotations for the new emoji, new scripts, etc.
  • Reflects the most recent updates to external standards and data sources, such as the language subtag registry, UN M49 macro regions, ISO 4217 currencies, etc.
  • New formatting options:
    • New date & time formatting options including:
      • Localized patterns for gluing date & timezoneAppend Items — e.g., Sept 3, EST
      • Ordinal days in dates — e.g., Sept 3rd
      • Customizing numeric separators for dates and times in patterns: 3-10-2031 → 3/10/2031
      • UTC timezone display patterns
      • Dual Standard/Daylight format — Sept 3, UTC+3/+2
      • Structure for preventing digit-digit merges — e.g., '2026/1/29 GMT-817时'
      • Additional skeleton-patterns added for flexible and interval date formats
    • Nested bracket replacement — for constructing locale names with parts that have parentheses, e.g., ”birmanês (Mianmar [Birmânia])”
    • Additional localized locale display name options - see keys
    • New and updated plural and ordinal rules
    • New and improved rule-based number formatting (spellout) rules for many new locales
    • New units — poundal, dyne, milliinch (US mil)
  • Many enhancements of the CLDR specification (LDML) will be in the CLDR 49 Beta (September 23rd).

Locale Coverage Levels

Level Count Regional
Variants
Usage
Modern 100 399 Suitable for full UI internationalization
Moderate 12 14 Suitable for “document content” internationalization, eg. in spreadsheet
Basic 73 99 Suitable for locale selection, eg. choice of language on mobile phone

This is the second release where the new CLDR Organization process is in place for DDL languages. As a result, several locales were able to reach higher levels or had substantial contributions:
  • Adyghe, Kabardian: Adyghe Language and Literature Association
  • Laz: Laz Instittue
  • Ligurian: Council for Ligurian Linguistic Heritage
  • Mara: Mara Language Preservation
  • Coptic: St. Shenouda Coptic Society
± New Level Locales
📈 Modern Akan
📈 Moderate Breton, Coptic
📈 Moderate* Romansh, Shan, Tigrinya
📈 Basic Adyghe, Central Kurdish, Colognian, Kabardian, Kikuyu, Kʼicheʼ, Ladin, Prussian, Qʼeqchiʼ, Sunwar (Sunuwar)
📉 Basic* Tajik, Bashkir, Interlingua, Sardinian, Faroese, Venetian

Note: Each release, the number of items needed for Modern and Moderate increases. So locales without active contributors may drop down in coverage level.

For the details, see the CLDR 49 release note page, which has information on accessing the data, reviewing charts of the changes, and — importantly — will cover Migration issues.