🎤 New Session Confirmed for Unicode Technology Workshop 2026 Thirty Years of Line Breaking - hosted by Robin Leroy In 1998, Michel Suignard gave a talk at the 12th Internationalization and Unicode Conference titled “Worldwide Typography and How to Apply JIS X 4051-1995 to Unicode”. This synthesis of Japanese and French typographic rules was the inception of the Unicode Line Breaking Algorithm, which is used nearly everywhere that text is displayed to determine where lines can end. This talk by the current maintainer of the line breaking algorithm will present the history of that algorithm, from the ancestor standard JIS X 4051-1995 and the first version in Unicode 3.0 with 29 line breaking classes and 23 rules, to the current state with nearly 50 classes and more than 40 rules, with more being added year after year to improve the handling of the typographic rules of Hebrew (2012), Balinese (2023), French (2023, 2026), Simplified Chinese (2025), etc. The talk will review the evolution of implementation strategies over the years, from a simple pair table for JIS X 4051-1995 to a space-aware one in Unicode 3, to the abandonment of pair tables in 2017, and finally to the state tables for finite automata which ICU has used for the past 23 years and which are being prepared for publication by the UTC. This historical approach is of practical interest for today’s implementers, as many of the approaches that have been abandoned in the past are rediscovered, and their shortcomings forgotten, leading to incorrect implementations. The talk will also discuss some questions in theoretical computer science that arise in the maintenance of the state-of-the-art implementation based on finite automata. Secure your seat NOW! 🎟️ https://lnkd.in/g5_jMSZ6 #Unicode #i18n #ICU
Unicode Consortium
IT Services and IT Consulting
South San Francisco, California 3,020 followers
Everyone in the world should be able to use their own language on phones and computers. Unicode makes it possible.
About us
The Unicode Consortium is the premier standards body for the internationalization of all software and services. Today, most people take for granted that their phones and computers can display text in any of hundreds of languages, handle dates, times, numbers, currencies in familiar local formats, and faithfully send emoji-laden messages to friends using any device. That wasn’t always the case, and it didn’t happen by accident. For 30 years, the Unicode Consortium has coordinated the efforts of a worldwide team of volunteer programmers and linguists to standardize, evolve, and maintain a global software foundation that allows virtually every computer system and service to help people connect using their native language. This has real world consequences. Today’s global economy runs on networks that reach billions of users around the globe. Databases, commerce engines, websites and shipping systems handle local names, addresses, and text in hundreds of languages from Latin to Cyrillic to Hindi to Japanese – all thanks to Unicode standards and code. The most compelling rationale was interoperability across the world. Unicode Consortium Quick Facts - Founded in 1988, incorporated in 1991 - Public benefit, 501(c)3 non-profit organization - Open source standards, data, and software development - Orchestrates the contributions of 100s of professionals, expert volunteers, and language experts - 30+ organizational members across corporate, academic, and governmental institutions - Funded by membership dues and donations Unicode was founded on the basis that: - Local solutions require global collaboration - Interoperability across platforms serves you – and the greater good - Transparency and open source ensure: Reliability — Security — Stability - Localization respects and empowers users Latest news and information: https://www.unicode.org/consortium/general-contact-signup.html
- Website
-
http://home.unicode.org
External link for Unicode Consortium
- Industry
- IT Services and IT Consulting
- Company size
- 2-10 employees
- Headquarters
- South San Francisco, California
- Type
- Nonprofit
- Founded
- 1991
Locations
-
Primary
Get directions
611 Gateway Blvd
Suite 120
South San Francisco, California 94080, US
Employees at Unicode Consortium
Updates
-
🗣️ New Tutorial Confirmed for Unicode Technology Workshop 2026 Macrolanguages, what is a language and not a language - hosted by Conrad Nied Often we group languages into macrolanguages like Chinese (zh), Malay (my) and Cree (cr). Depending on the context it is useful to have these groupings. However, their constituent languages are often not mutually intelligible and in can cause communities to reject translations and services translated too broadly. In this tutorial, we will discuss language families, dialects, and all that are in between and how it impacts your localization use cases. Secure your seat NOW! 🎟️ https://lnkd.in/g5_jMSZ6 #Unicode #i18n #l10n #ICU #CLDR
-
-
📚 New Session Confirmed for Unicode Technology Workshop 2026 Digital text, legacy tech and the many meanings of identity: hosted by Bríd-Áine Parnell For many internet users, the mission to digitally encode the world’s languages can appear complete. Yet, while Unicode has accomplished a robust foundation for digital text encoding of the vast majority of the world’s scripts, in practice, even languages with globe-spanning communities and prominent status are not always adequately represented. Users have to perform an anglicisation of their name in order to gain access to digital services, stripping out the accents so that their name is “acceptable” to the system. Grounded in a recent public campaign in Ireland advocating for the systemic acceptance of Irish-language names, this study presents survey data assessing the social, emotional and cultural impacts of this name misrepresentation, and considers the implications for citizens as governments scale digital identity systems using the same structures. For Unicode in the World, this is a timely illustration of how Unicode offers tools for multilingual representation, but implementation still requires key drivers – commercial, ideological, or regulatory – to ensure adoption in every situation. It allows policymakers, localisation professionals and product managers to understand Unicode as part of a broader infrastructure of standards and requirements, and invites scholars and advocates to consider what it might take to contest the soft-power coloniality of an Anglocentric digital world. Secure your seat NOW! 🎟️ https://lnkd.in/g5_jMSZ6 #Unicode #DigitallyDisadvantagedLangues #DigitalHumanities #scripts
-
-
⏰ New Tutorial Confirmed for Unicode Technology Workshop 2026 Times are changing: how CLDR handles time zone updates This presentation, hosted by Robert Bastian, will explore the practical intersection of code and linguistics by detailing how the Unicode Common Locale Data Repository (CLDR) models time zone data. We will review the structural mechanics of CLDR’s time zone architecture, including the relationships between metazones, generic location formats, and exemplary cities. Attendees will see how CLDR translates raw tz database identifiers into localized, human-readable strings, and why understanding this data mapping is highly beneficial for creating seamless multi-locale applications. We will conclude by looking ahead at impending legislative shifts, such as British Columbia’s transition to permanent daylight saving time in November, to illustrate how these best practices apply to real-world updates. Secure your seat NOW! 🎟️ https://lnkd.in/g5_jMSZ6 #Unicode #CLDR #ICU #i18n
-
-
⚡️ New Tutorial Confirmed for Unicode Technology Workshop 2026 Paths to Proliferation - hosted by Mark Jamra, Neil Patel, and Morgane Pierson Through our own experiences of working with neographies over the past decade, we have seen a multitude of ways in which they evolve, are developed and deployed. While the motivation behind the creation of these novel writing systems is similar, how each proliferates amongst a community is different. We will look at examples of various scripts, including Garay and others, and think together about the research, community outreach, design and development required, and the sociopolitical dynamics that can affect those efforts. The presentations will be followed by Q & A, open discussion and, if desired, critiques of tutorial participants’ projects in this vein. Secure your seat NOW! 🎟️ https://lnkd.in/g5_jMSZ6 #Unicode #DigitallyDisadvantagedLanguages #fonts #scripts #encoding #digitalhumanities
-
-
📢 New Plenary Confirmed for Unicode Technology Workshop 2026 Encoding as an artistic practice: The ABCC of CACB Visual Communication for a Contemporary Art Center might seem an unlikely entry point into encoding, but it is much more relevant than you may think. For ten years, Coline Sunier and Charles Mazé were artists in residence at the Contemporary Art Center of Brétigny (CAC Brétigny), in the Paris suburbs. During this period they developed the museum’s visual identity, and their graphic work is now part of its collections. Over those years, they progressively enriched a digital font, LARA, incorporating signs found in the museum, in its surroundings, and in the works of contemporary artists featured in its exhibitions. Each sign is adapted typographically and correctly encoded, whether it be letters, symbols, emojis, or hieroglyphs. Conducted in the art world, this project deals precisely with the act of encoding, and it has opened up the question of Unicode to an audience previously unaware of its existence. 📅 Early bird registration ends August 7! Secure your seat NOW! 🎟️ https://lnkd.in/g5_jMSZ6 #Unicode #scripts #encoding #fonts #typedesign
-
-
📣 New Session Confirmed for Unicode Technology Workshop 2026 The Tokenizer Comedy: A Catalog of Joyful Failures in LLM Unicode Processing As Large Language Models (LLMs) scale globally, the industry is fiercely focused on multilingual success. This talk, hosted by Tiffa Foster, offers the opposite: a meticulous, empirical, and highly humorous catalog of catastrophic failures. When you push LLMs past standard ASCII and into the deep waters of complex Unicode (Enclosed Alphanumerics, Mathematical Double-Struck, Zero-Width Non-Joiners), the ‘superintelligence’ shatters. This 45-minute presentation details a months-long, exhaustive stress-test of 33 major models via the Kaggle Benchmark SDK, documenting the exact vectors where tokenizers fail, hallucinate, or revert to mimicry. Rather than proposing a single theoretical ‘fix,’ this session serves as a practical “What Not To Do” guide for engineers, researchers, and typographers. We will explore the beautiful, frustrating reality that while we are encoding the world, the machines are currently just reading the fonts. 📅 Secure your seat NOW! 🎟️ https://lnkd.in/g5_jMSZ6 #Unicode #AI #i18n
-
-
🎙️ New Session Confirmed for Unicode Technology Workshop 2026 Ahead of the Glyph: Predictive Intelligence for Multilingual Text Engines Every layer of today’s internationalization stack waits to be told what to do. Developers pick an ICU data configuration at build time. The engine fails, then fetches, when it misses a data file. Software guesses a text’s encoding only after the bytes are already garbled. Language is detected one call at a time. The result is an i18n stack that is reactive, manually configured, and increasingly strained by the messy, mixed-script, AI-generated text users now paste into every application. This session, hosted by Abhi Bhatnagar and Nisarg Shah presents a unifying approach: a single intelligence layer that sits in front of ICU and Adobe’s Globalization Library and reads the incoming text once to get ahead of every downstream Unicode decision. Attendees will leave with a concrete architecture for a predictive i18n layer, the trade-offs it introduces, and a set of open problems to collaborate on. 📅 Early bird registration ends August 7! Secure your seat NOW! 🎟️ https://lnkd.in/g5_jMSZ6 #Unicode #AI #fonts #i18n
-
-
✨Unicode would like to congratulate Salvatore Rinchiera on becoming a Bronze Sponsor. 🥉 https://lnkd.in/eqtaC-Dv ❗Interested in adopting a character? Learn more: https://lnkd.in/eb5ER4C8 #UnicodeAAC #Unicode #TechForGood #CharacterAdoption
-