Now Open: Legesher's Language Canon ⚖️
Our second open source dataset, holding language-specific decisions for programming vocabulary, and where we need you
Hey y’all 👋,
Today, we’re releasing Legesher’s Language Canon ⚖️

Legesher's Language Canon ⚖️ exists as a centralized standard, created and curated by language communities, to retain the unique qualities of each language while also being able to collaborate with the rest of the world.
The Legesher Language Canon ⚖️ is our second open source dataset holding language-specific iterations of programming keywords, builtins, and exceptions, and documenting where each natural language stands in that process.
Legesher’s Language Corpus 📊 (released Monday) gathers all renderings and examples of core vocabulary across natural and programming languages. Legesher's Language Canon ⚖️ builds on top of the corpus and documents the words and translations that will actually be used by downstream applications (like Legesher).

How the Canon Works ⚖️
In its first instance, without other global examples to build from, each Legesher-supported language will have its canon drafted by machines and prior art, then marked as experimental. From there, native speakers help move each canon through its status from experimental to reviewed to official. Once a language reaches official status, its canon becomes a resource for the community it exists to serve. These canons, published on Legesher's Hugging Face, directly sync with the downloadable Legesher Language Packs used in Legesher’s newest version.
Where We Need You 🫶🏽
It’s well documented that the difficulty in translating programming languages exists in the gray areas, where words are borrowed, created, or derived from English (think “elif” or “lambda,” for example). The canon captures the iterations native reviewers suggest.
This will undoubtedly require much collaboration, and the voices of those who don’t just speak a language other than English, but who know its culture, including the stories often interwoven into a language.
We believe these datasets, while mostly exciting to linguists and data folks at this stage, will become foundational layers of the digital future.
That’s all for today! More releases next week + how YOU can contribute to a more multilingual future in programming.
Please subscribe to our newsletter to stay up to date with our upcoming releases! (You may check out the Legesher Language Canon here on Hugging Face)
All the best,
Josh @ Legesher