Llama — Field Evidence
paperUnverified
Llama · Llama · incident
Field note
Language Models are not Equally Robust to Non-Canonical Tokenization across Languages — Despite the existence of exponentially many valid tokenizations for a given string, language models operate on a single canonical sequence deterministically produced by the tokenizer, leaving
collected 2026-07-30original 2026-07-29
Does this shift the US–China race?
Be the first to call it
Impact on the index
Regulation · major
US -110China +30
Directional contribution — before recency decay and per-type diminishing returns. How it works →
Related Front
US vs ChinaLikely
U.S. FrontiervsChina Open-Weight
Likely — capability gap narrowing on common tasks
Letters from the Front
Be the first to file a report.