What Bengali costs a tokenizer
A paper and an open harness measuring how much more Bengali text costs under multilingual tokenizers.
5 to 9xmore tokens per Bengali word than a dedicated tokenizer
Dedicated Bengali tokenizers spend 1.47 to 1.75 tokens per word. Multilingual ones spend 7.32 to 13.69. The same sentence costs five to nine times more context and money.
Much of the waste traces to one character. The nukta, U+09BC, pushes byte-fallback from about 2 percent of tokens to about 20 percent, and accounts for 89.8 percent of the byte tokens. NFKC normalization cuts it about seventeenfold.
The repo ships the evaluation framework, the failure analysis and 13 trained tokenizers.