Shifat Santo

Kotha-1 pipeline

A from-scratch Bengali language-model pipeline, and the label-shift bug that taught it the wrong task.

2026Pipeline open source, fix public, no retrain

i+2the token the buggy model learned to predict instead of the next one

Data collection, deduplication, language ID, a 32K SentencePiece tokenizer and a custom training loop for a LLaMA-style 306M model. The labels were shifted twice, so the model learned to predict the token after next. On a synthetic check the buggy loop scores 0 percent next-token accuracy and the fixed loop 100 percent at identical loss. The fix is public; the model has not been retrained.