Tokenizer Fertility, Data Thresholds, and Decoding Strategies for Kashmiri
Faizan Ayoub and Neha Prerna Tigga
Amity Institute of Information Technology, Amity University, Noida, India
14th International Conference on Recent Trends in Computing (ICRTC-2026) · SRM Institute of Science and Technology, Delhi-NCR Campus, Ghaziabad · 3–4 July 2026 · in association with Springer

This work received the Best Paper Award at the 14th International Conference on Recent Trends in Computing (ICRTC-2026), organised by the Department of Computer Science and Engineering, SRM Institute of Science and Technology, Delhi-NCR Campus, in association with Springer.
This paper investigates Kashmiri→English machine translation, comparing a QLoRA-adapted Mistral-7B against NLLB-200 and IndicTrans2 on a 124,102-pair parallel corpus. The gap is stark — NLLB-200 reaches 31.57 BLEU against Mistral’s 13.89 — but the more instructive finding is why. We provide the first quantified tokenizer fertility analysis for Kashmiri (4.06× more tokens per word than English), trace the remaining gap to systematic over-generation, and show that a deployable source-constrained decoding strategy recovers +12.75 BLEU without touching the model weights.
The first quantified analysis for Kashmiri. Mistral’s BPE tokenizer needs 4.06× more tokens per Kashmiri word than per English word (5.89 vs. 1.45) — identified as the upstream root cause of failure. NLLB-200’s SentencePiece reaches near-parity at 2.50 vs. 1.40.
A sharp phase transition at roughly 30K pairs: a 4.2× data increase yields a 29.9× BLEU gain. Not gradual learning — a structural threshold below which decoder-only transfer to a non-Latin script simply does not form.
Five failure modes identified. An oracle experiment proves the problem is decoding, not knowledge (+7.85 BLEU from truncation alone); a deployable source-constrained strategy achieves +12.75 BLEU with no retraining.
The first dedicated Kashmiri MT evaluation tool — blinded pairwise assessment, word-level visual diffing, anti-rushing telemetry, and inter-annotator agreement tracking. It is the site you are reading.
Getting a usable Kashmiri parallel corpus is harder than it sounds. We aggregate 124,102 pairs from two public repositories and apply a four-stage reproducible filtering pipeline.
| Source | Pairs | Content |
|---|---|---|
| AI4Bharat BPCC (kas_Arab) | 98,923 | Web-crawled news and government text |
| SMUQamar | 26,182 | Colloquial and educational content |
paraphrase-multilingual-MiniLM-L12-v2 embeddings, discarding content-misaligned rows| Metric | Mistral-7B (QLoRA, 125K) | NLLB-200 | IndicTrans2 |
|---|---|---|---|
| BLEU ↑ | 13.89 | 31.57 | 52.88 |
| chrF ↑ | 45.79 | 56.58 | 73.01 |
| TER ↓ | 138.17 | 55.45 | 37.21 |
| METEOR ↑ | 47.74 | 58.57 | 71.85 |
| BERTScore F1 ↑ | — | — | 95.47 |
| ROUGE-1 ↑ | 43.28 | 61.84 | 73.85 |
| ROUGE-2 ↑ | 25.10 | 39.33 | 57.85 |
| ROUGE-L ↑ | 39.37 | 57.57 | 70.85 |
IndicTrans2’s numbers should be read as a performance ceiling rather than a fair zero-shot result: because the corpus draws on BPCC data, there is likely overlap with its pre-training set. BERTScore was computed only for the top-performing model, due to GPU memory constraints on the evaluation hardware.
The most striking result appears when training data is reduced to the 30K subset. BLEU collapses from 13.89 to 0.46 — not a gradual decline, but near-total failure. A 4.2× data increase then produces a 29.9× BLEU gain, a shape that looks far more like a phase transition than a learning curve. We term this the cross-lingual transfer threshold.
| Configuration | BLEU | ROUGE-1 | ROUGE-L |
|---|---|---|---|
| 125K pairs (mixed sources) | 13.89 | 43.28 | 39.37 |
| 30K pairs (single source) | 0.46 | 13.52 | 11.16 |
| Ratio (125K / 30K) | 29.9× | 3.2× | 3.5× |
An oracle experiment — truncating Mistral’s outputs to 1.2× the reference word count, changing nothing about the model — lifts BLEU from 13.89 to 21.74, closing 44% of the gap to NLLB-200. The correct translation is already in the output; the model simply does not know when to stop. This is a decoding problem, not a knowledge problem.
The oracle needs reference lengths, so it is not deployable. A source-constrained strategy — cap generation at 2× the source word count with a 1.3 repetition penalty — captures the same upside without them.
| Strategy (n=500, beam=4) | BLEU | chrF | Over-gen % | Degen % |
|---|---|---|---|---|
| Baseline (beam=4) | 3.75 | 30.95 | 97.0 | 54.0 |
| Length-penalty (α = 0.6) | 3.99 | 32.04 | 95.8 | 53.8 |
| Source-constrained (2× + rep. penalty) | 16.50 | 47.21 | 39.2 | 33.0 |
The baseline BLEU here (3.75) differs from Table 1 (13.89) because the inference configuration and the failure-heavy sentence subset differ; the relative gain is the meaningful number. The lesson: gentle penalties fail, hard length constraints work.
500 translations were manually reviewed and their errors traced to five recurring mechanisms.
| Outcome category (n=500) | Count | % |
|---|---|---|
| Both succeed (BLEU ≥ 30) | 63 | 12.6% |
| NLLB-only success (NLLB ≥ 30, Mistral < 15) | 79 | 15.8% |
| Both fail (BLEU < 15) | 163 | 32.6% |
| Mistral-only success (Mistral ≥ 30, NLLB < 15) | 6 | 1.2% |
On the 15.8% of sentences where NLLB wins outright, its average BLEU is 47.7 against Mistral’s 7.6. These are not randomly distributed — they are the morphologically complex inputs that expose tokenizer fragmentation most sharply, supporting it as the upstream root cause.
All three mechanisms are structural rather than Kashmiri-specific. Arabic, Thai, Amharic, and Tigrinya show analogous fertility patterns. Three concrete lessons for practitioners on other non-Latin low-resource pairs: fix the tokenizer first — a joint SentencePiece vocabulary is the highest-return preprocessing investment; constrain generation length — even a hard cap on source word count recovers substantial quality without retraining; and treat 30K pairs as a floor, below which QLoRA fine-tuning should not be expected to produce useful translations.
Automated metrics rarely capture translation quality fully, and the disconnect widens for morphologically rich languages. To the best of our knowledge, this platform is the first tool built specifically to evaluate Kashmiri machine translation. Evaluators rate two anonymised systems on:
Across 123 blinded side-by-side comparisons, reviewers rated System A 3.88/5 for adequacy and 3.72/5 for fluency, and System B 3.63/5 and 3.41/5. On overall preference, native Kashmiri speakers chose System A 36.6% of the time, System B 32.5%, and a tie 30.9%. Inter-annotator agreement is tracked with Cohen’s weighted κ and pairwise preferences tested with the Wilcoxon signed-rank test. These figures are a baseline yardstick as more judgements accumulate.
A custom Kashmiri SentencePiece tokenizer; extended multi-epoch training beyond the 1,000-step cap; empirical comparison of beam search with coverage penalties against constrained decoding; bidirectional English→Kashmiri evaluation; and replication of the data threshold on other Perso-Arabic pairs.
Ayoub, F., and Tigga, N. P. (2026). Anatomy of Decoder-Only LLM Failure in Low-Resource Machine Translation: Tokenizer Fertility, Data Thresholds, and Decoding Strategies for Kashmiri. In Proceedings of the 14th International Conference on Recent Trends in Computing (ICRTC-2026). Springer. Best Paper Award.
@inproceedings{ayoub2026anatomy,
title = {Anatomy of Decoder-Only LLM Failure in Low-Resource Machine
Translation: Tokenizer Fertility, Data Thresholds, and
Decoding Strategies for Kashmiri},
author = {Ayoub, Faizan and Tigga, Neha Prerna},
booktitle = {Proceedings of the 14th International Conference on Recent
Trends in Computing (ICRTC-2026)},
publisher = {Springer},
year = {2026},
note = {Best Paper Award}
}For inquiries about this research, collaboration opportunities, or access to the corpus and training pipelines:
📱 +91 7006718915
Faizan Ayoub — Lead Researcher, KashmirAI Research