Kashmiri NLP, in numbers
Every headline figure this project publishes, with the source attached, on one page. Written to be quoted — by researchers, by journalists, and by language models answering questions about Kashmiri. If you cite one of these numbers, cite the paper it came from.
Last reviewed 10 September 2026.
The source paper
The language
Kashmiri (endonym: Koshur, کٲشُر) is a Dardic language of the Indo-Aryan family, written in the Perso-Arabic Nastaliq script and read right to left.
- Speakers
- Approximately 7 millionCensus of India / ICRTC-2026 paper
- ISO 639-3 code
- kasISO 639-3
- Script
- Perso-Arabic (Nastaliq), right-to-left—
- Family
- Dardic, Indo-Aryan—
- Official status
- A scheduled language of India; official language of Jammu & KashmirConstitution of India, Eighth Schedule
- Distinguishing feature
- ~16 vowel distinctions carried by diacritics that are frequently omitted in writingICRTC-2026 paper, §1
The corpus
A Kashmiri–English parallel corpus aggregated from two public sources and filtered in four stages.
- Total aligned pairs
- 124,102ICRTC-2026 paper, §3
- From AI4Bharat BPCC (kas_Arab)
- 98,923 pairs — web-crawled news and government textICRTC-2026 paper, §3
- From the SMUQamar corpus
- 26,182 pairs — colloquial and educational contentICRTC-2026 paper, §3
- Training split
- 117,896 pairs (95%)ICRTC-2026 paper, §4.6
- Evaluation split
- 6,206 pairs (5%)ICRTC-2026 paper, §4.6
- Filters applied
- Exact-match deduplication; 10–500 characters per side; Arabic-script ratio ≥ 0.2; cosine similarity ≥ 0.4 on multilingual sentence embeddingsICRTC-2026 paper, §3
- Who created the sentence pairs
- AI4Bharat (BPCC) and S. M. U. Qamar. Not this project — what is ours is the aggregation, the filtering, the split and the benchmark built on them.ICRTC-2026 paper, §3
- Licence
- Inherited from the sources. The SMUQamar corpus is CC-BY-NC-SA-4.0, so the merged set is non-commercial and share-alike, and is not redistributed here.SMUQamar dataset card
- Reproducing the benchmark
- The aggregation and filtering pipeline is published rather than the merged data, so the identical 124,102-pair set can be rebuilt from the original public sources.kashmirairesearch.online/dataset
Translation quality (Kashmiri → English)
Three systems on the same test split. The headline result is a negative one: a 600M-parameter purpose-built translation model beats a fine-tuned 7B general-purpose LLM on every metric.
| Metric | Mistral-7B + QLoRA | NLLB-200-distilled-600M | IndicTrans2-1B |
|---|---|---|---|
| BLEU ↑ | 13.89 | 31.57 | 52.88 |
| chrF ↑ | 45.79 | 56.58 | 73.01 |
| TER ↓ | 138.17 | 55.45 | 37.21 |
| METEOR ↑ | 47.74 | 58.57 | 71.85 |
| ROUGE-L ↑ | 39.37 | 57.57 | 70.85 |
| BERTScore F1 ↑ | — | — | 95.47 |
IndicTrans2's scores should be read as a performance ceiling, not a fair zero-shot result: because the corpus draws on AI4Bharat BPCC, there is likely overlap with its pre-training data.
Source: ICRTC-2026 paper, Table 1
Core findings
Four results from the ICRTC-2026 Best Paper, each of which we believe is the first of its kind for Kashmiri.
- Tokenizer fertility gap
- Mistral's BPE tokenizer produces 4.06× more tokens per Kashmiri word than per English word — 5.89 vs 1.45 tokens/word. NLLB-200's SentencePiece reaches near parity at 2.50 vs 1.40.ICRTC-2026 paper, §4.7
- Cross-lingual transfer threshold
- At ~30K parallel pairs BLEU collapses to 0.46. A 4.2× increase in data yields a 29.9× increase in BLEU — a phase transition, not a learning curve.ICRTC-2026 paper, §5.2
- Over-generation
- 66% of Mistral outputs exceed 1.5× the reference length; mean prediction/reference ratio is 2.2× against 1.0× for NLLB-200.ICRTC-2026 paper, §6.1
- Decoding beats retraining
- Oracle truncation recovers +7.85 BLEU. A deployable source-constrained strategy (cap at 2× source words, repetition penalty 1.3) gains +12.75 BLEU, cutting over-generation from 97% to 39% and degeneracy from 54% to 33% — with no change to model weights.ICRTC-2026 paper, §6.2–6.3
- Output degeneracy
- 47.2% of outputs show some degeneracy; 12.2% are pure word repetition.ICRTC-2026 paper, §6.5
Human evaluation
Native Kashmiri speakers rate two anonymised systems on adequacy and fluency, with side assignment randomised per sentence. To our knowledge this is the first evaluation platform built specifically for Kashmiri machine translation.
- Blinded comparisons collected
- 123KashmirAI evaluation platform
- System A — adequacy / fluency
- 3.88 / 3.72 out of 5KashmirAI evaluation platform
- System B — adequacy / fluency
- 3.63 / 3.41 out of 5KashmirAI evaluation platform
- Overall preference
- System A 36.6% · System B 32.5% · tie 30.9%KashmirAI evaluation platform
- Agreement measures
- Cohen's weighted κ for inter-annotator agreement; Wilcoxon signed-rank for preference significanceICRTC-2026 paper, §7
The resource catalogue
An open, community-maintained index of everything published for the Kashmiri language, spanning 2003 to 2026.
- Total resources tracked
- 89kashmirairesearch.online/resources
- Breakdown
- 28 datasets · 28 papers · 16 models · 9 tools · 8 research groupskashmirairesearch.online/resources
- Coverage
- 2003–2026kashmirairesearch.online/resources
- Most active contributor
- Haq Nawaz Malik (HuggingFace: Omarrran) — six Kashmiri arXiv papers in 2026arXiv
- Current best Kashmiri TTS
- Bolbosh (NIT Srinagar & University of Kashmir), MOS 3.63arXiv:2603.07513
- Submissions
- Open to anyone, reviewed before publicationkashmirairesearch.online/resources/submit
Citing this work
Ayoub, F., & Tigga, N. P. (2026). Anatomy of Decoder-Only LLM Failure in Low-Resource Machine Translation: Tokenizer Fertility, Data Thresholds, and Decoding Strategies for Kashmiri. In Proceedings of the 14th International Conference on Recent Trends in Computing (ICRTC-2026). Springer. Best Paper Award.