کٲشُر

Seven million speakers.
Almost no machine that understands them.

Kashmiri is a scheduled language of India, written in Perso-Arabic Nastaliq, and almost absent from modern language technology. KashmirAI Research builds the missing pieces — a corpus, a benchmark, and a place for native speakers to judge the results.

7M+
Kashmiri speakersand counting
124,102
Sentence pairsaligned & filtered
89
Resources catalogued2003–2026
31.57
Best BLEU on our benchmarkNLLB-200

Best Paper

Anatomy of Decoder-Only LLM Failure in Low-Resource Machine Translation

Faizan Ayoub and Neha Prerna Tigga · ICRTC-2026, in association with Springer. The first quantified tokenizer fertility analysis for Kashmiri, a data threshold below which cross-lingual transfer collapses, and a decoding fix worth +12.75 BLEU without retraining a single weight.

What we are building

A corpus

124,102 Kashmiri–English sentence pairs, aggregated from AI4Bharat BPCC and the SMUQamar corpus, then filtered on script, length and semantic alignment. It is the training set behind the benchmark, and it is open.

See the dataset

A benchmark

A QLoRA-adapted Mistral-7B measured against NLLB-200 and IndicTrans2. The headline result is a negative one: a 600M translation model beats a fine-tuned 7B LLM, and we traced exactly why.

Read the paper

An evaluation platform

Automated metrics punish valid Kashmiri paraphrases. So native speakers rate translations here directly — blinded, randomised, with agreement tracked. As far as we know it is the first such tool for the language.

How it works

If you speak Kashmiri, you can do something no model can.

You do not need a technical background. Read a Kashmiri sentence, read two English translations, and say which one is closer. It takes ten minutes, and it produces the ground truth this research runs on.

Start evaluating

Free, no institution required.

Everything built for Kashmiri, tracked in one place

Kashmiri research is scattered across arXiv, HuggingFace, government portals and university PDFs. The resource hub catalogues all of it — 89 datasets, papers, models, tools and the groups behind them, filterable by task and year, and open to submissions.

Browse the catalogue