A corpus
124,102 Kashmiri–English sentence pairs, aggregated from AI4Bharat BPCC and the SMUQamar corpus, then filtered on script, length and semantic alignment. It is the training set behind the benchmark, and it is open.
See the datasetکٲشُر
Kashmiri is a scheduled language of India, written in Perso-Arabic Nastaliq, and almost absent from modern language technology. KashmirAI Research builds the missing pieces — a corpus, a benchmark, and a place for native speakers to judge the results.
Best Paper
Faizan Ayoub and Neha Prerna Tigga · ICRTC-2026, in association with Springer. The first quantified tokenizer fertility analysis for Kashmiri, a data threshold below which cross-lingual transfer collapses, and a decoding fix worth +12.75 BLEU without retraining a single weight.
124,102 Kashmiri–English sentence pairs, aggregated from AI4Bharat BPCC and the SMUQamar corpus, then filtered on script, length and semantic alignment. It is the training set behind the benchmark, and it is open.
See the datasetA QLoRA-adapted Mistral-7B measured against NLLB-200 and IndicTrans2. The headline result is a negative one: a 600M translation model beats a fine-tuned 7B LLM, and we traced exactly why.
Read the paperAutomated metrics punish valid Kashmiri paraphrases. So native speakers rate translations here directly — blinded, randomised, with agreement tracked. As far as we know it is the first such tool for the language.
How it worksYou do not need a technical background. Read a Kashmiri sentence, read two English translations, and say which one is closer. It takes ten minutes, and it produces the ground truth this research runs on.
Free, no institution required.
Kashmiri research is scattered across arXiv, HuggingFace, government portals and university PDFs. The resource hub catalogues all of it — 89 datasets, papers, models, tools and the groups behind them, filterable by task and year, and open to submissions.
Browse the catalogue