We are currently building and quality-filtering the corpus. Public release on Hugging Face is planned for Q2 2026.
None of these sentence pairs originate here. They were collected and aligned by others, and this benchmark exists on top of their work. If you use it, cite them.
Because the SMUQamar corpus is share-alike and non-commercial, the merged set inherits those terms — it is not ours to relicense, and we do not redistribute it. What we publish instead is the pipeline that rebuilds it: the aggregation and filtering code, the source manifests, and the split seed. Run it against the original sources and you get the identical 124,102-pair benchmark.
124,102 aligned sentence pairs after filtering — 98,923 from AI4Bharat BPCC, 26,182 from the SMUQamar corpus
Kashmiri (ks) → English (en), Perso-Arabic Nastaliq script
Exact-match deduplication, a 10–500 character bound per side, Arabic-script ratio ≥ 0.2, and semantic alignment on multilingual sentence embeddings
117,896 training and 6,206 evaluation pairs, at a fixed seed so the benchmark is reproducible
Inherited from the sources. The SMUQamar corpus is CC-BY-NC-SA-4.0, so any derivative including it is non-commercial and share-alike.
We publish the aggregation and filtering pipeline rather than redistributing the merged data, so you can rebuild the exact benchmark from the original sources
Kashmiri is spoken by over 7 million people, yet what parallel data exists is scattered across separate collections with different formats, quality levels and licences. That fragmentation is one reason Kashmiri is missing from multilingual NLP benchmarks, translation leaderboards and commercial APIs.
This benchmark does not add new sentence pairs. It takes two existing public corpora, removes duplicates and misaligned rows, verifies the script, and fixes a train/evaluation split — so that results reported against it are comparable between papers. That standardisation is the contribution, and it is what the ICRTC-2026 results are measured on.
Native Kashmiri speakers can contribute by evaluating translations on our platform.
Start Evaluating →Reach out to be notified when the dataset is publicly released on Hugging Face.
Contact Us →