📦 Open Dataset

Kashmiri Parallel Corpus

A 124,102-pair Kashmiri→English evaluation benchmark, aggregated from two existing public corpora and quality-filtered so results can be compared between papers. The sentence pairs are the work of AI4Bharat and S. M. U. Qamar; the filtering, split and benchmark are ours.

🔬
Dataset in active development

We are currently building and quality-filtering the corpus. Public release on Hugging Face is planned for Q2 2026.

Where the data comes from

None of these sentence pairs originate here. They were collected and aligned by others, and this benchmark exists on top of their work. If you use it, cite them.

AI4Bharat BPCC (kas_Arab)
98,923 pairs — web-crawled news and government text.AI4Bharat, IIT Madras · huggingface.co/datasets/ai4bharat/BPCC
SMUQamar Kashmiri-English Parallel Corpus
26,182 pairs — colloquial and educational content. Licensed CC-BY-NC-SA-4.0.S. M. U. Qamar · huggingface.co/datasets/SMUQamar/Kashmiri-English-Parallel-Corpus

Because the SMUQamar corpus is share-alike and non-commercial, the merged set inherits those terms — it is not ours to relicense, and we do not redistribute it. What we publish instead is the pipeline that rebuilds it: the aggregation and filtering code, the source manifests, and the split seed. Run it against the original sources and you get the identical 124,102-pair benchmark.

📊

Scale

124,102 aligned sentence pairs after filtering — 98,923 from AI4Bharat BPCC, 26,182 from the SMUQamar corpus

🌐

Languages

Kashmiri (ks) → English (en), Perso-Arabic Nastaliq script

🔍

Filtering

Exact-match deduplication, a 10–500 character bound per side, Arabic-script ratio ≥ 0.2, and semantic alignment on multilingual sentence embeddings

✂️

Split

117,896 training and 6,206 evaluation pairs, at a fixed seed so the benchmark is reproducible

📝

Licence

Inherited from the sources. The SMUQamar corpus is CC-BY-NC-SA-4.0, so any derivative including it is non-commercial and share-alike.

🔁

Reproducing it

We publish the aggregation and filtering pipeline rather than redistributing the merged data, so you can rebuild the exact benchmark from the original sources

Why This Dataset Matters

Kashmiri is spoken by over 7 million people, yet what parallel data exists is scattered across separate collections with different formats, quality levels and licences. That fragmentation is one reason Kashmiri is missing from multilingual NLP benchmarks, translation leaderboards and commercial APIs.

This benchmark does not add new sentence pairs. It takes two existing public corpora, removes duplicates and misaligned rows, verifies the script, and fixes a train/evaluation split — so that results reported against it are comparable between papers. That standardisation is the contribution, and it is what the ICRTC-2026 results are measured on.

🗣️

Help Build the Dataset

Native Kashmiri speakers can contribute by evaluating translations on our platform.

Start Evaluating →
📧

Get Notified at Release

Reach out to be notified when the dataset is publicly released on Hugging Face.

Contact Us →