Kashmiri research is scattered across arXiv, HuggingFace, government portals and university PDFs — which makes it hard to know what already exists, and easy to rebuild something twice. This is a single, tracked catalogue of 89 datasets, papers, models, tools and the groups behind them, spanning 2003–2026.
Resource count by task. The long tail is the interesting part — it shows which problems in Kashmiri NLP are still essentially untouched.
Best Paper at ICRTC-2026. First quantified tokenizer fertility analysis for Kashmiri (4.06x overhead), a sharp cross-lingual transfer threshold near 30K pairs, and a deployable source-constrained decoding strategy worth +12.75 BLEU without retraining. Compares QLoRA Mistral-7B against NLLB-200 and IndicTrans2 on a 124,102-pair corpus.
Script-aware flow-matching TTS purpose-built for Kashmiri, with a 272-grapheme vocabulary covering the language's diacritics and vowel distinctions. Current state of the art for Kashmiri speech synthesis at MOS 3.63.
A new benchmark for Kashmiri speech synthesis. Uses Optimal Transport Conditional Flow Matching within a Matcha-TTS architecture, expanding the vocabulary to 272 graphemes to model Kashmiri diacritics and fine-grained vowel distinctions explicitly. Achieves MOS 3.63 and MCD 3.73, substantially outperforming multilingual baselines.
The most prolific single contributor to Kashmiri NLP data in 2026, with six arXiv papers covering KS-LIT-3M, KS-PRET-5M, 600K-KS-OCR, Koshur Pixel, Koshur Diacritizer and SynthOCR-Gen. Publishes datasets openly on HuggingFace as 'Omarrran'.
Converts legacy InPage (.inp) Kashmiri documents into standardised Unicode. Decades of Kashmiri publishing sits locked in InPage files that are invisible to search engines and unusable by AI models; this unlocks that archive.
This platform. Kashmiri MT research plus the first open human evaluation infrastructure for the language, the InPage digitization engine, and this resource hub — built to give Kashmiri NLP a single home.
First large-scale synthetic OCR dataset for Kashmiri: 613,078 image-text pairs generated from the KS-PRET-5M corpus using the SynthOCR-Gen framework. Spans multiple fonts and granularities from single words to full-page documents, with 25+ augmentation strategies emulating real-world document degradation.
5 million words / 12 million tokens of Kashmiri text recovered from professionally typeset InPage Nastaliq archives. The largest and highest-fidelity Kashmiri pretraining corpus published to date, and the source corpus behind Koshur Pixel.
Successor to KS-LIT-3M. A 5-million-word, 12-million-token Kashmiri pretraining corpus recovered largely from professionally typeset InPage archives in Nastaliq script — the highest-quality stratum currently available for Kashmiri NLP.
Introduces a 602K-word synthetic OCR dataset for Kashmiri in three typefaces (Naskh, Nastaleeq, Nakash). Evaluates CRNN and TrOCR architectures for Perso-Arabic script.
Crowdsourced transcription of handwritten Kashmiri manuscripts. Native speakers verify and correct machine transcriptions, turning inaccessible physical archives into training data.
Byte-level sequence-to-sequence model that restores omitted diacritics in Kashmiri text — a necessary preprocessing step for TTS, and for any task where vowel distinctions carry meaning.
Byte-level sequence-to-sequence model for restoring diacritics in Kashmiri text. Diacritic restoration is critical for Kashmiri because the language's ~16 vowel distinctions are carried by diacritics that are frequently omitted in written text.
613,078 synthetic image-text pairs for Kashmiri OCR, spanning multiple fonts and granularities from individual words to full pages, with 25+ augmentation strategies simulating real document degradation.
Introduces a 3.1-million-word curated Kashmiri text dataset designed for pretraining LLMs. Addresses the scarcity of high-quality Kashmiri training data with CC-BY-4.0 licensing.
The generation framework behind Koshur Pixel. Produces synthetic OCR training data for low-resource scripts, breaking the annotation bottleneck that blocks OCR development for languages without large labelled document collections.
First manually labeled corpus of 15,036 Kashmiri news snippets across 10 domains. Benchmarks ParsBERT (F1=0.98), IndicBERTv2+GRU, BLOOM-560m, and Flan-T5 in zero-shot settings.
Introduces the first comprehensive NMT benchmark for Kashmiri with a 270K parallel corpus. Compares Encoder-Decoder, Attention-Enhanced, and Transformer architectures. Achieves BLEU-4 of 0.2965.
15,036 manually labelled Kashmiri news snippets across ten domains: Medical, Politics, Sports, Tourism, Education, Art & Craft, Environment, Entertainment, Technology and Culture. The first labelled Kashmiri text classification dataset.
The first standard WSD dataset for Kashmiri: a sense-annotated corpus covering 124 commonly used ambiguous Kashmiri words, drawn from a raw corpus of roughly 1 million tokens. Establishes the baseline resource for Kashmiri word sense disambiguation.
Survey of NLP for 650+ South Asian languages since 2020, covering data, models and tasks. Situates Kashmiri within the broader regional resource landscape and maps where the gaps remain.
Builds Indic-to-Indic parallel corpora and translation systems across Indian languages including Kashmiri. Notes that Kashmiri and Sindhi consistently yield the lowest scores even within the Indic family, quantifying how far behind the low-resource tail sits.
Evaluates coefficient-based acoustic features with bidirectional LSTM networks for recognising emotion in Kashmiri speech — one of the few affective-computing studies targeting the language.
RL-inspired framework using Direct Preference Optimization and Hypergeometric-Gamma Reward with mT5-Large model. Covers Kashmiri, Konkani, and Dogri translation.
A stacked ensemble learning system for Kashmiri sentiment analysis with manually labeled data. Compares SVM, Random Forest, XGBoost, mBERT, and LSTM architectures.
Multilingual text-to-speech model covering Indian languages with reported support for Kashmiri. Provides a general-purpose TTS baseline that Kashmiri-specific systems such as Bolbosh are measured against.
The first standard sense-annotated WSD corpus for Kashmiri, covering 124 commonly used ambiguous words drawn from a ~1M-token raw corpus.
Multi-style, professionally recorded speech corpus covering Urdu and Kashmiri, designed for both text-to-speech and automatic speech recognition.
Wide catalogue of text, speech and multimodal research across South Asian languages, useful for locating where Kashmiri work sits relative to neighbouring languages and which task areas remain untouched.
Documents the construction of a Kashmiri parallel corpus and the practical obstacles — script normalisation, source scarcity, alignment quality — that make corpus building for Kashmiri harder than for better-resourced Indic languages.
A translation ecosystem spanning 36 Indian subcontinent languages including Kashmiri, covering model training, evaluation and deployment across scripts and resource levels.
Domain-specific MT system for tourism using Encoder-Decoder architectures. Addresses challenges of Kashmiri language in practical tourism applications.
Document-level monolingual corpora spanning 22 Indian languages including Kashmiri. Useful for pretraining and for document-level rather than sentence-level modelling.
Hybrid CNN-gMLP model for spoken Kashmiri word recognition. Addresses phonetic diversity and dialectal variations with data augmentation strategies.
Commercial ASR corpus: 115.34 hours of Kashmiri speech from 218 speakers (102 male, 116 female) across 77,255 utterances recorded in quiet environments. The largest Kashmiri speech collection known, though it is licensed rather than open.
Open-source project for automatic speech recognition of the Kashmiri language, providing code and pipeline scaffolding for building Kashmiri ASR systems.
First Lexical Sample WSD dataset for Kashmiri with 50+ ambiguous words. Implements SVM, k-NN, Naïve Bayes, and Decision Tree classifiers achieving 75-89% accuracy.
Foundational Kashmiri-to-English MT system using LSTM architecture. Discusses challenges of Perso-Arabic script morphology and limited digital resources.
Develops POS taggers using CRF models achieving 80-94% accuracy. Proposes hierarchical tagsets for Kashmiri's V2 word order and inflectional complexity.
Proposes a translator system based on ML algorithms and parallel corpora to bridge communication gaps for Kashmiri speakers in digital spaces.
Survey of the Kashmiri NLP resource landscape — dictionaries, WordNet, the EMILLE monolingual corpus, parallel corpora (NLLB-200, BPCC), spell checkers and speech tools. A useful map of what existed before the 2025–2026 wave.
Government-funded raw text and speech corpora released for Kashmiri alongside Assamese, Dogri, Gujarati, Odia and Tamil, built under India's national language-resource programme.
The most significant producer of open Indic language technology. Behind IndicTrans2, the BPCC corpora, IndicBERT, IndicConformer ASR and Indic Parler-TTS — most of the general-purpose models that currently support Kashmiri at all.
Natural raw speech corpus for Kashmiri covering multiple dialects, age groups and both genders. Speech was collected across the Valley from Pulwama, Srinagar and Anantnag, comprising words, sentences and running text.
Paradigm-based morphological analyser and generator for Kashmiri, with separate paradigms constructed for nouns, verbs, adjectives and adverbs covering their possible inflections.
Early Kashmiri ASR work recognising isolated spoken digits zero to nine across male and female speakers, using LPC feature extraction and artificial neural networks. Reports 92% accuracy on 350 speech samples.
One of the earliest digital Kashmiri monolingual corpora, built by the EMILLE project as part of a wider effort covering South Asian languages. Historically important as a starting point for Kashmiri corpus linguistics.
India's national repository for language-technology resources, established in 2003 and fully government-funded. Distributes Kashmiri text and speech corpora.
Coordinates IndoWordNet, the linked lexical knowledge base covering 18 scheduled Indian languages. Kashmiri WordNet with its 29,466 synsets was built and linked here.
India's national programme for building linguistic resources and language technology across scheduled languages, running since 1991. The distribution point for many government-funded Kashmiri resources.
The primary academic home of Kashmiri linguistic research. Runs the 'Development of Language Tools and Linguistic Resources for Kashmiri' project, and produced the Kashmiri POS tagset, morphological analyser and WordNet work.
Engineering research base in the Valley. Contributed the Bolbosh Kashmiri TTS system in collaboration with the University of Kashmir, currently the strongest Kashmiri speech synthesis result.
Conformer-Large (120M params) ASR model with hybrid CTC-RNNT decoder specifically for Kashmiri speech. Requires 16kHz mono WAV input. Uses AI4Bharat NeMo framework.
State-of-the-art multilingual NMT model supporting 22 scheduled Indian languages including Kashmiri. Uses script unification for cross-lingual transfer. The top-performing model for Indian MT.
A 124,102-pair Kashmiri→English evaluation benchmark aggregated from AI4Bharat BPCC and the SMUQamar corpus, deduplicated and quality-filtered. The sentence pairs are the work of their original authors; the filtering, split and human evaluation are ours.
Human-in-the-loop evaluation platform for Kashmiri MT. Native speakers rate translations using MQM framework to create gold-standard quality benchmarks.
Largest open Kashmiri-English parallel corpus with ~270,000 sentence pairs built from digitized literary texts, manually authored dialogues, and filtered legacy data. Foundation of the first NMT benchmark for Kashmiri.
Meticulously cleaned 3.1-million-word (16.4M characters) Kashmiri text corpus specifically designed for pretraining LLMs from scratch. The largest open Kashmiri text resource.
~602,000 synthetic word-level images for Kashmiri OCR in three typefaces (Naskh, Nastaleeq, Nakash). Includes ground-truth transcriptions compatible with CRNN and TrOCR.
Open-source rule-based MT platform. Potential for building Kashmiri morphological analyzers and translation modules using paradigm-based approach.
Official Indian government AI platform with Kashmiri translation, ASR, and TTS APIs. Free to use for developers. Includes models for all 22 scheduled languages.
Official Government of India platform for Indian language datasets and models under the National Language Translation Mission. Central repository for Kashmiri digital resources.
Open multilingual LLM pretrained on ROOTS corpus (incl. Persian/Urdu). Used for zero-shot Kashmiri text classification. Effective for cross-lingual transfer to Kashmiri.
Multilingual MT evaluation benchmark covering 200+ languages including Kashmiri. Human-translated sentences from Wikipedia for standardized evaluation.
India-specific multi-domain evaluation benchmark (IN22-Gen + IN22-Conv) for 22 scheduled Indian languages. The standard for evaluating Indian MT models.
Indic-focused variant of mBART, specifically designed for Indian language generation and translation tasks including Kashmiri. Optimized for Indic script families.
Multilingual BERT model pretrained on Indian languages including Kashmiri. Used as embedding backbone in news classification benchmarks (F1=0.98 with GRU).
Comprehensive catalog of NLP resources for Indian languages including Kashmiri. Tracks datasets like NLLB-Seed parallel data and connects to FLORES-200 benchmarks.
Reverse-direction model of IndicTrans2 for translating from Indian languages (including Kashmiri) to English. Essential for bidirectional MT research.
Open-source database of Kashmiri words with English meanings. Installable via pip (`pip install kashmiri`). Great for lexicon building and dictionary apps.
1,955 segmented speech samples (16kHz, 16-bit, mono WAV) derived from OpenSLR-122. Ready-to-use format for ASR training pipelines.
BERT-base model trained specifically for the Kashmiri language. Useful for downstream tasks like text classification, NER, and feature extraction.
Transcribed audio recordings from native Kashmiri speakers for Automatic Speech Recognition (ASR) development. The foundational Kashmiri speech dataset. GPL-3.0 licensed.
Multilingual dictionary dataset with entries in English, Kashmiri, Urdu, Chinese, and Turkish — including example sentences for cross-lingual research.
5,000 samples of Kashmiri text images paired with text labels for OCR model training and testing.
Processed audio data of 12 frequently spoken Kashmiri words. Designed for spoken word recognition and voice command research.
Cleaned and processed Kashmiri text dataset designed for linguistic research and model training. Preprocessed and deduplicated for quality.
Data and tools to collect Kashmiri text from various online sources and dictionaries. Includes word pronunciations, PDFs, HTML files, and CSV data.
Fine-tuned Llama 3 8B model for Kashmiri question answering and English→Kashmiri translation. Research-grade prototype for Kashmiri text generation.
Lexical database organizing Kashmiri words into synsets with semantic relationships. Essential resource for WSD, morphological analysis, and NLP preprocessing.
30,000+ Kashmiri→English sentence pairs organized into raw, cleaned, and processed directories. A widely-used open parallel corpus for Kashmiri MT research.
AI-powered Kashmiri language assistant accepting Roman and Perso-Arabic input. Provides responses in Kashmiri script, Roman transliteration, and English for cultural preservation.
Curated community collection grouping Kashmiri-specific LLMs, text datasets, and ASR models in one browsable hub. Great starting point for discovery.
Tools for processing audio and text data from OpenSLR Kashmiri Data Corpus. Includes preprocessing pipelines, data loaders, and segmentation utilities.
Multilingual sequence-to-sequence model pretrained with denoising objective across 50 languages. Base model for fine-tuning Kashmiri translation systems.
Full-size NLLB model with 3.3B parameters. Highest quality translation for 200+ languages including Kashmiri. Research-grade multilingual translation.
No Language Left Behind — supports 200+ languages including Kashmiri (kas_Arab). Designed for high-quality low-resource translation with distilled efficiency.
Standardized MT evaluation toolkit used in all Kashmiri NMT papers. Recommended for reproducible BLEU, ChrF++, and TER scoring on Kashmiri translation output.
Self-supervised speech model pretrained on 436K hours of audio in 128 languages. Fine-tunable for Kashmiri ASR tasks with limited labeled data.
This hub is only useful if it stays current. If you have built a Kashmiri dataset, model, tool, or published a paper — or you know of one that is not listed — send it over. Submissions are reviewed before they appear.
Showing the bundled catalogue — the live database is not reachable right now.