📚 The Kashmiri NLP Resource Hub

Everything built for the Kashmiri language, in one place

Kashmiri research is scattered across arXiv, HuggingFace, government portals and university PDFs — which makes it hard to know what already exists, and easy to rebuild something twice. This is a single, tracked catalogue of 89 datasets, papers, models, tools and the groups behind them, spanning 20032026.

+ Submit a resource
89
Resources
tracked & verified
28
Datasets
corpora & benchmarks
28
Papers
published research
16
Models
pretrained & fine-tuned
9
Tools
software & platforms
26
Since 2025
the field is accelerating

Where the effort is going

Resource count by task. The long tail is the interesting part — it shows which problems in Kashmiri NLP are still essentially untouched.

89 of 89 resources
paper2026

Anatomy of Decoder-Only LLM Failure in Low-Resource Machine Translation

Faizan Ayoub, Neha Prerna Tigga · ICRTC-2026 (Springer) — Best Paper Award

Best Paper at ICRTC-2026. First quantified tokenizer fertility analysis for Kashmiri (4.06x overhead), a sharp cross-lingual transfer threshold near 30K pairs, and a deployable source-constrained decoding strategy worth +12.75 BLEU without retraining. Compares QLoRA Mistral-7B against NLLB-200 and IndicTrans2 on a 124,102-pair corpus.

Best PaperTokenizer FertilityQLoRA
Open Access
Featured
model2026

Bolbosh — Kashmiri TTS

Ashraf, Zargar, Muizz, Mushtaq, Mehdi, Gillani, Kak, Bashir · NIT Srinagar · University of Kashmir

Script-aware flow-matching TTS purpose-built for Kashmiri, with a 272-grapheme vocabulary covering the language's diacritics and vowel distinctions. Current state of the art for Kashmiri speech synthesis at MOS 3.63.

SOTAMOS 3.63Script-Aware
See paper
Featured
paper2026

Bolbosh: Script-Aware Flow Matching for Kashmiri Text-to-Speech

Tajamul Ashraf, Burhaan Rasheed Zargar, Saeed Abdul Muizz, Ifrah Mushtaq, Nazima Mehdi, Iqra Altaf Gillani, Aadil Amin Kak, Janibul Bashir · arXiv:2603.07513

A new benchmark for Kashmiri speech synthesis. Uses Optimal Transport Conditional Flow Matching within a Matcha-TTS architecture, expanding the vocabulary to 272 graphemes to model Kashmiri diacritics and fine-grained vowel distinctions explicitly. Achieves MOS 3.63 and MCD 3.73, substantially outperforming multilingual baselines.

TTSMOS 3.63Flow Matching
arXiv
Featured
org2026

HNM Research — Haq Nawaz Malik

Independent

The most prolific single contributor to Kashmiri NLP data in 2026, with six arXiv papers covering KS-LIT-3M, KS-PRET-5M, 600K-KS-OCR, Koshur Pixel, Koshur Diacritizer and SynthOCR-Gen. Publishes datasets openly on HuggingFace as 'Omarrran'.

Independent ResearcherHuggingFaceProlific
Featured
tool2026

InPage → Unicode Converter

Faizan Ayoub · KashmirAI Research

Converts legacy InPage (.inp) Kashmiri documents into standardised Unicode. Decades of Kashmiri publishing sits locked in InPage files that are invisible to search engines and unusable by AI models; this unlocks that archive.

InPageUnicodeLegacy Formats
Free
Featured
org2026

KashmirAI Research

Founded by Faizan Ayoub

This platform. Kashmiri MT research plus the first open human evaluation infrastructure for the language, the InPage digitization engine, and this resource hub — built to give Kashmiri NLP a single home.

Research PlatformHuman EvaluationOpen
Featured
paper2026

Koshur Pixel: A Large-Scale Synthetic OCR Dataset for Kashmiri

Haq Nawaz Malik, Faizan Iqbal, Nahfid Nissar · arXiv:2606.23144

First large-scale synthetic OCR dataset for Kashmiri: 613,078 image-text pairs generated from the KS-PRET-5M corpus using the SynthOCR-Gen framework. Spans multiple fonts and granularities from single words to full-page documents, with 25+ augmentation strategies emulating real-world document degradation.

613K PairsSynthetic DataPerso-Arabic
arXiv
Featured
dataset2026

KS-PRET-5M — 5M Word Kashmiri Pretraining Corpus

Haq Nawaz Malik, Nahfid Nissar · HNM Research

5 million words / 12 million tokens of Kashmiri text recovered from professionally typeset InPage Nastaliq archives. The largest and highest-fidelity Kashmiri pretraining corpus published to date, and the source corpus behind Koshur Pixel.

5M Words12M TokensInPage
See paper
Featured
paper2026

KS-PRET-5M: A 5 Million Word, 12 Million Token Kashmiri Pretraining Dataset

Haq Nawaz Malik, Nahfid Nissar · arXiv:2604.11066

Successor to KS-LIT-3M. A 5-million-word, 12-million-token Kashmiri pretraining corpus recovered largely from professionally typeset InPage archives in Nastaliq script — the highest-quality stratum currently available for Kashmiri NLP.

5M Words12M TokensInPage
arXiv
Featured
paper2026

600K-KS-OCR: A Synthetic Corpus for Kashmiri Script Recognition

HNM Research Group · arXiv

Introduces a 602K-word synthetic OCR dataset for Kashmiri in three typefaces (Naskh, Nastaleeq, Nakash). Evaluates CRNN and TrOCR architectures for Perso-Arabic script.

OCR602K ImagesPerso-Arabic Script
tool2026

Community OCR Hub — Manuscript Transcription

Faizan Ayoub · KashmirAI Research

Crowdsourced transcription of handwritten Kashmiri manuscripts. Native speakers verify and correct machine transcriptions, turning inaccessible physical archives into training data.

CrowdsourcedManuscriptsCommunity
Free
model2026

Koshur Diacritizer

Haq Nawaz Malik, Nahfid Nissar, Faizan Iqbal · HNM Research

Byte-level sequence-to-sequence model that restores omitted diacritics in Kashmiri text — a necessary preprocessing step for TTS, and for any task where vowel distinctions carry meaning.

Byte-LevelPreprocessing
See paper
paper2026

Koshur Diacritizer: A Byte-Level Seq2Seq Model for Kashmiri Diacritic Restoration

Haq Nawaz Malik, Nahfid Nissar, Faizan Iqbal · arXiv:2606.15883

Byte-level sequence-to-sequence model for restoring diacritics in Kashmiri text. Diacritic restoration is critical for Kashmiri because the language's ~16 vowel distinctions are carried by diacritics that are frequently omitted in written text.

Byte-LevelSeq2SeqDiacritics
arXiv
dataset2026

Koshur Pixel — 613K Synthetic OCR Pairs

Haq Nawaz Malik, Faizan Iqbal, Nahfid Nissar · HNM Research

613,078 synthetic image-text pairs for Kashmiri OCR, spanning multiple fonts and granularities from individual words to full pages, with 25+ augmentation strategies simulating real document degradation.

613K PairsSyntheticMulti-Font
See paper
paper2026

KS-LIT-3M: A Curated Kashmiri Text Dataset for LLM Pretraining

HNM Research Group · arXiv

Introduces a 3.1-million-word curated Kashmiri text dataset designed for pretraining LLMs. Addresses the scarcity of high-quality Kashmiri training data with CC-BY-4.0 licensing.

LLM PretrainingDataset Paper3.1M Words
paper2026

SynthOCR-Gen: A Synthetic OCR Dataset Generator for Low-Resource Languages

Haq Nawaz Malik, Kh Mohmad Shafi, Tanveer Ahmad Reshi · arXiv:2601.16113

The generation framework behind Koshur Pixel. Produces synthetic OCR training data for low-resource scripts, breaking the annotation bottleneck that blocks OCR development for languages without large labelled document collections.

Data GenerationLow-ResourceFramework
arXiv
paper2025

Dataset Creation & Benchmarking for Kashmiri News Snippet Classification

Various · Scientific Reports (Nature) · PMC12630696

First manually labeled corpus of 15,036 Kashmiri news snippets across 10 domains. Benchmarks ParsBERT (F1=0.98), IndicBERTv2+GRU, BLOOM-560m, and Flan-T5 in zero-shot settings.

Text Classification15K SnippetsParsBERT
Featured
paper2025

Deep Neural Architectures for Kashmiri-English Machine Translation

Qamar, S.M.U., Azim, M., Quadri, S.M.K., et al. · Scientific Reports (Nature)

Introduces the first comprehensive NMT benchmark for Kashmiri with a 270K parallel corpus. Compares Encoder-Decoder, Attention-Enhanced, and Transformer architectures. Achieves BLEU-4 of 0.2965.

NMT270K CorpusTransformer
Featured
dataset2025

Kashmiri News Snippet Classification Dataset (15K)

— · Published in Scientific Reports

15,036 manually labelled Kashmiri news snippets across ten domains: Medical, Politics, Sports, Tourism, Education, Art & Craft, Environment, Entertainment, Technology and Culture. The first labelled Kashmiri text classification dataset.

15K Snippets10 DomainsLabelled
Open Access
Featured
paper2025

Word Sense Disambiguation Corpus for Kashmiri

Tawseef Ahmad Mir, Aadil Ahmad Lawaye, et al. · Natural Language Processing (Cambridge), 31, 631–654

The first standard WSD dataset for Kashmiri: a sense-annotated corpus covering 124 commonly used ambiguous Kashmiri words, drawn from a raw corpus of roughly 1 million tokens. Establishes the baseline resource for Kashmiri word sense disambiguation.

124 Words1M TokensSense-Annotated
Cambridge Core
Featured
paper2025

Bhaasha, Bhāṣā, Zaban: A Survey for Low-Resourced Languages in South Asia

Sampoorna Poria, Xiaolei Huang · Findings of EMNLP 2025 · arXiv:2509.11570

Survey of NLP for 650+ South Asian languages since 2020, covering data, models and tasks. Situates Kashmiri within the broader regional resource landscape and maps where the gaps remain.

SurveySouth AsiaEMNLP
ACL Anthology
paper2025

CorIL: Enriching Indian Language to Indian Language Parallel Corpora and MT Systems

Soham Bhattacharjee, Mukund K Roy, Yathish Poojary, Bhargav Dave, Mihir Raj, et al. · arXiv:2509.19941

Builds Indic-to-Indic parallel corpora and translation systems across Indian languages including Kashmiri. Notes that Kashmiri and Sindhi consistently yield the lowest scores even within the Indic family, quantifying how far behind the low-resource tail sits.

Indic-IndicParallel Corpus
arXiv
paper2025

Emotion Recognition in Kashmiri Speech Using Bidirectional LSTM Networks

— · Procedia Computer Science (ScienceDirect)

Evaluates coefficient-based acoustic features with bidirectional LSTM networks for recognising emotion in Kashmiri speech — one of the few affective-computing studies targeting the language.

EmotionBiLSTMAcoustic Features
Elsevier
paper2025

Enhancing Low-Resource Indian Language MT Using LLMs with Preference Optimization

Various · IEEE

RL-inspired framework using Direct Preference Optimization and Hypergeometric-Gamma Reward with mT5-Large model. Covers Kashmiri, Konkani, and Dogri translation.

DPOmT5-LargeReinforcement Learning
paper2025

EnsembleSenti-Kash: Stacked Ensemble for Kashmiri Sentiment Classification

Various · ResearchGate

A stacked ensemble learning system for Kashmiri sentiment analysis with manually labeled data. Compares SVM, Random Forest, XGBoost, mBERT, and LSTM architectures.

Sentiment AnalysisEnsemble LearningmBERT
model2025

Indic Parler-TTS

AI4Bharat · AI4Bharat

Multilingual text-to-speech model covering Indian languages with reported support for Kashmiri. Provides a general-purpose TTS baseline that Kashmiri-specific systems such as Bolbosh are measured against.

TTSMultilingualAI4Bharat
Apache-2.0
dataset2025

Kashmiri Word Sense Disambiguation Corpus

Tawseef Ahmad Mir, Aadil Ahmad Lawaye, et al. · Cambridge University Press

The first standard sense-annotated WSD corpus for Kashmiri, covering 124 commonly used ambiguous words drawn from a ~1M-token raw corpus.

124 WordsSense-Annotated1M Tokens
Cambridge Core
dataset2025

UrduSpeech — Urdu & Kashmiri Speech Corpus

humairawan · HuggingFace

Multi-style, professionally recorded speech corpus covering Urdu and Kashmiri, designed for both text-to-speech and automatic speech recognition.

Multi-StyleTTSASR
See dataset card
paper2024

A Breadth-First Catalog of Text, Speech and Multimodal Research in South Asian Languages

Pranav Gupta · arXiv:2501.00029

Wide catalogue of text, speech and multimodal research across South Asian languages, useful for locating where Kashmiri work sits relative to neighbouring languages and which task areas remain untouched.

CatalogMultimodalSouth Asia
arXiv
paper2024

Addressing the Data Gap: Building a Parallel Corpus for Kashmiri

— · Research publication

Documents the construction of a Kashmiri parallel corpus and the practical obstacles — script normalisation, source scarcity, alignment quality — that make corpus building for Kashmiri harder than for better-resourced Indic languages.

Corpus ConstructionData Gap
ResearchGate
paper2024

BhashaVerse: Translation Ecosystem for Indian Subcontinent Languages

Vandan Mujadia, Dipti Misra Sharma · arXiv:2412.04351

A translation ecosystem spanning 36 Indian subcontinent languages including Kashmiri, covering model training, evaluation and deployment across scripts and resource levels.

36 LanguagesEcosystem
arXiv
paper2024

English-Kashmiri MT System for the Tourism Domain

Various · INDIACom Conference

Domain-specific MT system for tourism using Encoder-Decoder architectures. Addresses challenges of Kashmiri language in practical tourism applications.

Domain-Specific MTTourismEncoder-Decoder
dataset2024

IITB-IndicMonoDoc

CFILT · CFILT, IIT Bombay

Document-level monolingual corpora spanning 22 Indian languages including Kashmiri. Useful for pretraining and for document-level rather than sentence-level modelling.

22 LanguagesDocument-LevelMonolingual
See dataset card
paper2024

Spoken Kashmiri Recognition Using Hybrid CNN-gMLP

Various · Preprints.org

Hybrid CNN-gMLP model for spoken Kashmiri word recognition. Addresses phonetic diversity and dialectal variations with data augmentation strategies.

Speech RecognitionCNN-gMLPPhonetics
dataset2023

Indian Kashmiri Speech Recognition Corpus (Mobile)

Speechocean · Speechocean

Commercial ASR corpus: 115.34 hours of Kashmiri speech from 218 speakers (102 male, 116 female) across 77,255 utterances recorded in quiet environments. The largest Kashmiri speech collection known, though it is licensed rather than open.

115 Hours218 SpeakersCommercial
Commercial
tool2023

Kashmir-ASR

kamrandar · GitHub

Open-source project for automatic speech recognition of the Kashmiri language, providing code and pipeline scaffolding for building Kashmiri ASR systems.

Open SourceGitHubASR
See repository
paper2023

Lexical Sample Word Sense Disambiguation for Kashmiri

Mir, T.A., Lawaye, A.A., et al. · Indian J. Science & Technology

First Lexical Sample WSD dataset for Kashmiri with 50+ ambiguous words. Implements SVM, k-NN, Naïve Bayes, and Decision Tree classifiers achieving 75-89% accuracy.

WSDDisambiguationSVM
paper2023

Machine Intelligence for Language Translation: Kashmiri to English

Various · J. Information & Knowledge Management

Foundational Kashmiri-to-English MT system using LSTM architecture. Discusses challenges of Perso-Arabic script morphology and limited digital resources.

Machine TranslationLSTMNastaliq Script
paper2023

POS Tagging for Kashmiri Using Conditional Random Fields

Various · Academic Conferences

Develops POS taggers using CRF models achieving 80-94% accuracy. Proposes hierarchical tagsets for Kashmiri's V2 word order and inflectional complexity.

POS TaggingCRF80-94% Accuracy
paper2023

Tarjama: The Kashmiri Translator

Various · IJRASET

Proposes a translator system based on ML algorithms and parallel corpora to bridge communication gaps for Kashmiri speakers in digital spaces.

Translation SystemMLParallel Corpus
paper2022

Natural Language Processing Resources for the Kashmiri Language

Aadil Ahmad Lawaye, et al. · Indian Journal of Science and Technology

Survey of the Kashmiri NLP resource landscape — dictionaries, WordNet, the EMILLE monolingual corpus, parallel corpora (NLLB-200, BPCC), spell checkers and speech tools. A useful map of what existed before the 2025–2026 wave.

Resource SurveyKashmiri NLP
Open Access
dataset2021

CIIL / LDC-IL Kashmiri Text & Speech Corpora

CIIL Mysore · Central Institute of Indian Languages

Government-funded raw text and speech corpora released for Kashmiri alongside Assamese, Dogri, Gujarati, Odia and Tamil, built under India's national language-resource programme.

GovernmentTextSpeech
LDC-IL Licence
org2020

AI4Bharat

IIT Madras

The most significant producer of open Indic language technology. Behind IndicTrans2, the BPCC corpora, IndicBERT, IndicConformer ASR and Indic Parler-TTS — most of the general-purpose models that currently support Kashmiri at all.

Research LabOpen SourceIndic NLP
Featured
dataset2020

Kashmiri Raw Speech Corpus (LDC-IL)

Linguistic Data Consortium for Indian Languages · LDC-IL, CIIL Mysore

Natural raw speech corpus for Kashmiri covering multiple dialects, age groups and both genders. Speech was collected across the Valley from Pulwama, Srinagar and Anantnag, comprising words, sentences and running text.

Raw SpeechDialectsGovernment
LDC-IL Licence
paper2019

Developing a Morphological Analyser/Generator for Kashmiri

Aadil Ahmad Lawaye, Tawseef Ahmad Mir, et al. · University of Kashmir — Journal of Linguistics

Paradigm-based morphological analyser and generator for Kashmiri, with separate paradigms constructed for nouns, verbs, adjectives and adverbs covering their possible inflections.

MorphologyParadigm-BasedGenerator
Open Access
paper2017

Kashmiri Speech Recognition Using Linear Predictive Coding and Neural Networks

— · Conference paper

Early Kashmiri ASR work recognising isolated spoken digits zero to nine across male and female speakers, using LPC feature extraction and artificial neural networks. Reports 92% accuracy on 350 speech samples.

LPCANNIsolated Digits
ResearchGate
dataset2003

EMILLE / CIIL Corpus — Kashmiri Monolingual

EMILLE Project · Lancaster University · CIIL

One of the earliest digital Kashmiri monolingual corpora, built by the EMILLE project as part of a wider effort covering South Asian languages. Historically important as a starting point for Kashmiri corpus linguistics.

MonolingualHistoricalEMILLE
Academic
org2003

LDC-IL — Linguistic Data Consortium for Indian Languages

LDC-IL · CIIL Mysore

India's national repository for language-technology resources, established in 2003 and fully government-funded. Distributes Kashmiri text and speech corpora.

RepositoryGovernmentSpeech
LDC-IL Licence
org2000

CFILT — Center for Indian Language Technology, IIT Bombay

IIT Bombay

Coordinates IndoWordNet, the linked lexical knowledge base covering 18 scheduled Indian languages. Kashmiri WordNet with its 29,466 synsets was built and linked here.

IndoWordNetLexicalIIT Bombay
org1991

TDIL Data Centre

Technology Development for Indian Languages · MeitY, Govt. of India

India's national programme for building linguistic resources and language technology across scheduled languages, running since 1991. The distribution point for many government-funded Kashmiri resources.

GovernmentNational ProgrammeRepository
Varies
org1970

Department of Linguistics, University of Kashmir

University of Kashmir

The primary academic home of Kashmiri linguistic research. Runs the 'Development of Language Tools and Linguistic Resources for Kashmiri' project, and produced the Kashmiri POS tagset, morphological analyser and WordNet work.

UniversityLinguisticsKashmir
Featured
org1960

NIT Srinagar — Language & Systems Research

National Institute of Technology Srinagar

Engineering research base in the Valley. Contributed the Bolbosh Kashmiri TTS system in collaboration with the University of Kashmir, currently the strongest Kashmiri speech synthesis result.

UniversityEngineeringKashmir
model

IndicConformer ASR — Kashmiri

AI4Bharat · HuggingFace

Conformer-Large (120M params) ASR model with hybrid CTC-RNNT decoder specifically for Kashmiri speech. Requires 16kHz mono WAV input. Uses AI4Bharat NeMo framework.

ASRConformer120M Params
Featured
model

IndicTrans2 (En→Indic 1B)

AI4Bharat · HuggingFace

State-of-the-art multilingual NMT model supporting 22 scheduled Indian languages including Kashmiri. Uses script unification for cross-lingual transfer. The top-performing model for Indian MT.

NMT1B Params22 Languages
Featured
dataset

KashmirAI Parallel Corpus

KashmirAI Research · KashmirAI

A 124,102-pair Kashmiri→English evaluation benchmark aggregated from AI4Bharat BPCC and the SMUQamar corpus, deduplicated and quality-filtered. The sentence pairs are the work of their original authors; the filtering, split and human evaluation are ours.

Parallel CorpusHuman-EvaluatedMQM
Featured
tool

KashmirAI Research — Evaluation Platform

Faizan Ayoub · KashmirAI

Human-in-the-loop evaluation platform for Kashmiri MT. Native speakers rate translations using MQM framework to create gold-standard quality benchmarks.

EvaluationHuman-in-the-LoopMQM
Featured
dataset

Kashmiri-English Dataset 270K

SMUQamar · HuggingFace

Largest open Kashmiri-English parallel corpus with ~270,000 sentence pairs built from digitized literary texts, manually authored dialogues, and filtered legacy data. Foundation of the first NMT benchmark for Kashmiri.

Parallel Corpus270K PairsNMT Benchmark
Featured
dataset

KS-LIT-3M — 3.1 Million Word Pretraining Dataset

Omarrran (HNM) · HuggingFace

Meticulously cleaned 3.1-million-word (16.4M characters) Kashmiri text corpus specifically designed for pretraining LLMs from scratch. The largest open Kashmiri text resource.

LLM Pretraining3.1M WordsCC-BY-4.0
Featured
dataset

600K-KS-OCR Dataset

Omarrran (HNM) · HuggingFace

~602,000 synthetic word-level images for Kashmiri OCR in three typefaces (Naskh, Nastaleeq, Nakash). Includes ground-truth transcriptions compatible with CRNN and TrOCR.

OCR602K ImagesPerso-Arabic
tool

Apertium — Rule-Based MT Platform

Open Source · GitHub

Open-source rule-based MT platform. Potential for building Kashmiri morphological analyzers and translation modules using paradigm-based approach.

Rule-Based MTMorphologyOpen Source
tool

Bhasini — National Language Translation Mission

MeitY, Govt. of India · Government

Official Indian government AI platform with Kashmiri translation, ASR, and TTS APIs. Free to use for developers. Includes models for all 22 scheduled languages.

APIGovernmentTranslation
dataset

Bhasini / ULCA Portal

MeitY, Govt. of India · Government

Official Government of India platform for Indian language datasets and models under the National Language Translation Mission. Central repository for Kashmiri digital resources.

OfficialIndian LanguagesGovernment
model

BLOOM-560M

BigScience · HuggingFace

Open multilingual LLM pretrained on ROOTS corpus (incl. Persian/Urdu). Used for zero-shot Kashmiri text classification. Effective for cross-lingual transfer to Kashmiri.

LLM560M ParamsZero-Shot
dataset

FLORES-200 Benchmark

Meta AI (FAIR) · GitHub

Multilingual MT evaluation benchmark covering 200+ languages including Kashmiri. Human-translated sentences from Wikipedia for standardized evaluation.

Benchmark200+ LanguagesEvaluation
dataset

IN22 Benchmark (IndicTrans2)

AI4Bharat · GitHub

India-specific multi-domain evaluation benchmark (IN22-Gen + IN22-Conv) for 22 scheduled Indian languages. The standard for evaluating Indian MT models.

BenchmarkIndian LanguagesMulti-Domain
model

IndicBARTSS

AI4Bharat · HuggingFace

Indic-focused variant of mBART, specifically designed for Indian language generation and translation tasks including Kashmiri. Optimized for Indic script families.

IndicBARTSeq2Seq
model

IndicBERTv2 (ai4bharat)

AI4Bharat · HuggingFace

Multilingual BERT model pretrained on Indian languages including Kashmiri. Used as embedding backbone in news classification benchmarks (F1=0.98 with GRU).

BERTEmbeddingsIndian Languages
dataset

IndicNLP Catalog (incl. NLLB-Seed)

AI4Bharat · GitHub

Comprehensive catalog of NLP resources for Indian languages including Kashmiri. Tracks datasets like NLLB-Seed parallel data and connects to FLORES-200 benchmarks.

Indic NLPNLLBCatalog
model

IndicTrans2 (Indic→En 1B)

AI4Bharat · HuggingFace

Reverse-direction model of IndicTrans2 for translating from Indian languages (including Kashmiri) to English. Essential for bidirectional MT research.

NMT1B ParamsIndic→English
dataset

Kaeshir Database

izan-majeed · GitHub

Open-source database of Kashmiri words with English meanings. Installable via pip (`pip install kashmiri`). Great for lexicon building and dictionary apps.

DictionaryLexiconpip install
dataset

Kashmiri Audio Corpus (Segmented)

programindz · HuggingFace

1,955 segmented speech samples (16kHz, 16-bit, mono WAV) derived from OpenSLR-122. Ready-to-use format for ASR training pipelines.

Speech1.9K Segments16kHz
model

Kashmiri BERT (kashmiri-llm-bert-base)

Omarrran (HNM) · HuggingFace

BERT-base model trained specifically for the Kashmiri language. Useful for downstream tasks like text classification, NER, and feature extraction.

BERTKashmiri-SpecificNLU
dataset

Kashmiri Data Corpus (OpenSLR-122)

OpenSLR · OpenSLR

Transcribed audio recordings from native Kashmiri speakers for Automatic Speech Recognition (ASR) development. The foundational Kashmiri speech dataset. GPL-3.0 licensed.

SpeechASRAudio
dataset

Kashmiri Multilingual Dictionary

Omarrran · HuggingFace

Multilingual dictionary dataset with entries in English, Kashmiri, Urdu, Chinese, and Turkish — including example sentences for cross-lingual research.

DictionaryMultilingual5 Languages
dataset

Kashmiri Sample Text Recognition (OCR)

Omarrran · HuggingFace

5,000 samples of Kashmiri text images paired with text labels for OCR model training and testing.

OCR5K SamplesImage-Text
dataset

Kashmiri Spoken Words Dataset

Omarrran · HuggingFace

Processed audio data of 12 frequently spoken Kashmiri words. Designed for spoken word recognition and voice command research.

SpeechWord RecognitionAudio
dataset

Kashmiri Text Corpus Cleaned (2025)

Omarrran (HNM) · HuggingFace

Cleaned and processed Kashmiri text dataset designed for linguistic research and model training. Preprocessed and deduplicated for quality.

Text CorpusCleaned2025
dataset

Kashmiri Text Dataset Collection

mzmmoazam · GitHub

Data and tools to collect Kashmiri text from various online sources and dictionaries. Includes word pronunciations, PDFs, HTML files, and CSV data.

Text CorpusWeb ScrapingData Collection
model

Kashmiri Text Generation — Llama3 8B Instruct

MISHANM · HuggingFace

Fine-tuned Llama 3 8B model for Kashmiri question answering and English→Kashmiri translation. Research-grade prototype for Kashmiri text generation.

LLMLlama 38B Params
tool

Kashmiri WordNet

University of Kashmir · Academic

Lexical database organizing Kashmiri words into synsets with semantic relationships. Essential resource for WSD, morphological analysis, and NLP preprocessing.

WordNetLexical DBSynsets
dataset

Kashmiri-English Parallel Corpus (30K)

SMUQamar · HuggingFace

30,000+ Kashmiri→English sentence pairs organized into raw, cleaned, and processed directories. A widely-used open parallel corpus for Kashmiri MT research.

Parallel CorpusMachine Translation30K+ Pairs
tool

KashmiriGPT — AI Language Assistant

Saqlain Yousuf · Web App

AI-powered Kashmiri language assistant accepting Roman and Perso-Arabic input. Provides responses in Kashmiri script, Roman transliteration, and English for cultural preservation.

ChatbotCultural AIMultilingual I/O
model

kashurai — Kashmiri Dataset Collection

Community · HuggingFace

Curated community collection grouping Kashmiri-specific LLMs, text datasets, and ASR models in one browsable hub. Great starting point for discovery.

CollectionCommunityCurated Hub
dataset

kscp — Kashmiri Speech Corpus Processing

erstan · GitHub

Tools for processing audio and text data from OpenSLR Kashmiri Data Corpus. Includes preprocessing pipelines, data loaders, and segmentation utilities.

Speech ProcessingToolsASR Pipeline
model

mBART-large-50

Meta AI · HuggingFace

Multilingual sequence-to-sequence model pretrained with denoising objective across 50 languages. Base model for fine-tuning Kashmiri translation systems.

Seq2Seq50 LanguagesDenoising
model

NLLB-200 (3.3B Full)

Meta AI · HuggingFace

Full-size NLLB model with 3.3B parameters. Highest quality translation for 200+ languages including Kashmiri. Research-grade multilingual translation.

3.3B ParamsFull Model200+ Languages
model

NLLB-200 Distilled 600M

Meta AI · HuggingFace

No Language Left Behind — supports 200+ languages including Kashmiri (kas_Arab). Designed for high-quality low-resource translation with distilled efficiency.

200+ LanguagesMetakas_Arab
tool

sacreBLEU — MT Evaluation Toolkit

Open Source · GitHub

Standardized MT evaluation toolkit used in all Kashmiri NMT papers. Recommended for reproducible BLEU, ChrF++, and TER scoring on Kashmiri translation output.

EvaluationBLEUChrF++
model

XLS-R 300M (Wav2Vec2)

Meta AI · HuggingFace

Self-supervised speech model pretrained on 436K hours of audio in 128 languages. Fine-tunable for Kashmiri ASR tasks with limited labeled data.

ASRSelf-Supervised128 Languages

Something missing?

This hub is only useful if it stays current. If you have built a Kashmiri dataset, model, tool, or published a paper — or you know of one that is not listed — send it over. Submissions are reviewed before they appear.

Showing the bundled catalogue — the live database is not reachable right now.