Why we are doing this
Our goal is to keep Kashmiri alive in the digital age by helping AI learn it: tools that translate Kashmiri into English and English into Kashmiri, so the language can be used on phones, computers and the internet like any other.
Today's AI still translates Kashmiri badly, and no one has checked its mistakes with native speakers. This study is the first step. You mark those mistakes; your markings show where the programs go wrong, in a research paper and a public benchmark (KashEval), and guide the work on better Kashmiri translation tools.
What you will do
- You get a pack of about 25 Kashmiri sentences. Each has an English reference translation and three computer translations, marked A, B and C. You are not told which program wrote which.
- For each translation you give a score from 1 to 5, and mark five kinds of error as none, minor or major.
- A pack takes about 45 to 60 minutes. You may do more than one. Some sentences are in two teachers' packs, so we can check how often people agree.
- The sentences come from SMOL, a public dataset published by Google. They contain no personal information.
What we collect and how we use it
- Your markings, filed under a code (T01, T02…). Your name appears only on this form.
- Optional: your age range and home district, because Kashmiri differs by region.
- The markings are used in a research paper (planned for the LoResMT workshop at EACL 2027) and in the KashEval benchmark. Results are reported as totals, never by name.
- This form is kept securely until 31 December 2030 and then deleted.
Your rights
- Taking part is voluntary. You can stop at any time without giving a reason.
- You can ask for your markings to be deleted until 31 December 2030. Totals already published in a paper cannot be changed.
- Under the Digital Personal Data Protection Act 2023, you can ask to see, correct or delete your personal data by writing to faizan@kashmirairesearch.online.
- No risks are expected beyond your time. Thank you: major contributors receive a certificate.