Transcribe Tajik audio with WhisperAI
WhisperAI transcribes Tajik (Тоҷикӣ) audio for the around 8 million speakers who use the language across Tajikistan, Uzbekistan and Afghanistan. Built on OpenAI Whisper and tuned for the realities of Tajik speech, the service produces publication-grade transcripts in Cyrillic with additional letters (ғ, қ, ҳ, ҷ, ӣ, ӯ).
How Tajik speakers actually use WhisperAI
- Dushanbe government documentationTajik ministries and government offices transcribe Tajik-language meetings.
- Tajik media and broadcastingDushanbe-based Tajik broadcasters transcribe content for caption use.
- Tajik-diaspora workTajik diaspora communities (especially in Russia) document oral histories.
- Tajik academic and cultural workUniversities transcribe Tajik qualitative research.
Tajik as Persian written in Cyrillic
Tajik is essentially Persian (mutually intelligible with Iranian Farsi and Afghan Dari) but written in Cyrillic with six additional letters. ASR trained on Russian Cyrillic produces broken text; ASR trained on Persian script can't read it. Whisper handles Tajik in Cyrillic with the full alphabet.
Standard Northern Tajik and Southern dialects are supported in Cyrillic.
Who actually needs Tajik transcription
Tajik demand comes from Dushanbe government and media, Tajik-diaspora work, and academic Central-Asian-studies.
Dialects and varieties handled
- Standard Tajik (Northern, Dushanbe)
- Southern
- Pamiri varieties (separate languages)
How to transcribe Tajik audio
Upload your Tajik recording
Drop in MP3, WAV, M4A, MP4 or similar — files up to 5GB. Long-form Tajik interviews, lectures and meetings work without splitting.
Whisper transcribes the audio
The model recognises Tajik speech across the dialects above and outputs Cyrillic with additional letters (ғ, қ, ҳ, ҷ, ӣ, ӯ). Code-switched English and other languages in the same recording are handled inline.
Edit, label and export
Open the transcript in the editor, fix any names or terms, label speakers, and export to PDF, DOCX, TXT or SRT. Optional AI summary surfaces the key points and decisions.
What Tajik transcripts include
Native Cyrillic with additional letters (ғ, қ, ҳ, ҷ, ӣ, ӯ)
Output in Cyrillic with additional letters (ғ, қ, ҳ, ҷ, ӣ, ӯ) with all diacritics, tone marks, special characters and script-specific conventions preserved — never transliteration.
Speaker labels
Diarisation that separates and labels speakers in Tajik interviews, panels and multi-party meetings.
Timestamps and SRT
Word-level timestamps and SRT subtitle export for Tajik video captioning, with proper line-breaking for the script.
Editor and exports
In-browser editor with PDF, DOCX, TXT and SRT exports. Edit the transcript without leaving the browser, then download the format your downstream workflow expects.
Transcription guides and best practices
Transcribe in 100+ languages
WhisperAI supports Tajik alongside over 100 other languages with the same accuracy and editor experience.
Other widely-used pages: English, Spanish, French, German, Portuguese, Japanese, Korean, Chinese, Arabic, Russian, Hindi, Italian, and 80+ more languages including Swedish, Norwegian, Danish, Finnish, Greek, Malay, Filipino and beyond.
See every supported language, with transcription, realtime and translation coverage.
Start transcribing Tajik today
Sign up free, drop in your first Tajik file, and have a usable transcript in minutes.