Dialect-faithful transcription
Automatic detection by default. Dialect wording is preserved instead of quietly rewritten into formal Arabic.
Accurate dialect transcription, word-perfect timing, real RTL styling, and a publish-ready export in minutes.
Join the beta See supported dialects
One free export to prove it. Then US$7.99/month.
The comparison video walks through one real clip, in the order that matters:
ASSET SLOT
Replace this block with the real comparison video (see DEMO_VIDEO_ASSET). Ship a captioned poster frame and a text transcript alongside it.
A real creator speaking their own dialect, unedited audio.
The same clip through a mainstream multilingual captioning tool.
Dialect words kept as spoken, foreign words kept in their own script.
Counted live, both workflows, same clip: CORRECTION_TIME_DELTA and EDIT_COUNT_DELTA.
Word highlighting, real RTL layout, 1080p, no watermark.
| What matters for Arabic | Generic multilingual tool | Kalemio |
|---|---|---|
| Dialect handling | Often normalises spoken dialect into formal Arabic | Keeps the dialect wording as spoken by default |
| Arabic + English in one sentence | Frequently drops or transliterates one of the two | Keeps each word in the script it was spoken in |
| Timing | Segment-level timing, retimed by hand | Word-level timing, preserved through text edits |
| RTL typography | Arabic displayed, but punctuation and mixed text often break | Unicode bidirectional layout, shaped and measured with the real text engine |
| Where the video lives | Full video uploaded to a server | Video stays on device; only audio is uploaded |
| Correction time per minute of audio | Your current baseline | CORRECTION_TIME_DELTA |
| Edits required per minute of audio | Your current baseline | EDIT_COUNT_DELTA |
Competitor prices and behaviour change. Everything in the last two rows comes from our own evaluation set of BENCHMARK_CLIP_COUNT clips, measured BENCHMARK_DATE, and is published only when it is measured.
Automatic detection by default. Dialect wording is preserved instead of quietly rewritten into formal Arabic.
Every word carries its own start and end time, so highlighting lands on the syllable, not near it.
Low-confidence words are flagged first. Fixing text reflows the cue without shifting the timing you already trust.
Proper bidirectional text with numbers, hashtags, usernames, Latin words and emoji in the same line.
1080p/30 with no watermark, plus SRT and VTT. Up to 5 minutes per project at launch.
Only the audio needed for the transcript is uploaded, and it is deleted within 24 hours.
It is not a general video editor. It does one job: Arabic captions on a finished clip.
Kalemio is built for Arabic dialects, not for every language. We publish a dialect's status only after it passes our private benchmark, and every dialect starts as "not yet benchmarked."
We do not claim to be the world's best captioning tool, and we do not quote accuracy numbers we have not measured ourselves on clips we are allowed to use.
Your original video stays on your device. We upload only the audio needed to create the transcript. Audio is automatically deleted within 24 hours. We do not train on your content unless you separately opt in.
We are onboarding a limited group of Arabic creators and agencies before the App Store launch. Tell us the dialect you publish in.