Source
Audio, Video · Required
Return to the catalogue and choose an available model.
Back to the catalogueTurn multilingual speech into accurate, production-ready transcripts.
What is said in a clip, word by word, speaker by speaker.
The model defines every slot it reads and every result it returns. Required inputs must be present before a run can begin.
Prompts and source material
Audio, Video · Required
Files and structured results
Metadata
| Control | Available values | Requirement |
|---|---|---|
| Language | af, am, ar, as, az, ba, be, bg, bn, bo, br, bs, ca, cs, cy, da, de, el, en, es, et, eu, fa, fi, fo, fr, gl, gu, ha, haw, he, hi, hr, ht, hu, hy, id, is, it, ja, jw, ka, kk, km, kn, ko, la, lb, ln, lo, lt, lv, mg, mi, mk, ml, mn, mr, ms, mt, my, ne, nl, nn, no, oc, pa, pl, ps, pt, ro, ru, sa, sd, si, sk, sl, sn, so, sq, sr, su, sv, sw, ta, te, tg, th, tk, tl, tr, tt, uk, ur, uz, vi, yi, yo, zh | Optional |
| Speakers | 1–… | Optional |
Works especially well for
Know the trade-offs
The generation returns to the scene, asset folder and timeline it belongs to.