Skip to content
Latest Effects
Captions / Transcribe speech
  • ElevenLabs
  • Standard
  • Tools

Scribe v2

Turn multilingual speech into accurate, production-ready transcripts.

What is said in a clip, word by word, speaker by speaker.

  • Transcribe speech

Curated output references.

Current API configuration

What goes in. What comes back.

The model defines every slot it reads and every result it returns. Required inputs must be present before a run can begin.

Inputs

Prompts and source material

Source

Audio, Video · Required

Outputs

Files and structured results

Captions

Metadata

Available controls

Configure the run.

Control Available values Requirement
Language af, am, ar, as, az, ba, be, bg, bn, bo, br, bs, ca, cs, cy, da, de, el, en, es, et, eu, fa, fi, fo, fr, gl, gu, ha, haw, he, hi, hr, ht, hu, hy, id, is, it, ja, jw, ka, kk, km, kn, ko, la, lb, ln, lo, lt, lv, mg, mi, mk, ml, mn, mr, ms, mt, my, ne, nl, nn, no, oc, pa, pl, ps, pt, ro, ru, sa, sd, si, sk, sl, sn, so, sq, sr, su, sv, sw, ta, te, tg, th, tk, tl, tr, tt, uk, ur, uz, vi, yi, yo, zh Optional
Speakers 1–… Optional

Works especially well for

  • Transcribe speech
  • Productions that need captions
  • Keeping generated results beside the edit and references

Know the trade-offs

  • Generation results are probabilistic and need editorial review.
  • Availability, speed and final price can change with the provider.
  • The public catalogue does not include vendor benchmarks or model-specific artwork.

Configure the model.
Keep the result in the cut.

The generation returns to the scene, asset folder and timeline it belongs to.