Skip to content
Latest Effects
Captions / Transcribe speech
  • local
  • Local
  • Tools

Local Whisper

Transcribe multilingual audio locally with the open Whisper model.

Writes down what is said in a clip, word by word with a time on each, and fills the clip with the captions. Runs on this device.

  • Transcribe speech

Curated output references.

Current API configuration

What goes in. What comes back.

The model defines every slot it reads and every result it returns. Required inputs must be present before a run can begin.

Inputs

Prompts and source material

Source

Audio, Video · Required

Outputs

Files and structured results

Captions

Metadata

Available controls

Configure the run.

Control Available values Requirement
Language en, zh, de, es, ru, ko, fr, ja, pt, tr, pl, ca, nl, ar, sv, it, id, hi, fi, vi, he, uk, el, ms, cs, ro, da, hu, ta, no, th, ur, hr, bg, lt, la, mi, ml, cy, sk, te, fa, lv, bn, sr, az, sl, kn, et, mk, br, eu, is, hy, ne, mn, bs, kk, sq, sw, gl, mr, pa, si, km, sn, yo, so, af, oc, ka, be, tg, sd, gu, am, yi, lo, uz, fo, ht, ps, tk, nn, mt, sa, lb, my, bo, tl, mg, as, tt, haw, ln, ha, ba, jw, su Optional

Works especially well for

  • Transcribe speech
  • Productions that need captions
  • Keeping generated results beside the edit and references

Know the trade-offs

  • Generation results are probabilistic and need editorial review.
  • Availability, speed and final price can change with the provider.
  • The public catalogue does not include vendor benchmarks or model-specific artwork.

Configure the model.
Keep the result in the cut.

The generation returns to the scene, asset folder and timeline it belongs to.