Adobe Speech to Text for Premiere Pro
Version 2.2.5 · Adobe Inc. · Windows 10 / 11 (64-bit) · ranked #226 in the library
The transcription component for the timeline editor, running the speech models locally so an interview turns into a searchable, editable transcript without anything leaving the machine.
Read the installation section below before you start. The order of the steps matters on this release.
About this release
Overview
This package installs the speech recognition component that the transcript panel in Premiere Pro depends on, together with the language models it runs against. With it in place, selecting a clip or a sequence and running a transcription produces a time aligned transcript with speaker labels, which then drives text based editing, caption generation and search across a whole project.
The point of the offline models is that nothing is uploaded. The standard behaviour sends audio to a service for processing, which is unacceptable for embargoed material, legal work, medical content or anything under a non-disclosure agreement, and impossible on a machine that is not connected. With the models installed locally the transcription runs on the workstation using the processor and, where supported, the graphics card.
Once a transcript exists the rest of the workflow opens up. Delete a sentence in the transcript and the matching footage leaves the sequence. Search across every transcribed clip in the project to find a phrase and jump to it. Generate captions from the transcript with automatic line breaking and timing, style them as a group, and burn them in or export as a sidecar file. Speaker detection separates voices and labels them, and the labels can be renamed once and applied throughout.
What it does
Feature breakdown
Local transcription with no upload
Speech models run on the workstation, so audio never leaves the machine and transcription works on a system with no connection at all.
Time aligned transcripts with speakers
Every word carries a timecode and voices are separated into labelled speakers, which makes long interviews navigable in seconds.
Text based sequence editing
Delete paragraphs or sentences from the transcript and the corresponding clips leave the sequence, including on multicam and synced audio.
Caption generation
Build captions from a transcript with automatic segmentation and timing, styled as a track so a change applies to every caption at once.
Project wide transcript search
Search across every transcribed clip in a project to find a spoken phrase and jump straight to that frame.
Multiple language models
Models for a broad set of languages, installed selectively so you only carry the ones you actually work in.
Changed in this build
Version 2.2.5
- Additional language models added and existing ones retrained for better accuracy on accented speech.
- Graphics acceleration used for transcription where a supported card is present, cutting runtime substantially.
- Speaker separation improved on recordings with overlapping dialogue.
- Filler word detection extended to more languages.
- Transcript panel now handles clips over three hours without the scroll performance degrading.
Installation
Follow these in order
- Close Premiere Pro completely before installing.
- Extract the archive to a folder with a few gigabytes free.
- Run the installer as administrator and select the languages you want models for.
- Let it write the models into the application support folder.
- Open Premiere, go to the transcript panel, and confirm the language list matches what you installed.
- Run a short test transcription to confirm it completes without asking for a connection.
Release notes
From whoever packed it
- This is a component for an existing Premiere Pro install. It does nothing on its own.
- Each language model is a large download. Install only the ones you need or the footprint grows quickly.
- If the panel still tries to reach a service, clear the transcription cache once and restart the application.
System requirements
Minimum and recommended
| Operating system | Windows 10 version 22H2 or Windows 11, 64-bit |
|---|---|
| Processor | Six cores or more recommended, transcription is processor bound without a supported GPU |
| Memory | 16 GB minimum |
| Graphics | Optional but strongly recommended, a card with 6 GB VRAM cuts transcription time dramatically |
| Storage | 8 GB with several language models installed |
Questions about this release
Answered before you ask
- Does it work with any version of Premiere?
- It targets the current release generation. Older versions use a different component layout and will not pick the models up.
- How accurate is it compared to the online service?
- Very close on clean dialogue. Heavy accents, crosstalk and poor recordings widen the gap somewhat.
- Can I export the transcript?
- Yes, as plain text, as a caption sidecar file, or as a formatted document with timecodes and speaker labels.
Download
Pick a mirror