To extract text from video, first decide which text you need. Spoken words come from a transcript: open the "Show transcript" panel on YouTube, or run the file through a speech-to-text tool. Words that only appear on screen, like slides, code, prices and captions burned into the picture, need video OCR, which reads the frames instead of the audio. Most marketing videos need both.
You are probably here because the useful part of a video is stuck inside it. A client hands you a 50-minute webinar and wants a blog post. A competitor's YouTube video outranks your article and you want to see what it covers. A product demo shows pricing on a slide that nobody says out loud.
Video is hard to copy from and search engines read it far less well than they read a page of text. Here are five ways to get the text out, what each one costs, and where each one falls short.
Key takeaways
Spoken text and on-screen text need different tools. Check which one you are missing before you pick.
YouTube's own transcript is free and is the fastest route for speech.
A screenshot plus Live Text or Google Lens is enough for one or two frames.
A video OCR tool reads a whole video's frames and gives you timestamps.
For hundreds of videos, an API is cheaper per minute but needs a developer.
Spoken text vs on-screen text
A transcript is the audio turned into words. It gets you everything the speaker says and nothing else.
On-screen text is whatever the picture shows: slide headings, bullet points, a terminal command, a name in a lower third, a chart label, a price. A transcript misses all of it unless the speaker reads it aloud. Optical character recognition (OCR) is what reads it.
This matters for SEO work because slides often carry the structure of a talk. The headings on screen are the outline, and the outline is what you need for a brief.
Method
Speech
On-screen text
Price
YouTube transcript
Yes
No
Free
Screenshot + Live Text or Google Lens
No
One frame at a time
Free
Scribiz
Yes
Yes
Free tier, Pro $10 a month
Descript
Yes
No
Free tier, paid from $24 a month
Google Cloud Video Intelligence
Yes
Yes
1,000 free minutes a month, then per minute
Prices checked on each vendor's pricing page on October 5, 2026.
1. YouTube's built-in transcript
If the video is on YouTube and has captions, you already have a transcript.
Open the video on desktop.
Expand the description under the video.
Click Show transcript.
Select the text in the side panel and copy it. You can switch timestamps off from the panel's menu.
Where it falls short: it only works on YouTube and only when captions exist. Auto-generated captions have no punctuation to speak of and they get brand names, people's names and technical terms wrong, so proofread before you publish anything. It reads no on-screen text at all.
2. Screenshot plus Live Text or Google Lens
For a single slide or one line of code, do not open a tool. Pause the video, then:
Mac and iPhone: Live Text lets you select text in a paused video frame in Safari, QuickTime and Photos, or in a screenshot.
Android and Chrome: take a screenshot and open it in Google Lens, then choose the text.
Windows: the Snipping Tool has a Text actions button that copies text from a capture.
All three are free and built in.
Where it falls short: it is one frame at a time. Ten slides is fine. A 40-slide webinar is an afternoon you will not get back, and you have to find each frame yourself.
3. A video OCR tool for the whole video
When you need every slide from a full video, use a tool that reads the frames for you. Scribiz video OCR takes a public video link, reads the text shown on screen, groups frames that show the same thing into scenes, and gives each scene a timestamp. It picks up slides and titles, code and terminal commands, burned-in captions and lower thirds. The same run can return the spoken transcript, a summary and chapters, so you get both kinds of text in one place.
It is free on the website without an account, for videos up to 15 minutes and about 10 minutes of frame reading a day. Pro is $10 a month or $84 a year and raises the limit to videos up to 6 hours. There is also a Mac app, a command-line tool (npm install -g scribiz) and an MCP server, which is handy if you want an AI agent to pull the text into a content brief for you.
Where it falls short: by its own account it misses small type, handwriting, text at an angle and text that flashes by quickly. The website only takes public links, so a file on your laptop needs the Mac app or the command-line tool. If you only need the speech from a YouTube video that already has captions, method 1 is quicker.
4. Descript, when you also edit the video
Descript transcribes the audio and then lets you edit the video by editing the text. Delete a sentence in the transcript and the clip is cut from the timeline.
If your job is to trim the webinar, cut clips for social and export captions, Descript is the better pick and a plain extraction tool is the wrong one. The free plan includes 60 media minutes a month. The Hobbyist plan is $24 a month billed monthly, or $16 a month billed yearly, with 10 media hours.
Where it falls short: the transcript comes from the audio, so text that only appears on a slide is not in it. It is also more tool than you need if all you want is the words.
5. Google Cloud Video Intelligence for bulk work
If a client has a library of hundreds of videos, pay per minute through an API instead of clicking through a web app. Google Cloud Video Intelligence has a text detection feature that returns each piece of on-screen text with its timestamps and its position in the frame, as JSON. The first 1,000 minutes a month are free, then text detection is $0.15 a minute. Speech transcription is a separate feature with its own price.
For speech alone in bulk, OpenAI's open-source Whisper model runs on your own machine at no charge beyond the hardware.
Where it falls short: both need a developer. You have to set up a cloud project, upload files, write code against the API and clean up raw output that repeats the same text across frames. For a one-off video it is far more work than methods 1 to 3.
What SEO teams do with the extracted text
Getting the text out is step one. Here is where it pays off.
Put a transcript under embedded videos. A page that is only a video player gives search engines very little to read. A cleaned-up transcript gives the page text to rank with.
Turn one webinar into many pieces. The transcript is the raw material for an article, a newsletter and social posts. See our content distribution guide for the full workflow.
Research competitor videos. Pull the slide headings from the videos that rank for your keyword. They show you which subtopics the winning content covers.
Add chapters. Timestamps from the extraction become YouTube chapters, which help viewers jump to the part they came for.
Feed AI search. ChatGPT, Claude and Perplexity quote text. If your best explanation only exists as a video, there is much less for them to quote.
If you are building a wider stack, we keep a list of the SEO tools we use.
Which method should you pick?
One YouTube video, speech only: YouTube's transcript panel.
One or two slides: screenshot and Live Text or Google Lens.
A full video with slides, code or captions on screen: a video OCR tool.
You need to edit the video too: Descript.
Hundreds of videos and a developer on hand: Google Cloud Video Intelligence, or Whisper for speech.
Conclusion
Start with the free route that matches the text you are missing. Speech is nearly always a few clicks away. On-screen text is the part people forget, and it is often where the outline, the numbers and the code live.
Then put the text to work. A transcript sitting in a doc ranks for nothing. If you want help turning a video library into pages that bring in signups, take a look at how we work with SaaS teams.