Glasser Team8 min read
How to Get YouTube Transcripts with Python
Fetch YouTube transcripts in Python, choose language tracks, preserve timestamps, and diagnose missing captions or blocked requests with current examples.

How can you get a YouTube transcript with Python?
The youtube-transcript-api Python package retrieves available captions for a YouTube video and returns text with start times and durations. Its current interface uses YouTubeTranscriptApi().fetch(video_id), which is worth checking before copying an older get_transcript example.
This is an independent open-source library. It is separate from Google's official YouTube Data API, and it does not generate speech-to-text output for a video that has no available captions.
Want the transcript without working through the Python setup below? Start with Glasser to use a hosted YouTube transcript endpoint, check its price, and retrieve the text when available.
Install the package version used in this guide
This guide targets youtube-transcript-api 1.2.4, the version listed on PyPI when checked on October 8, 2026. Use a virtual environment so that your terminal and editor load the same package.
python -m venv .venv
source .venv/bin/activate
python -m pip install 'youtube-transcript-api==1.2.4'
On Windows, activate .venv\Scripts\activate instead. Save this as fetch_transcript.py:
import sys
from youtube_transcript_api import YouTubeTranscriptApi
video_id = sys.argv[1] # The ID from v=... or the path of a youtu.be URL.
fetched = YouTubeTranscriptApi().fetch(video_id, languages=["en"])
for segment in fetched:
print(f"{segment.start:.2f}\t{segment.duration:.2f}\t{segment.text}")
Run python fetch_transcript.py YOUR_VIDEO_ID, replacing the final argument with a video ID. Supply the ID rather than the full watch URL. The maintainer's documentation defines the returned FetchedTranscript object and language-selection behavior.
The printed columns are start time in seconds, duration in seconds, and caption text. For example, the following is illustrative output, not a transcript fetched for this article:
0.00 2.40 Welcome to this example.
2.40 3.10 We will keep the timing information.
The examples in this guide were checked against version 1.2.4 and local fixtures. No live video-retrieval success rate or cloud-hosting benchmark is claimed.
How do you get the right language and caption track?
List the video’s available caption tracks
A video can have manual captions, automatically generated captions, multiple languages, or no available track. List the tracks before deciding that a failed English request means the entire video is unusable.
import sys
from youtube_transcript_api import YouTubeTranscriptApi
api = YouTubeTranscriptApi()
tracks = api.list(sys.argv[1])
for track in tracks:
print(track.language_code, track.is_generated, track.is_translatable)
The languages argument is an ordered preference list. A request with languages=["de", "en"] tries the German option before English; it does not ask the library to translate English into German.
Choose manual, generated, or translated captions explicitly
By default, the package prefers manually created captions when both types are available for the requested language. You can require a particular kind of track:
# Continue after creating `tracks` in the preceding example.
manual_track = tracks.find_manually_created_transcript(["en"])
fetched = manual_track.fetch()
Use find_generated_transcript(["en"]) when you specifically want the generated track. Requiring a manual English track can raise NoTranscriptFound even when an automatically generated English track exists.
For a summarization workflow, a reasonable author policy is to prefer manual captions, accept generated captions when needed, and preserve is_generated in the stored record. Generated captions may mishear names or technical terms, so a downstream answer should remain traceable to the relevant video segment.
Translation is another explicit choice. A track exposes is_translatable and translation_languages; inspect those before calling track.translate("fr").fetch(). Store the source and requested language if your application produces a translated version so it cannot be mistaken for the original caption track.
Why should you keep timestamps when saving a transcript?
Store the video ID, language, and timed segments
For a search index or summarizer, retain the video ID, language, caption origin, and segment timing alongside the text. The following is an author-defined storage envelope around the library's documented output:
import json
from datetime import datetime, timezone
from pathlib import Path
# `fetched` is the FetchedTranscript returned by fetch().
record = {
"video_id": fetched.video_id,
"language_code": fetched.language_code,
"is_generated": fetched.is_generated,
"retrieved_at": datetime.now(timezone.utc).isoformat(),
"segments": fetched.to_raw_data(),
}
Path("transcript.json").write_text(
json.dumps(record, ensure_ascii=False, indent=2),
encoding="utf-8",
)
Each raw segment contains text, start, and duration. The retrieval timestamp belongs to your pipeline; it is not a caption field returned by YouTube or proof that the video content changed at that time.
If you need plain text as an additional representation, join the segment text while retaining the timed record as the source. Flattening everything into one string makes it harder to show where a generated summary came from.
Export subtitles as SRT or WebVTT
The package includes SRT and WebVTT formatters:
from pathlib import Path
from youtube_transcript_api.formatters import SRTFormatter, WebVTTFormatter
Path("transcript.srt").write_text(
SRTFormatter().format_transcript(fetched), encoding="utf-8"
)
Path("transcript.vtt").write_text(
WebVTTFormatter().format_transcript(fetched), encoding="utf-8"
)
Choose JSON for structured processing, SRT for subtitle workflows that expect it, and WebVTT for compatible web players. Conversion changes the format; it does not correct transcription errors or add missing speech.
For chunking, keep the first and last segment times for each chunk. If a user follows a summary back to the source, your application can point to the right part of the video instead of only its homepage. Chunk size and overlap are application decisions, so test them against the retrieval task you actually need.
Why does transcript retrieval fail, and what can you check?
Separate missing captions from blocked requests
| Symptom | What it indicates | Next step |
|---|---|---|
NoTranscriptFound | Requested language or track type is unavailable | Inspect list(video_id) and your fallback policy |
TranscriptsDisabled | Captions are unavailable through this retrieval path | Record the condition; use an authorized alternative if needed |
VideoUnavailable | The requested video cannot be accessed | Recheck the ID and the video's availability |
RequestBlocked or IpBlocked | YouTube blocked the request or network | Stop repeated attempts and review deployment access |
Missing get_transcript attribute | Old code against the newer interface | Use instance fetch()/list() and verify the installed version |
The package README describes cloud-provider IP blocking. A successful call from a laptop therefore does not establish that the same workload will run from your server.
Do not convert these errors into an empty successful transcript. A downstream summarizer might otherwise interpret a retrieval failure as evidence that the video contains no speech.
If maintaining transcript retrieval on your server is taking time away from your application, explore Glasser’s YouTube endpoints. Check each transcript endpoint’s language support, timestamp handling, and availability before trying your sample videos.
Record a success or failure status for each video
This wrapper illustrates a per-video outcome record. It handles known retrieval conditions; unexpected programming or network errors remain visible to the caller.
from youtube_transcript_api import (
YouTubeTranscriptApi,
NoTranscriptFound,
TranscriptsDisabled,
VideoUnavailable,
RequestBlocked,
IpBlocked,
)
api = YouTubeTranscriptApi()
def fetch_record(video_id):
try:
result = api.fetch(video_id, languages=["en"])
return {
"video_id": video_id,
"status": "ok",
"language_code": result.language_code,
"is_generated": result.is_generated,
"segments": result.to_raw_data(),
}
except NoTranscriptFound:
status = "requested_track_unavailable"
except TranscriptsDisabled:
status = "captions_unavailable"
except (RequestBlocked, IpBlocked):
status = "request_blocked"
except VideoUnavailable:
status = "video_unavailable"
return {"video_id": video_id, "status": status, "segments": None}
The status strings above are application labels introduced in this guide. They are not a response schema supplied by the package or Glasser. For a batch, count successful transcripts separately from the number of video IDs submitted, and review the failed categories before increasing volume.
When should you use the official API or a managed service?
Use the official captions API for videos you can edit
Google's captions.list returns caption-track metadata, not the caption text, and requires authorization. captions.download retrieves a track and requires the authorized user to have permission to edit the video.
Consequently, a YouTube Data API key alone does not give you a general transcript download for any public video. For your own channel or another video you are authorized to edit, evaluate the official captions workflow and its quota costs before using a separate extraction library.
Consider speech-to-text when captions are unavailable
If no usable caption track exists, speech-to-text may be an option for audio you are authorized to process. It introduces separate acquisition, transcription, timing, and quality decisions. It can also produce words that differ from the creator's edited captions, so label that output's origin explicitly.
Check a managed endpoint’s output before integrating it
Glasser’s search, inspect, run model lets you review a transcript endpoint’s accepted input, price, and outcome clauses before calling it.
Before choosing any managed service, check whether it returns available captions or performs new transcription, whether timing and language metadata are included, and how unavailable videos are represented. Those distinctions determine whether its output can replace the segments record used by your application. A managed service should be evaluated with your sample videos; this guide makes no uptime or universal-coverage claim.
Ready to turn a video into research material? Open Glasser or connect a compatible agent with
set up https://glasser.ai/SKILL.md, then ask it to inspect a transcript endpoint and retrieve the available text for your video.
Frequently asked questions
Does youtube-transcript-api need a YouTube API key?
The open-source package does not require a YouTube Data API key. It uses a separate retrieval path and can still encounter unavailable captions or blocked requests.
Why does get_transcript fail in an older example?
The current interface uses YouTubeTranscriptApi().fetch(video_id). Check the installed package version and use the matching maintainer documentation.
Can I request another language?
Yes, pass language codes as an ordered list or inspect available tracks with list(). Fetching an existing language and translating a track are separate operations.
Can it transcribe a video with no captions?
The library retrieves caption tracks; it does not create a new speech recognition transcript. Process authorized audio with a separate transcription service if that is your task.
Related posts
View more
6 min read
Dify × Glasser: Bringing Business Data into Your AI Workflows
Bring business data into your Dify workflows with Glasser. Find contacts, research companies, and retrieve search insights without leaving your workflow.

8 min read
Free Web Search API: Choose by Output, Free Limits, and Task Cost
Compare free web search API tiers for AI agents and Google SERP data. See which return links or content, then budget search and extraction together.
Make your agent
work with real data.
Search, inspect, and run data APIs through one Glasser key. Turn your next question into a result you can use.
Get started with Glasser