Chunking YouTube transcripts for RAG: timestamps, citations and missing captions
Most RAG demos over video content start with "get the transcript", and that step breaks more often than the embedding step. Three things matter: getting timestamps, chunking so every chunk can be cited, and not paying fo
Most RAG demos over video content start with "get the transcript", and that step breaks more often than the embedding step. Three things matter: getting timestamps, chunking so every chunk can be cited, and not paying for videos with no captions.
1. Keep timestamps through chunking
If you split a transcript by characters you lose where each piece came from. Keep the caption segments (text, start, duration) and build chunks by walking segments until you reach your size limit, ending on a sentence boundary. Store the start of the first segment with the chunk.
def chunk(segments, max_chars=1000):
out, cur, start = [], "", None
for s in segments:
if start is None:
start = s["start"]
cur += " " + s["text"].strip()
if len(cur) >= max_chars and cur.rstrip()[-1:] in ".?!":
out.append({"start": start, "text": cur.strip()})
cur, start = "", None
if cur.strip():
out.append({"start": start, "text": cur.strip()})
return out
2. Make citations clickable
Store a deep link with each chunk: https://www.youtube.com/watch?v=VIDEO_ID&t=SECONDS. When your assistant answers, it can link to the exact moment in the video instead of the whole hour-long talk.
3. Handle missing captions without failing the batch
Some videos have no captions or have them disabled. Collect those IDs in an errors list and move on; do not retry them forever. Prefer the requested language track, then auto-generated, then any track.
Skipping the plumbing
I packaged all of the above as an Apify actor: paste video URLs, IDs or a playlist, set chunkChars, and get text, timestamped segments and sentence-aligned chunks with deep links. No API key or cookies, and you pay only for transcripts actually delivered.
YouTube Transcript Extractor on Apify
from apify_client import ApifyClient
client = ApifyClient("<APIFY_TOKEN>")
run = client.actor("quiethand098/youtube-transcript-rag-extractor").call(run_input={
"videos": ["https://www.youtube.com/watch?v=dQw4w9WgXcQ"],
"chunkChars": 1000,
})
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
for c in item["chunks"]:
print(c["link"], c["text"][:80])
Source and notes: https://github.com/quiethand098/youtube-transcript-extractor
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.