Dev.to AI 🤖 Ai 👁 0 📖 2 min read

Chunking YouTube transcripts for RAG: timestamps, citations and missing captions

Most RAG demos over video content start with "get the transcript", and that step breaks more often than the embedding step. Three things matter: getting timestamps, chunking so every chunk can be cited, and not paying fo

Most RAG demos over video content start with "get the transcript", and that step breaks more often than the embedding step. Three things matter: getting timestamps, chunking so every chunk can be cited, and not paying for videos with no captions.

1. Keep timestamps through chunking

If you split a transcript by characters you lose where each piece came from. Keep the caption segments (text, start, duration) and build chunks by walking segments until you reach your size limit, ending on a sentence boundary. Store the start of the first segment with the chunk.

def chunk(segments, max_chars=1000):
    out, cur, start = [], "", None
    for s in segments:
        if start is None:
            start = s["start"]
        cur += " " + s["text"].strip()
        if len(cur) >= max_chars and cur.rstrip()[-1:] in ".?!":
            out.append({"start": start, "text": cur.strip()})
            cur, start = "", None
    if cur.strip():
        out.append({"start": start, "text": cur.strip()})
    return out

2. Make citations clickable

Store a deep link with each chunk: https://www.youtube.com/watch?v=VIDEO_ID&t=SECONDS. When your assistant answers, it can link to the exact moment in the video instead of the whole hour-long talk.

3. Handle missing captions without failing the batch

Some videos have no captions or have them disabled. Collect those IDs in an errors list and move on; do not retry them forever. Prefer the requested language track, then auto-generated, then any track.

Skipping the plumbing

I packaged all of the above as an Apify actor: paste video URLs, IDs or a playlist, set chunkChars, and get text, timestamped segments and sentence-aligned chunks with deep links. No API key or cookies, and you pay only for transcripts actually delivered.

YouTube Transcript Extractor on Apify

from apify_client import ApifyClient
client = ApifyClient("<APIFY_TOKEN>")
run = client.actor("quiethand098/youtube-transcript-rag-extractor").call(run_input={
    "videos": ["https://www.youtube.com/watch?v=dQw4w9WgXcQ"],
    "chunkChars": 1000,
})
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    for c in item["chunks"]:
        print(c["link"], c["text"][:80])

Source and notes: https://github.com/quiethand098/youtube-transcript-extractor

📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.