Dev.to AI ๐Ÿค– Ai ๐Ÿ‘ 0 ๐Ÿ“– 7 min read

Add AI meeting summaries to any app with Zoom AI Services in 15 minutes

Add AI meeting summaries to any app with Zoom AI Services in 15 minutes If you have meeting recordings (standups, sales calls, support calls) and you want summaries + action items out of them without running your own s

Add AI meeting summaries to any app with Zoom AI Services in 15 minutes

If you have meeting recordings (standups, sales calls, support calls) and you want summaries + action items out of them without running your own speech models, two REST calls are all it takes:

  1. Scribe โ€” turn the audio file into a transcript (with speakers and timestamps).
  2. Summarizer โ€” turn the transcript into a recap, a detailed summary, and action items.

Both are part of Zoom AI Services, Zoom's developer API platform. The developer portal lives at zoom.ai (verified October 2026 โ€” that's Zoom's own portal, not a third party). No SDK needed โ€” it's plain REST.

Total code: about 60 lines of Python.

Prerequisites

  • Python 3.9+ with requests (pip install requests).
  • A Zoom AI Services developer portal account. Sign in at zoom.ai, open the API keys section, and create a key with the Scribe and Summarizer scopes. The key is shown once โ€” copy it immediately and store it server-side.

That's it for credentials. One important rule: API keys are server-to-server only. Never put the key in client-side code or ship it in a mobile/web app bundle. All calls in this tutorial run from a backend.

Try it with zero setup first

Not ready to create a key? The portal has an interactive playground with sample data โ€” no setup required. I'd suggest poking it for two minutes so you know what the responses look like before wiring up the script.

Step 1: Transcribe audio with Scribe (Fast mode)

Scribe has three modes: Live (real-time transcription over a secure WebSocket at wss://api.zoom.us/v2/aiservices/scribe/live), Fast (synchronous โ€” POST your audio bytes, the response returns when transcription completes), and Batch (async jobs API for long files, with webhook callbacks). For this tutorial we use Fast mode, the simplest. Fast mode handles files up to 5 minutes; longer recordings go through Batch (up to 6 hours per file).

Input is the raw audio bytes: WAV, Opus, and ยต-law encodings (audio/wav, etc. via Content-Type). 9 locales: en-US, de-DE, es, fr-FR, it-IT, ja-JP, ko-KR, pt-BR, zh-CN.

Endpoint: POST https://api.zoom.us/v2/aiservices/scribe/transcribe

Auth: one header โ€” x-api-key: <your API key>. That's the whole auth story per the portal docs.

Options go as query parameters (model, language, diarization, text_polishing, profanity_filter, word_time_offsets, channel_separation, custom_terms):

curl -X POST "https://api.zoom.us/v2/aiservices/scribe/transcribe?model=zoom-scribe-en&language=en-US&diarization=true&text_polishing=true" \
  -H "x-api-key: $ZOOM_API_KEY" \
  -H "Content-Type: audio/wav" \
  --data-binary @meeting.wav

Setting diarization=true adds a speaker label (Speaker 1, Speaker 2, โ€ฆ) to each segment โ€” worth it for meetings. Set Content-Type to match your file.

Response (200) โ€” synchronous; it returns after transcription completes:

{
  "request_id": "req_123",
  "duration_sec": 181.04,
  "result": {
    "text_display": "So, let's have a good conversation, my friend. Do you think...",
    "segments": [
      {
        "id": "seg_1",
        "start": 0.0,
        "end": 5.2,
        "channel": 1,
        "speaker": "Speaker 1",
        "text": "So, let's have a good conversation, my friend.",
        "words": []
      }
    ]
  },
  "model": "zoom-scribe-en",
  "usage": { "input_units": 181.04, "unit_type": "seconds" }
}

With word_time_offsets=true the words arrays carry per-word timestamps; channel_separation=true splits stereo channels.

In Python:

import os, requests

API_BASE = "https://api.zoom.us/v2/aiservices"
API_KEY = os.environ["ZOOM_API_KEY"]

def transcribe(audio_path: str, language: str = "en-US") -> dict:
    with open(audio_path, "rb") as f:
        audio_bytes = f.read()
    resp = requests.post(
        f"{API_BASE}/scribe/transcribe",
        headers={"x-api-key": API_KEY, "Content-Type": "audio/wav"},
        params={"model": "zoom-scribe-en", "language": language,
                "diarization": "true", "text_polishing": "true"},
        data=audio_bytes,
        timeout=300,
    )
    resp.raise_for_status()
    return resp.json()["result"]

Tip: feed the Summarizer a speaker-labeled transcript โ€” it's much better at action-item ownership. Build it from the segments:

def segments_to_transcript(result: dict) -> str:
    lines = []
    for seg in result.get("segments", []):
        speaker = seg.get("speaker") or "Speaker"
        lines.append(f"{speaker}: {seg.get('text', '').strip()}")
    text = "\n".join(lines)
    return text or result.get("text_display", "")

Step 2: Summarize the transcript (Summarizer Fast mode)

The Summarizer API takes transcript text (from any system โ€” it doesn't have to come from Scribe) and returns rendered, display-ready text in result.text. Fast mode accepts one inline transcript (up to 96 KB) and returns synchronously.

Endpoint: POST https://api.zoom.us/v2/aiservices/summarizer/summarize

The request is flat JSON โ€” no config wrapper. transcript is a JSON-encoded string of speaker/text turns:

curl -X POST https://api.zoom.us/v2/aiservices/summarizer/summarize \
  -H "x-api-key: $ZOOM_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "summary_type": "CONVERSATION",
    "task": "full_summary",
    "transcript": "[{\"speaker\": \"Speaker 1\", \"text\": \"Let us finalize the timeline.\"}, {\"speaker\": \"Speaker 2\", \"text\": \"We need one more week for validation.\"}]",
    "language": "en-US"
  }'

Four tasks to choose from, set in task:

task What result.text contains
recap A concise Recap: paragraph
summary A detailed Summary: section
action_items An Action Items: section grouped by owner
full_summary All three โ€” recap + summary + action items

summary_type is "CONVERSATION" (default; also "GENERIC", "MEETING_NOTES"), language is a BCP-47 output locale (14 supported, en-US default).

Response (200) โ€” the exact shape, from the official docs:

{
  "request_id": "req_123",
  "task": "full_summary",
  "result": {
    "text": "Recap:\n\nThe team agreed to delay the launch by one week.\n\nSummary:\n\n..."
  }
}

For full_summary, result.text is rendered text, ready to display โ€” recap, summary, and action items grouped by owner.

In Python:

import json

def summarize(transcript: str, task: str = "full_summary") -> str:
    turns = []
    for line in transcript.splitlines():
        if ":" in line:
            speaker, text = line.split(":", 1)
            turns.append({"speaker": speaker.strip(), "text": text.strip()})
    resp = requests.post(
        f"{API_BASE}/summarizer/summarize",
        headers={"x-api-key": API_KEY, "Content-Type": "application/json"},
        json={"summary_type": "CONVERSATION", "task": task,
              "transcript": json.dumps(turns), "language": "en-US"},
        timeout=120,
    )
    resp.raise_for_status()
    data = resp.json()
    result = data.get("result") or {}
    return result.get("text") or data.get("text") or data.get("summary") or ""

Step 3: Put it together

One script: audio file in, meeting summary out. Save as summarize_meeting.py:

#!/usr/bin/env python3
"""Audio file -> transcript -> meeting summary, via Zoom AI Services (Scribe + Summarizer).

Usage:
    ZOOM_API_KEY=<your key> python summarize_meeting.py meeting.wav

The key is server-side only -- never expose it to clients.
"""
import json, os, sys, requests

API_BASE = "https://api.zoom.us/v2/aiservices"
API_KEY = os.environ["ZOOM_API_KEY"]

def transcribe(audio_path: str, language: str = "en-US") -> dict:
    with open(audio_path, "rb") as f:
        audio_bytes = f.read()
    resp = requests.post(
        f"{API_BASE}/scribe/transcribe",
        headers={"x-api-key": API_KEY, "Content-Type": "audio/wav"},
        params={"model": "zoom-scribe-en", "language": language,
                "diarization": "true", "text_polishing": "true"},
        data=audio_bytes,
        timeout=300,
    )
    resp.raise_for_status()
    return resp.json()["result"]

def segments_to_transcript(result: dict) -> str:
    lines = [
        f"{seg.get('speaker') or 'Speaker'}: {seg.get('text', '').strip()}"
        for seg in result.get("segments", [])
    ]
    return "\n".join(lines) or result.get("text_display", "")

def summarize(transcript: str, task: str = "full_summary") -> str:
    turns = [{"speaker": s.strip(), "text": t.strip()}
             for s, _, t in (line.partition(":") for line in transcript.splitlines()) if t.strip()]
    resp = requests.post(
        f"{API_BASE}/summarizer/summarize",
        headers={"x-api-key": API_KEY, "Content-Type": "application/json"},
        json={"summary_type": "CONVERSATION", "task": task,
              "transcript": json.dumps(turns), "language": "en-US"},
        timeout=120,
    )
    resp.raise_for_status()
    data = resp.json()
    result = data.get("result") or {}
    return result.get("text") or data.get("text") or data.get("summary") or ""

def main() -> None:
    audio_path = sys.argv[1]
    print("Transcribing...", flush=True)
    result = transcribe(audio_path)
    transcript = segments_to_transcript(result)
    print(f"Transcript ready ({len(transcript)} chars). Summarizing...\n", flush=True)
    print(summarize(transcript))

if __name__ == "__main__":
    main()

Run it:

ZOOM_API_KEY=your_key_here python summarize_meeting.py meeting.wav

That's the whole pipeline: audio bytes โ†’ diarized transcript โ†’ recap + summary + action items. (Fast mode caps at 5 minutes of audio; longer files go through the Batch jobs API.)

On pricing: Zoom AI Services are usage-based via prepaid credits. I don't have verified per-minute rates, so check the current numbers at zoom.us/pricing/developer before you estimate costs.

Troubleshooting

401 Unauthorized. Most common causes: the key is wrong or expired (remember, it's shown only once at creation โ€” regenerate it in the portal if you lost it), or the wrong header name. Per the portal docs, auth is the x-api-key header only.

Unsupported format / fetch errors. Scribe Fast mode takes the raw audio bytes (--data-binary), not a URL โ€” set Content-Type to match (audio/wav, audio/mpeg, โ€ฆ). Keep Fast-mode files under 5 minutes; longer audio goes through Batch.

429 rate limits. Back off and retry with exponential backoff. Exact rate limits and quotas depend on your plan โ€” check the portal rather than guessing.

Diarization quality. Heavy background noise, overlapping speech, and very short clips all degrade speaker separation. If labels matter to you (they should โ€” the Summarizer groups action items by owner), a quick pass with a noise-reduction tool before upload pays off.

Transcript too long for Summarizer Fast mode. Inline input caps at 96 KB. For longer meetings, split the transcript into chunks and summarize each, or move to Batch mode.

Where to go next

  • Batch mode (both APIs): submit async jobs, poll GET /aiservices/summarizer/jobs/{jobId} or .../scribe/jobs/{jobId}, or get webhook callbacks. Webhooks are signed โ€” verify with the x-zm-signature (HMAC-SHA256, sha256= prefix) and x-zm-request-timestamp headers against your secret.
  • Live mode: real-time transcription by streaming audio over wss://api.zoom.us/v2/aiservices/scribe/live โ€” one completed segment per detected speech turn.
  • Translator API: localize your summaries with another single call.
  • Reference implementations: ai-services-quickstart (Express + React playground), scribe-quickstart, and the skills repo runbooks for Scribe/Summarizer, webhooks, and troubleshooting.
  • Docs: AI Services overview and the Scribe API docs.

Disclosure: this article was drafted with AI assistance; all API calls and code were checked against the official Zoom AI Services portal docs (October 2026), and the Scribe flow was run live in the portal playground.

๐Ÿ“ฐ Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes โ€” full credit and traffic to the original publisher.