Dev.to WebDev 🛠 Dev 👁 0 📖 7 min read

Building a zero-dependency Python and Node client for a speech-to-text API

Most speech-to-text and text-to-speech APIs ship a fat SDK: a package manager, a dependency tree, a version matrix. We went the other way. Our transcription and synthesis service speaks plain HTTP with JSON payloads and

Most speech-to-text and text-to-speech APIs ship a fat SDK: a package manager, a dependency tree, a version matrix. We went the other way. Our transcription and synthesis service speaks plain HTTP with JSON payloads and exactly one multipart upload, and we wrote the reference clients as a proof: they talk to the API with nothing but the standard library. No requests, no axios, no node-fetch — Python's urllib and Node's built-in fetch.

This is a walkthrough of how those clients work and which parts of the API forced the design. The code is on GitHub at VavilkinAlex/bowhard-speech; the endpoint reference is at bowhard.ru/api.

No API key means the cookie jar is the client

The service does not issue API keys. It sets a signed cookie (tools_id) on your first request and counts your daily free volume — 15 minutes of transcription and 5,000 characters of synthesis — by that cookie. Your jobs are bound to the same identity.

That single decision shapes both clients: the client is a cookie jar plus a request wrapper.

In Python that is one line of standard library:

import urllib.request
from http.cookiejar import CookieJar

self._opener = urllib.request.build_opener(
    urllib.request.HTTPCookieProcessor(CookieJar())
)

Every call goes through self._opener, so Set-Cookie on the first response is replayed on all the following ones without any code from us.

Node is where it gets interesting. fetch is built in and good enough, but it is deliberately stateless about cookies — there is no jar. So the Node client carries its own, a plain Map:

this.cookies = new Map();

// when sending
if (this.cookies.size) {
  headers.cookie = [...this.cookies]
    .map(([k, v]) => `${k}=${v}`)
    .join('; ');
}

// when receiving
for (const raw of response.headers.getSetCookie?.() || []) {
  const [pair] = raw.split(';');
  const [name, ...rest] = pair.split('=');
  this.cookies.set(name.trim(), rest.join('='));
}

Two details are easy to get wrong. First, getSetCookie() — the plural accessor — is what you need; the single headers.get('set-cookie') flattens multiple cookies into one string and breaks parsing. Second, you must keep only the name=value part: the attributes after the first semicolon (Path, HttpOnly, Max-Age) are server-side instructions, not something you echo back.

One consequence worth stating plainly: one client instance is one job owner. Creating a new Client discards the cookie, and with it the ability to read the jobs it just started. If you build a worker that fans out work, keep one instance alive.

Building multipart by hand in Python

There is exactly one place where a JSON library is not enough: POST /api/tools/stt wants multipart/form-data with a file field. urllib has no multipart encoder, so the Python client assembles the body itself:

boundary = uuid.uuid4().hex
parts = []

for key, value in (extra or {}).items():
    parts.append(
        f'--{boundary}\r\n'
        f'Content-Disposition: form-data; name="{key}"\r\n\r\n'
        f'{value}\r\n'.encode('utf-8')
    )

mime = mimetypes.guess_type(filename)[0] or 'application/octet-stream'
parts.append(
    f'--{boundary}\r\n'
    f'Content-Disposition: form-data; name="{field_name}"; '
    f'filename="{os.path.basename(filename)}"\r\n'
    f'Content-Type: {mime}\r\n\r\n'.encode('utf-8')
)
parts.append(content)
parts.append(f'\r\n--{boundary}--\r\n'.encode('utf-8'))

body = b''.join(parts)
content_type = f'multipart/form-data; boundary={boundary}'

The rules that actually matter:

  • The boundary must not appear in the payload. A random 32-character hex string from uuid4() makes that a non-issue.
  • CRLF, not LF. The spec says \r\n and some parsers check. Python's email module would handle this, but then you are back to assembling anyway.
  • Extra fields come before the file part. Cheap to respect, annoying to debug when a proxy-side parser assumes order.
  • Send the basename, not the path. filename="/home/you/secret/meeting.mp3" leaks your directory layout to the server and to any log in between.
  • Always send a Content-Type per part. mimetypes.guess_type() covers the formats the service accepts (mp3, wav, ogg/opus, m4a, mp4, mov, mkv — anything ffmpeg reads); application/octet-stream is the fallback.

The same request in Node is four lines, because Node 18+ ships FormData, Blob and fetch together:

const form = new FormData();
form.append('file', new Blob([await readFile(filePath)]), basename(filePath));
const response = await this.#request('/stt', { method: 'POST', body: form });

fetch writes the boundary and the part headers itself. The interesting part is that the server cannot tell the two apart — a hand-rolled body and a FormData body are the same bytes on the wire.

Timeouts and retries

The two clients make opposite choices here, on purpose.

Python uses a single attempt with a 120-second timeout: urllib raises, the client translates the exception into its own error type, and the caller decides. Retrying a request that uploads a 200 MB recording automatically is not obviously a kindness.

Node retries — up to three attempts, with a linear backoff of 500 * (attempt + 1) milliseconds — because the failure it most often hits is a dropped connection on a long upload, and fetch gives it a clean AbortController to bound each attempt:

const controller = new AbortController();
const timer = setTimeout(() => controller.abort(), this.timeout);

The timeout is per attempt, not per operation, which is the behaviour you want: three tries of 120 seconds each, not one budget of 120 seconds split three ways.

Polling, because transcription is not instant

Uploading a file does not return text. It returns a job:

{"id": "abc123", "kind": "stt", "status": "queued"}

Transcription happens in a queue — the service processes two files at a time and the rest wait — so the time to completion depends on what else is running and on the length of your recording. There is no honest way to predict it, which is why the API is poll-based rather than webhook-based: a webhook would need a publicly reachable URL from you, and most client code does not have one.

GET /api/tools/job/{id} returns everything you need to decide what to do next:

field meaning
status queued, processing, done, error
stage what the worker is doing right now (used for progress output)
error human-readable failure reason when status is error
durationSec length of the audio
chars number of characters for synthesis jobs
costRub, free what this job cost and whether it came out of the free daily allowance
files the result files that are ready to download

Both clients wrap that in a wait() helper that polls every 3 seconds, prints progress through a callback, and gives up after an hour by default:

job = client.transcribe("meeting.mp3")
job = client.wait(job.id, on_progress=lambda j: print(j.status, j.stage))
const job = await client.transcribe('meeting.mp3');
await client.wait(job.id, { onProgress: (j) => console.log(j.status, j.stage) });

Three seconds is a deliberate compromise. A tighter loop gains nothing — the queue position changes on the scale of seconds — and every poll is a request against the same cookie-counted account.

Downloading the results

GET /api/tools/job/{id}/file/{kind} serves the artefacts, where kind is one of four:

  • txt — the plain transcript, one paragraph flow, no timecodes;
  • srt — subtitles, already wrapped to a maximum of 42 characters per line, which is the usual readability norm and the same limit our subtitle converter uses;
  • json — the same transcript with a timecode for every single word, which is what you need if you are cutting video against the speech rather than just captioning it;
  • mp3 — the synthesised audio, for text-to-speech jobs.

In Python the helpers return strings or bytes (client.text(job.id), client.srt(job.id), client.words(job.id), client.audio(job.id)); in Node they return strings or a Buffer. Both have a save() that writes straight to a path. One operational note: results live for 24 hours, and the job statistics are kept for 90 days. If a transcript matters, download it in the same run that produced it.

Errors that carry the server's own message

An HTTP client that hides the response body costs you an hour of guessing. Both clients parse the error payload before throwing, and fall back to a truncated raw body when the response is not JSON:

try:
    ...
except urllib.error.HTTPError as error:
    body = error.read()
    try:
        payload = json.loads(body.decode('utf-8'))
    except ValueError:
        payload = {"error": body.decode('utf-8', 'replace')[:300]}
    raise BowhardError(payload.get('error') or f'HTTP {error.code}',
                       error.code, payload)

The resulting BowhardError carries status and payload, so except BowhardError as e: print(e.status, e.payload) tells you whether you hit a rate limit (more than ten link imports per hour), a bad file type, or an expired job — without parsing strings.

The whole thing

Both clients are a couple of hundred lines each, under MIT, with no dependencies to install and no lockfile to audit. Installation for the Python one is a single wheel from the repository's releases; the Node one installs from GitHub Packages. If you would rather not write HTTP by hand at all, there are CLI entry points in both:

bowhard-speech transcribe soveschanie.mp3 -o protokol.txt --srt protokol.srt
bowhard-speech say "…" -o privet.mp3 --voice alena --speed 1.1
npx bowhard-speech transcribe meeting.mp3 --srt meeting.srt

The interesting lesson is not "write your own HTTP client". It is that the API was small enough to make that a reasonable afternoon's work: one cookie, one multipart endpoint, one job resource, four file kinds. If your client library needs a dependency tree, it is usually because the API asked for one.

📰 Read the original article on Dev.to WebDev

Originally published by Dev.to WebDev. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.