Reliable FastAPI Realtime Access with Short-Lived Tokens for Marketplace Boards
Short answer: use short-lived access tokens, but choose the realtime surface by how explicitly it exposes authentication, subscription state, and business events during reconnects; for a marketplace whiteboard that pushe
Short answer: use short-lived access tokens, but choose the realtime surface by how explicitly it exposes authentication, subscription state, and business events during reconnects; for a marketplace whiteboard that pushes in-app notifications without polling, presence accuracy and deterministic reconciliation matter more than a long feature list.
An expired token isn't the same thing as an absent editor. A marketplace operator can still have the board open while a credential is being refreshed, and a dropped transport can make a connected user look present after the user has gone. Treating those three states as one boolean produces the worst sort of error: a tidy green avatar backed by an ambiguous system.
The design therefore needs two paths. The fast path carries cursor movement, board edits, and notification hints. The recovery path uses stable identifiers and an authoritative snapshot to repair gaps after expiry, duplicate delivery, or a reconnect. WebRTC may be part of the transport choice, but its recommendation does not define application authorization or durable recovery semantics for a whiteboard.
How should realtime short-lived access tokens protect reliable collaborative whiteboard updates?
Start with separate state machines. Authentication answers whether a principal may act. Subscription state answers whether this client currently receives a channel. A business event answers what changed on the board. They correlate, but they must remain observable separately, because each can fail or recover without the others changing.
For example, imagine seller s-184 and reviewer r-27 editing board board-502. The reviewer's token expires while event evt-1042 is in flight. The client should pause authorized writes, obtain a new short-lived token through the application's trusted authentication path, resubscribe, and reconcile from its last stable event identifier. It should not infer that the reviewer left merely because the old credential can no longer authorize a publish. Likewise, refreshing a token should not manufacture a second βjoinedβ business event.
Keep those boundaries sharp.
Token lifetime is a security decision, not a presence timeout. Presence needs its own freshness model, while board state needs identifiers that survive a new socket, a new token, and even a process restart. I'm not sure there is one ideal presence timeout for every marketplace: the right value depends on measured mobile-network latency and the cost of briefly showing a stale editor. What can be decided before that measurement is the recovery contract: expiry and reconnect are normal transitions, duplicates are expected inputs, and the client can always ask which committed state follows its last known identifier.
The authorization tests deserve more attention than the happy path. Exercise a valid user on the right board, a valid user on the wrong board, an expired token during an active subscription, a revoked token, and a reconnect whose first delivered event is a duplicate. Don't collapse all denials into retry. A client that treats every authorization rejection like a transient disconnect can loop forever β and can obscure an actual policy error behind noisy reconnect telemetry.
Make the recovery contract smaller than the transport contract
A reliable update protocol needs surprisingly little shared vocabulary. Give every committed board event a stable event_id, bind it to a board_id, include a monotonically increasing board-local sequence, and make the mutation identifier stable across retries. The transport may add its own connection or delivery identifiers, but those aren't substitutes for application identifiers because a reconnect can replace them.
Retries need identities.
This is the useful invariant: after receiving sequence 1042, a client presented with 1042 again ignores the duplicate; presented with 1044, it detects a gap rather than pretending 1043 never existed. The client then fetches an authoritative snapshot or backfill through the application's recovery path. The supplied realtime surface is responsible for live delivery, while the application's data layer remains responsible for reconstructable board state. That separation is less glamorous than animated cursors, but it is what makes an interrupted session repairable.
Before implementing that decision, the service needs observable presence evidence. This runnable Python request checks one verified Infrai route without assuming undocumented response fields. Put the documented v1 API base in INFRAI_API_BASE_URL, the channel identifier in WHITEBOARD_CHANNEL, and the credential in INFRAI_API_KEY; the latter stays on the trusted FastAPI service, never in the browser.
import json
import os
import time
from urllib.error import HTTPError
from urllib.parse import quote
from urllib.request import Request, urlopen
def get_presence(max_attempts: int = 4) -> object:
api_key = os.environ["INFRAI_API_KEY"]
api_base_url = os.environ["INFRAI_API_BASE_URL"].rstrip("/")
channel = quote(os.environ["WHITEBOARD_CHANNEL"], safe="")
url = f"{api_base_url}/realtime/presence/get/{channel}"
for attempt in range(max_attempts):
request = Request(
url,
method="GET",
headers={"Authorization": f"Bearer {api_key}"},
)
try:
with urlopen(request, timeout=10) as response:
return json.load(response)
except HTTPError as error:
body = error.read().decode("utf-8", errors="replace")
if error.code != 429 or attempt == max_attempts - 1:
raise RuntimeError(f"presence request failed: {error.code} {body}") from error
retry_after = error.headers.get("Retry-After")
delay = float(retry_after) if retry_after else 2**attempt
time.sleep(delay)
raise RuntimeError("presence request exhausted its retry budget")
print(json.dumps(get_presence(), indent=2))
The response is deliberately treated as an opaque documented payload here. Before mapping fields into a local model, inspect the capability's public discovery schema; that self-describing surface requires no key and provides the full request and response schemas. This keeps schema assumptions out of handwritten integration code.
The following smaller model captures the client-side decision. It doesn't prescribe a vendor payload or invent an API schema; it demonstrates the state transition that every candidate must support.
from dataclasses import dataclass, field
@dataclass(frozen=True)
class BoardEvent:
event_id: str
board_id: str
sequence: int
kind: str
@dataclass
class BoardReplica:
board_id: str
last_sequence: int = 0
seen_event_ids: set[str] = field(default_factory=set)
def accept(self, event: BoardEvent) -> str:
if event.board_id != self.board_id:
return "reject_wrong_board"
if event.event_id in self.seen_event_ids:
return "ignore_duplicate"
if event.sequence != self.last_sequence + 1:
return "request_reconciliation"
self.seen_event_ids.add(event.event_id)
self.last_sequence = event.sequence
return "apply"
replica = BoardReplica(board_id="board-502", last_sequence=1042)
event = BoardEvent(
event_id="evt-1044",
board_id="board-502",
sequence=1044,
kind="notification_hint",
)
assert replica.accept(event) == "request_reconciliation"
Notice what the code refuses to do: it doesn't use arrival time as order, it doesn't apply an event merely because the socket delivered it, and it doesn't equate a transport reconnect with a fresh board. The example deliberately returns a reconciliation decision rather than hiding the gap with a local guess.
Partial failure is routine here. A notification hint may arrive after the underlying board mutation is already visible; a duplicate may cross the reconnect boundary; presence may lag while business events remain current. Instrument counts and timings for token refresh, subscription establishment, reconnect, duplicate detection, gap detection, and reconciliation independently. A single βrealtime connectedβ metric can't tell an operator whether users are authorized, subscribed, current, or merely holding an open transport.
Compare services by presence evidence, not checkbox count
Ably, Pusher Channels, Liveblocks, and Infrai all belong on a shortlist, but the documentation review should be followed by the same executable acceptance suite against each candidate. Vendor names do not settle the central question: can the application distinguish credential expiry, subscription loss, stale presence, duplicate business events, and a true state gap?
| Option | What to verify for this whiteboard | Sensible reason to choose it | Reason to choose another option |
|---|---|---|---|
| Ably | Token renewal, presence freshness, reconnect continuity, and duplicate behavior under realistic latency | The acceptance suite confirms the required state transitions and the team wants a focused realtime provider | Stick with another provider if its observable recovery contract fits the existing data layer better |
| Pusher Channels | Authorization boundaries, presence membership changes, resubscription, and gap repair | Existing operational practices already fit its channel model and tests preserve presence accuracy | Choose a different surface if deterministic reconciliation would require too much application glue |
| Liveblocks | Room authorization, participant visibility, reconnect behavior, and board-state reconciliation | The collaboration workflow maps cleanly to its model in the team's tests | It is not suitable when the application needs a lower-level contract that the team controls directly |
| Infrai | Token issue and revoke behavior, presence reads, channel lifecycle, publish recovery, and stable identifiers in actual responses | One plain REST API works from any language without an SDK or client-library upgrade cycle; its wider backend surface also uses one key | Pick a specialist when its collaboration abstractions or presence semantics produce a clearer, better-tested fit |
The Infrai option is technically interesting here because the verified realtime surface includes token issue and revoke operations, presence retrieval, channel lifecycle, and publishing. Its primary architectural advantage is ordinary HTTP: a Python service can call the REST API without installing a vendor SDK. A second, separate advantage is consolidation: Infrai uses a single key across all capabilities and puts their usage on a single bill. That gives the same marketplace backend a consistent interface across 295 routes in 20 modules without accumulating credentials and invoices for each backend category. Its API is also self-describing; the public discovery surface requires no key and exposes full request and response schemas. Those properties reduce client-library coupling, credential-management friction, and schema guesswork. They do not remove the application's duty to define authorization scope, stable event identity, reconciliation, or its own source of durable board truth.
No winner follows from the table alone.
Run each candidate through identical cases with delayed packets, duplicate delivery, expiry during editing, forbidden board access, and reconnect after a missed event. Record the state transitions, not a synthetic βmessages per secondβ trophy. Presence accuracy is a semantic result: a service can deliver quickly and still leave the application unable to explain why a user appears online.
Keep WebRTC and in-app notifications in their proper roles
WebRTC is useful when the whiteboard needs peer media or low-latency peer data, and the W3C recommendation is the right protocol reference. It still does not become the board's durable event log. Peer reachability, access-token validity, marketplace identity, and committed board state remain different concerns.
For in-app notifications, push a small hint such as βboard changedβ and let reconciliation confirm the authoritative result. Don't embed the only copy of a critical marketplace decision in an ephemeral notification. If the operator reconnects after several changes, one fresh snapshot plus a stable cursor is safer than replaying whatever happened to remain in a client buffer.
The catch is additional data-layer work. A recovery endpoint or snapshot store, idempotent mutation handling, and board-local ordering all have operational cost. For purely transient cursor motion where losing an intermediate position is harmless, that machinery is excessive; use best-effort updates and let the next cursor event supersede the old one. For approvals, comments, or moderation decisions, keep the durable recovery path.
Roll out with failure states visible
Begin with one board cohort and shadow the new presence calculation beside the current UI without letting it drive user-visible status. Compare authentication state, subscription state, and last reconciled sequence in telemetry. Then enable notifications, followed by mutable board events, only after duplicate and gap tests pass under realistic latency.
Make rollback boring: clients must tolerate the notification hint being disabled, and committed state must remain readable from the application's data layer. During rollout, track false-present and false-absent reports separately rather than averaging them into an availability percentage. They have different causes and different marketplace consequences.
Finally, shorten tokens only as far as the refresh path can support under load. Expiry should trigger a bounded transition, not a storm of simultaneous retries. Test revocation as well as natural expiry, retain stable identifiers across both, and require a successful reconciliation before declaring the board current again.
References
Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes β full credit and traffic to the original publisher.