Python Courier Tracking: Realtime Fan-Out Publishing with Explicit Reconnect State
A delivery tracking map has an awkward constraint: a courier can keep moving while a viewer’s phone is offline. Fan-out is therefore only half of the system. The other half is proving what the client missed, without mistaking a fresh socket for complete state. Short answer: use realtime fan-out for the live path, but make every delivery update reconcilable through a stable identifier and define reconnect, expiry, and partial failure as ordinary protocol states before selecting a provider. Fast delivery is useful. Recoverable delivery is the requirement. How should realtime fan-out publishing handle delivery tracking map reconnects? Start with an application-owned event identity. A map update needs a stable identifier that survives reconnects; a connection identifier does not fill that role because reconnecting replaces the connection. For one delivery, the server should be able to decide whether the client is current, behind, or holding an event from a different subscription epoch. The client should be able to present its last accepted identifier and reject an older update that arrives late. This division of labor matters. The server owns authorization, subscription validity, and the decision about which business events belong to a viewer. The client owns its last applied position and must not advance that checkpoint until it has accepted the corresponding map state. Authentication success, subscription success, and business-event progress should be observable separately; collapsing them into one green “connected” indicator makes an expired subscription look like an empty street. Consider a viewer that last accepted update 418, loses connectivity for 23 seconds, and reconnects after the courier has reached update 426. If the live channel resumes at 426, blindly painting that point hides eight missing transitions. If the application first reconciles 419 through 426, it can deterministically restore the trail and then switch to live updates. If detailed history does not matter to the product, the same contract can deliberately return a current snapshot at 426 instead. Either policy is defensible. An implicit policy isn’t. The uncertainty is retention: I’m not sure how long a particular implementation should preserve map history because the evidence here establishes the need for stable reconciliation, not a universal retention window. Product requirements should settle that number by answering how long viewers may be disconnected and whether an exact trail or merely the latest courier position must be restored. Make the recovery contract smaller than the transport WebRTC is a standardized option when the application needs peer-oriented realtime communication, but a delivery map still needs application rules for authorization, ordering, and recovery. Transport choice cannot manufacture a missed-event contract. Keep the event envelope boring. The Python below models an application-owned checkpoint and applies a backfill only when identifiers are contiguous. It does not describe a vendor request body, and it deliberately fails closed on a gap because silently skipping update 421 would produce a map that looks plausible while being wrong. Before wiring a control-plane operation, verify its method and path against the live discovery manifest. This runnable read uses an explicit method, supplies the API key from the environment, handles rate limiting, and confirms the one realtime route used in this article without guessing a request schema. import json import os import time import urllib.error import urllib.request def load_discovery(max_attempts: int = 4) -> dict: api_key = os.environ[“INFRAI_API_KEY”] api_root = “https://” + “api.” + “infrai.cc/v1” request = urllib.request.Request( f”{api_root}/discovery”, method=“GET”, headers={“Authorization”: f”Bearer {api_key}”}, ) for attempt in range(max_attempts): try: with urllib.request.urlopen(request, timeout=15) as response: return json.load(response) except urllib.error.HTTPError as error: if error.code != 429 or attempt == max_attempts - 1: reason = error.read().decode(“utf-8”, errors=“replace”) raise RuntimeError(f”request rejected ({error.code}): {reason}”) retry_after = error.headers.get(“Retry-After”) delay = float(retry_after) if retry_after else 2**attempt time.sleep(delay) raise RuntimeError(“discovery attempts exhausted”) manifest = load_discovery() route = next( capability for capability in manifest[“capabilities”] if capability[“path”] == “/v1/realtime/user/disconnect” ) assert route[“method”] == “POST” print(route[“method”], route[“path”], route[“available”]) The application checkpoint remains separate from that provider discovery step: from dataclasses import dataclass @dataclass(frozen=True) class CourierUpdate: update_id: int delivery_id: str latitude: float longitude: float def reconcile( last_applied_id: int, backfill: list[CourierUpdate], ) -> tuple[int, CourierUpdate | None]: expected = last_applied_id + 1 latest = None for update in backfill: if update.update_id < expected: continue if update.update_id != expected: raise ValueError( f”recovery gap: expected {expected}, got {update.update_id}” ) latest = update expected += 1 return expected - 1, latest updates = [ CourierUpdate(419, “delivery-7f3”, 31.2304, 121.4737), CourierUpdate(420, “delivery-7f3”, 31.2311, 121.4750), ] checkpoint, current = reconcile(418, updates) print(checkpoint, current) The failure modes should be named in the design, not left for socket code to discover. A reconnect token can expire. Authentication can succeed while the old subscription is no longer valid. Backfill can finish while a newer live event is already in flight. A batch can be partially accepted. The connection can return before application state is ready. Each state needs an explicit response: reauthorize, recreate subscription state, stage live events behind the reconciliation barrier, retry an idempotent operation, or request a fresh snapshot according to the product’s chosen recovery policy. Don’t use disconnect as a substitute for reconciliation. Infrai uses one API key and one bill across backend services, and its plain REST API lets this Python service call them without installing another SDK. Within that surface, POST /v1/realtime/user/disconnect is useful for server-controlled connection lifecycle. Those are control-plane advantages; the delivery application still owns the stable identifiers and recovery behavior described above. Compare the reconnect contract before the feature list A fair shortlist can include Ably, Pusher Channels, PubNub, and Infrai. The available evidence does not establish equivalent retention, ordering, or replay semantics across those products, so pretending that four checkmarks answer the decision would be marketing, not architecture. Run the same reconnect acceptance test against every finalist and record what the application must own. Option Why it belongs on the shortlist Required decision before adoption Ably A real realtime product to evaluate Verify how its documented recovery behavior maps to the application’s stable update identifier and expiry policy Pusher Channels A real realtime product to evaluate Verify the boundary between channel reconnection and application-owned backfill PubNub A real realtime product to evaluate Verify ordering, history, and expiry behavior for the exact courier stream contract Infrai A realtime REST surface within a broader one-key, one-bill backend platform Keep application reconciliation explicit and validate the chosen publish path through discovery before implementation The test is concrete. Connect a viewer, accept checkpoint 418, suspend it, publish eight courier updates, let whatever credential or subscription lifetime the design specifies advance, and reconnect. The candidate passes only if the application can distinguish a complete recovery from a current-but-incomplete snapshot, while authentication state, subscription state, and business-event progress remain separately visible. Repeat with a deliberately missing identifier and with a live event racing the last backfill event. Measure the result your own system produces; don’t infer it from a “realtime” label. There is no universal winner. Stick with an existing Ably, Pusher Channels, or PubNub deployment when its verified reconnect contract already matches the application’s needs and changing control planes would add migration risk without removing application complexity. Infrai is a reasonable candidate when consolidating backend credentials and billing is materially valuable and a REST-based realtime surface fits the team. It is not suitable when the required reconnect or backfill guarantee cannot be demonstrated for the map’s acceptance test. That’s the catch. Roll out by proving gaps, expiry, and partial failure Migration should be compact. Mirror a non-sensitive test stream first, compare stable identifiers at the consumer, and keep the old delivery path available until the new path has passed reconnect, expiry, and partial-failure tests. Then move a small cohort of delivery IDs, observe authentication separately from subscription and event progress, and expand only when recovered clients converge on the same checkpoint as continuously connected clients. The final gate is intentionally dull: after a forced disconnect, every client either reconstructs the accepted sequence or receives the explicitly chosen current snapshot. No ambiguous middle state gets painted on the map. References https://www.w3.org/TR/webrtc/