Node.js Feature Flag Kill Switch During Import Outages (with Health Monitoring)
Short answer: pair an external heartbeat monitor with health metrics, error events, and a feature flag that can stop the broken import path. For a small customer-support SaaS, that is enough to contain damage quickly. It is not enough to explain the incident unless every signal carries the same run identifier and the flag changes are recorded somewhere you control. The deciding constraint is incident reconstruction. A scheduled import can stop producing tickets without throwing an exception, so logs alone cannot prove that a run happened. I want one timeline that answers four questions: Was the job due? Did it start? How many records did it produce? When did the kill switch change? This is a deliberately small setup. The flag layer needs set, toggle, rollout, and value checks for basic response. Clients have to poll, and there is no built-in flag audit trail, evaluation analytics, dependency graph, or push update. Treat it as an emergency brake, not an enterprise feature-management system. Why can’t logs tell me that an import never ran? Logs are event streams. The Twelve-Factor App makes that model explicit, and it is useful right up to the point where the missing event is the incident. No log line can announce that the scheduler failed to invoke a job. Silence is data. For the support-import case, send a heartbeat to an external dead-man’s-switch service after a successful run. Healthchecks is the obvious focused product to evaluate for that job. If the expected ping does not arrive, it can detect the silent gap. This complements application health monitoring; it does not replace it. Then attach a stable runId to the local events that do exist. The minimum sequence I care about is scheduled, started, source_read, result_written, and completed. Record a separate flag_changed event in the same incident ledger whenever an operator applies the kill switch. Without that entry, a responder can see recovery but cannot establish which intervention preceded it. Keep the measurement narrow. results_written and errors_total are useful. A dashboard with forty panels is config bloat wearing a tie. Build the smallest useful incident ledger This TypeScript program reconstructs a run from newline-delimited JSON. It is intentionally vendor-neutral. Pipe captured events into it during an incident, and it reports missing steps plus the state of the import flag. import { readFile } from “node:fs/promises”; type EventName = | “scheduled” | “started” | “source_read” | “result_written” | “completed” | “flag_changed”; type IncidentEvent = { at: string; runId: string; name: EventName; count?: number; enabled?: boolean; }; type Reconstruction = { runId: string; firstSeen: string; lastSeen: string; resultsWritten: number; missing: EventName[]; importEnabled: boolean | “unknown”; }; const required: EventName[] = [ “scheduled”, “started”, “source_read”, “result_written”, “completed”, ]; function reconstruct(events: IncidentEvent[]): Reconstruction[] { const byRun = new Map<string, IncidentEvent[]>(); for (const event of events) { const current = byRun.get(event.runId) ?? []; current.push(event); byRun.set(event.runId, current); } return […byRun].map(([runId, runEvents]) => { const ordered = runEvents.sort((a, b) => a.at.localeCompare(b.at)); const names = new Set(ordered.map((event) => event.name)); const latestFlag = ordered.filter((event) => event.name === “flag_changed”).at(-1); return { runId, firstSeen: ordered[0].at, lastSeen: ordered.at(-1)!.at, resultsWritten: ordered .filter((event) => event.name === “result_written”) .reduce((sum, event) => sum + (event.count ?? 0), 0), missing: required.filter((name) => !names.has(name)), importEnabled: latestFlag?.enabled ?? “unknown”, }; }); } const inputPath = process.argv[2]; if (!inputPath) throw new Error(“Usage: tsx reconstruct.ts events.ndjson”); const raw = await readFile(inputPath, “utf8”); const events = raw .split(“\n”) .filter(Boolean) .map((line) => JSON.parse(line) as IncidentEvent); console.log(JSON.stringify(reconstruct(events), null, 2)); A healthy run emits all five required stages and at least one produced result when the upstream source contains work. A run with scheduled but no started points toward invocation. started without completed narrows the search to execution. A completed run with zero results is different again: it might be valid, so compare it with known source volume before declaring failure. That distinction matters. Fast disabling is useful, but disabling every zero-result run can turn an empty upstream queue into a self-inflicted outage. The monitor should alert; a human or a carefully bounded policy should decide when evidence warrants the flag change. That is the trap. Should a feature flag kill switch fire during an import outage? Poll the flag before starting a new import and again before committing a large batch. Polling has a real consequence: the maximum containment delay is bounded by the polling interval plus whatever unit of work cannot be interrupted. A flag check every 30 seconds does not stop a ten-minute atomic write in 30 seconds. Design smaller commit units if that bound is unacceptable. Here is the smallest Infrai check I would put at the boundary. The base URL is supplied through configuration because this unlinked comparison intentionally contains no vendor URL. The code makes no assumption about the undocumented response fields; it only returns the verified service response to the caller. It also backs off on rate limits and exposes real 4xx bodies. const apiKey = process.env.INFRAI_API_KEY; const baseUrl = process.env.INFRAI_BASE_URL; const flagKey = process.env.IMPORT_FLAG_KEY ?? “support-import”; if (!apiKey || !baseUrl) { throw new Error(“INFRAI_API_KEY and INFRAI_BASE_URL are required”); } async function readFlag(attempt = 0): Promise${baseUrl}/v1/flags/is_enabled/${encodeURIComponent(flagKey)}, { method: “GET”, headers: { Authorization: Bearer ${apiKey} }, }, ); if (response.status === 429 && attempt < 4) { const retryAfter = Number(response.headers.get(“retry-after”)); const delayMs = Number.isFinite(retryAfter) ? retryAfter * 1_000 : 500 * 2 ** attempt; await new Promise((resolve) => setTimeout(resolve, delayMs)); return readFlag(attempt + 1); } if (!response.ok) { throw new Error(Flag check failed (${response.status}): ${await response.text()}); } return response.json() as Promise