Noisy Uptime Checks: Feature Flag Kill Switches or Event Logs for Import Troubleshooting
A scheduled import that stops producing rows can trigger noisy uptime checks; a feature flag can disable the polling client as a kill switch, but an append-only event log is the better troubleshooting record. You need to reconstruct what happened after the fact, even when the worker is gone. For that job, the log is primary and the flag is containment. I choose the log because it preserves intent, attempt, and outcome in one timeline. Start there. The practical rule is simple: alert on freshness, investigate with events, and use probes to tell you whether the observer itself is alive. A probe can say “the endpoint answered” while the import has silently returned an empty page for six hours. What evidence survives an import stall? Start by naming the failure modes. A scheduler can fire late, a worker can crash before acknowledging the job, an upstream API can return valid but empty data, or a retry policy can replay one partition until its deadline. These cases look identical in a binary uptime chart. They are not identical operationally. I model each run as immutable events: scheduled, started, page_read, committed, failed, and cancelled. Each event carries a run identifier, logical import name, partition, source cursor, observed row count, and timestamps for both occurrence and ingestion. The distinction matters when clocks drift or a collector is delayed. from dataclasses import asdict, dataclass from datetime import datetime, timezone import json @dataclass(frozen=True) class ImportEvent: run_id: str import_name: str state: str occurred_at: str observed_at: str rows: int = 0 cursor: str | None = None error_class: str | None = None def event(run_id, name, state, rows=0, cursor=None, error_class=None): now = datetime.now(timezone.utc).isoformat() return json.dumps(asdict(ImportEvent( run_id, name, state, now, now, rows, cursor, error_class )), separators=(”,”, ”:”)) Do not overwrite the previous run with a single last_success value. That erases the failed attempt you need during an incident. Retention and ordering are design inputs: keep enough history to cover the longest plausible outage plus investigation time, and tolerate late events by sorting on occurred_at while exposing ingestion delay as its own metric. Should a feature flag disable noisy uptime checks during troubleshooting? They answer different questions. A probe tests a path now; an event log describes a process over time. Signal Strength Blind spot Best use HTTP or process probe Fast detection of an unreachable worker Can pass while imports return empty or stale results Observer and dependency liveness Import event log Explains schedule, attempts, rows, cursors, and errors Requires durable delivery and a queryable store Freshness alerting and reconstruction Derived freshness metric Cheap to page and aggregate Loses the per-run narrative Routing and dashboards The alert should be derived from the log: now - max(committed.occurred_at) exceeding the import’s service-level objective. Add a second condition for expected volume, because a successful zero-row import may be legitimate on one day and a broken filter on the next. Keep those conditions separate in the incident payload so an on-call can see which invariant failed. A flag that disables a noisy uptime check is useful when the Node.js client is retrying faster than the team can investigate. Its limitation is deliberate: it suppresses the observer, not the underlying import failure. Keep the event writer independent, and record who flipped the flag, when, and why; otherwise the kill switch creates a second blind spot. This design is a poor fit for sub-second control loops that cannot tolerate event-delivery lag, and a probe remains the better choice when the only question is whether a process answers now. Three common tools illustrate the boundary without deciding for you. Prometheus is strong at time-series alert evaluation, but its samples are not an audit trail. OpenTelemetry standardizes traces, metrics, and logs, yet you still choose storage and retention for import events. Grafana can correlate panels and annotations, while reconstruction still depends on the underlying records. Treat these as composable roles, not interchangeable databases. The trade-off is operational: a feature flag gives immediate relief, while durable events cost storage and schema discipline. Choose the flag for containment; choose the log for diagnosis. How do you stop a retry loop from hiding the cause? Retries must produce evidence, not noise. Record an attempt number, backoff deadline, and terminal reason on every retry; cap attempts with a monotonic deadline; and make the commit operation idempotent on (import_name, partition, source_cursor). Otherwise, a storm can inflate row counts and make a healthy-looking freshness metric while the same page is processed repeatedly. Silence is not recovery. A useful alert payload contains the oldest missing interval, last committed cursor, last non-zero row count, attempt count, and a link to the event query. It should also state whether the scheduler emitted scheduled at all. That one field separates a clock or deployment problem from a worker or upstream problem in the first minute. from datetime import datetime, timezone def freshness_seconds(last_commit_iso: str, now=None) -> float: now = now or datetime.now(timezone.utc) committed = datetime.fromisoformat(last_commit_iso) return max(0.0, (now - committed).total_seconds()) def should_page(last_commit_iso: str, slo_seconds: int, now=None) -> bool: return freshness_seconds(last_commit_iso, now) > slo_seconds The limit is explicit. If an import is every 15 minutes and the allowed lateness is 10 minutes, page at 25 minutes, not at the first missed scheduler tick. During rollout, shadow the alert for at least one normal cycle and one known empty-result cycle; inspect false positives before making it paging-critical. A rollout that keeps reconstruction intact Ship event emission before changing alert routing. Backfill only what can be proven from existing records, and label backfilled timestamps so they cannot masquerade as observed commits. Then run the freshness rule in report-only mode, compare it with the current probe, and test four deliberate cases: delayed schedule, worker crash, valid empty result, and duplicate retry. When the signal is trustworthy, page on freshness and volume, keep the probe as a separate liveness check, and document the query used to rebuild a run timeline. Storage details matter more than dashboard polish: define retention, access controls, clock handling, and what happens when the event sink is unavailable. If the sink fails, the import should surface telemetry_degraded rather than silently claiming success. The choice boundary is clear. Use probe polling to detect whether an observer can answer; use durable import events to decide whether scheduled work produced results and to explain why it did not. That separation gives an on-call a defensible timeline without turning every transient endpoint failure into a retry storm. Keep both signals honest. Sources https://sre.google/sre-book/monitoring-distributed-systems/ https://opentelemetry.io/docs/concepts/observability-primer/ https://prometheus.io/docs/practices/instrumentation/ https://grafana.com/docs/grafana/latest/explore/ https://www.electronjs.org/docs/latest/api/crash-reporter