Attributing Changed DNS Records in Edtech Zone Drift Reconciliation
TL;DR: For edtech onboarding, treat DNS as current state, not a change log. Store an intent ledger before every write, capture a normalized observation afterward, and require matching authorization evidence before accepting drift. That is the least complex design that can prove domain control without confusing a visible record with proof of who created it. Choice Attribution strength Deliverability evidence Operational weight Compare live DNS with desired state Low Shows current publication Low Add periodic signed snapshots Medium Shows when an observed value changed Medium Pair an intent ledger with snapshots and control-plane audit events High Connects approval, publication, and later drift Higher Recommendation: use the third approach for onboarding gates. A diff can detect an unexpected ownership or mail-policy record, but it cannot name the actor. Attribution needs evidence from the system where the mutation was authorized and executed. This distinction matters in education platforms. A school may need to prove control of a domain before onboarding completes, while mail from that domain must remain aligned with its published policy. One stray TXT mutation can look harmless and still invalidate the evidence bundle that an automated gate relies on. Fast setup is useful. Unexplained state is not. How do you find who changed DNS records in a zone? A DNS answer tells a resolver what data is currently available. It does not carry the identity of the person, deployment job, registrar workflow, or API credential that changed the zone. Asking a recursive lookup to identify the writer is asking the data plane for control-plane history. The information is absent. This is the first debugging fork. If the question is “what value is published?”, query DNS from more than one vantage point and save the complete answer. If the question is “who changed it?”, inspect the authoritative change path: approval records, infrastructure runs, provider audit events, credential identity, and timestamps. Do not collapse those questions into one green check. For email, DMARC makes the boundary concrete. RFC 7489 defines DMARC policy discovery through a DNS TXT record and describes reporting that gives domain owners feedback about message authentication. The policy record is publication evidence. Aggregate or failure reports are delivery-related evidence. Neither is an author ledger for the zone. That leads to a blunt rule: never infer authorship from a DNS value alone. Matching text proves equality, not provenance. The diff is only a clue. The two criteria that decide whether the evidence is useful The first criterion is causal attribution. Every automation write needs a durable intent entry created before mutation. It should identify the domain, record key, normalized previous and desired values, an actor or workload identity, a request identifier, and the authorization that allowed the change. After the write, attach the control-plane event identifier when one exists. A human edit needs the same shape of evidence, even if its approval originated in a ticket rather than a deployment. The second criterion is deliverability relevance. An ownership challenge and a mail policy record can both be TXT data, but they do different jobs. The onboarding gate should isolate its exact owner name and expected token. Mail checks should independently observe the relevant policy owner name and retain the raw result used for evaluation. RFC 7489 specifies the DMARC record location and its tag-based content; parsing that record as an opaque substring is brittle. Parse fields, preserve the raw value, and report malformed or multiple-policy observations instead of silently choosing one. I would benchmark this pipeline by time-to-first-trustworthy-verdict, not raw lookup latency. A fast resolver followed by a manual hunt through 3 dashboards is slow DX. The useful clock starts when onboarding requests verification and stops when the system can emit one of 3 explicit outcomes: verified with evidence, pending because observations have not converged, or blocked by unauthorized drift. Keep those outcomes separate. A snapshot is trustworthy only when normalization is deterministic. DNS names are compared case-insensitively, while application tokens may have their own comparison rules. Record ordering should not manufacture drift. TTL changes may be operationally relevant but should not masquerade as an ownership-token change. Define these decisions once, version them, and include the normalizer version in each snapshot. Config sprawl starts when every worker invents its own comparison logic. A small reconciliation core The core does not need a provider-shaped SDK object. Give it desired records, observed records, and independent evidence about authorized writes. The example below deliberately refuses to assign an actor when evidence is missing. type RecordKind = “TXT” | “CNAME” | “MX”; type RecordKey = { name: string; kind: RecordKind; }; type RecordState = RecordKey & { values: string[]; ttl: number; }; type WriteEvidence = RecordKey & { actor: string; requestId: string; approvedAt: string; before: string[]; after: string[]; }; type Drift = { key: RecordKey; expected: string[]; observed: string[]; attribution: “authorized” | “unattributed”; evidence?: WriteEvidence; }; const normalizeName = (name: string): string => name.trim().replace(/.$/, "").toLowerCase(); const normalizeValues = (values: string[]): string[] => […values].map((value) => value.trim()).sort(); const sameValues = (left: string[], right: string[]): boolean => JSON.stringify(normalizeValues(left)) === JSON.stringify(normalizeValues(right)); const sameKey = (left: RecordKey, right: RecordKey): boolean => normalizeName(left.name) === normalizeName(right.name) && left.kind === right.kind; export function reconcile( desired: RecordState[], observed: RecordState[], writes: WriteEvidence[], ): Drift[] { const drift: Drift[] = []; for (const expected of desired) { const actual = observed.find((item) => sameKey(item, expected)); const observedValues = actual?.values ?? []; if (sameValues(expected.values, observedValues)) continue; const evidence = writes.find( (item) => sameKey(item, expected) && sameValues(item.after, observedValues), ); drift.push({ key: { name: expected.name, kind: expected.kind }, expected: normalizeValues(expected.values), observed: normalizeValues(observedValues), attribution: evidence ? “authorized” : “unattributed”, evidence, }); } return drift; } This function is intentionally incomplete at the network edge. The collector must query authoritative data, record collection time and response details, and avoid claiming convergence from one cached view. The adapter that loads audit events must verify identity and time bounds. Those concerns belong outside reconciliation because they fail differently. A lookup timeout is not drift; a malformed value is not absence; an unexplained mutation is not automatically malicious. The missing-record case deserves special care. During onboarding, deletion of the challenge record may mean a legitimate cleanup, a premature manual edit, or observation through stale cache. Return “pending” until the observation policy is satisfied. Return “blocked” only when the evidence rules justify it. False certainty is worse than a slower verdict. Debug the timeline, not just the diff Start with the earliest known-good snapshot and the first bad snapshot. Bound the mutation window. Then correlate authorized intents and control-plane events inside that window. If an audit event matches the key and resulting value, attribution is supported. If a deployment claims success but no matching observed state appears, keep the write and observation as separate events; retries and delayed visibility must not rewrite history. For a concrete trace, suppose the stored intent says the onboarding worker planned value A, the next snapshot still shows A, and a later snapshot shows value B. An audit event for a human identity inside that bounded interval records B. Reconciliation can attach that event and report an attributed change. Remove the audit event, and the same pair of snapshots supports only a chronology: B appeared between 2 observations. It does not support a person or workload name. Remove the earlier snapshot too, and even the useful time bound disappears. Each missing layer weakens a different claim, so the verdict should expose which layer is absent rather than flattening all 3 cases into “drift detected.” Next, check for competing writers. Infrastructure automation, an onboarding service, and a human administration path should not share an anonymous credential. Distinct identities turn a vague timestamp correlation into useful evidence. They also make revocation surgical. Identity is the expensive part. Finally, replay reconciliation from stored inputs. The same desired snapshot, observed snapshot, evidence set, and normalizer version should produce the same verdict. This gives a support engineer something better than screenshots: a compact evidence bundle that can be inspected and rerun. It also keeps the hot path small. No sprawling provider configuration belongs in the decision function. Observability should follow the same boundaries. Count lookup failures separately from drift findings. Track time from requested verification to a stable verdict. Alert on unattributed mutations of protected names, repeated oscillation between values, and evidence records that never acquire a matching observation. Avoid high-cardinality record values in metric labels; keep sensitive tokens in access-controlled evidence storage instead. When the lighter choices are better Desired-versus-live comparison is enough for a disposable preview domain or a check whose only consequence is asking a developer to retry. It is cheap to reason about and quick to ship. Its limitation is plain: it detects difference but does not attribute change. Periodic signed snapshots are the runner-up when control-plane audit events are unavailable or fragmented across administrators. The trade-off is weaker identity evidence. Snapshots can narrow the change window and show that retained evidence was not altered later, but they cannot prove which human or workload caused the mutation unless identity-bearing authorization records are joined to them. There is also a point where automatic repair is the wrong default. For an onboarding ownership token, restoring the declared value may be safe after the domain holder explicitly requests verification. For mail-policy data, an unexplained change deserves review before a reconciler overwrites it. The policy may encode an intentional operational decision, and RFC 7489 gives receivers behavior to apply from that published policy. Blind repair destroys evidence and can reverse a legitimate change. The final decision is narrow: use diffs for detection, snapshots for chronology, and authorization plus control-plane events for attribution. Gate edtech onboarding on evidence you can replay. Treat deliverability as its own observed outcome, not as a promise inferred from a successful DNS write. References RFC 7489, Domain-based Message Authentication, Reporting, and Conformance (DMARC): https://datatracker.ietf.org/doc/html/rfc7489