85k or 90k? My Agent Won’t Guess
This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content What I Built Ask most award-travel tools “what’s the cheapest way to fly SFO to Tokyo in business with my points?” and you get a confident number. Ask where the number came from and things get vague. Award travel is a domain built on sources that disagree: a printed award chart says one price, a devaluation notice from last week says another, a transfer partner page says a third thing. Pick the wrong one and you move 90,000 points into a program that wanted 95,000. Safari is an agent that answers that question the way a careful person would: It constructs the candidate routings by walking a typed graph in Sanity: points currency → transfer partner → loyalty program → award chart entry. One bounded GROQ query, no free-text guessing. It stops when its sources disagree. The ANA chart entry for SFO→NRT business carries a contradiction: the printed chart says 85,000 points, a devaluation notice says 90,000 effective 2026-09-25. Until that’s resolved, Safari refuses to price anything. Not “prices it with a caveat”. It returns NOT_COMPUTED. It resolves the conflict against the authoritative source, cites it, and writes the decision back to Sanity so the next run carries it forward. It proves the answer. A deterministic TypeScript solver returns the single cheapest valid routing and a minimality proof: every other valid routing, priced higher, in order. The model never does arithmetic. The challenge prompt literally lists “an award-travel agent that untangles which transfer partner story is current.” I took that, then aimed past it: Safari doesn’t only work out which story is current, it builds the answer from the structure and proves nothing cheaper exists. The data is synthetic and labeled that way everywhere. Real award data is a scraping and terms-of-service minefield, and the demo needs contradictions I control. It has five programs (ANA, Virgin Atlantic, Air Canada Aeroplan, British Airways, Avianca LifeMiles), three points currencies (Amex Membership Rewards, Chase Ultimate Rewards, Capital One Miles), eight transfer partners with their ratios, fourteen award chart entries across four routes, eight cited sources and two engineered contradictions. Demo Live: https://safari-tan.vercel.app There are two ways in, and both run the same pipeline against the same Sanity data: /solver: a model-free form. No LLM, no cost, works every time. Start here. /agent: Amazon Nova Pro chains the tools and narrates. It’s capped at 50 runs a day and 5 per visitor per hour, because every question costs real money. If the cap is hit, the page says so and points you to /solver. To see the gate fire, press Reset demo first. The demo state is shared, so if someone before you resolved a contradiction, you’ll see “resolved earlier, decision carried forward” instead. That’s the persistence working, but the gate firing is the better show. Try these. The chips above the /solver form fill it in: Trip What it shows Answer SFO→NRT business · Amex + Chase The gate: printed chart vs devaluation notice ANA, 90,000 JFK→LHR business · Capital One Fewer program points can still lose BA, 60,000 JFK→LHR economy · Amex A points blog says 20k, the official chart says 25k Virgin, 25,000 LAX→SYD business · Chase + Capital One Transfer ratios decide it again Aeroplan, 90,000 SFO→NRT first · Chase No program you hold flies it No valid routing Here’s what /solver shows for SFO → NRT, business, holding both currencies: 1. The contradiction, side by side with its sources. 2. The Knowledge Base agrees, independently. This panel comes from a Sanity Context Knowledge Base, read through Context MCP. More on why that’s interesting below. 3. The gate fires. No price yet. 4. Resolution, cited. The devaluation notice has higher authority and a later effective date, so it supersedes the printed chart. 5. The proof. Amex → ANA at 90,000 wins. Every other valid routing is listed, cheapest first. 6. Where the structure earns its keep. JFK→LHR business, holding only Capital One miles. Avianca wants 52,000 LifeMiles, fewer than British Airways’ 60,000. But Capital One transfers to Avianca at 2:1.5, so those 52,000 LifeMiles cost you 69,334 of your own miles. British Airways wins. A search over chart prices would pick the wrong one; the answer only exists once you follow the transfer-partner edge and its ratio. 7. The precedence rule isn’t “take the newest”. On JFK→LHR economy, a points blog published after the official Virgin chart says 20,000. The chart says 25,000. Safari resolves to the chart: source authority outranks recency, so a newer but weaker source can’t move the price. The agent path renders the same cards, interleaved with the model’s narration: Code App (Next.js 16, Vercel AI SDK v7): https://github.com/lewisawe/safari Studio (Sanity v6 schema): https://github.com/lewisawe/studio-safari The parts worth reading: solver/: the pure, deterministic solver. It returns either COMPUTED { chosen, proof[], minimalityHolds } or NOT_COMPUTED { reason, message }, which has no price field at all. It minimizes points in your currency (ceil(points / transferRatio)). Taxes are shown but labeled informational and left out of the optimization. lib/groq.ts: the one traversal query. lib/context.ts and lib/traverse.ts: the Context MCP client, the GROQ literal binder, and the fallback to @sanity/client. app/api/chat/route.ts: the agent, its tools, and the cost guard. 160 Vitest tests cover the solver, the resolution rule, the fail-closed routes, the Context client (against a mocked MCP server), the cost guard and the offline end-to-end pipeline. Fail-closed is the core design rule. Every price on screen comes from the solver’s typed output. It’s enforced in four places: The gate. The traversal embeds each chart entry’s contradiction. A contradiction with no committedResolution blocks pricing for the whole request, not just its own route. The solver throws, the routes degrade. An invariant violation (a resolved claim with no number, a zero transfer ratio, an unknown source authority) throws a typed error. Every caller turns it into NOT_COMPUTED, never a fallback price. The UI’s structural price guard. On /agent, price cells render only from the runSolver tool result. The model’s text is narration and can never fill a price slot. The model can’t feed the solver. runSolver takes only currencies, origin, destination and cabin, then re-runs the traversal on the server. I learned this one the hard way (see below). How I Used Sanity The content model Seven document types in the Studio repo: source, pointsCurrency, program, transferPartner, awardChartEntry, contradiction (two inline claim objects plus an optional inline resolution), and userDecision. Two choices made the structure load-bearing instead of decorative: Contradictions are first-class documents, not a flag on a chart entry. Each one points at the chart entry it blocks and holds both claims, each citing its own source. The query folds them into the routing rows, so the gate is decided by the same read that finds the routes. Source authority is one shared object. AUTHORITY_RANK (devaluation notice > transfer partner > official program > aggregator) lives in a dependency-free file that the solver and the Studio schema both import. The schema’s dropdown options are derived from it, so the two can’t drift, and an unknown authority throws instead of defaulting. Sanity Context, both modes I set up one Context MCP endpoint with two sources attached: the production dataset and a Knowledge Base built from it. The app opens it in both modes. GROQ mode (groq_query). The routing traversal runs through Context MCP’s groq_query tool. It’s the same fixed four-hop query the app would otherwise send through @sanity/client, and it returns the same rows. The UI shows a “read via Context MCP (GROQ mode)” badge, and if Context is unreachable it falls back to the client and says so. Knowledge Base mode (initial_context, knowledge_base_read). This is the part that surprised me. I pointed the Knowledge Base at the dataset with five document types: sources, chart entries, programs, currencies and transfer partners. I deliberately left out the contradiction documents. Each source has a short prose excerpt: the printed ANA chart says business SFO→NRT is 85,000; the devaluation notice says 90,000 from 2026-09-25. The build flagged it on its own: The SFO–NRT Business Class Award Prices entry lists the current ANA Mileage Club price as 85,000 points one-way, but the Devaluation Notices entry states that this price is being increased to 90,000 points effective 2026-09-25, superseding the 85,000-point price. (Conflict, Critical) I resolved it in favor of the devaluation notice, which became a standing instruction for future builds. So Context’s Knowledge Base and Safari’s own gate found the same disagreement independently, one from prose and one from typed references, and reached the same answer. On every run, the app reads the outline through initial_context, picks the entry about the contradicted route, and reads it with knowledge_base_read. The keyless /solver page picks the entry by plain matching, with no LLM. On /agent, the server attaches the entry to the traversal result, because Nova kept skipping the tool otherwise. Either way it’s evidence only: it’s cited on screen, but it never gates and never supplies a price. What stays out of Context: writing the resolution. Context MCP is read-only, so resolveContradiction is a small server-side write that does one patch on one document type. What the agent does with it On /agent, Nova Pro gets five tools: traverseRoutings (through Context MCP, with the Knowledge Base entry attached), readContradictions, readKnowledgeBase, resolveContradiction and runSolver. A typical run calls them in that order, narrates the contradiction with both sources, resolves it, and quotes the price only after the solver returns. Sanity Project Details Project ID: 62hh3v9t Dataset: production Schema: https://github.com/lewisawe/studio-safari (7 document types; also deployed to the Content Lake) Sanity Context: one MCP endpoint (safari-award-routing) serving the dataset in GROQ mode, plus the Knowledge Base Safari award-travel facts, built from the same dataset Seed: 40 synthetic documents with fixed IDs (5 programs, 3 currencies, 8 transfer partners, 14 chart entries, 8 sources, 2 contradictions), plus one demoUsage counter document per day for the agent’s cost cap Every document is synthetic. Nothing here is real award pricing. My Build Process I built this with Kiro CLI, mostly by orchestrating workflows: a design agent wrote the spec, a reviewer pushed back on it until it held up, a planner split it into features, and coder/reviewer loops built them. I steered, verified and made the calls. The parts that broke are the most useful bits of this post. The model invented its own inputs. My first runSolver accepted the traversal rows from the model. Nova Pro sometimes passed rows it had made up. The fix wasn’t a stricter prompt. runSolver now takes only the trip parameters and re-runs the traversal on the server, so there’s nothing for the model to fabricate. Nova Pro ignores prompt rules, so I stopped relying on them. In one test run it skipped the Knowledge Base tool, opened its answer with a literal