Claude Haiku 5.5 Is Cheap Until 100k Tokens. Build a Tiny Price Gate in TypeScript.
On October 7, 2026, Anthropic released Claude Haiku 5.5. The model id is claude-haiku-5-5. For prompts up to 100,000 tokens, input is $0.10 per million tokens and output is $0.50. That is 90% below Haiku 4.5’s $1.00 and $5.00. Cross 100,000 tokens and the rates become $0.50 and $2.50. That is a 50% cut, not a 90% cut. Anthropic also says the new tokenizer uses about 30% more tokens than Haiku 4.5 for the same text. A cheaper token is not automatically a cheaper task. And Haiku 5.5 is the first Haiku with an effort dial. The API default is medium. low, high, xhigh, and max are real settings. They change how much the model thinks, which changes the bill. The list price does not tell you that. The useful question is what your code does before the request leaves. The gate This toy does not call Claude. You hand it token counts you already have, a task name, and an effort. It answers ALLOW, REVIEW, or REFUSE, and it prints the list-price bill in USD. Prompt size is input tokens plus cache reads plus cache writes. At or under 100,000, the short rates apply. Over that line, the long rates apply and the decision is REVIEW. ALLOW is only for five narrow jobs: classify, extract, route, summarize, and compact. Effort must be low or medium. If you omit effort, the gate stamps medium, because that is the Claude API default for this model. agentic-coding and computer-use are REVIEW even on a short prompt. Anthropic’s own Terminal-Bench 4.0 numbers are 39.2% for Haiku 5.5 and 70.6% for Sonnet 5.5. A lower price is not the same job. high, xhigh, and max are also REVIEW. This program has no eval that says those levels pay for themselves. Bad counts and unknown labels are REFUSE. A negative token count should not become a discount. Run it The full source is below. Save it as src/gate.ts. Node 22 or newer. No packages. No key. node —experimental-strip-types src/gate.ts /** * Price gate for Claude Haiku 5.5. * No network. No API key. Rates are list prices in microdollars per million tokens. * * Sources, read 2026-10-08: * - https://www.anthropic.com/claude-haiku-5-5 (October 7, 2026) * - https://platform.claude.com/docs/en/models/haiku-5-5/overview * - https://platform.claude.com/docs/en/build-with-claude/effort */ const MILLION = 1_000_000n; const PROMPT_LIMIT = 100_000n; const RECOUNT_NUM = 130n; const RECOUNT_DEN = 100n; const ALLOW_TASKS = new Set([ “classify”, “extract”, “route”, “summarize”, “compact”, ]); const REVIEW_TASKS = new Set([“agentic-coding”, “computer-use”]); const ALLOW_EFFORT = new Set([“low”, “medium”]); const ALL_EFFORT = new Set([“low”, “medium”, “high”, “xhigh”, “max”]); type Effort = “low” | “medium” | “high” | “xhigh” | “max”; type Ttl = “5m” | “1h”; type Tier = “short” | “long”; type Rates = { input: bigint; output: bigint; cacheRead: bigint; cacheWrite5m: bigint; cacheWrite1h: bigint; }; const HAIKU_55: Record<Tier, Rates> = { short: { input: 100_000n, output: 500_000n, cacheRead: 10_000n, cacheWrite5m: 125_000n, cacheWrite1h: 200_000n, }, long: { input: 500_000n, output: 2_500_000n, cacheRead: 50_000n, cacheWrite5m: 625_000n, cacheWrite1h: 1_000_000n, }, }; const HAIKU_45_5M: Rates = { input: 1_000_000n, output: 5_000_000n, cacheRead: 100_000n, cacheWrite5m: 1_250_000n, cacheWrite1h: 0n, }; const SONNET_55_CACHE_READ_NEW = 100_000n; const SONNET_55_CACHE_READ_OLD = 200_000n; export type Job = { name: string; task: string; effort?: string; inputTokens: number; outputTokens: number; cacheReadTokens: number; cacheWriteTokens: number; cacheWriteTtl?: Ttl; recountFromHaiku45?: boolean; }; export type Decision = “ALLOW” | “REVIEW” | “REFUSE”; export type Quote = { name: string; decision: Decision; reason: string; model: “claude-haiku-5-5”; effort: Effort | null; assumedDefaultEffort: boolean; tier: Tier | null; promptTokens: bigint | null; haiku55Usd: string | null; haiku45Usd: string | null; savingsBps: number | null; sonnetCacheReadNewUsd: string | null; sonnetCacheReadOldUsd: string | null; }; function isWhole(value: number): boolean { return Number.isFinite(value) && Number.isInteger(value) && value >= 0; } function mulDiv(tokens: bigint, perMillion: bigint): bigint { return (tokens * perMillion) / MILLION; } function usd(micro: bigint): string { const whole = micro / MILLION; const frac = (micro % MILLION).toString().padStart(6, “0”); return ${whole}.${frac}; } function bill(tokens: { input: bigint; output: bigint; cacheRead: bigint; cacheWrite: bigint; }, rates: Rates, ttl: Ttl): bigint { const writeRate = ttl === “1h” ? rates.cacheWrite1h : rates.cacheWrite5m; return ( mulDiv(tokens.input, rates.input) + mulDiv(tokens.output, rates.output) + mulDiv(tokens.cacheRead, rates.cacheRead) + mulDiv(tokens.cacheWrite, writeRate) ); } function savingsBps(before: bigint, after: bigint): number | null { if (before <= 0n) return null; return Number(((before - after) * 10_000n) / before); } export function quoteJob(job: Job): Quote { const base = { name: job.name, model: “claude-haiku-5-5” as const, effort: null, assumedDefaultEffort: false, tier: null, promptTokens: null, haiku55Usd: null, haiku45Usd: null, savingsBps: null, sonnetCacheReadNewUsd: null, sonnetCacheReadOldUsd: null, }; const counts = [ job.inputTokens, job.outputTokens, job.cacheReadTokens, job.cacheWriteTokens, ]; if (!counts.every(isWhole)) { return { …base, decision: “REFUSE”, reason: “token counts must be non-negative integers”, }; } const ttl = job.cacheWriteTtl ?? “5m”; if (ttl !== “5m” && ttl !== “1h”) { return { …base, decision: “REFUSE”, reason: “cache write ttl must be 5m or 1h” }; } let effort = job.effort; let assumedDefaultEffort = false; if (effort === undefined) { effort = “medium”; assumedDefaultEffort = true; } if (!ALL_EFFORT.has(effort)) { return { …base, decision: “REFUSE”, reason: unknown effort ${job.effort} }; } const knownTask = ALLOW_TASKS.has(job.task) || REVIEW_TASKS.has(job.task); if (!knownTask) { return { …base, decision: “REFUSE”, reason: unknown task ${job.task} }; } let input = BigInt(job.inputTokens); let cacheRead = BigInt(job.cacheReadTokens); let cacheWrite = BigInt(job.cacheWriteTokens); const output = BigInt(job.outputTokens); if (job.recountFromHaiku45) { input = (input * RECOUNT_NUM) / RECOUNT_DEN; cacheRead = (cacheRead * RECOUNT_NUM) / RECOUNT_DEN; cacheWrite = (cacheWrite * RECOUNT_NUM) / RECOUNT_DEN; } const promptTokens = input + cacheRead + cacheWrite; const tier: Tier = promptTokens <= PROMPT_LIMIT ? “short” : “long”; const haiku55 = bill( { input, output, cacheRead, cacheWrite }, HAIKU_55[tier], ttl, ); let haiku45: bigint | null = null; if (ttl === “5m”) { const oldInput = BigInt(job.inputTokens); const oldRead = BigInt(job.cacheReadTokens); const oldWrite = BigInt(job.cacheWriteTokens); haiku45 = bill( { input: oldInput, output, cacheRead: oldRead, cacheWrite: oldWrite }, HAIKU_45_5M, “5m”, ); } const reasons: string[] = []; if (!ALLOW_TASKS.has(job.task)) { reasons.push(${job.task} stays off the Haiku allow-list); } if (!ALLOW_EFFORT.has(effort)) { reasons.push(effort ${effort} needs an eval this toy does not have); } if (tier === “long”) { reasons.push(“prompt is over 100000 tokens, so the long-tier rates apply”); } if (job.recountFromHaiku45) { reasons.push(“prompt tokens were recounted at 130/100 before the Haiku 5.5 bill”); } if (assumedDefaultEffort) { reasons.push(“missing effort was stamped medium, the API default”); } const decision: Decision = reasons.some((reason) => reason.startsWith(“prompt is over”) || reason.includes(“allow-list”) || reason.includes(“needs an eval”) ) ? “REVIEW” : “ALLOW”; if (decision === “ALLOW”) { reasons.unshift(“short prompt, allow-listed task, effort is low or medium”); } if (ttl === “1h”) { reasons.push(“one-hour cache writes are quoted for Haiku 5.5 only”); } const sonnetNew = mulDiv(cacheRead, SONNET_55_CACHE_READ_NEW); const sonnetOld = mulDiv(BigInt(job.cacheReadTokens), SONNET_55_CACHE_READ_OLD); return { …base, decision, reason: reasons.join(”; ”), effort: effort as Effort, assumedDefaultEffort, tier, promptTokens, haiku55Usd: usd(haiku55), haiku45Usd: haiku45 === null ? null : usd(haiku45), savingsBps: haiku45 === null ? null : savingsBps(haiku45, haiku55), sonnetCacheReadNewUsd: usd(sonnetNew), sonnetCacheReadOldUsd: usd(sonnetOld), }; } function line(quote: Quote): string { const parts = [ quote.name, quote.decision, quote.model, quote.effort === null ? “effort=none” : effort=${quote.effort}, quote.tier === null ? “tier=none” : tier=${quote.tier}, quote.promptTokens === null ? “prompt=none” : prompt=${quote.promptTokens}, quote.haiku55Usd === null ? “haiku55=none” : haiku55=${quote.haiku55Usd}, quote.haiku45Usd === null ? “haiku45=none” : haiku45=${quote.haiku45Usd}, quote.savingsBps === null ? “savings_bps=none” : savings_bps=${quote.savingsBps}, quote.reason, ]; return parts.join(” | ”); } const fixtures: Job[] = [ { name: “classify-short”, task: “classify”, effort: “medium”, inputTokens: 2_000, outputTokens: 200, cacheReadTokens: 0, cacheWriteTokens: 0, }, { name: “compact-under-limit”, task: “compact”, effort: “low”, inputTokens: 4_000, outputTokens: 500, cacheReadTokens: 80_000, cacheWriteTokens: 0, }, { name: “compact-over-limit”, task: “compact”, effort: “medium”, inputTokens: 20_000, outputTokens: 400, cacheReadTokens: 90_000, cacheWriteTokens: 0, }, { name: “coding-short”, task: “agentic-coding”, effort: “medium”, inputTokens: 8_000, outputTokens: 1_200, cacheReadTokens: 0, cacheWriteTokens: 0, }, { name: “classify-max-effort”, task: “classify”, effort: “max”, inputTokens: 2_000, outputTokens: 200, cacheReadTokens: 0, cacheWriteTokens: 0, }, { name: “extract-default-effort”, task: “extract”, inputTokens: 1_500, outputTokens: 300, cacheReadTokens: 0, cacheWriteTokens: 0, }, { name: “recount-30”, task: “summarize”, effort: “medium”, inputTokens: 10_000, outputTokens: 500, cacheReadTokens: 0, cacheWriteTokens: 0, recountFromHaiku45: true, }, { name: “one-hour-write”, task: “route”, effort: “low”, inputTokens: 1_000, outputTokens: 100, cacheReadTokens: 0, cacheWriteTokens: 4_000, cacheWriteTtl: “1h”, }, { name: “bad-tokens”, task: “classify”, effort: “low”, inputTokens: -1, outputTokens: 10, cacheReadTokens: 0, cacheWriteTokens: 0, }, { name: “unknown-effort”, task: “classify”, effort: “turbo”, inputTokens: 100, outputTokens: 10, cacheReadTokens: 0, cacheWriteTokens: 0, }, ]; const expected: Record<string, { decision: Decision; haiku55Usd: string | null; savingsBps: number | null }> = { “classify-short”: { decision: “ALLOW”, haiku55Usd: “0.000300”, savingsBps: 9000 }, “compact-under-limit”: { decision: “ALLOW”, haiku55Usd: “0.001450”, savingsBps: 9000 }, “compact-over-limit”: { decision: “REVIEW”, haiku55Usd: “0.015500”, savingsBps: 5000 }, “coding-short”: { decision: “REVIEW”, haiku55Usd: “0.001400”, savingsBps: 9000 }, “classify-max-effort”: { decision: “REVIEW”, haiku55Usd: “0.000300”, savingsBps: 9000 }, “extract-default-effort”: { decision: “ALLOW”, haiku55Usd: “0.000300”, savingsBps: 9000 }, “recount-30”: { decision: “ALLOW”, haiku55Usd: “0.001550”, savingsBps: 8760 }, “one-hour-write”: { decision: “ALLOW”, haiku55Usd: “0.000950”, savingsBps: null }, “bad-tokens”: { decision: “REFUSE”, haiku55Usd: null, savingsBps: null }, “unknown-effort”: { decision: “REFUSE”, haiku55Usd: null, savingsBps: null }, }; function main(): void { let failed = 0; for (const job of fixtures) { const quote = quoteJob(job); const want = expected[job.name]; const ok = want.decision === quote.decision && want.haiku55Usd === quote.haiku55Usd && want.savingsBps === quote.savingsBps; if (!ok) { failed += 1; console.error(FAIL ${job.name}); console.error(quote); } console.log(line(quote)); } const cacheRead = 100_000n; const sonnetNew = usd(mulDiv(cacheRead, SONNET_55_CACHE_READ_NEW)); const sonnetOld = usd(mulDiv(cacheRead, SONNET_55_CACHE_READ_OLD)); console.log( sonnet-cache-read | QUOTE | claude-sonnet-5-5 | cache_read_tokens=100000 | new=${sonnetNew} | old=${sonnetOld} | October 7 cache-read cut, not a routing decision, ); if (sonnetNew !== “0.010000” || sonnetOld !== “0.020000”) { failed += 1; console.error(“FAIL sonnet-cache-read”); } if (failed > 0) { console.error(${failed} check(s) failed); process.exit(1); } console.log(checks=11 passed); } main(); Checked on October 8, 2026, Node.js 25.6.0. Eleven checks passed: classify-short | ALLOW | claude-haiku-5-5 | effort=medium | tier=short | prompt=2000 | haiku55=0.000300 | haiku45=0.003000 | savings_bps=9000 | short prompt, allow-listed task, effort is low or medium compact-under-limit | ALLOW | claude-haiku-5-5 | effort=low | tier=short | prompt=84000 | haiku55=0.001450 | haiku45=0.014500 | savings_bps=9000 | short prompt, allow-listed task, effort is low or medium compact-over-limit | REVIEW | claude-haiku-5-5 | effort=medium | tier=long | prompt=110000 | haiku55=0.015500 | haiku45=0.031000 | savings_bps=5000 | prompt is over 100000 tokens, so the long-tier rates apply coding-short | REVIEW | claude-haiku-5-5 | effort=medium | tier=short | prompt=8000 | haiku55=0.001400 | haiku45=0.014000 | savings_bps=9000 | agentic-coding stays off the Haiku allow-list classify-max-effort | REVIEW | claude-haiku-5-5 | effort=max | tier=short | prompt=2000 | haiku55=0.000300 | haiku45=0.003000 | savings_bps=9000 | effort max needs an eval this toy does not have extract-default-effort | ALLOW | claude-haiku-5-5 | effort=medium | tier=short | prompt=1500 | haiku55=0.000300 | haiku45=0.003000 | savings_bps=9000 | short prompt, allow-listed task, effort is low or medium; missing effort was stamped medium, the API default recount-30 | ALLOW | claude-haiku-5-5 | effort=medium | tier=short | prompt=13000 | haiku55=0.001550 | haiku45=0.012500 | savings_bps=8760 | short prompt, allow-listed task, effort is low or medium; prompt tokens were recounted at 130/100 before the Haiku 5.5 bill one-hour-write | ALLOW | claude-haiku-5-5 | effort=low | tier=short | prompt=5000 | haiku55=0.000950 | haiku45=none | savings_bps=none | short prompt, allow-listed task, effort is low or medium; one-hour cache writes are quoted for Haiku 5.5 only bad-tokens | REFUSE | claude-haiku-5-5 | effort=none | tier=none | prompt=none | haiku55=none | haiku45=none | savings_bps=none | token counts must be non-negative integers unknown-effort | REFUSE | claude-haiku-5-5 | effort=none | tier=none | prompt=none | haiku55=none | haiku45=none | savings_bps=none | unknown effort turbo sonnet-cache-read | QUOTE | claude-sonnet-5-5 | cache_read_tokens=100000 | new=0.010000 | old=0.020000 | October 7 cache-read cut, not a routing decision checks=11 passed savings_bps is basis points against a Haiku 4.5 bill at the original token counts. 9000 is 90.00%. 5000 is 50.00%. 8760 is 87.60%. Read three of those lines closely. classify-short is the brochure case. 2,000 input tokens and 200 output tokens, effort medium. Haiku 5.5 is $0.000300. Haiku 4.5 is $0.003000. Same tokens, 90% less. compact-over-limit is the line you can miss. 20,000 fresh input tokens plus 90,000 cache-read tokens is a 110,000-token prompt. The job is still compact, which is on the allow-list, and effort is still medium. The decision is REVIEW. Haiku 5.5 costs $0.015500. Haiku 4.5 costs $0.031000. The saving is 50%, because the long-tier rates are five times the short-tier rates, and Haiku 4.5’s flat rate is ten times the short-tier rate. recount-30 starts from a Haiku 4.5 count of 10,000 input tokens and 500 output tokens. The gate multiplies the input by 130/100 before it prices Haiku 5.5, so the new bill uses 13,000 input tokens. Haiku 5.5 is $0.001550. Haiku 4.5 is $0.012500. That is 87.60% less, not 90%. The last quote is not a route. On October 7, Anthropic cut Sonnet 5.5 cache reads from $0.20 to $0.10 per million tokens. 100,000 cache-read tokens are $0.010000 now and were $0.020000. The gate prints that. It does not send the job to Sonnet. What this is not There is no API call, no key, and no latency measurement. Effort is not turned into a token multiplier. Anthropic tells you that higher effort thinks more. It does not publish a single multiplier this program could apply honestly. max on a classify job still shows the short-tier list price, then REVIEW. Prompt size here is three counts added together. Anthropic says “prompts.” If your invoice counts a different subset, the tier can be wrong until you change the sum. The 30% recount is the phrase “approximately 30%” from the model overview, applied as the fraction 130/100. It is not a tokenizer run on your text. Five-minute cache writes are the rate in the October 7 table ($0.125 and $0.625 per million). One-hour writes are the model overview’s other rate ($0.20 and $1.00). Haiku 4.5’s one-hour write is not in this comparison, so that fixture leaves haiku45 empty. The Batch API is 50% off input and output. These figures are standard rates. Customer quotes on the launch page, including Asana’s latency note and HubSpot’s 92.8% CRM score, are Anthropic’s report of early tests. They are not reproduced here. Benchmarks above are Anthropic’s: OSWorld 2.1 offline subset 72.4% versus 15.7% for Haiku 4.5 and 48.9% for GPT-6 Luna, GDPval-AA v2.1 Elo 1620 versus 735 and 1437, Terminal-Bench 4.0 39.2% versus 0.0% and 16.4%. Rates used, USD per million tokens: Rate Haiku 5.5, prompt ≤100k Haiku 5.5, prompt >100k Haiku 4.5 Input 0.10 0.50 1.00 Output 0.50 2.50 5.00 Cache read 0.01 0.05 0.10 Cache write, 5 min 0.125 0.625 1.25 Cache write, 1 hour 0.20 1.00 not compared Sources Introducing Claude Haiku 5.5, October 7, 2026 Claude Haiku 5.5 model overview, including the model id, the 100k price split, the one-hour cache write, the medium default, and the tokenizer note Effort, including the five levels and the Haiku 5.5 default I’m building Roster, AI employees that do real work. A smaller model still needs a gate that knows which price it is about to pay.