Toby Allen

Six reviews caught the architecture and missed the data

· Auth0, Vercel, AI-Assisted Development, Identity Security

This is part eight of building the live demo for my apidays Australia 2026 workshop with Claude Code. Part seven closed out a run of small, scattered gaps - a tool the model didn't know it had, a page title nobody had ever looked at. This part is the opposite shape: one feature, six adversarial reviews before a line of code shipped, and a genuine production incident in between. It's also the clearest example this whole project has produced of a distinction worth having a name for - the difference between reviewing an argument and reviewing a system.

The feature, and the incident

I asked Claude to build a live monitoring dashboard for the workshop - signup and login volume in real time, and alerts if Auth0 or the app itself started reporting problems during the talk. It shipped as a Vercel Cron function polling Auth0's tenant log API every 60 seconds into Firestore, with a dashboard polling a snapshot route every 10 seconds. Within hours of going live it exhausted Firestore's free-tier daily read quota and took down /access-log, a completely unrelated citizen-facing page that happened to share the same Firestore project.

Firestore bills per document returned, not per query call. The dashboard's snapshot route was returning up to 220 full documents every 10 seconds, per open browser tab. A single tab left open all day was enough to burn the entire daily budget in under an hour. The pre-build adversarial review had checked Auth0's own Management API rate limit thoroughly and confirmed it was nowhere close to a problem. Nobody had checked Firestore's separate, much lower quota, because the review that would have caught it was busy verifying a different system - one that turned out not to be the one that broke.

We disabled the feature the same day and started again properly.

Round one: the caching claim didn't survive the docs

The redesign dropped Firestore's event store entirely and called Auth0 directly from the dashboard's own route, cached with Next.js's export const revalidate = 10 - the idea being that ten open tabs polling every ten seconds would share one underlying fetch instead of ten. I asked for an adversarial review before writing any more code, and it found the whole mechanism was built on a claim that wasn't actually true for this route.

The route's very first line calls auth0.getSession() from the Auth0 Next.js SDK, which internally reads the request's cookies - a documented Request-time API in Next.js's caching model. Anything downstream of a Request-time API runs on every request, no matter what revalidate says. revalidate was never actually going to fire for this route, and nobody had checked that against the framework's own documentation before writing the design around it.

Round two: the fix repeated the exact same mistake

The next attempt replaced the framework-level cache with a small in-memory Map, keyed to a time bucket, on the theory that concurrent requests inside the same window would share one cached result. A second review checked this against Vercel's own Fluid Compute documentation and found the same shape of problem, one layer down. Fluid Compute reuses warm instances to cut cold starts - it doesn't promise that concurrent requests land on the same instance. Under real load, Vercel's documented behaviour is to spin up additional instances, each with its own empty cache. Redone honestly, a single tab left open all day was still 73% over the daily quota. The fix for an unverified caching assumption had been a second, different unverified caching assumption.

Round three: the math only ever counted one person

At that point we dropped caching altogether in favour of a manual "click Refresh" button - a human can only ever fire one request at a time, which removes the concurrency question rather than trying to bound it. True, and also not the whole answer. A review of that design found the entire worst-case table only ever assumed exactly one person had the dashboard open, despite the plan's own narrative naming a presenter and a co-organiser as the realistic audience, and despite the live Auth0 tenant already having two real accounts with the role required to see the page. At two people refreshing every 15 seconds or so, the feature alone used 77% of the daily budget. At three, it went over.

Round four: the fix to that broke a different alert

Round four documented an explicit ceiling - four concurrent operators, sized by shrinking a Firestore query's limit from ten documents to three. That held the quota, and quietly disabled one of the five alert rules in the process. The FGA-denial alert fired on a threshold of fifteen or more matching records in the window. A query capped at three documents can never return more than three documents. The threshold could never be reached, by construction, and the plan's own justification cited a completely different alert rule - one that reads a different collection and was never affected by the change - as evidence the fix was safe.

Round five: reviewing a rewrite that had never been applied

The fifth review found something categorically different from a quota or a threshold bug. It went looking for the query shape the plan described, in the actual file, and it wasn't there. The real alerts.ts had no query limit of any size, anywhere - and the alert-evaluation function containing that query was only ever called from a disabled cron route, never from the dashboard's own snapshot route the way every plan since round two had assumed. The real snapshot route was still sitting behind the original MONITORING_DISABLED flag, reading the pre-computed Firestore collections from the very first version. Every plan from round two through round five - and every review of those plans - had been reasoning with complete rigor about a rewrite that had never touched a single real file. The reviews had correctly stress-tested the arguments they were handed. Nobody, across four rounds, had checked whether the code those arguments described actually existed.

Adversarial review, done properly, checks whether a plan's own logic holds together. It does not automatically check whether the plan is describing reality, and those are different jobs that happen to look identical from the outside until one of them is quietly missing.

Round six: read the file first, then write the diff

The recovery was almost embarrassingly mechanical once we saw the actual shape of the problem: stop describing a target state in prose and write the plan as an explicit diff against the real, freshly re-read file contents - current code quoted directly, new code shown against it, no paraphrasing in between. We also cut the two Firestore-backed alert rules entirely, since every correctness bug across four rounds had traced back to them specifically, and they were always the lower-priority signal. What survived was three rules sourced directly from Auth0's own logs - an elevated failed-login rate, an Attack Protection block, an Action execution failure - plus a manual refresh button and zero Firestore reads for the feature at all. Not a smaller version of the quota risk. The complete removal of the resource that caused it.

The sixth review's job, explicitly, was to check whether the plan had actually grounded itself this time - more important than any individual code-quality finding. It re-read all nine referenced files independently and found zero discrepancies against the plan's own quotes. It checked the proposed new TypeScript against the real, not described, types and function signatures. It found exactly one issue, and a minor one: a stale comment in an unrelated file that could have misled a future reader into deleting code a completely different feature still depended on. First clean pass in six rounds.

Four rounds of adversarial review, each one correctly stress-testing the plan's own reasoning, none of them checking whether the code being reasoned about actually existed in the repository - the fifth round finally read the files first

The data was still wrong

Deployed, working, quota-safe - and then a plain question from whoever was actually going to use the thing: are all the different Auth0 log event types translated for display? They weren't; only nine of Auth0's hundred-plus event codes were mapped at all, and two of the nine were mapped to the wrong meaning entirely. sepft had been labelled a critical Attack Protection alert. Auth0's public docs page for event codes wouldn't render its content for a fetch, so I pulled the source schema directly from Auth0's own docs-content repository on GitHub instead of trusting a secondary source. The schema says sepft means "Successful Exchange of Password for Access Token" - a routine success event with nothing to do with attack protection.

The second one mattered more. fce fed the alert rule specifically meant to catch a silently broken Auth0 Action - the exact failure mode this project had already hit once for real, in an earlier part of this series, when a setCustomClaim call turned out not to persist the way anyone assumed. The same schema says fce means "Failed to change user email," an unrelated event, and separately confirms the real code for an Action execution failure is actions_execution_failed. The alert this feature existed partly to provide had never once been capable of firing on a real Action failure, because it had been filtering for a code that Action failures don't emit. Six rounds of review, and not one of them had any reason to open the vendor's own API reference and check whether a short string constant meant what its variable name and inline comment both claimed.

Final Thoughts

Six rounds of good adversarial review, each one catching something real about caching, concurrency, or a query shape - and none of them exercised whether a single Auth0 event code was correct, because "is this architecture sound" and "is this constant's value true" are different questions that don't get tested by the same kind of scrutiny. The fifth round's finding is the one to keep: checking a decision's stated risk thoroughly is not the same thing as checking the risk nobody thought to name, and a review that's rigorous about the wrong layer looks exactly like due diligence right up until the wrong thing runs out first. The fix, in the end, wasn't a smarter reviewer. It was reading the file before describing it, and opening the schema before trusting the comment above the constant.