The Token Dad

Labwebsite-challenge-v1closed 2026-08-19

Did four tools converge on the same typeface?

Yes. Six independent readings of the same brief, one display face. The tools differed in process, not in taste, and the cost spread between them was 6x.

Four AI design tools (frontend-design, Impeccable, Taste, and gstack) each built the same website from one frozen brief, with typography and colour deliberately left open because those were the two axes expected to diverge most. They did not diverge. All four arms chose Archivo for display type. Arms 1 and 2 landed on the identical pair, Archivo paired with Faustina, with no contact between them. Inside arm 4, three further independent voices (the runner, an OpenAI model, and a separate Claude subagent) each proposed Archivo without seeing one another's work. Six independent readings of the same brief, one typeface.

Why was typography left open in the first place?

The brief instructed each tool: "I don't want to restrict, I want to be surprised," and named "an obvious serif font selected by AI" as the specific embarrassment to avoid. Type and colour were withheld from the answer sheet on purpose, precisely because a well-run comparison needs at least one place where the tools are free to disagree. This was the test of whether four separate design processes, given the same latitude, would actually use it differently.

What fonts did each arm choose?

Font choices by arm
ArmDisplayBodyThird
1 · frontend-designArchivoFaustinanone
2 · ImpeccableArchivoFaustinanone
3 · TasteArchivoLiterataGeist Mono
4 · gstackArchivoPiazzollaMartian Mono

How was this checked, not assumed?

The convergence was verified rather than taken on faith. The baseline lockfile going into the run contained no Archivo. Arm 1's installed font packages were deleted from disk before the parallel arms started, so no arm could inherit a leftover choice from another. Each arm ran in its own isolated worktree. No runner was given the playbook, the run log, or any other arm's output. The four Archivo choices were made independently, under conditions designed to rule out contamination.

Does this prove the tools agree, or just that they share training data?

This is unresolved, and it is the most interesting thing the run produced. The generous reading is that four tools independently found the correct answer to a well-specified brief: that Archivo really is what "confident, not obvious, not a serif" cashes out to, and four separate processes converged on a real answer. The skeptical reading is that four tools trained on overlapping data reached for the same non-obvious-but-safe font, and "surprise me" was answered with consensus rather than range. The run itself cannot distinguish these two readings. Both are worth publishing, and neither is proven false by the other.

What else converged beyond the typeface?

Two further agreements showed up without being asked for. All four arms independently built a wordless glyph system, which the shared vision document had named as its recurring idea. And all four arms rationed saturated colour against a dark ground, holding red in deep reserve rather than spending it freely. Typography was the axis this page tracks in detail, but it was not the only place the four arms agreed.

If the outputs converged, where did the four tools actually differ?

Not in taste. In process. The visual outputs were more alike than expected; the working methods that produced them were not alike at all. Each tool operated on a different theory of what "finished" means.

frontend-design
Asks nothing up front, self-fills every gap in the brief, runs one visual self-review, and commits. The cheapest arm by roughly six times.
Taste
Runs a templated six-slot intake, one self-review, and applies the strictest copy rules of the four. It also costs the most to port elsewhere, since it hands over React that has to be translated for static delivery.
gstack
Runs a structured review that produced thirteen findings, brings in an external model as a second voice, and commits once per finding. This process caught defects no screenshot could have reached.
Impeccable
Runs an internal adversarial system: a builder, a reviewer sitting behind an explicit fix-or-ship gate, and a documenter that re-derives its output from source rather than summarizing claims. It ran two full rounds before the gate would let it pass itself, at roughly six times arm 1's token cost.

That divergence in process, not the colour palettes, is the finding this page is built to publish.

What was the single most valuable defect found during review?

gstack's structured review surfaced a CSS specificity defect on .stamp--alert: the surrounding context rules carried a specificity of 0,2,0 against the alert modifier's 0,1,0, and the three files that were supposed to consume the modifier never actually applied it. The one place in the whole design that used a second hue would never have rendered red anywhere a reader could see it. No build gate would have flagged this, no screenshot would have shown it, and no purely visual review could have caught it, because the page looks correct: the wrong colour looks like a deliberate choice until someone reads the cascade.

Impeccable's two-round fix-or-ship gate produced the same class of result on its own build: it caught a hard-banned kicker style it had shipped itself, an h3 heading set quieter than its own body copy, and a full quarter of its own mark system left uncoded. The principle both findings point at is the same one: a verification step that is capable of contradicting the thing it verifies is worth more than a review that can only confirm it.

What did each tool cost in the design phase?

Design-phase metrics by arm
Metric1 frontend-design2 Impeccable3 Taste4 gstack
Intake questions0211823
Answered from the sheetnone15912
Answered "your call"none668
Tool calls71258116151
Wall clock~26 min~54 min~48 min~46 min
Tokens~169k1,064,198284,287409,632
Review rounds1 self-review2 (fix/ship gate)1 self-review13-finding review
External model usednononoyes (2 codex exec)
Modelled cost$1.12$7.02$1.88$2.70

What did deployment cost, and why did one arm take 38 minutes?

Only three arms went through a separate deployment phase.

Deployment-phase metrics by arm
MetricImpeccableTastegstack
Tool calls154075
Wall clock2.5 min7.1 min38.6 min
Tokens48,97168,648100,079
Archive size432K748K848K
Modelled cost$0.32$0.45$0.66

The gstack deployment's 38.6 minutes of wall clock is not 75 slow tool calls; the deploy script itself ran in 9 seconds. The gap was a Cloudflare control-plane success that the data plane did not follow: the DNS record never appeared until the binding was deleted and re-attached after a backoff. The token and tool-call counts reflect the work; the wall clock reflects waiting on infrastructure that said yes before it meant it.

What was the total cost across all four arms?

Tool calls, all phases
726
Total tokens
2,144,815
Modelled total cost
~$14.16

How was cost modelled, and how reliable is the number?

Token counts in every table above are measured, taken directly from the agent sessions. Cost is not measured; it is modelled, and the model's assumption is the whole ballgame. Anthropic reports each of these sessions as a single total-token figure with no input/output split, so the split has to be assumed, and it dominates the result. Opus 5 bills input at $5 and output at $25 per million tokens: read every token as input and the total is $10.72; read every token as output and the total is $53.62. This page uses 92% input / 8% output, which is typical for agent sessions that resend a growing history on every turn, and that assumption is what produces the $14.16 figure above. It is an estimate carrying a stated assumption, not an invoice.

Two factors push the real number down from there. Prompt caching bills cache reads at roughly a tenth of the standard input rate, and these sessions resend large stable prefixes on nearly every turn, which caching is built for. And the run itself was executed under a Claude Code subscription rather than per-token API billing, so no per-token invoice exists for this run at all: the $14.16 is a hypothetical replication cost, not a bill anyone paid. An honest range for an API-billed replication is $11 to $54, most likely landing near the low end once caching is accounted for.

The number worth remembering is not the total. Arm 1 produced a complete, shipped, reviewed design for a modelled $1.12. Whether the roughly six-times spread up to arm 2 buys a better website is not something this page can answer. That is what the exhibits are for.

For how these numbers were captured, see the methodology. To see what each arm actually shipped, visit the exhibits.