The tools passed three times. The page scored 40, 80, and 85

Assessment course, unit 00
What a green checkmark leaves out

Webflow built an eval harness that runs Claude Code and Codex CLI against disposable sites and scores what they build. On one landing-page task the deterministic tool checks passed every run while design quality came in at 40, 80, then 85 percent. Plus Vercel gives eve agents a memory and puts Sandbox in every region.

  1. Run 1 of 3 40% Tool checks 1.00
  2. Run 2 of 3 80% Tool checks 1.00
  3. Run 3 of 3 85% Tool checks 1.00
Design-quality score from the harness's visual judge, one story run three times on a fresh site each time. The deterministic checks passed 1.00 on every run. Figures from Webflow's post.
A two-colour risograph diagram of a web page drawn as blocks in cobalt keyline and poster-red fills, with a round grading loupe over one block and a small rubric grid of filled and empty squares in the corner
Exam plate 1. The page under the loupe: the harness screenshots what the agent built and grades the render, then reads the markup underneath it.
Unit 01

Tooling

Unit 1 of 6

Webflow's MCP eval harness grades the page the agent built, not the calls it made

Webflow Engineering, 10 September 2026

Gil Levin describes a harness that runs real agents, headlessly, against disposable Webflow sites. Each story is a plain-English task with assertions, such as "build a full SaaS landing page and publish it." A fresh site is provisioned, Claude Code or Codex CLI runs the task, deterministic checks confirm the expected tool use and that every claimed write actually landed, an LLM judge reads the transcript, and page-building stories add a visual judge over the rendered result.

The number that matters for designers is the spread. In one three-run diagnostic the tool checks passed 1.00 every time while the design-quality score landed at 40, 80, and 85 percent. The console had summarised that as a perfect mean with zero deviation because it only aggregated the deterministic score. Webflow now reports the two dimensions separately, and every run leaves a punch list of findings even when the story passes. It continues the direction from Webflow Source, covered in issue 116, where agents and designers share one canvas.

Tool execution was stable while output quality was not.

Gil Levin, Webflow Engineering

Vercel Sandbox runs in all 20 compute regions, with failover you can order

Vercel changelog, 10 September 2026

Vercel Sandbox, the disposable environment agents use to build and run code, is now available in all 20 Vercel compute regions, up from four. Region selection is on every plan, Pro and Enterprise teams can list failover regions in order, and a region passed at creation overrides the project default. For a studio whose clients sit under data-residency rules, an agent's build box can now stay inside approved geographies.

Unit 02

Technique

Unit 2 of 6

Score the structure separately from the screenshot

Webflow Engineering, 10 September 2026

Two pages can render pixel-identical while one is built from nav, section, and button elements and the other is nested wrappers all the way down. A visual judge cannot see the difference, so Webflow added a semantic HTML ratio: a deterministic signal computed from the rendered markup it already captured for the screenshot, the fraction of semantic tags against generic wrappers, combined across pages by markup volume. It is a trend line rather than a gate, since a visual builder is naturally wrapper-heavy, but it gives a second reading of the same page.

The findings list is the other technique to borrow. One real entry: Webflow renames a CSS class on collision, so the agent's intended hero class ships with a numeric suffix and no warning from the MCP server. Nothing failed, the story passed, and the class name a designer would search for no longer exists. A pass/fail number would have erased that.

Unit 03

Workflow

Unit 3 of 6

eve agents keep a memory across sessions, scoped per signed-in user

Vercel changelog, 9 September 2026

Vercel's eve agents can now retain context across sessions through named memory slots. Each slot declares a provider that stores and retrieves the memory and a scope that decides who shares it, so one command adds a file-backed slot scoped to each authenticated caller. Before every turn eve pulls the relevant memory into context; on Vercel the built-in provider writes to a private Blob store that survives restarts and deploys. For a site agent that means the brand rules a client explained on Monday are still in play on Friday.

Test against the agent host that is actually failing

Webflow Engineering, 10 September 2026

OpenAI rejected one of Webflow's ChatGPT app submissions after an agent reached for a Designer-canvas tool that needs a live Designer tab. The harness never reproduced it, because the harness had only ever been driven by Claude, which never took the bait. Webflow added a second driver that shells out to Codex CLI with the same provisioning, scoring, and screenshots, and a five-repeat comparison of the original prompt against a hardened one showed materially different behaviour per host. If your MCP server or design system is meant to work in more than one agent, the eval coverage has to include each of them.

Unit 04

Design move

Unit 4 of 6

Borrow this pattern: the criterion bar

Give each section a header that carries its own scale. The bar is a three-column grid: a mono unit label, the section name, and a row of small blocks that shows position or score. On this page the blocks show where you are in six units. On a course page they can show module progress, on a comparison page a rating per criterion, on an audit report the score a section earned. The reader gets orientation and a number from the same ruled line, and the page loses its decorative dividers.

To keep it honest, the blocks must mean something you can state in the caption, the meter is aria-hidden with the value written out for screen readers, and the bar never becomes a fake control. If a section has no scale, use the same rule without the blocks.

Criterion

Visual judge

4 of 5
Criterion

Semantic ratio

2 of 5
Criterion

Tool checks

5 of 5
A risograph diagram of two web page thumbnails with the same silhouette, the left built from a few clean nested rectangles in cobalt, the right from dozens of tangled nested rectangles in red
Exam plate 2. Same render, different markup: the pair the screenshot judge cannot tell apart and the semantic ratio can.
Unit 05

Prompt Lab

Unit 5 of 6

Answer sheet

Paste this into an AI page builder to reproduce The Rubric as a course or assessment landing page. Swap the course, the scores, and the plates for your own.

Build a course landing page for a design-school assessment course, called The Rubric. One page, semantic HTML and CSS only, no JavaScript.

Ground: the whole page is poster red (#A61D2B) with a drafting grid drawn in CSS, 32px squares, 1px lines of cream at 8 percent opacity. No gradients, no glow, no shadows except one soft shadow under the framed plate.

Masthead: a breadcrumb rail, a ruled single row across the top with small uppercase monospace crumbs separated by a thin chevron: Issues, Course library, then the course number in a filled cream chip, then the date and reading time pushed to the right edge. No boxed metadata cells.

Hero: the headline runs full width in an extended grotesque (Syne 800), 6vw, line height 0.98, cream. Below it a 12-column grid: a two-line mono unit label in the left five columns, the deck in EB Garamond 23px in the right six columns. Then a grade strip: an ordered list of three tiles inside one cream keyline, each with a small run label, a very large tabular number (Syne 800, 80px) with a small percent sign, a two-line mono note on the right, and a thin cream bar whose fill width equals the score. Then a framed plate: a cream frame with 14px padding holding a 3:2 image, with a mono caption in pale rose below.

Section markers: every section opens with a criterion bar, a three-column grid ruled with a 1px cream line above and a 34 percent cream line below: mono unit number on the left, the section name in Syne 800 at 30px in the middle, and on the right a meter of six 22px blocks with 1px cream borders, filled cream for the units already passed. Sections: Tooling, Technique, Workflow, Design move, Prompt Lab, Field note, then Sources.

Items: a Syne 700 subhead at 26px, a mono source line in pale rose, then one or two paragraphs of EB Garamond at 20px with line height 1.7 and a 66 character measure, cream on red, never justified. One pullquote only, EB Garamond italic at 34px between thin rules.

Course card: a cream (#F3E7CF) card in two columns with dark ink text, a cobalt (#1F3FBF) mono kicker, the course name in Syne 800 at 44px, an italic one-line summary in cobalt, one paragraph, a row of keyline chips with 3px corners, and on the right a definition list of course details with a cobalt rule above and thin cobalt rules between rows. Cobalt appears only on cream surfaces, never on the red ground.

Design move: two columns, text left and a square framed plate right, with a small demo of three criterion bars used as a five-block rating scale.

Prompt block: a cream answer sheet with a cobalt kicker, one line of instructions, and the prompt in a monospace pre at 15.5px that wraps.

Images: risograph prints, two-colour overprint of poster red and cobalt on cream uncoated paper, diagrams of web page blocks with visible misregistration and ink grain, no legible text.

Guardrails: body text at least 18px with line height 1.6 or more; small labels at least 13.5px and WCAG AA against their real surface; no accent left border stripes; no pill radii except none; hover and focus states on every link with a 150ms transition; a prefers-reduced-motion guard; no em dashes; no fake search fields, toggles, or filter pills; the meter is aria-hidden with its value written for screen readers.

Runs in v0, Lovable, Bolt, Framer AI, Claude Code, and Beaver Builder AI. Swap the fonts for any wide grotesque paired with a Garamond-style serif.

Unit 06

Field note

Unit 6 of 6

The scores designers should care about were the ones the summary hid. When a vendor grades its own agent output on rendering and structure, and publishes the spread, the page becomes the test.

Reading list

Sources

  1. Your tools work. Will the agent use them right? Inside Webflow's MCP eval harness Webflow Engineering, Gil Levin, 10 Sep 2026
  2. Webflow MCP server Webflow Developer Documentation
  3. Vercel Sandbox is now available in all regions Vercel changelog, 10 Sep 2026
  4. Persistent memory for eve agents Vercel changelog, 9 Sep 2026
  5. Issue 116: Webflow stops requiring Webflow Art Direction Daily, 2 Sep 2026