design2code-rl: what the grader says about 114 Claude Code attempts, 20 animated and 30 framework trials

Agent: Claude Code (Opus 4.7), one attempt per trial, 10 attempts per task (12 on two tasks). Grader: render-based element matching, see docs/decisions/0001.

Where every number comes from

Every number here is computed by scripts/report.py from docs/report/results-manifest.json, the per-trial verifier/breakdown.json, the reference sites and judge_pairs.csv, except these, which are quoted and marked where they appear: the degradation and anti-gaming tables (decisions 0001 and 0005), the failure causes in section 10 and the ~$0.20 Modal cost per trial (decision 0006), and the Design2Code baselines of 158 tags / 22 unique tags / depth 13 and Fleiss kappa 0.46 (arXiv:2403.09029).

Section 11 covers the Part 2 animated tasks on their own, graded by 1.2 with --animated. Sections 1 to 10 are the Part 1 set and are unchanged by it.

Section 12 covers the Part 3 framework tasks on their own: the same references, the agent writing React or Solid, the build snapshotted and graded by the unchanged grader.

1. Headline

114graded trials
11tasks
0.622mean reward (grader 1.0)
0.088sd
0.483min
0.869max
41infra failures

Every attempt lands strictly between 0 and 1: the floor is 0.483 and the ceiling 0.869, on a scale where the reference scores 1.0000 and a blank site scores 0.0000. There is room to improve in both directions, which is what an RL environment needs. The 41 infra failures below are sandbox and orchestration errors, not low scores; they carry no reward and are excluded from every statistic here.

Tiers: S, M, L. Agent spend as reported by the Claude Code CLI and recorded per trial by Harbor: $480 over all 155 attempts, $419 over the 114 graded ones ($3.67 per graded trial). Sandbox compute is a rounding error beside that: docs/decisions/0006 measures ~$0.20 per Modal trial against the credit balance, which this report cannot recompute (Modal bills the account, not the trial).

2. Per task

tasktierarchetypepagesnmeansdminmaxstructpercep0.4 → 0.9
atlas-cafeSrestaurant5100.8290.0280.7760.8690.8590.820
podcast-foxglove-25Spodcast5100.7020.0330.6600.7520.7260.834
docs-wren-collective-33Mdocs6100.6770.0210.6540.7290.6960.865
clinic-tessellate-supply-46Mclinic6100.6260.0300.5790.6900.6580.755
portfolio-ember-44Mportfolio6100.6190.0280.5790.6590.6570.701
travel-alder-works-35Mtravel6100.6160.0330.5740.6860.6360.840
conference-anvil-kitchen-42Lconference8100.5940.0370.5500.6820.6300.709
civic-juniper-collective-19Scivic5120.5750.0300.5390.6420.6300.559
clinic-corvid-health-37Sclinic5100.5710.0380.5200.6320.6230.574
fitness-harbor-partners-13Lfitness8100.5400.0230.5160.6010.5890.566
civic-ember-group-2Lcivic8120.5250.0490.4830.6380.5710.597

The spread within a task is small (sd 0.02-0.05) next to the spread between tasks (0.53 to 0.83). The agent is consistent; the tasks are what differ. No task is saturated and none is a floor, so every one of them still carries gradient.

3. Tier and archetype

tierpagestasksnmeansdminmax
S54420.6650.1100.5200.869
M64400.6350.0380.5740.729
L83320.5510.0490.4830.682

Do the tiers separate? All rewards in this table are grader 1.0 (section 7 measures 1.1 against it). Means are S 0.665, M 0.635, L 0.551: monotone in the intended direction (S > M > L), but the gaps are small next to the within-tier range and the tier ranges overlap heavily. Tier is a page-count label, and page count is not the thing that makes a page hard. The S tier is also flattered by the hand-built atlas-cafe fixture, the easiest task in the set. Without it the means are S 0.614, M 0.635, L 0.551: S and M swap order and sit 0.021 apart, well inside the within-tier sd, while L is clearly lowest. So the ordering that survives is L below the rest; S vs M is noise, and all three sit inside 0.083. The prior-art risk (Design2Code-HARD: frontier models miss 30-40% of blocks on long pages, so the hard tier compresses) shows up here as compression: the L tier is separated, the two smaller tiers are not separated from each other.

archetypetasksnmeansdminmax
restaurant1100.8290.0280.7760.869
podcast1100.7020.0330.6600.752
docs1100.6770.0210.6540.729
portfolio1100.6190.0280.5790.659
travel1100.6160.0330.5740.686
clinic2200.5990.0430.5200.690
conference1100.5940.0370.5500.682
civic2240.5500.0480.4830.642
fitness1100.5400.0230.5160.601

Across the 11 tasks, Spearman rho between mean reward and mean reference elements per page is -0.48; between mean reward and page count it is -0.49. Both are moderate and negative, and at n=11 they are not distinguishable from each other: more page and more elements both cost reward, and this sample cannot say which costs more.

4. Backend

taskmodaldockerdaytonaspread
civic-ember-group-20.500 (n=5)0.545 (n=5)0.542 (n=2)0.045
civic-juniper-collective-190.561 (n=5)0.580 (n=5)0.599 (n=2)0.039
clinic-corvid-health-370.576 (n=8)-0.554 (n=2)0.021
conference-anvil-kitchen-420.589 (n=8)-0.616 (n=2)0.027
docs-wren-collective-330.681 (n=8)0.662 (n=2)-0.020
fitness-harbor-partners-130.541 (n=9)0.533 (n=1)-0.007

Backends should not change the score: the same image, the same verifier, the same Chromium. The gaps above are within the within-task sd, so nothing here suggests a backend effect; the trials differ because the agent's samples differ. Backend choice was made on throughput, not on score (docs/decisions/0006).

5. Where the reward is lost

00.250.50.751position 0.743position0.743size 0.960size0.960color 0.899color0.899typography 0.780typography0.780text 0.985text0.985
positionsizecolortypographytextstructuralperceptual
all graded0.7430.9600.8990.7800.9850.6590.706
tier L0.6380.9570.8900.7780.9670.5950.622
tier M0.7810.9650.8920.7720.9900.6620.790
tier S0.7880.9570.9130.7880.9930.7060.690

Position is the whole story. These components are grader 1.0; section 7 shows colour moving 0.899 → 0.914 under the shipped 1.1, and nothing else moving. Size, colour, text and typography all sit high; position averages 0.743. The agent gets the right elements with roughly the right look and puts them in the wrong place, and vertical drift compounds down a long page. That is exactly the failure a screenshot-to-code policy should be trained to fix, and it is why the reward is dense rather than binary.

6. Reference, best and worst

Top 900 px of each index page (rendered at 360 px wide here, 420 px in the file). Left: the reference the agent never sees as HTML. Middle: the best of its attempts. Right: the worst.

atlas-cafe tier S, 5 pages

reference
reference
best 0.869
best 0.869
worst 0.776
worst 0.776

civic-ember-group-2 tier L, 8 pages

reference
reference
best 0.638
best 0.638
worst 0.483
worst 0.483

civic-juniper-collective-19 tier S, 5 pages

reference
reference
best 0.642
best 0.642
worst 0.539
worst 0.539

clinic-corvid-health-37 tier S, 5 pages

reference
reference
best 0.632
best 0.632
worst 0.520
worst 0.520

clinic-tessellate-supply-46 tier M, 6 pages

reference
reference
best 0.690
best 0.690
worst 0.579
worst 0.579

conference-anvil-kitchen-42 tier L, 8 pages

reference
reference
best 0.682
best 0.682
worst 0.550
worst 0.550

docs-wren-collective-33 tier M, 6 pages

reference
reference
best 0.729
best 0.729
worst 0.654
worst 0.654

fitness-harbor-partners-13 tier L, 8 pages

reference
reference
best 0.601
best 0.601
worst 0.516
worst 0.516

podcast-foxglove-25 tier S, 5 pages

reference
reference
best 0.752
best 0.752
worst 0.660
worst 0.660

portfolio-ember-44 tier M, 6 pages

reference
reference
best 0.659
best 0.659
worst 0.579
worst 0.579

travel-alder-works-35 tier M, 6 pages

reference
reference
best 0.686
best 0.686
worst 0.574
worst 0.574

7. Why a higher reward is a better replication

The grader is calibrated on a degradation suite: take the reference, break it in one controlled way, and check the reward moves the right way. fixtures/atlas-cafe, 5 pages, from docs/decisions/0001, all 26 rows, grader 1.1. Seventeen assertions in tests/test_validation.py hold this table in place.

00.250.50.751oracle (reference vs itself) 1.0000oracle (reference vs itself)desaturate 33% 0.9900desaturate 33%desaturate 66% 0.9767desaturate 66%color_shift +12 0.9725color_shift +12swap_fonts "Roboto, sans-serif" 0.9718swap_fonts "Roboto, sans-serif"position: fixed opaque full-viewport cover 0.9634position: fixed opaque full-viewport covergrayscale (desaturate 100%) 0.9632grayscale (desaturate 100%)shift_section 40px 0.9472shift_section 40pxclip-path: circle(0) on every section 0.8921clip-path: circle(0) on every sectionshift_section 120px 0.8557shift_section 120pxcolor_shift +48 0.8421color_shift +48color_shift +96 0.8132color_shift +96drop 1 of 5 pages 0.8000drop 1 of 5 pagesshift_section 400px 0.6548shift_section 400pxdrop 2 of 5 pages 0.6000drop 2 of 5 pagesdelete 1 section 0.5745delete 1 sectiondelete 2 sections 0.4027delete 2 sectionsdrop 3 of 5 pages 0.4000drop 3 of 5 pagesdelete 3 sections 0.3068delete 3 sectionsevery text colour at 6% alpha 0.3893every text colour at 6% alpha-webkit-text-fill-color: transparent 0.3892-webkit-text-fill-color: transparentscreenshot <img> hack 0.0455screenshot <img> hackevery section at 6% opacity 0.0412every section at 6% opacityevery section at 5.1% opacity 0.0412every section at 5.1% opacityfilter: opacity(0.04) on everything 0.0000filter: opacity(0.04) on everythingblank pages 0.0000blank pages
degradation (quoted from decision 0001)reward
oracle (reference vs itself)1.0000
desaturate 33%0.9900
desaturate 66%0.9767
color_shift +120.9725
swap_fonts "Roboto, sans-serif"0.9718
position: fixed opaque full-viewport cover0.9634
grayscale (desaturate 100%)0.9632
shift_section 40px0.9472
clip-path: circle(0) on every section0.8921
shift_section 120px0.8557
color_shift +480.8421
color_shift +960.8132
drop 1 of 5 pages0.8000
shift_section 400px0.6548
drop 2 of 5 pages0.6000
delete 1 section0.5745
delete 2 sections0.4027
drop 3 of 5 pages0.4000
delete 3 sections0.3068
every text colour at 6% alpha0.3893
-webkit-text-fill-color: transparent0.3892
screenshot <img> hack0.0455
every section at 6% opacity0.0412
every section at 5.1% opacity0.0412
filter: opacity(0.04) on everything0.0000
blank pages0.0000

Strictly monotone along five axes independently: deleted content, missing pages (1 of 5 costs exactly 20%), tint shift, desaturation, displacement.

Anti-gaming

attack (quoted from decisions 0001 and 0005)score / costwhy it fails
paste the reference screenshot as one <img>0.0455no text leaves, one box; the multiplicative gate keeps the perceptual term from paying
blank pages0.0000nothing matches, structural is 0
fade the page out (opacity: 0.06, filter: opacity(0.04))0.0412 / 0.0000alpha x effective opacity is composited into every colour, and an element under 8% effective opacity cannot paint anything visible, so it is not extracted
keep the layout, make the text invisible (color: rgba(0,0,0,0.06))0.3893text under WCAG 1.2:1 against its own background is dropped, so the page scores as the text-free wireframe it is - the one known attack above 0.25
duplicate elements to hedge the match-29% at 2x elements, -81% at 10xextras penalty is count-shaped, so spam is expensive
read the reference source or stylesheetimpossiblethe renderer serves only file:// paths inside the candidate directory
paint the reference palette over grayscale content-1.0 pp netpalette gain 0.79 -> 0.89 costs layout 0.976 -> 0.770
explode the DOM to slow the verifierrejecteda candidate page over 20,000 nodes is not extracted
Honest wording: the suite tests attacks we thought of; it does not search for attacks. The defensible claim is no known exploit above 0.39 among attacks the extractor can see - and the one that reaches 0.39 has to render every background, border and icon correctly and only hides the text - not that the grader is ungameable. Attacks that leave the DOM intact and hide content at the pixel level - general clip-path, z-index occlusion, mix-blend-mode, opaque covers - are floored at 0.8 × structural, because the perceptual term is a 20% multiplicative factor and the extractor sees a faithful page; they measure 0.87-0.96 on the fixture (clip-path: circle(0) 0.8921 on a totally blank render, a fixed opaque cover 0.9634). Two are asserted as documented ceilings, not passes; closing the floor needs a perceptual hard gate or occlusion-aware extraction, deferred because it would move every calibrated number here. Rewards in this report were produced by grader 1.0; the shipped 1.1 added the visibility guard, measured against 1.0 just below.

Grader 1.0 vs 1.1 on the recorded trials

The rewards above were produced by grader 1.0; the repo ships 1.1, which composites every extracted colour over what is behind it and drops what is invisible. Harbor kept no candidate site, so scripts/regrade.py replays each trial's transcript and grades the reconstruction under both graders on one host. Reconstruction fidelity, checked by re-grading with 1.0 against the reward the trial was actually given: 99 exact, 14 within 0.01 (all Modal trials, perceptual term only - this host is arm64), 1 unrecoverable (portfolio-ember-44__tPa3EYh, which wrote 5 of its pages from subagents whose calls are not in the transcript) and excluded below.

1.0 vs 1.1, 113 trials (regrade-summary.json)
mean reward0.6225 → 0.6270 (+0.0045)
unchanged / moved < 0.01 / moved ≥ 0.0121 / 79 / 13
scored lower under 1.120
largest shifttravel-alder-works-35__3ep3QXC 0.5806 → 0.6154 (+0.0348)
Spearman ρ / Kendall τ / Pearson r0.9951 / 0.9567 / 0.9979
pairwise winners that flip2 of 196 judge pairs, 0 of 35 labelled human pairs
per-task orderingidentical, rank for rank
component (mean over trials)1.01.1delta
position0.74280.7430+0.0002
typography0.77960.7796+0.0000
color0.89890.9142+0.0153
size0.95980.9605+0.0007
text0.98450.9845+0.0000
recall0.76960.7691-0.0005
extras_ratio0.18170.1816-0.0001
structural0.65930.6641+0.0048
perceptual0.70750.7075+0.0000

Only colour moves: 1.1 reports the background a border-only or gradient box actually paints instead of black, so both sides agree more often. Pages scoring under 0.85 on colour: 34 of 700 under 1.0, 7 under 1.1 - the one figure in section 5 that the grader change moves. Everything else in this report is unaffected.

8. Does the grader agree with people?

The degradation suite says the reward falls when a page is broken in a known way. It does not say the reward orders two independent replications the way a viewer would. That needs a blind pairwise check: show the reference and two candidates with no scores, in randomised order, and ask which is closer. 40 pairs were put to a person: 30 spread across tasks at a reward gap of 0.03 or more, and 10 cut at a gap of 0.10 or more. 35 carry a verdict and 5 were left blank as genuinely uncallable. A Claude judge makes the same forced choice on all 30 of the first set and on 200 more random pairs, which is the only way to get an interval this side of a paid annotation run.

Both checks compare index pages; the reward does not. A contact sheet and a judge call show one page per trial, while the reward is the mean over 5 to 8 pages. So a disagreement can mean the reward is wrong, or simply that the pages nobody looked at decided the ranking. The rows below quantify that directly.
comparisonagreerateWilson 95% CI
Claude judge vs grader, whole-task reward (random pairs)97/20048%42% - 55%
Claude judge vs the index page's own reward (same pairs)112/20056%49% - 63%
Claude judge vs grader, on pairs where index and total agree89/16953%45% - 60%
Human vs grader (25 of 30 blind pairs labelled, index page)18/2572%52% - 86%
Human vs Claude judge (25 of 30 blind pairs labelled, index page)15/2560%41% - 77%
Claude judge vs grader (the same 30 pairs)16/3053%36% - 70%
Human vs grader, clear-gap set (10 blind pairs, gap >= 0.10)10/10100%72% - 100%
Human vs grader, all 35 labelled pairs (both sets pooled)28/3580%64% - 90%

The page-scope confound is real and measurable: in 31 of 200 random pairs the index page alone ranks the two trials the opposite way from the whole-task reward. Scored against the page it was actually shown, the judge does better (112/200, 56%) than against the whole-task reward (97/200, 48%) - some of the disagreement is scope, not grader error. The same caveat applies to every human row in this table, which is why the human rows are labelled as index-page judgements.

Human protocol. 40 pairs, full-page contact sheets at 500 px per column (docs/report/pairs/, docs/report/pairs_cleargap/), no scores shown, A/B order randomised. The labeller is the author of the grader: blind to the scores on these pairs, but not an independent annotator, and there is only one of them, so there is no inter-annotator agreement to report. 35 of 40 rows carry a verdict, which puts the 95% interval on the pooled human row at about +/-13 pp.

Blanks. 5 of the 40 rows were left blank: on those the labeller reported the two replications as good and bad in different ways, with no honest preference to state. Their reward gaps run 0.033 to 0.070, inside the 0.033 to 0.130 range of the rows that were called. Blanks are excluded from every rate above rather than counted as a miss.

Human agreement is not one number, it is a function of how far apart the grader put the two replications. Both label sets pooled, split by that gap:

grader reward gappairs shownhuman vs graderrateWilson 95% CIhuman vs judge
< 0.05126/967%35% - 88%5/9
0.05 - 0.101610/1471%45% - 88%8/14
≥ 0.101212/12100%76% - 100%7/9

The first set is dominated by near-ties, so the second set of 10 pairs was cut at the other end of the scale: same task, reward gap 0.10 to 0.13 (docs/report/pairs_cleargap/). That is the claim that matters - that the ordering holds where the difference is large enough to see - and it is the band where the judge agrees most often too. The reading is that the reward is trustworthy for coarse orderings and noisy for fine ones, which is what a dense shaping signal is usually allowed to be.

The same random pairs, split by how far apart the grader put the two replications. If the reward tracks something a viewer can see, agreement should climb with the gap.

grader reward gappairsagreerateWilson 95% CI
0.02 - 0.0510143/10143%33% - 52%
0.05 - 0.108340/8348%38% - 59%
0.10 and up1614/1688%64% - 97%

Pooling pairs hides the shape of the disagreement, so here is each trial's judge win rate against its grader reward, per task. Read the rho column as a direction, not a measurement: a win rate built from a handful of comparisons is a noisy quantity.

tasktrialscomparisons per trial (median, range)Spearman rho, judge win rate vs reward
civic-ember-group-2123 (1-8)+0.64
clinic-tessellate-supply-46104 (2-9)+0.59
podcast-foxglove-25104 (3-6)+0.50
travel-alder-works-35104 (1-6)+0.29
atlas-cafe104 (3-6)+0.26
fitness-harbor-partners-13102 (2-9)-0.04
clinic-corvid-health-37104 (2-9)-0.18
conference-anvil-kitchen-42105 (2-6)-0.39
portfolio-ember-44104 (2-6)-0.68
civic-juniper-collective-19124 (1-6)-0.75
docs-wren-collective-33103 (2-9)-0.89

The honest reading. On the widest gaps the two agree (the last bucket above). Below that they are indistinguishable from a coin flip - the pooled 97/200 is not anti-correlation, it is no signal. The mean per-task rho is -0.06, i.e. about zero; individual tasks run from -0.89 to +0.64, but each of those is fitted on a median of 4 comparisons per trial, so the spread is what noise looks like at this sample size and no single task's sign should be believed. Three readings fit the data and it cannot separate them: near-ties dominate the pair set and two replications that close are hard for a viewer to rank either (the human band table above shows the same slope, and every blank row sits in this band); or the judge is ranking the index page while the reward ranks 5-8 pages, which is measured above; or the reward genuinely weights fine layout geometry a viewer does not. What the data does not show is that the reward is arbitrary: the degradation suite, which compares a page against a controlled damage of itself, is monotone on every axis.

Read these intervals with care. The Wilson intervals treat pairs as independent; they are not. The random pairs come from 114 trials, so one wrong ranking inside a task flips every pair drawn from it - the effective sample is nearer 11 rankings than 200 comparisons. The widest gaps are also not evenly spread: 16 of the 40 pairs above 0.08 come from civic-ember-group-2. Resolution is not the differentiator: the judge encoder scales each screenshot by min(900/w, 1400/h, 1), which over the 114 candidate index pages here is a median effective width of 568 px (p25 457, p75 644, range 313-900) against the human sheets' 500 px columns - comparable, and wider than the human sheet on more than half the pages. Height is not the differentiator either: the sheets' 2600 px crop is applied after the resize to 500 px, and 0 of 114 candidate index pages are tall enough to reach it (the tallest sheet actually built is 2223 px), so not one sheet was cropped and the two views are near-equivalent on both axes. And judge_pairs.csv as run does not record which candidate was shown first, so these 230 rows cannot be audited for position bias; the column exists in the writer for future runs.

Judge: claude-opus-5, 230 forced-choice vision calls (750,642 input / 48,577 output tokens), estimated $4.97. A/B order randomised per call; the judge never sees a score.

The Claude judge is evidence about the reward, never part of it. No judge output enters grader/, the verifier, or any reward path; the reward is deterministic and the judge is not. A pairwise check is used rather than a rating because Design2Code's own study shows raters agree on comparisons (Fleiss kappa 0.46) far better than on absolute scores.

9. Site complexity against the Design2Code baseline

Design2Code criticises WebSight for generating pages simpler than real ones. The same criticism applies to us until we measure it. Source complexity of the reference sites against their 484 real scraped pages (Appendix Table 4). Depth is counted inside <body>: add 2 for the html/body wrapper to compare. Rendered elements are the verifier's own reference-element count.

sitetierpagestags/pageof which SVGunique tagsbody depthrendered elements/page
atlas-cafe fixtureS51111324.64.483
civic-ember-group-2L863811340.89.1505
civic-juniper-collective-19S53246839.88.0222
clinic-corvid-health-37S53225444.68.6275
clinic-tessellate-supply-46M64287943.58.5318
conference-anvil-kitchen-42L84868242.88.4378
docs-wren-collective-33M63486142.58.7278
fitness-harbor-partners-13L854512642.67.6409
podcast-foxglove-25S53867143.07.4309
portfolio-ember-44M644311845.07.5314
travel-alder-works-35M63885045.88.2292
mean, the 10 generated sites4318243.08.2330
atlas-cafe fixture, on its own1111324.64.483
Design2Code real pages158-2213-

Read honestly, and over the 10 generated sites only - the hand-built fixture is a different animal and is kept out of the mean. On tag count they are 2.7x the real-page baseline (2.2x if every inline SVG tag is thrown away), and ahead on variety (43 unique tags vs 22) - though inline SVG is the only graphics these sites may use, so part of both numbers is drawing, not structure. The gap that remains is depth: 8.2 inside <body>, so about 10 against their 13. That follows from the generator's own header + 4-8 sections + footer constraint and is the one axis on which these pages are simpler than the web. The fixture, at 111 tags and depth 4.4, is the small S-tier page the note in docs/decisions/0005 measured; the generated sites are several times denser. Nothing gates on any of this.

10. Infrastructure failures

backendexceptiontrialswhat it is (docs/decisions/0006)
modalCancelledError21job killed after the containers went silent (alive, no output)
dockerCancelledError7duplicate top-up jobs launched by two sessions, then killed
modalInternalError5Modal dropped the exec stdio stream mid-run
daytonaDaytonaNotFoundError2our own auto_stop_interval_mins=15 stopped the sandbox during the agent run
daytonaRuntimeError2same auto-stop, wrapped
dockerNetworkConnectionError2transient DNS/egress on the claude-code bootstrap curl
modalRuntimeError2same stdio-stream drop, surfaced as RuntimeError
backendtrialsgradedfailedfailure rate
daytona128433%
docker2213941%
modal121932823%

41 of 155 trials (26%) died before the verifier ran and wrote no reward and no breakdown (the manifest collector asserts that reward present <=> breakdown present). Per docs/decisions/0006, none is a task, grader or agent defect: 28 are Modal dropping the exec stdio stream or hanging silently 20+ minutes into a run, 4 are our own auto_stop_interval_mins=15 flag stopping a Daytona sandbox mid-agent, and 9 are local Docker jobs killed as duplicates or lost to a bootstrap curl. They are orchestration losses, not zero scores: scoring them as zeros would move the mean by a third and would be measuring Modal, not the agent. The fix is a retry around a trial rather than around a job.

11. Part 2: animated tasks

These tasks carry animated = true and are graded by 1.2 with --animated, which renders both sides at 8 fixed instants (0, 150, 300, 500, 750, 1000, 1500, 2000 ms) with reduced-motion: no-preference and scores the matched elements' delta trajectories. A page with reference motion scores page = 0.7 × static + 0.3 × motion; a page with none scores exactly what 1.1 gave it. A static misplacement is the static term's business and cancels out of the delta comparison, so the two structural halves do not double count; the motion perceptual term compares whole viewport frames, so it does see static misplacement, at 20 percent of the motion term, as 0007 states. See docs/decisions/0007 and docs/design-part2.md. Sections 1 to 10 above are the Part 1 set and exclude these trials entirely, cost and infrastructure failures included, so every Part 1 number there is the number it was before Part 2 existed.

20graded trials
2animated tasks
0.489mean reward (1.2)
0.661mean static
0.089mean motion
2cancelled or failed

Agent spend on the Part 2 attempts, as recorded per trial by Harbor and counted separately from section 1: $86 over all 22 attempts, $85 over the 20 graded ones ($4.24 per graded trial).

tasktierarchetypepagesnmeansdminmaxstaticmotionpositionsizecoloropacity0.4 → 0.9
agency-bramble-139-animSagency5100.5150.0250.4790.5730.7000.0840.6070.9410.4110.617
ecommerce-wren-studio-101-animSecommerce5100.4640.0220.4430.5080.6220.0940.1960.4220.4330.566

The two terms are worth reading apart. The static half scores 0.661 and the motion half 0.089; mixed at 0.3 on the pages that carry reference motion, that lands the mean reward at 0.489. Motion is the harder of the two here, which is where the headroom is. Section 1's Part 1 mean is 0.622, and the two figures are not comparable: different tasks, grader 1.0 there against 1.2 here, and 20 graded animated trials. Both are printed; no delta is quoted.

motion, 20 Part 2 trialsmean
position0.401
size0.681
color0.422
opacity0.591
trials that animated nothing (no candidate element moves)0 of 20
motion extras per trial (candidate moves, reference does not)579.45

The index page at each of the 8 instants, whole viewport scaled to 240x150 per cell: reference row over candidate row. These are the frames the motion term is computed from.

agency-bramble-139-anim tier S, 5 pages

best 0.573 (static 0.714, motion 0.245)
best 0.573 (static 0.714, motion 0.245)
worst 0.479 (static 0.664, motion 0.048)
worst 0.479 (static 0.664, motion 0.048)

ecommerce-wren-studio-101-anim tier S, 5 pages

best 0.508 (static 0.630, motion 0.225)
best 0.508 (static 0.630, motion 0.225)
worst 0.443 (static 0.603, motion 0.069)
worst 0.443 (static 0.603, motion 0.069)

Part 2 attempts that were cancelled or failed, so they carry no reward: 2 of 22.

backendexceptiontrialscancelled or failed: what it was
daytonaCancelledError2cancelled by the user (subscription usage limit)
What this section does not claim. 2 animated tasks is a sample, not a spread: nothing here supports a claim about archetypes, and the tier column is carried over for reference only. The blind human check in section 8 covers the static term alone, on Part 1 pairs; no person has been asked to rank two animations, so the motion term is checked by the degradation suite in docs/decisions/0007 and by nothing else. The motion perceptual term inherits the Part 1 floor: an attack that leaves the DOM intact and hides content at the pixel level is still bounded below by the structural term, not by the pixels. And 0.3 was chosen before the degradation suite ran and confirmed by it, unchanged (decision 0007).

12. Part 3: frameworks

These tasks carry framework in task.toml. The agent is given the same screenshots as the HTML task of the same site and a ready Vite project; the verifier runs npm run build, opens each built page with JavaScript on, freezes the rendered DOM and grades it with the unchanged grader (docs/design-part3.md). A compliance violation returns 0 and keeps the grader's number as graded_reward; both are shown. 1 site: one site cannot separate framework effects from site effects, so every comparison below is within a site.

30graded framework trials
3tasks
8zeroed by compliance
0build failures
0ended on an API error
1infra failure

podcast-foxglove-25

0.50.60.70.80.91HTML + CSSHTML + CSS 0.663HTML + CSS 0.669HTML + CSS 0.670HTML + CSS 0.686HTML + CSS 0.700HTML + CSS 0.708HTML + CSS 0.719HTML + CSS 0.743HTML + CSS 0.744HTML + CSS 0.764React + CSSReact + CSS 0.645React + CSS 0.659React + CSS 0.670React + CSS 0.678React + CSS 0.706React + CSS 0.714React + CSS 0.719React + CSS 0.723React + CSS 0.731React + CSS 0.743React + TailwindReact + Tailwind 0.636React + Tailwind 0.653React + Tailwind 0.669React + Tailwind 0.670 (zeroed by compliance, plotted at its graded reward)React + Tailwind 0.678 (zeroed by compliance, plotted at its graded reward)React + Tailwind 0.679React + Tailwind 0.684React + Tailwind 0.684 (zeroed by compliance, plotted at its graded reward)React + Tailwind 0.690React + Tailwind 0.704Solid + TailwindSolid + Tailwind 0.648Solid + Tailwind 0.657 (zeroed by compliance, plotted at its graded reward)Solid + Tailwind 0.658 (zeroed by compliance, plotted at its graded reward)Solid + Tailwind 0.660 (zeroed by compliance, plotted at its graded reward)Solid + Tailwind 0.666Solid + Tailwind 0.673 (zeroed by compliance, plotted at its graded reward)Solid + Tailwind 0.681Solid + Tailwind 0.682Solid + Tailwind 0.702 (zeroed by compliance, plotted at its graded reward)Solid + Tailwind 0.713

One dot per trial at the reward the grader scored; a hollow dot is a trial the compliance check zeroed, drawn where the grader scored it. A build failure has no dot.

columnnzeroedmean rewardsdminmaxmean graded rewardpositionsizecolortypographytext
HTML + CSS1000.7070.0330.6630.7640.8640.9700.9270.7790.991
React + CSS1000.6990.0320.6450.7430.6990.8710.9730.9000.8120.989
React + Tailwind1030.4710.3090.0000.7040.6750.8780.9690.9290.8250.990
Solid + Tailwind1050.3390.3390.0000.7130.6740.8760.9680.9320.8300.990

HTML + CSS column: 10 Part 1 trials, rewards from docs/report/regrade-summary.json (grader 1.1, identical to 1.2 on static pages). Framework columns: grader 1.2 on the snapshot. mean graded reward ignores the compliance zero.

Per page, paired against the HTML column

pageHTML + CSSReact + CSS (delta)React + Tailwind (delta)Solid + Tailwind (delta)
contact0.7590.750 (-0.009)0.736 (-0.022)0.726 (-0.032)
episode0.6910.673 (-0.018)0.694 (+0.003)0.688 (-0.003)
episodes0.7490.742 (-0.007)0.623 (-0.126)0.628 (-0.121)
hosts0.6440.644 (-0.000)0.656 (+0.012)0.662 (+0.018)
index0.6910.686 (-0.005)0.664 (-0.027)0.665 (-0.025)

Compliance violations

ruletrialsentries
inline-style821

React + CSS

best 0.743
podcast-foxglove-25-react-css best index
worst 0.645
podcast-foxglove-25-react-css worst index

React + Tailwind

best 0.704
podcast-foxglove-25-react-tailwind best index
worst 0.000
podcast-foxglove-25-react-tailwind worst index

Solid + Tailwind

best 0.713
podcast-foxglove-25-solid-tailwind best index
worst 0.000
podcast-foxglove-25-solid-tailwind worst index
What this section does not claim. One site per framework is a sample, not a spread. The compliance checks catch the bypasses they name (raw markup in an entry file or public/, changed manifests, hand-written CSS or inline styles in the Tailwind variants); a determined circumvention could pass. Behaviour numbers (builds run, snapshot tool used, npm install attempts, arbitrary-value share) are in patterns.json.

Generated by scripts/report.py. Pairwise material: docs/report/human_pairs.csv, docs/report/pairs/, docs/report/judge_pairs.csv.