Every card below is a recorded run, not a screenshot. The steps are the
tool calls the agent actually made, in order; the verdict is the agent's own words; the figure is what the task used. Nothing here is a claim you cannot check —
the file that produced this page is the same one the bench harness writes.
39tasks completed
145tool calls
622stotal wall clock
89%of a Free day, all 39
social
It built the page, then looked at it
make us a wordle that only uses words from our server's inside jokes: BONK, YEET, SUSSY, GRIND, VIBED. give me a playable page
Write→inside-joke-wordle.html
BrowserNavigate→inside-joke-wordle.html
BrowserScreenshot→—
Read→screenshot.png
Attach→inside-joke-wordle.html
What it made, as it saw it
Now let me check how it actually looks: Here's your Inside Joke Wordle!
does this actually do what the comment says?
# returns the average score
def average(scores):
return sum(scores) / len(scores)
no tools needed — it read the question and answered it
Yes, it does exactly what the comment says — for valid input.
0 calls10.7s0.3% of a Free day
social
A page you can pin, in twelve seconds
make us a page with the server rules, numbered, that i can pin: be kind, no spoilers without tags, keep clips under a minute, no self-promo on weekdays, ping mods not everyone
Write→server-rules.html
BrowserNavigate→server-rules.html
BrowserScreenshot→—
Attach→server-rules.html
What it made, as it saw it
Here's your server rules page — clean, numbered, and ready to pin: 1.
Model openrouter/qwen/qwen3.7-flash. Every figure is a share of one Free day's
budget — the same percentage /usage reports. Bars are scaled to
that run's share against the heaviest in the set, so a heavier turn reads
heavier. Runs scoring below 1.0 say so on the card rather than being dropped.