The lab

From here on: build once, reuse for months

You have now done the hard part. The thinking about how to evaluate an AI is finished, and it is saved in one document. From here on, a new release costs you no thinking at all.

From here on: build once, reuse for months

When something new comes out, you do Part 2 again with the new option and what you use today: the same saved requests, unchanged, names hidden, pick, decide. That is the whole routine. You do it as often as releases interest you, and you never start from a blank page again.

Refresh the benchmark itself only rarely. Three moments call for it:

  • Your work has changed. A new role, a new kind of task filling your week. Run Prompt 1 again and swap in what is missing.
  • A step change in what AI can do. When a new kind of ability arrives, add one wish-list task that tests it.
  • A task stopped telling you anything. If every option handles it perfectly, replace it with a harder one from your real work.

Between those moments, leave the benchmark alone. That is what makes results comparable across releases. Built well, a benchmark serves you for a good few months, and often longer.

For a new agent tool the routine is the same, with the tasks that use files and tools going first.

Done: one saved document, a lineup you trust, and a routine you can repeat at every release.

Keep going with us

Where to go from here

Cohort 6 · starts October 5

Executive Agent Leadership

Go deep on building and leading agents, and design your organization's agentic operating system.

Sign up

Cohort 4 · starts October 12

Executive Catch-Up

Need to catch up first? Get fluent in using AI and building with AI, for yourself.

Sign up
More from The AI Daily Brief