Cohort 6 · starts October 5
Executive Agent Leadership
Go deep on building and leading agents, and design your organization's agentic operating system.
Sign upReference card
Nufar Gaspar · Frontier Lab · October 2026
Your own real requests, given to two to four AI options side by side, names hidden. An option is a model, a setting inside one product, or an AI agent tool.
Pick work where you can spot a weak answer fast. Add one matter-of-taste task and one you can check. Write what good looks like first. Made-up material only.
A friend who keeps the key. A chat in a tool you are not testing. A spreadsheet sorted by a random number. The blind-shuffle script, also for images and websites.
Finding your lineup: your AI lineup and default. Testing something new: Switch (new default), Split (joins your lineup for some tasks), or Stay (test again later).
A clear win on tasks that matter moves you. A close call goes to the tie-breakers.
Rerun the same requests, unchanged. Hide the names, pick, decide. Refresh the benchmark only when your work changes, a new kind of AI ability arrives, or a task stops separating the options.
Optional: the scored version adds a scoring guide and a strict AI judge; your scores decide. The automated version runs it all.
Prompts 1 to 7 are in the lab.
Keep going with us
Cohort 6 · starts October 5
Go deep on building and leading agents, and design your organization's agentic operating system.
Sign upCohort 4 · starts October 12
Need to catch up first? Get fluent in using AI and building with AI, for yourself.
Sign up