Cohort 6 · starts October 5
Executive Agent Leadership
Go deep on building and leading agents, and design your organization's agentic operating system.
Sign upResources
Links we mention in the session, plus a few good reads on testing AI on your own work.
One account and one key for models from most AI companies. It reports what each answer cost and how long it took, which the automated version shows after the reveal.
Open ↗Send one message to several models and read the answers side by side. The names stay visible, so hide them yourself before you pick.
Open ↗Built for teams testing prompts inside software they ship. On all four, the names of the options stay visible.
An AI judge, review queues for people, and side-by-side comparison of two answers.
Open ↗A playground you can use without code, an AI judge, and review scores from people. The easiest of the four to start with.
Open ↗Open source, with a hosted version. Runs experiments on a saved set of prompts from its own page, with an AI judge.
Open ↗Free and open source, for builders. A settings file and one command run your prompts across models, with an AI judge, your own scores, cost, and speed.
Open ↗The previous Frontier Lab. Start with the AI Daily Brief episode behind it, "What the Heck is Graph Engineering?"
Open ↗Our experiment for getting a feel for tokens: what they are and what you get for spending them.
Open ↗The AI Daily Brief Operator's Cut on agents that work for a whole team.
Open ↗Why public benchmarks tell part of the story, and the case for testing AI on the work it will really do.
Open ↗One person's benchmark: the same odd drawing request, given to each new model.
Open ↗For builders: how teams test AI inside a product.
Open ↗Keep going with us
Cohort 6 · starts October 5
Go deep on building and leading agents, and design your organization's agentic operating system.
Sign upCohort 4 · starts October 12
Need to catch up first? Get fluent in using AI and building with AI, for yourself.
Sign up