Start here

Start here

The short version. Everything on this page comes from the lab, where each step has its full instructions.

A new AI model or tool arrives every few weeks, with test scores and a week of hot takes. None of it tells you whether it is better at your work. A personal AI benchmark does. It is a small set of your own real requests. You give them to several AI options side by side, with the names hidden, and see which answer you would use.

The whole process on one page: find your tasks, choose your options, run the tryout, hide the names and pick, decide and change your habits

The taste test in six steps

  1. Choose two to four options.
  2. Run every request in each, in a fresh chat.
  3. Hide the names.
  4. Read all, then pick the one you would use, or score each from 1 to 5 (few 5s). One line on why: clear or close?
  5. Reveal and count.
  6. Decide.

Begin with Prompt 1: Find my tasks

Paste this into the AI tool that knows you best. If it has memory of you or access to your files, it starts from what it knows. If it knows nothing about you, it interviews you from the start. Answer in your own words, and correct it when it guesses wrong.

I want to build my personal AI benchmark. That is a small saved set of real requests from my own life, which I will give to several AI tools side by side, with the names hidden, to see which one does my work best. I will reuse the same requests, unchanged, every time a new AI model or tool comes out. Help me choose those requests by interviewing me.

A good set covers the mix of what I really do. Use these six questions as your map of my work:

1. What kind of work is it? Writing; analysis of numbers or documents; research; decisions and advice; planning and admin; making things such as images, slides, or web pages.
2. How often does it happen? Daily habits, and rare moments that matter a lot when they come.
3. What is at stake? Quick, low-risk requests, and ones where a weak answer would cost me.
4. Where does the AI start? From a blank page, or from my own material such as notes, a draft, a spreadsheet, or a long document.
5. Which part of my life is it? Work, or personal life.
6. Is it on my wish list? Things I have wanted AI to do for me that it could not do, or could not do well.

How to run the interview:

- Start with what you already know. If you have memory of me, our past conversations, or access to my files, begin by summarizing in a few lines what you think I use AI for, including the requests I seem to repeat or go back and forth on. Then ask me what you got wrong. If you know nothing about me, say so and start asking.
- Ask one question at a time, in plain language, and wait for my answer. Keep each question short.
- Start broad: my role, what a normal week looks like, and what I already use AI for. Then go deeper on the parts of the map my answers have not reached.
- Every few answers, tell me in one sentence which parts of the map are covered so far and which are still empty. Then ask about an empty one.
- Useful questions to draw on: Which request do you make most often? Where do you usually rewrite the answer, or ask several times before you are happy? Which rare task matters a lot when it comes, such as a board update or a difficult conversation? Where would a wrong answer embarrass you or cost money? What do you use AI for outside work, if anything?
- Near the end, always ask these three, one at a time: What have you always wanted AI to do for you that it could not do, or could not do well? What has frustrated you with AI recently? Which parts of your work do you still not trust AI with, and what would change your mind?
- Ask twelve questions at most. Stop earlier if the map is well covered or I say I am done.

Then propose my set: three to five everyday tasks (things I really do, regularly or at important moments) plus one or two wish-list tasks, with six tasks at most in total. Choose them so the set looks like my real life, and prefer tasks where I know the work well enough to tell quickly when an answer is off, because I will be the judge: several kinds of work, at least one rare but important moment, at least one where a weak answer would cost me, at least one that starts from my own material, and at least one where the best answer is a matter of taste, such as a message in my own voice. Weight the set toward the work that fills my week. Include a personal task only if I use AI outside work. If I use AI tools that can work with my files or take several steps on their own, include one task of that kind.

For each task, give me:

- A number and a short name.
- One line on why it earns its place, and which parts of the map it covers.
- The complete request I will paste, ready to use as it is. Where the request needs background material, such as meeting notes, a draft, figures, or an email thread, write realistic made-up material directly into it, as messy and specific as the real thing, so I never need to paste anything private or confidential. Keep each request short enough to paste into any chat window.
- A short scoring guide: two or three plain lines on what a great answer must include and what would make me reject it. To write it, ask me one question per task: how would I tell a great answer from a weak one here? Use my own words. If I cannot answer, tell me the task may not belong in the set, because I will be the judge.
- If the task asks for an image, slides, or a web page, a reminder that I will need to save those results in a folder to compare them.

Finish with a short table showing each task against the six map questions. Name any part of the map that is still missing, and say whether that matters for someone like me.

Then save it in the shape that fits where we are. If you can create files, make a folder called my-ai-benchmark with three things in it: a README.md that lists my tasks by number and name and leaves space for my lineup; a tasks folder with one file per task, named like 01-weekly-update.md, each with three headings in this order: "Why it is here" (one line), "Request" (the complete request to paste, with any made-up material written into it), and "Scoring guide" (the two or three lines we agreed); and an empty answers folder for the answers I collect later. If you cannot create files, give me everything in one document, numbered, with the requests complete, and I will save it myself. Before you save anything, ask me to review the set: for each request, is this something I would really send, in these words, with material like this? Ask what to swap, what to sharpen, and what is missing from my real week. Revise until I say the set is mine. Only then save it, and remind me to keep these exact requests unchanged from that point on, because I will paste them every time something new comes out.

Then continue in the lab

Keep going with us

Where to go from here

Cohort 6 · starts October 5

Executive Agent Leadership

Go deep on building and leading agents, and design your organization's agentic operating system.

Sign up

Cohort 4 · starts October 12

Executive Catch-Up

Need to catch up first? Get fluent in using AI and building with AI, for yourself.

Sign up
More from The AI Daily Brief