The lab

Part 3 (optional). The scored version

For the extra diligent. You keep the same tasks and add three things.

  1. A scoring guide for each task. A few written lines that say what a great answer must include, what earns top marks, and what counts as an automatic fail. A good guide is specific enough that two people, or a person and an AI, would give the same answer the same score. For example, "sounds like me" is a wish. A scoring guide turns it into lines anyone can check: "no sentence over 25 words, opens with a concrete example, no hedging phrases such as 'it may be worth considering'."
  2. Scores from 1 to 5, which you give yourself, with the names hidden.
  3. An AI judge. A second AI, from a company that made none of the options you are testing, scores the same hidden answers against your scoring guide. You then compare its scores with yours. An AI judge left to itself gives top marks to almost everything, so Prompt 6 asks it to act as a tough editor. Your own scores decide.

Picking by taste is fast and builds your judgment. A scoring guide makes results comparable from one release to the next, and once you trust the judge, it can do the scoring for you.

Three kinds of saved task

  • Everyday tasks are the regular work you chose in Part 1. Keep them word for word between runs, so this month's results compare with last month's. If you do change one, save it as a new, dated version.
  • Wish-list tasks are things you want AI to do and it has not managed yet. Refresh them at each new release. When testing something new, add one or two built around what the release says it can now do.
  • Matter-of-taste tasks are the ones where the best answer is simply the one you prefer, such as advice on a tricky conversation. Judge them on four points: pushes back when it should, right length, sounds as sure as it should be, and usable as it is. These results give you a read on personality. Take them with a pinch of salt.

Also mark each task as a chat task or an agent task. A chat task works in any chat window with the material pasted in. In an agent task, the AI has to open files, use tools, or take several steps on its own. When you compare agent tools, run the agent tasks first, because that is where the tools differ most.

You and the judge

Score first, then read the judge's scores, then reveal the names. When you and the judge disagree, one of three things happened:

Judging taste, fairly: two judges, you or an AI judge

  • Your scoring guide is missing something you care about. Add it, and date the change.
  • The judge misread. Try a different judge, or score that task yourself from now on.
  • Length or polish swayed you. Reread the scoring guide before you score next time.

Check the judge now and then. When you first set it up, and again every few releases, score a few answers yourself and compare. If you mostly agree on a kind of task, you can let the judge score that kind alone. The benchmark checks your judge, too.

The decision rule

Write it down before you see any answers, so a polished surprise cannot talk you into anything. A good default:

  • Switch when the new option matches or beats what you use today on every everyday task, and wins at least one wish-list task.
  • Split when it wins some tasks and loses others. List the tasks it wins.
  • Stay in every other case.

When finding your lineup, the option with the most wins becomes your default, and each other option takes the kinds of task it won. Matter-of-taste tasks shape your notes on personality and stay out of the verdict. Cost, speed, and how you like each tool settle the close calls, exactly as in the taste test.

The scored version, step by step

  1. Testing something new? Research the release with Prompt 4, in a tool with web search. Skip this step when finding your lineup. Done: a short brief on what the company claims, what people report, and which tasks to watch.
  2. Build your scored benchmark with Prompt 5. Paste your task list from Prompt 1, and the research brief if you have one. Done: one document with every task, its material, its scoring guide, an empty scorecard, and your decision rule.
  3. Run every task in every option, and hide the names, exactly as in the taste test. Done: lettered answers for every task.
  4. Score each answer yourself from 1 to 5, rereading the scoring guide first. Write your scores down. Done: your scores, recorded before you see anyone else's.
  5. Bring in the judge with Prompt 6, in a fresh chat with an AI from a company you are not testing. Done: a filled scorecard with both sets of scores, how often you agreed, and a verdict or a lineup.
PROMPT

Prompt 4: Research a new release

A new AI model or AI tool has just come out, and I want to decide whether it matters for my work. Help me research it. Use web search throughout and cite a source for every claim.

Before you start, ask me three questions, one at a time, and wait for each answer: what was released and by which company (it may be a model, a new option inside a product, or a tool that works on its own with files and apps); how many days ago it came out; and which AI tools I can open today.

If the release is a tool, read its release notes and help pages, which play the part a model announcement plays for a model. Its claims will be features, such as memory, connections to other apps, or the ability to work on its own for longer. Tell me which kinds of work those features would change.

Timing: if it came out less than five days ago, label everything people online are saying as EARLY AND UNRELIABLE, because the first days are mostly reposts of the announcement and showy demos. Tell me to run this research again in a week.

Then give me four short sections, in plain language:

1. What the company claims. A table with the claim, the evidence offered, whether that evidence comes from a test the whole industry uses or one the company created itself, and the kind of everyday task where I would notice the difference. Add a few lines on changes to cost, to speed, and to how much material it can work with at once. Add one line on anything the company is quiet about, such as tests it stopped reporting.
2. What people report. Search X, LinkedIn, Reddit, Hacker News, Substack, and independent blogs. Give more weight to posts that show the actual request and answer, to people who work in a named field such as law, finance, or marketing, and to detailed reports of failures. Leave out posts that only repeat the company's numbers, and do not let any single popular account decide. A table with what people say got better, what got worse, which reports show evidence, and which field they come from.
3. Tasks to watch. Eight to twelve kinds of task where this release is most likely to be noticeably better or worse than what came before, each with one line on why, and whether the evidence comes from the company, from users, or both.
4. Personality. One short paragraph on what people say about its style: how long its answers are, whether it pushes back, how often it hedges, and how it handles unclear requests. Label it "impressions only".

End with three questions I should ask myself about my own work before I decide what to test.
PROMPT

Prompt 5: Build my scored benchmark

Help me build a scored personal AI benchmark. A personal AI benchmark is a saved set of my own real requests that I give to several AI options side by side, with the names hidden, to see which one does my work best. An option can be an AI model, a setting inside one product (such as fast or thinking), or an AI tool that works through several steps on its own. The scored version adds a written scoring guide to every task, scores from 1 to 5, and an AI judge whose scores I compare with my own.

Interview me first, one question at a time, and wait for each answer:

- Ask me to paste my saved task list. If I do not have one, tell me to create it first with a task-finding interview, and stop there.
- Ask which situation I am in. Finding my lineup means comparing a few options to choose my default and which kinds of task go elsewhere. Testing something new means comparing a new release with what I use today. If I am testing something new, ask me to paste any research I have on the release.
- Ask me to name the options, and which one I use today, if any.
- Ask whether my tools show me the cost or usage of each conversation.

Then build one document I will still understand a year from now. Explain every part inside the document in plain words.

Part 1, everyday tasks. Use my regular tasks, five to eight in total. If I have fewer than five, suggest more from what you have learned and ask me to approve them. For each task, write down: the exact request, word for word; any material it needs, made up and realistic so nothing private is included; what a great answer must include; the common ways an answer fails; a scoring guide from 1 to 5 that says what earns a 5 and what is an automatic fail, specific enough that two people would give the same answer the same score; and a tag, either chat task (works in any chat window with the material pasted in) or agent task (the AI must open files, use tools, or take several steps on its own). Tell me these tasks stay word for word between runs, and that any change becomes a new dated version.

Part 2, wish-list tasks. Two to four tasks I have wanted AI to do and it has not done well. Take them from my list. If I am testing something new, build one or two around what the release claims it can now do, shaped so that the claimed ability has to show up for the task to succeed. Write the same fields as for the everyday tasks, plus one line on where each task came from. Tell me to refresh these at each release.

Part 3, matter-of-taste tasks. Choose up to three tasks where the best answer is simply the one I prefer, such as writing in my voice or advice on a tricky conversation. Judge them on four points: pushes back when it should; right length; sounds as sure as it should be; usable as it is. If I am comparing agent tools, add two more points: asks before taking actions that matter, and reports clearly what it did. Label this part "a read on personality; take with a pinch of salt".

The scorecard. An empty table with one row per task. For each option, two columns: my score and the judge's score. Then columns for the winner, whether the judge and I picked the same winner, the cost per task where I can see it, speed, and notes.

The decision rule, written now, before any answers exist. If I am testing something new, one verdict per new option: Switch (it becomes my default) when it matches or beats what I use today on every everyday task and wins at least one wish-list task; Split (it joins my lineup for specific kinds of task, which I list) when it wins some tasks and loses others; Stay (keep my lineup and test again next release) in every other case. If I am finding my lineup, the option with the most wins becomes my default, and each other option takes the kinds of task it won. Matter-of-taste tasks shape my notes on personality and stay out of the verdict. Cost, speed, and how much I like each tool settle close calls.

A short run checklist: a fresh chat for every option; the same request and material; the same settings for web search, memory, and thinking; the same level of model when comparing companies (top against top, fast against fast) unless comparing levels is the point; names hidden before any scoring; I score first; the judge is an AI from a company that made none of the options; the names are revealed only after scoring. For agent tools: the same model inside each tool where the tool allows it, the finished result is what gets judged, and agent tasks run first. Say which tasks to run first if I am short on time.

Refer to AI options by company and level, such as "the top model from company X", so the document still makes sense after products are renamed.
PROMPT

Prompt 6: The AI judge

Score the answers yourself first. Then run this in a fresh chat with an AI from a company that made none of your options.

I ran my personal AI benchmark: the same saved requests given to several AI options, with the names hidden and the answers labeled with letters. I want you to act as a tough, independent editor who has to pick what actually gets used, and who has seen a great deal of competent, forgettable work. I will paste my benchmark document first, so you have the tasks, the scoring guides, and my decision rule. Then I will paste the lettered answers one task at a time. Do not guess which AI wrote which answer, and do not ask.

For each task, read every letter first and rank them from best to worst. Then score every letter from 1 to 5 against the scoring guide, and keep the scores in the order of your ranking. Competent work earns a 3. A 5 means I would use it exactly as it is and it clearly beats the others. Give at most one 5 per task, and give none if no answer earns it. Write one line for each score that names the answer's main weakness, including for the winner, and name the winner. Judge the finished result, and do not reward length or polish that the request did not ask for. If an answer is a web page, remember that you can read only its code and cannot see the page, so leave the look of it to me. If I also paste a record of the steps an agent tool took, use it only for the points about asking before acting and reporting what it did. For matter-of-taste tasks, rate the four points in the document and add one line on each letter's personality: what kind of work it seems suited to, and what I should watch out for.

After each task, ask me for my own scores, which I gave before reading yours. Tell me where we picked the same winner and where we differ. Where we differ, help me work out which of three things happened: my scoring guide is missing something I care about; you misread the answer or the guide; or I was swayed by length or polish. Push back if I try to move the goalposts or explain away a failure.

When all tasks are done, fill in the scorecard, including any cost and speed figures I give you, and tell me how often we agreed. For each task, say whether the win was clear or close. Ask me whether anything beyond the answers matters to my choice: whether I am allowed to use each option at work, its privacy terms, my plan, cost, speed, a feature I love, or the effort of switching. Then apply the decision rule from my document to my own scores. My scores decide, and yours are a second opinion. Let clear wins on tasks that matter move me, and let what I told you settle the close calls. If I am testing something new, give one verdict per new option: Switch, Split, or Stay, as defined in my document. If I am finding my lineup, give my AI lineup: my default, and which kinds of task go elsewhere. Add three notes on what I seem to value, written as qualities that will still make sense after products are renamed.

I will reveal which letter was which only after your decision. Then help me update my lineup and note what to retest next time.