---
name: benchmark-builder
description: >-
  Build a personal AI benchmark: a few of the user's own real requests, given
  to two or more AI options (models, settings inside one product, or AI agent
  tools) side by side with the names hidden. Starts with the Work Mix interview
  to find tasks that cover what the user really does, then guides the taste
  test with four ways to hide the names, then the decision (must-haves, the
  answers, tie-breakers), ending in an AI lineup or a Switch, Split, or Stay
  verdict. Offers the scored version, with a scoring guide and an AI judge, to
  users who want more rigor. Use when the user says "build my benchmark,"
  "which model should I use," "is this new model worth it for me," or "test
  this release on my work."
---

<!-- BUILD YOUR PERSONAL AI BENCHMARK · Frontier Lab giveaway · AIDB × Superintelligent × Nufar Gaspar
INSTALL: Claude Code → save as .claude/skills/benchmark-builder/SKILL.md in your project
(or ~/.claude/skills/benchmark-builder/SKILL.md for all projects), restart your session.
Cursor → .cursor/skills/ the same way. Codex → .codex/skills/ the same way.
Any chat tool → paste the whole file into a conversation and say "act as this skill." -->

# Benchmark Builder

Help the person find out, on their own work, which AI option to use and for what. Leave them with a saved set of requests they can rerun every time something new comes out.

## How to talk to the person

The person hears every word you say, and most people who use this skill have never tested AI tools before. So:

- Use plain, everyday words and short sentences. Ask one question at a time and wait for the answer.
- Use only the names in the table below. Define each one in a single sentence the first time you use it. If a technical idea comes up, describe what it does in everyday words.
- Treat the person as smart and busy. Explain the reason for a step in one line, then move on.

| Name to use | What to say the first time |
|---|---|
| **your personal AI benchmark** | A small set of your own real requests that you give to several AI options side by side, with the names hidden, to see which answer you would use. |
| **options** | The things you compare: an AI model, a setting inside one product (fast, thinking, or top), or an AI agent tool that works through several steps on its own, with files and apps. |
| **finding your lineup** | You have no clear favorite yet, so you compare two to four options. |
| **testing something new** | A new model or tool just came out, and you compare it with what you use today. |
| **your AI lineup** | A short card naming your default AI and any kinds of task you send to another one. |
| **Switch, Split, Stay** | Switch: it becomes your new default. Split: it joins your lineup for specific kinds of task. Stay: keep your lineup and test again next time. |
| **the Work Mix** | Six questions that check your set of tasks covers what you really do. |
| **everyday tasks, wish-list tasks, matter-of-taste tasks** | Things you really do, regularly or at important moments. Things you wanted AI to do and it couldn't. Tasks where the best answer is simply the one you prefer, such as a message in your voice. |
| **the taste test** | Run each request in every option, hide the names, pick the answer you would use, and write one line on why. |
| **must-haves, the answers, tie-breakers** | Must-haves: allowed at work, privacy terms, on your plan. The answers: who won, clear or close. Tie-breakers: cost, speed, the features you enjoy, plan limits, the effort of changing habits. |
| **scoring guide, AI judge** | Scored version only. A few written lines per task on what a great answer includes, what earns top marks, and what is an automatic fail. A second AI, from a company that made none of the options, that scores the hidden answers against that guide. |

## What to say on first use

Say this, briefly, in your own words: "Published test scores describe the average user. Your personal AI benchmark describes you. We will pick a few of your real requests, run them in the options you want to compare, hide the names, and see which answer you would actually use. The most important part is choosing requests that look like your real work, so we start there."

**Where this skill runs.** Installed as a skill in an agent tool such as Claude Code, Cursor, or Codex, you can read the person's working folders and notes. Pasted into a chat tool, you can use memory when it is on. With neither, you interview from the start.

## Phase 1: Find the tasks (the heart of this skill)

A benchmark can only tell the person about the work it contains. If every task is an email, they learn which AI writes the best email, and that is all they learn. Give this phase the most care.

**Step 1. Start with what you already know.** If you have memory of the person, past conversations, or access to their files, summarize in a few lines what you think they use AI for. Include the requests they seem to repeat, and the ones where they go back and forth or rewrite the answer. In an agent tool, look through their working folders and notes for repeated kinds of work before you ask anything. Then ask what you got wrong. If you know nothing about them, say so and start asking.

**Step 2. Interview with the Work Mix as your map.** Start broad: their role, what a normal week looks like, and what they already use AI for. Then go deeper on the parts of the map their answers have not reached. The six questions:

1. **What kind of work?** Writing, analysis of numbers or documents, research, decisions and advice, planning and admin, making things such as images, slides, or web pages.
2. **How often?** Daily habits, and rare moments that matter a lot when they come.
3. **What's at stake?** Quick, low-risk requests, and ones where a weak answer would cost them.
4. **Where does the AI start?** A blank page, or their own material such as notes, a draft, a spreadsheet, or a long document.
5. **Which part of life?** Work, or personal life.
6. **What's on their wish list?** Things they wanted AI to do and it could not, or could not do well.

Useful questions to draw on: Which request do you make most often? Where do you usually rewrite the answer, or ask several times before you are happy? Which rare task matters a lot when it comes, such as a board update or a difficult conversation? Where would a wrong answer embarrass you or cost money? What do you use AI for outside work, if anything?

**Step 3. Track coverage.** Keep a running note of which of the six questions the answers have covered. Every few answers, tell the person in one sentence what is covered and what is still empty, then ask about an empty part. For example: "So far we have your writing and your weekly routines. We have nothing yet on rare, high-stakes moments. Which one comes to mind?"

**Step 4. Always ask the three wish-list questions,** near the end, one at a time:

1. What have you always wanted AI to do for you that it could not do, or could not do well?
2. What has frustrated you with AI recently?
3. Which parts of your work do you still not trust AI with, and what would change your mind?

Ask twelve questions at most in the whole interview. Stop earlier if the map is well covered or the person says they are done.

**Step 5. Propose the set.** Prefer tasks the person knows well enough to tell quickly when an answer is off, because they will be the judge. Three to five everyday tasks plus one or two wish-list tasks, six tasks at most. Choose them so the set looks like the person's real life:

- several kinds of work
- at least one rare but important moment
- at least one where a weak answer would cost them
- at least one that starts from their own material
- one matter-of-taste task, such as a message in their own voice
- one task they can check, where an answer is plainly right or wrong, such as figures pulled from a table
- weighted toward the work that fills their week
- a personal task only if they use AI outside work
- one agent task (the AI opens files, uses tools, or takes several steps on its own) if they use agent tools

For each task, give:

- A number and a short name.
- One line on why it earns its place, and which Work Mix questions it covers.
- The complete request, ready to paste as it is. Where it needs background material, such as meeting notes, a draft, figures, or an email thread, write realistic made-up material directly into it, as messy and specific as the real thing. Keep each request short enough to paste into any chat window.
- For an image, slides, or a web page: a reminder that they will save those results in a folder to compare them.

Then show a short coverage table: one row per task, one column per Work Mix question, a mark where the task covers it. Name any part of the map that is still empty, and say whether that matters for someone like them. If the set came out lopsided, say so with the reason. Five writing tasks from their own material, all low stakes, would end in a tie that tells them very little.

**Step 5a. Ask what good looks like.** For each proposed task, ask one question: how would you tell a great answer from a weak one here? Turn the reply into two or three plain lines: what a great answer must include, and what would make them reject it. If they cannot answer for a task, say it may not belong in the set, because they will be the judge.

**Step 5b. Make it theirs.** Before saving, ask the person to review the set as if they were about to send each request for real: would they really ask this, in these words, with material like this? What would they swap, sharpen, or add from their real week? Revise until they say the set is theirs. A set that tests work they never do is worthless, however well it covers the map.

**Step 6. Save it.** If you can create files, build a folder called `my-ai-benchmark` with three things: `README.md` (the tasks by number and name, today's date, and a blank "My lineup" section), a `tasks` folder with one file per task named like `01-weekly-update.md`, each with three headings in order: "Why it is here", "Request" (the complete request with the made-up material written in), and "Scoring guide" (two or three lines in their own words), and an empty `answers` folder for the answers they collect. If you cannot create files, print the whole set in one document, numbered. Tell the person to keep these requests word for word, because they will paste the same ones every time something new comes out, and that the folder also feeds the blind-shuffle script and the automated version later.

**Done:** four to six numbered requests ready to paste, each with its reason, plus the coverage table, saved in one place.

## Phase 2: The situation and the options

Ask these one at a time:

1. **Which situation?** Finding your lineup, or testing something new?
2. **Which options?** Two to four they can open today. Four is the most anyone can judge well, even when a program runs the requests. Examples to offer: the fast, thinking, and top options in the product they already use, or the top option from three different companies.
3. **If testing something new:** what do they use today? It joins the comparison.

Then say there are three depths, and recommend the first. *The taste test* is complete on its own and sharpens their sense of what a good answer looks like far better than one quick try. *The scored version* is optional, for people who want more rigor. *The automated version* is optional, for builders. Offer the scored version only if the person asks for more rigor or wants results they can compare from one release to the next.

## Phase 3: The taste test

**Explain the run.** For each saved request, open a fresh chat in each option and paste the request exactly as saved. Keep the settings the same everywhere: web search on or off, memory on or off, and the same thinking setting where there is one. Copy each full answer into one document under the option's name. Paste quickly and try not to read yet. If one option searched the web and another did not, note it, because it can explain a surprise later.

**Offer the four ways to hide the names.** Brand preference is louder than most people expect, so the names come off before anyone picks.

- **A friend.** A colleague gets the answers with the names on, sends them back shuffled and labeled A, B, C, and keeps the key until the person has picked.
- **A chat.** Paste the named answers into a fresh chat in a tool that is not one of the options. It removes the names, shuffles, relabels, and keeps the key until the person types "reveal". No setup needed.
- **A spreadsheet.** One row per answer: the option's name in the first column, the answer in the second, and the formula `=RAND()` in the third. Sort by the third column, then hide the first. Unhide it to reveal.
- **A script,** for people comfortable running a small program. Save each answer as a file in a folder, one subfolder per request and one file per option, named after the option. The blind-shuffle script, `blind-shuffle/anonymize.py`, makes a copy with the answers relabeled and shuffled separately for each request, keeps the key in a separate file, and can tally the picks. Agent tools can save their answers straight into the folder.

**Things that cannot be pasted.** Images, slides, and websites go in a folder: one file per option, or one folder per option for a website. The script or a friend relabels them A, B, C. Judge a website from a full-page screenshot, because a live link shows which tool made it in the web address. After the reveal, the person opens the real thing and checks what needs no hiding: does it work, does it match the request, and how many follow-up requests did it take? If a tool left a visible mark or a recognizable style, note it next to the pick.

**Offer to be the chat, only when that is fair.** First check whether you, the AI running this skill, are one of the options, or come from the same company as one of them. If so, say plainly that you should not hide the names for this test, because you would be handling your own answers, and point to the other three ways or a different tool. If you are outside the options, offer this: the person pastes the named answers for one request at a time. You remove the names and any line where an answer mentions who wrote it, shuffle afresh for every request, and show the answers in full as A, B, C, with no comment, ranking, or hint. You record each pick. You show the key, the picks with the real names, and a count of wins only when the person types "reveal".

**Pick.** Ask the person to read every answer to a request before judging any, because judging is a comparison. Then they ask one question: which of these would I use as it is? They choose one letter, write one line on why in concrete words ("sounded like me", "caught the risk I missed", "half the length and nothing lost"), and note whether the win was clear or close. For a finer read, they can score each answer from 1 to 5 instead. Tell them to be stingy with top marks, because if everything gets a 4 or a 5, the scores stop telling them anything.

**Reveal and count.** Count the wins for each option. If an option they were rooting for lost, mention that the feeling is the reason the names were hidden.

**Done:** for every request, one letter, one line on why, "clear" or "close", and a count of wins for each option.

## Phase 4: Before you decide

The taste test judges one thing: the finished answer. Cost, speed, and how much the person enjoys a tool are real reasons to choose, and they come in now, with the names visible. Ask these one at a time, and wait for each answer:

1. For each request, if you do not have it yet: which option won, the one line on why, and was the win clear or close?
2. **Must-haves.** Are you allowed to use each option at work? Are you comfortable with its data and privacy terms? Is it on your plan? An option that fails here is out, however good its answers were.
3. **Tie-breakers.** Does anything else matter to your choice: cost, speed, a feature you love (voice, connections to your other apps, memory, a good mobile app), plan limits, or the effort of changing your habits?

Decide with one rule: **a clear win on tasks that matter moves the person. A close call goes to the tie-breakers.** Keep it to those two plain questions, clear or close, and whether anything outside the answers outweighs a close call. If it helps, give the example: a new model comes a close third, the person loves the voice mode in its app, and staying in that app is a sound decision. Had it come a clear first on one kind of task, the verdict would be Split.

Then give the person:

1. **Their AI lineup:** their default, and which kinds of task, if any, go to another option.
2. **If testing something new:** one verdict per new option, Switch, Split, or Stay, with one sentence of reasoning. For Split, name the kinds of task that move.
3. **Three notes on what they seem to value** in an answer, written as qualities that will still make sense after products are renamed, such as "gets to the point in the first line".
4. The result that surprised them most, and one thing to retest the next time something new comes out.

Close with the line: *the benchmark judges the answers, you judge the experience, and the decision uses both.* Save the result under today's date next to the task list.

**Done:** a lineup, plus a verdict for each new option when testing something new, saved next to the task list.

## Phase 5 (optional): The scored version

Offer this only to people who ask for more rigor. It keeps the same tasks and adds a scoring guide for each task, the person's own scores from 1 to 5 with the names hidden, and an AI judge whose scores they compare with their own.

**If testing something new, research the release first,** with web search, or send the person to Prompt 4 in the lab, in a tool that has web search. Separate what the company claims from what people report after real use. If the release is less than five days old, label online reactions as early and unreliable, and suggest checking again in a week. End with a short list of tasks to watch.

**Build one document** the person will still understand a year from now, with every part explained inside it in plain words:

- **Everyday tasks.** Five to eight in total, from the Phase 1 set. If there are fewer than five, suggest more from what you learned and ask the person to approve them. For each task: the exact request, word for word; its made-up material; what a great answer must include; the common ways an answer fails; a scoring guide from 1 to 5 that says what earns a 5 and what is an automatic fail, specific enough that two people would give the same answer the same score; and a label, chat task (works in any chat window with the material pasted in) or agent task. These stay word for word between runs, and any change becomes a new, dated version.
- **Wish-list tasks.** Two to four, from the wish-list answers. When testing something new, add one or two built around what the release claims it can now do, shaped so the claimed ability has to show up for the task to succeed. Same fields, plus one line on where each came from. Refresh these at each release.
- **Matter-of-taste tasks.** Up to three, judged on four points: pushes back when it should, right length, sounds as sure as it should be, and usable as it is. For agent tools, add two more: asks before acting, and reports clearly what it did. Label this part "a read on personality; take it with a pinch of salt".
- **An empty scorecard.** One row per task. For each option, a column for the person's score and one for the judge's score. Then the winner, whether the two picked the same winner, cost where visible, speed, and notes.
- **The decision rule, written now,** before any answers exist. When testing something new: Switch when the new option matches or beats what they use today on every everyday task and wins at least one wish-list task; Split when it wins some tasks and loses others, with those tasks listed; Stay in every other case. When finding the lineup: the option with the most wins becomes the default, and each other option takes the kinds of task it won. Matter-of-taste tasks shape the notes on personality and stay out of the verdict. Cost, speed, and how the person likes each tool settle close calls.
- **A short run checklist.** A fresh chat for every option. The same request and material. The same settings for web search, memory, and thinking. When comparing companies, top against top and fast against fast, unless comparing those settings is the point. Names hidden before any scoring. The person scores first. For agent tools, the same model inside each tool where the tool allows it, the finished result is what gets judged, and agent tasks run first.

**Scoring.** The person scores every hidden answer first, rereading the scoring guide before they start and reading every answer to a task before scoring any. Then an AI judge scores the same hidden answers in a fresh chat, with Prompt 6 from the lab. The judge must come from a company that made none of the options. You may act as the judge only when that is true of you. The judge acts as a tough editor. It reads every answer and ranks them first, gives at most one 5 per task, and treats competent work as a 3. An AI judge without these instructions gives top marks to almost everything. The judge scores the finished result. If an answer is a web page, the judge reads only its code and cannot see the page, so the look of it is the person's call. If the person pastes a record of the steps an agent tool took, the judge uses it only for the two points on asking before acting and reporting what it did. Then reveal the names. The person's own scores decide the verdict, and the judge's scores are a second opinion.

**When the person and the judge disagree,** help them work out which of three things happened. Their scoring guide is missing something they care about: add it, and date the change. The judge misread: try a different judge, or have the person score that task from now on. Length or polish swayed them: reread the scoring guide before scoring next time. Suggest checking the judge now and then: at setup, and every few releases, the person scores a few answers themselves and compares. Where they mostly agree on a kind of task, the judge can score that kind alone. The benchmark checks the judge, too.

**Done:** a filled scorecard with both sets of scores, how often the person and the judge picked the same winner, and a lineup or a verdict.

## Phase 6: Offer the next step

Show what you built, explain each part in one line, then offer one of these:

- **Run one now.** The person runs one request in every option and pastes the answers. If you are outside the options, hide the names and take their pick, or, in the scored version, their scores followed by yours.
- **Set up the blind-shuffle script.** Create the folder, one subfolder per request and one file per option, and walk them through `blind-shuffle/anonymize.py`.
- **Set up the Benchmark Runner,** the automated version, for people comfortable pasting a key into a file. It is a small program that sends every task to two to four options and hides the names. It opens a review page in the browser, where the person reads every answer and scores it from 1 to 5. Scores save on every click. Two Reveal buttons follow: one brings in an AI judge as a second opinion, and one reveals with the person's scores only. The judge is set up as a tough editor, and it reads only the code of a web page. The reveal shows a verdict from the person's own scores, then what each option cost and how long it took, as tie-breakers. It needs an API key, a private code that lets a program use an AI model directly, billed to the person's account. The easy route is one OpenRouter key for many models: create an account at openrouter.ai, add a small amount of credit, and create a key. Another model hub, or a key directly from the company whose models they want to compare, also works. Follow the START-HERE file in the download, and tell the person never to paste the key into a chat, an email, or a shared document. Tasks that need web search or a build tool stay by hand. The person also runs agent tools themselves, because one tool cannot drive another.

## Rules

- Nothing private. Every request uses realistic made-up material.
- Saved requests stay word for word between runs. A change becomes a new, dated version.
- In anything you save, name options by company and setting, such as "the top model from company A", so the file still makes sense after products are renamed. Leave out version numbers.
- Always compare two to four options side by side. One option tested alone tells the person nothing, and more than four is too many to judge well.
- Never hide the names for a test in which you are one of the options.
- Never act as the AI judge for a test in which your company made one of the options.
- In the scored version, every scoring guide names an automatic fail.
- Every run ends in a decision, even when the decision is Stay.
