You test a new AI model by building a personal benchmark from four to six of your real tasks, then scoring anonymized outputs side by side.
NLW, host of The AI Daily Brief, made the case in an October 7 episode. That episode was built from a recorded operator webinar with Nufar Gaspar, who walked through the method she uses on her own work. The central argument is simple to state and easy to skip. Published benchmarks cannot tell you how a model handles your customer emails, your estimates, or your weekly schedule, because those tests measure different work than yours.
A small business owner can answer that question only by running the model on tasks that matter. The owner then compares the results without knowing which tool produced them, and decides deliberately whether anything should change. This post explains that method, the side points the episode added, and how a small team in Northeast Indiana can apply it without a data science staff.
Key Takeaways
- Published benchmarks increasingly sit inside the data that models learn from, so high scores say less about your work than they once did.
- A useful personal benchmark uses four to six real tasks, including one wish-list task that AI has handled poorly before.
- Compare your current setup against one new candidate, or at most four options, and always keep your current baseline in the group.
- Hide model names, score side by side, and run important requests more than once before trusting a result.
- A better score supports a switch but does not require one, because cost, plan limits, company policy, speed, and habit all carry weight.
Why do published AI benchmarks fall short for your business?
Published benchmarks fall short because most of them have become part of the material that models learn from, which means a high score tells you less than it once did. NLW put the point plainly in the episode: labs are not necessarily lying, but benchmarks get absorbed into the body of examples that goes into training sets, and the longer a test circulates, the less it reveals about what a model can do today.
Even perfect benchmarks would leave a gap, because model choice now depends as much on fit and feel as on rank. NLW warned that the model at the top of a chart can underperform a lower-ranked model on your particular work. He said this is especially likely with writing, where quality is subjective and depends on your voice, your customers, and your standards.
That leaves the decision with the person doing the work. The most useful evidence about a model comes from your own tasks, and collecting that evidence takes a deliberate process rather than a quick glance at the launch announcement. Launch week is a poor source of evidence in any case. NLW described a familiar arc: an announcement, a benchmark table, then about a week of demonstrations and hot takes.
Real signal tends to arrive later, once people use the model on genuine work. One example from the episode was a model that quietly dropped a critical condition from a contract review, a failure that a flashy demonstration would never reveal.
Which tasks belong in your personal benchmark?
Your benchmark should include four to six tasks that you already perform, that you perform regularly, and that you would trust a model's output to influence. Nufar recommended that range, and NLW agreed that this selection matters more than any other step in the process. Diversity matters, so pick tasks drawn from different parts of your work rather than six variations of one email. Frequency matters too, because the largest gains usually come from tasks you repeat every week.
High stakes matter as well, since you should test only tasks where a good result would change a real decision. The episode also suggested including at least one task where you know what a costly mistake looks like, because a bad output is often easier to recognize than an excellent one.
Two more filters help sharpen the list. Ask whether the task starts from a blank page or from an existing document, because polishing a draft and generating something new reveal different strengths in a model. Then add at least one task that you have tried with AI and been unhappy with. NLW called this a wish-list task, and it is among the most informative signals you can collect. When a new model unlocks something you could not previously accomplish, that task usually shows the change first.
Nufar also used a discovery step that is worth copying. She ran a structured prompt inside an AI tool with a deep memory of her work. After roughly eight questions, the tool proposed six use cases, wrote a benchmark prompt for each one, and drafted a scoring rubric for the judge.
She warned against trusting that output unread. If the proposed tasks look generic, rewrite them until they reflect how you actually work. A benchmark built on generic tasks measures generic ability, and generic ability will not help you decide anything about your business.
How many models should you compare, and how do you run them?
Compare your current setup against one new candidate, or against no more than four options in total, and always keep your current baseline in the group. Nufar tested seven models across six tasks and called the work "extremely tedious," which is a fair warning about how quickly a large field wears down your judgment.
A smaller field also makes the decision clearer. If the new option beats your baseline, you can see exactly what you gained. If it does not, you can see that switching would cost effort without a clear payoff.
Run every request in a fresh chat. A long conversation carries accumulated context that can distort the result, and a clean start keeps each answer comparable to the others. Send the identical prompt to each model, with the same wording and the same attached material, so that any difference in output reflects the model rather than the setup. Where possible, use similar settings for each model, though NLW and Nufar both noted that effort settings do not translate perfectly between vendors.
If you want to be thorough, run several settings and see whether the ranking changes. Many small teams will not need that level of rigor, but they should record the settings they used so the comparison remains honest later.
Choose how you will run the process: by hand, with a script, or through a platform. The episode described three routes. The manual route works well for one person and a handful of tasks. The script route, which Nufar shared, sends one request to many models through a single API key, collects the outputs, and serves a local scoring page.
Platforms such as LangSmith, Braintrust, and Langfuse suit teams that evaluate models on a regular schedule and need monitoring built in. For a single owner, the manual or script route is more than enough.
Why should you hide which model produced each answer?
Hiding model names keeps your preferences from steering the result, because people tend to favor brands they already trust. Nufar admitted she is biased toward tools she uses every day, and she expected GPT to win the LinkedIn task. When the reveal came, it did not. Kimi, Fable, and Opus topped that task, and GPT did not appear at the top.
Blind scoring makes that kind of surprise possible, and surprises are where useful learning happens. Anonymizing is simple. You can ask a colleague to relabel the outputs, use a spreadsheet random number to assign letters, or let a script shuffle and label them for you automatically.
After scoring, reveal the names and compare the results against your original guesses. The episode treated this as the enjoyable part of the process, and it also has a practical use. If you consistently guess wrong about which tool wrote the strongest output, your assumptions about your own stack need updating. Write down a one-line reason for each score as well. Those notes make the next benchmark easier to build, and they capture the vague preferences that numbers often miss entirely.
Keep sensitive information out of the process altogether. The episode warned that sending company data through an aggregator can expose that data to parties you never intended to involve. If the work involves client records, financial details, or patient information, use mock data and mock use cases instead. The benchmark still shows how a model behaves on that kind of task, and nothing confidential leaves your control while you run it.
Should an AI judge score the outputs for you?
An AI judge can help, but it should supplement your scores rather than replace them. Nufar used Gemini Pro as the judge because none of her candidate models came from Google. Models often favor outputs from their own family, so the judge should come from outside the group being tested whenever you can arrange it.
Even with an outside judge, disagreement can be large. Nufar's judge and her own scores agreed on only one task, the pricing pushback, and everywhere else they diverged. The judge's top pick also cost far more than the option she would have chosen herself.
One structural reason for that divergence is that an API-based judge may see only the code or text behind an output, not the rendered result a person would see. Her website test made this clear, because the judge evaluated the underlying code rather than the page a visitor would actually encounter. For visual work, you need to look at the output yourself.
Test the judge too, because its value depends on whether its rankings track your taste over several tasks. If they do not, the guidance is straightforward: discard the judge. Nufar concluded that she would trust her own scores over the judge's.
Side-by-side comparison also works better than scoring each output in isolation. Ranking several answers from most to least liked is considerably easier than assigning each one an absolute number, so put the candidates next to each other whenever you can. Give the top and bottom picks a clear reason, and let your taste do most of the work. Preference often includes a factor you cannot fully explain, and that unexplained part is still useful information about what you value.
Why run the same request more than once?
Run important requests more than once, because the same model can produce very different answers to an identical prompt. Nufar generated a website she liked on the first attempt, but a second run of the same request came out much worse, with something getting stuck in the middle. If the decision carries real money or real customer impact, repeat each request several times per model and compare the spread of results rather than a single output.
That step reveals reliability, which one lucky output can easily hide. A model that is excellent once and unreliable the next time may still be the wrong choice for a daily task.
What else was in the October 7 episode?
The episode was mostly a process walkthrough, but several side points matter for planning. Speed emerged as a criterion that is growing in importance. Sam Altman said he did not appreciate how much speed mattered until he had access to OpenAI's ultra-fast mode, which is currently offered on a $500 a month plan.
Nufar felt the same effect after watching the demonstration, and waiting about 20 minutes for a single website generation suddenly felt unbearable to her. Your criteria will change as the tools improve, so a benchmark should be rerun when they do.
Cost came up as well, and the numbers were revealing. The table summarizes what the episode reported about four of the models Nufar tested.
| Model | What the episode reported |
|---|---|
| Fable | Topped the judge's average, but its total cost sat far above the others |
| GPT Sol | Tied with Fable on Nufar's own scores at roughly one tenth the cost |
| Kimi | Won on both cost and duration |
| Groq | Inexpensive, but extremely slow on average |
The episode also warned that generous flat-rate subscriptions can hide the true token cost. Per-task price comparisons may therefore miss the full story, especially if you rarely hit usage limits.
Open-weight models came up too. For email drafting and basic research, the guests saw no noticeable difference between strong open models and commercial frontier models. That held especially inside a well-built harness. The gap appeared mainly on the most demanding work.
What does an AI teammate do with a benchmark like this?
An AI teammate can run the repeatable parts of a benchmark, but a person should make the final decision. The most useful setup keeps the workflow under draft-and-approve rules. The teammate prepares the same prompt for each candidate, collects the outputs, removes labels where needed, and presents the results for scoring.
Nothing is sent to a customer, and nothing about the business configuration changes until you approve it. That separation matters because the benchmark is fundamentally about judgment, and judgment is the part a small business owner should keep for themselves.
Hey Button works this way. Hey Button runs on OpenClaw, an open automation framework, with models like Claude and GPT answering through it. Our post on how model routing works for an AI teammate explains the separation. A model switch does not change the teammate's instructions, the business facts it was configured with, or the access boundaries you set. A fallback model can take over when the first choice is slow or unavailable, and the boundaries stay the same.
For a sense of what a teammate handles day to day, read what an AI teammate can actually do for a small business today. Access levels are read-only, draft-and-approve, or full access, and you choose the level that fits each kind of task.
Consider an illustrative example. A landscaping office receives a message asking for a spring cleanup quote. A teammate could organize the request, flag missing details such as the address and photos, and draft a reply for approval. A benchmark would test which candidate model writes the cleanest draft for that exact message, using a fresh chat for each candidate and blind scoring. Nothing reaches the customer until a person approves it.
A front desk at a Fort Wayne clinic could run the same kind of test on appointment-request replies. It should use mock patient details rather than real records, so no sensitive information reaches an outside model hub.
When should you stay with your current tool?
You should stay with your current tool in three situations.
The new result may not clearly beat your baseline. Switching may cost more than the gain. Company policy or plan limits may block the change. The episode described the decision as switch, split, or stay. A better benchmark score does not automatically require a change, because habits are expensive and NLW noted that a switch needs a significant boost to justify the effort. Terms of service, data policies, speed, price, and features you enjoy all count in the calculation too.
If you work at a company that limits you to one tool, the benchmark still matters.
For a Fort Wayne or Auburn shop juggling one subscription, mock-data testing before buying a second license matters more than chasing every frontier release. The guests said you can compare fast and thinking modes or other settings within the tool you already have. You can also run a benchmark on mock data to show decision-makers what a second license would buy. Either way, keep a written record of the results, the criteria, and the date, because a benchmark is most useful when you can rerun it later and see what has changed.
Here is a practical plan. Choose four tasks this week, write the prompts, and run your current tool against one new candidate. Score the outputs blind, reveal the names, and write one decision line. Repeat the process when a major release changes the picture. If you want that routine built around an AI teammate that works under your approval rules, start the Hey Button questionnaire. Tell us which task you would hand off first.
Setup and ongoing support are quoted around your workflows and integrations, and any software subscriptions or usage charges are identified before you commit.
Sources
- The AI Daily Brief, "The Best Way to Test New AI Models," October 7, 2026 (episode page, show notes, and transcript): https://aidailybrief.ai/e/2026-10-07
- Allowed figures from the episode: about 11 days average between frontier releases (NLW); seven models tested and a recommendation of up to four candidates; four to six benchmark tasks; about eight discovery questions; a $500 a month ultra-fast tier (as relayed by NLW); about 20 minutes for one website generation (Nufar Gaspar).
Frequently Asked Questions
- How often do new frontier models arrive, and how often should I rerun a benchmark?
- NLW said the average gap between frontier model releases was about 11 days when the episode aired. Given that pace, the guests suggested rerunning a saved benchmark when a release could change your work, rather than on a fixed weekly schedule.
- How many wish-list tasks should a benchmark include?
- The episode says to include at least one wish-list task inside your set of four to six. Pick work you most want AI to handle but have given up on. If a new release improves on it, that is the clearest sign a switch may be worth testing.
- Should my benchmark include a visual task?
- Include one only if you regularly use AI for visual work, such as website design, app design, or image generation. Visual outputs are quicker to compare side by side. Store them in a dedicated folder so you can open each one and compare them properly before scoring.
- What should I save with each benchmark run?
- Keep the prompt, the date, the settings, each anonymized output, your score, and your one-line reason for it. Those records let you rerun the same benchmark later and see what has changed. They also give a manager a clear record of what you tested and why you decided.
- What is the fastest way to start if I only have an hour?
- Write four tasks from your real work and paste each prompt into your current tool, saving each output. Ask a colleague to relabel the results, score them without names, and reveal the names afterward. Write one decision line. Rerun the process when a major release changes the picture.
Hey Button
Want a teammate that handles this for you?
Answer a few questions about your business and Hey Button will show you what an AI teammate would take off your plate, with you approving anything that goes out.




