Test AI Models Against Real Work, Not Vendor Benchmarks

The Best Way to Test New AI Models

As models multiply and switching carries hidden costs, teams need evaluations that reflect actual tasks, constraints, and changing expectations.

3 key takeaways
  1. 1Repeatable evaluations start with realistic tasks, measurable prompts, explicit constraints, and clear rejection criteria.
  2. 2Model choices should account for quality, speed, effort settings, security, cost, and the practical burden of switching.
  3. 3Teams can switch models, split usage across tools, or stay put as capabilities and user expectations evolve.

Don't miss

Nufar explains why the final evaluation should produce a practical choice—switch, split usage, or stay—rather than a simplistic leaderboard ranking.

The brief

With frontier, efficient, and open-weight models arriving quickly, vendor benchmarks offer an incomplete guide to choosing what belongs in real workflows.

Nufar Gaspar proposes a repeatable five-step evaluation process built around standardized tasks, measurable prompts, consistent settings, and evidence from actual work.

The decision is not simply which model scores highest: teams must weigh quality against speed, effort, security, cost, and the friction of switching.

A particularly useful distinction is between switching entirely, splitting usage across tools, and staying with the current setup when gains do not justify disruption.

The framework also has to evolve, because exposure to faster systems changes expectations and makes speed itself part of perceived model quality.

What was said on this episode

20 statements · 10 positive · 4 negative · 1 mixed · 5 neutral

  1. Nathaniel Whittemoreon AI model benchmarksNegative0:29

    Vendor benchmarks provide limited information about a model’s personal relevance.

    “benchmarks don't really tell us much about how it's going to be relevant for us personally”

    Listen at 0:29

  2. Nathaniel Whittemoreon Published AI benchmarksNegative2:37

    Many published AI benchmarks are included in model training datasets.

    “Most of these benchmarks are now in the training datasets”

    Listen at 2:37

  3. Nathaniel Whittemoreon AI model performanceNegative3:18

    Benchmark-leading models can perform worse on a user’s specific task than lower-ranked models.

    “there are going to be times, I guarantee, where the thing that is technically state of the art on some benchmark is worse at the version of the thing that you're doing that's lower on the charts”

    Listen at 3:18

  4. Nathaniel Whittemoreon New AI modelsNeutral3:36

    Testing a new model in one’s own process is necessary to assess its fit.

    “there's kind of no shortcut solution to figuring out how a new model is or isn't going to fit into your process other than using it”

    Listen at 3:36

  5. Nufar Gasparon Blind AI model evaluationPositive6:15

    Blind evaluation should hide model and tool identities to reduce evaluator bias.

    “we'll be advised to hide the names of which models or which tools created each result because we are very biased towards our beloved tools”

    Listen at 6:15

  6. Nufar Gasparon AI model switchingNeutral6:31

    A single benchmark win does not automatically justify changing the current AI setup.

    “the fact that in a specific benchmark run, one tool or one model outperformed your current setup does not necessarily mean that automatically you need to make changes”

    Listen at 6:31

  7. Nufar Gasparon Personal AI benchmarksPositive13:16

    Personal benchmarks should include one or two previously unsuccessful wishlist tasks.

    “I would strongly recommend that you have one or two of those”

    Listen at 13:16

  8. Nufar Gasparon Personal AI benchmark use casesPositive13:34

    A personal AI benchmark should contain roughly four to six use cases.

    “My recommendation would be like four to six type of such use cases”

    Listen at 13:34

  9. Nufar Gasparon AI model candidatesPositive15:11

    Evaluations should compare the existing model with no more than four candidates.

    “I would recommend to run the existing model versus a new one or up to four different candidates”

    Listen at 15:11

  10. Nufar Gasparon AI benchmark executionPositive15:23

    Each benchmark request should run in a fresh chat to avoid context effects.

    “we want to run every request in a fresh chat such that the context window will not send us on a tailspin”

    Listen at 15:23

  11. Nufar Gasparon High-stakes AI benchmarksPositive21:17

    High-stakes benchmarks should repeat each request several times per model.

    “if you're running this personal benchmark on something that is very critical and there is perhaps a monetary implication or other implications to you making changes to how you work, ideally you should run the same request several times per model”

    Listen at 21:17

  12. Nufar Gasparon AI tool switchingPositive26:25

    Users should change AI tools when benchmark results and practical considerations justify it.

    “if all of these tests justify the change, then make the change”

    Listen at 26:25

  13. Nufar Gasparon AI model evaluation criteriaNeutral43:11

    Evaluation criteria and quality standards for AI models change frequently.

    “the considerations of what makes you change your mind or what good looks like, those change quite frequently”

    Listen at 43:11

  14. Nufar Gasparon AI model effort settingsNeutral43:59

    Testing multiple effort settings can reveal whether configuration changes evaluation results.

    “if you want to be extra diligent, you can always run several models with different effort setting and see if it changes the picture”

    Listen at 43:59

  15. Nufar Gasparon Operational AI tasksNeutral45:11

    Operational AI tasks are likely strongly influenced by the surrounding system setup.

    “for operations, I would guess it's going to be most influenced by the overall setup”

    Listen at 45:11

  16. For basic tasks, capable open-weight models may perform indistinguishably from commercial models.

    “for them, you will probably see that any decent model, let's call it Kimia and above or Mistrel and GLM and above, you will not be able to see any noticeable difference and then open source will be as good as it gets”

    Listen at 46:11

  17. Nufar Gasparon Commercial versus open-weight AI modelsMixed46:23

    Commercial models can retain an advantage on sophisticated tasks.

    “For the more aggressive things or more sophisticated stuff, sometimes you will still see the gap between commercial and open source”

    Listen at 46:23

  18. Nufar Gasparon Open-source AI modelsPositive46:35

    Open-source models are currently adequate for most knowledge-work tasks with a good harness.

    “as of now, open source is good enough for most knowledge work tasks, especially if the harness in which it's working is a good one”

    Listen at 46:35

  19. Nufar Gasparon Agentic AI toolsPositive47:16

    Agentic tools require measurable goals rather than vague quality requests.

    “Agentic tools need goals. Saying create a good blog post is not measurable.”

    Listen at 47:16

  20. Nufar Gasparon AI subscription plansNegative48:29

    AI subscriptions may become less generous in the future.

    “a good prediction to make is that it might not stay as generous in the future”

    Listen at 48:29

Statements are attributed to the speaker as said on the episode and reflect their view at the time, not PodLume's. They are not advice.

Listen to the full episode and explore every guest, topic, and moment on PodLume.

Test AI Models Against Real Work, Not Vendor Benchmarks · PodLume