Back to blogApprofondimenti

    Choosing the model for the task, not from the leaderboard

    Leaderboards measure general capability on standardised tests. Thirty real cases from your own workflow tell you in half a day what no benchmark can.

    ZeroFive.AI October 11, 2026 6 min

    In short. Public leaderboards measure general capability on standardised tests, while a company needs to know whether a model solves a specific task on its own documents, in its own language and within its own cost and latency constraints. The gap between those two criteria is why many model choices get redone within months, with the integration cost paid twice.

    The question that comes up most often is which model is best. It is the wrong question, and the useful answer starts elsewhere: which task it has to perform, with what acceptable margin of error, and who notices when it gets it wrong.

    Building an internal evaluation set

    It takes less work than it sounds. Collect thirty to fifty real cases, drawn from the workflow you want to support, each with the answer an experienced professional would consider correct, including the awkward ones where the right answer is that the information isn't there.

    Run that set across every candidate model, with the same instruction and without knowing which model produced which answer during scoring. Half a day of work yields information no leaderboard can give, because it is built on your documents and your edge cases.

    The set stays useful after the choice, because it becomes the reference for checking a supplier update or a version change, both of which can shift behaviour without notice.

    The criteria that actually weigh

    CriterionWhy it matters in practice
    Accuracy on the specific taskThe only measure concerning your work
    Behaviour when information is missingA model that invents costs more than a duller one that stops
    LatencyInside a process with an operator on screen, two seconds change adoption
    Cost at real volumeUnit price says little until multiplied by actual volumes
    Data residency and terms of useDetermines whether the use case is feasible, before it is affordable
    Supplier documentationYou need it for your own documentation, and it has to be asked for upfront
    Version deprecation policySets how much notice you get when the model you use is retired

    The last row is the one companies discover late. A model retired with three months' notice forces the use case validation to be redone, and if that validation wasn't documented you start from scratch.

    Which framework does your company actually need?

    AI Rating measures maturity across the four areas of the model and shows where to start, with priorities and estimated effort.

    Start your AI Rating

    The model is a component, not the project

    In projects that reach production the model matters less than people assume. What matters more is the quality of the data fed to it, the precision of the instructions, the downstream controls and the way the output enters the process.

    That has a practical consequence for the choice: decide quickly, with an honest evaluation set, and spend the time saved on integration. A decent model inside a well-built process beats an excellent one bolted onto an improvised flow.

    The documentation to keep

    Whoever uses a general-purpose model inside their own system remains responsible for that system. The model provider has obligations of its own and must make technical documentation and information on capabilities and limitations available, and that material feeds your own system documentation.

    So keep the model version in use, the internal evaluation results, the system instructions in force and the date they changed. It is the material that lets you explain a decision a year later, when nobody remembers why it was made that way.

    Where to start

    The general criteria between proprietary and open models are in proprietary and open source LLMs, while benchmarks are covered in AI model benchmarks and evaluation.

    Where to run is covered in on premise, cloud or API, and dependency risk in AI vendor lock-in.

    To set up an evaluation on your own real cases you can book a session.

    Want to discuss this for your company?

    30 minutes with us to figure out where to start, or an AI Rating to measure your starting point.

    #model selection#evaluation#LLM#cost#GPAI
    Share

    Keep reading