Key takeaways
- Choose on the criteria that change slowly, not on this quarter's benchmark scores.
- Admin controls, data handling and commercial terms outlast any model release.
- If a provider is already bundled into a suite you pay for, that fact deserves real weight.
- Two providers on a shortlist is enough. Four is a symptom of an unmade decision.
Every few months a different provider tops the leaderboards, a wave of comparison articles follows, and somewhere a procurement committee decides to wait for more information. That committee will be waiting indefinitely, because it has adopted a decision rule that cannot converge. Here is a framework that can.
The premise is simple. Rank your selection criteria by how fast they change. Weight the slow-moving ones heavily and the fast-moving ones lightly. Model quality moves fast. Almost everything else you actually care about moves slowly.
Why benchmark comparisons mislead
Public benchmarks measure general capability on standardised tasks. Your business does not run on standardised tasks. It runs on summarising your meeting notes, drafting your kind of client email, checking your kind of document against your kind of requirement.
At the frontier, the leading models are close enough on this sort of work that the difference is usually swamped by other variables: how well the tool is configured, whether staff know what to ask for, and whether the relevant document is actually available to it. Choosing a provider on a two-point benchmark gap while ignoring whether your team will get trained is optimising the wrong variable.
There is a second problem. Benchmark leadership is temporary by construction. Whichever provider is ahead when you sign is unlikely to still be ahead at renewal, and if your decision was based on that lead, you have committed yourself to a switching debate every year.
A better framing: assume all shortlisted providers will be roughly equally capable for your work over the contract term. Then ask which one you would rather administer, argue with about an invoice, and explain to your auditor.
The seven criteria that actually last
1. Admin controls
Can you see every user from one console? Can you enforce settings centrally rather than trusting each person to configure their own? Can you control whether conversations can be shared outside the organisation? Can you deprovision in one action? This is the difference between a governed tool and a pile of subscriptions, and it changes very little between releases.
2. Data handling and retention
At the tier you are buying, is your content excluded from model training by default? How long is history retained, and can you change it? Where is data processed, and does that satisfy the obligations you actually have rather than the ones you vaguely worry about? Get the answers in writing from the contract, not from a marketing page.
3. Identity integration
If you run an identity provider, single sign-on turns offboarding from a checklist item into an automatic consequence. Check which tier includes it. This is the criterion most likely to change your effective price, and it is frequently discovered late.
4. Ecosystem fit
If your organisation lives in one productivity suite, an assistant embedded in that suite has a structural advantage that has nothing to do with model quality: it already knows where your documents are, it inherits your permissions, and staff do not have to change tools to use it. That advantage is durable. Weight it accordingly, but do check that the embedded version is actually good at the jobs your people need, because bundling is not the same as fitness.
5. Commercial terms
Contract length, notice period, whether you can reduce seats mid-term, what happens to your data on exit, and whether pricing is per-seat, consumption-based, or both. These terms determine how expensive it is to change your mind, which matters more in a fast-moving category than in a slow one.
6. Support and accountability
When something breaks or a bill looks wrong, is there a named route to a human, and on what timeframe? For a business without in-house AI expertise, this is worth more than a marginal capability advantage.
7. Model range
Does the provider offer a spread of models at different price points, so you can route routine work to a cheaper tier? A provider with only one expensive model gives you no cost lever at all. This is a capability criterion, but it is about economics rather than intelligence, which makes it more stable than benchmark position.
A scoring approach that ends the debate
Weighted scoring gets a bad reputation because it is often used to justify a decision already made. Used honestly, before opinions harden, it works well. Assign weights before you look at any vendor.
| Criterion | Suggested weight | Why |
|---|---|---|
| Admin controls | 20% | Determines whether you can govern the tool at all. |
| Data handling and retention | 20% | The part your legal, audit or client obligations will test. |
| Ecosystem fit | 15% | Drives adoption more strongly than most people expect. |
| Commercial terms | 15% | Sets the cost of changing your mind later. |
| Identity integration | 10% | Offboarding risk and the real price you pay. |
| Model range and pricing | 10% | Your only structural lever on running cost. |
| Support | 10% | Matters most when you have no in-house expertise. |
Notice that raw model quality does not appear as its own line. That is deliberate. Screen for it as a threshold instead: any provider on the shortlist must be capable enough to do your top five tasks acceptably. Once a provider passes that bar, extra capability is a bonus rather than a differentiator.
The one week selection process
- Day 1. Write down the five tasks you most want AI to help with, in the words the people who do them would use. This document is the whole basis of the evaluation.
- Day 2. Shortlist two providers. Two, not four. Usually one is whatever is bundled into your existing suite and the other is the strongest standalone option.
- Day 3. Check the unglamorous facts: tier that includes SSO, retention settings, training exclusion, contract length, exit terms.
- Day 4. Run the five tasks through both, using real content, judged by the people who normally do them. Not a demo. Their actual work.
- Day 5. Score against the weights you set on day one, write a one page recommendation with the reasoning, and circulate it.
The written recommendation matters as much as the choice. Six months later, when a new model launches and someone asks whether you picked the wrong vendor, the document is what stops the debate reopening from scratch.
Red flag in any trial: if the evaluation is being run by the most enthusiastic person in the building rather than by people with ordinary workloads, you are measuring enthusiasm, not fit.
When more than one provider makes sense
Running two providers is defensible when they serve genuinely different jobs. One assistant for general staff use, plus a second provider for a specific technical workload where its models or tooling are a better match, is a common and sensible arrangement.
What rarely works is two assistants doing the same job for different departments. It splits your admin surface, doubles the training burden, halves your negotiating position, and guarantees that half the organisation is using a tool nobody has configured properly.
Frequently asked questions
Which AI provider is best for business?
There is no permanent answer, and treating it as if there were is what stalls decisions. The durable question is which provider's admin controls, data handling, commercial terms and ecosystem fit suit your organisation, because those are what you live with daily.
Should we run a bake-off?
Run a short trial on your own real tasks, timeboxed to a few days, judged by ordinary users. Avoid open-ended comparisons, which consume months and rarely produce a clearer answer than the first week did.
What if we choose wrong?
Keep the initial contract term short, avoid building deep dependencies on one vendor's proprietary features in the first release, and document why you chose. Switching an assistant is genuinely disruptive only once you have entangled your processes with it.
Does it matter that our staff already prefer one tool?
Yes, more than most evaluation criteria. A tool people already reach for has cleared the adoption hurdle that kills most rollouts. Take it seriously, then check it also clears the governance bar.
End the provider debate in one session
The Clarity Package delivers a provider and model direction with the reasoning written down, so the decision is made and defensible. If you would rather skip straight to the build, the Implementation Package includes the choice and the configuration.