Every AI support tool runs the same models. So why do results differ so much?

Anna Hordiienko
Anna Hordiienko

At some point in every sales call, a founder asks us the same question: “Do you have your own proprietary model?”

The honest answer is no. And the honest answer is the one we like giving, because the question is built on the most successful piece of misdirection in this category.

Under the hood, nearly every AI support tool on the market – ours included, the expensive enterprise ones included, the one your competitor just bought included – runs on the same small handful of frontier models. OpenAI, Anthropic, Google. That’s the list. When a vendor says “proprietary model,” they almost always mean one of two things: a fine-tuned version of a model you could name, or a branded wrapper around several of them. What they rarely mean is a from-scratch model that competes with the frontier labs, because building one of those costs more than the entire AI support category makes in a year.

So if everyone is renting the same engines, why does one tool resolve most of your tickets while another confidently invents a refund policy you don’t have?

That difference – the whole difference – lives in the system built around the model. We call ours the Quidget harness. Here’s what’s in it, and why it’s the part worth interrogating in any demo, from any vendor.

What “proprietary model” actually buys you

Start with what the claim hides, because it hides two things.

First, a fine-tune is a snapshot. A model tuned on support data last year is frozen at last year, while the frontier models it forked from keep improving every few months. Vendors don’t retrain on every release – retraining is expensive and breaks things – so a “proprietary model” spends most of its life competing against a newer, smarter version of its own parent. The pitch says moat. The mechanics say time capsule.

Second, you can’t look inside it. When a black-box model gives a wrong answer, there is no “why.” You can’t see what it knew, what it drew on, or what its training data taught it about refunds. You’re asked to trust the brand instead of the mechanism – which is exactly backwards from how you’d evaluate a human hire.

None of this makes fine-tuning useless (more on that below). It makes it the wrong thing to buy on faith.

The four stages that actually produce a resolution rate

A frontier model out of the box is a brilliant generalist with no knowledge of your business and no sense of its own limits. Ask it your customers’ questions and you get fluent, confident, occasionally fictional answers. That’s a demo.

Turning it into an agent you can leave alone with customers takes four systems. Together, they’re the harness.

1. Retrieval – the model answers from your docs, or not at all. Before the model writes a word, the harness finds the relevant passages from your help center, your docs, your past conversations – and constrains the answer to them. Your actual refund policy, your actual shipping cutoff. The model supplies the language; your content supplies the facts. This single stage kills most of what people call “hallucination,” which was never a model defect so much as a missing leash.

2. Guardrails – scope and tone are set by you, not the model. The agent answers inside boundaries you define: these topics, this voice, these things it may never promise. An off-scope question doesn’t get a creative answer. It gets stage four.

3. Evaluation – quality is measured, not vibed. Answers get checked against your real content and real past conversations – did the agent say what your best human agent would have said? This is the unglamorous stage every vendor skips in the demo, and it’s where resolution rates are actually made. A support agent that isn’t evaluated against real tickets is a support agent whose accuracy is a rumor.

4. The handoff rule – knowing when not to answer. When confidence drops or a question crosses scope, the agent stops and hands the conversation to a human – with the full transcript and context attached, so your customer never repeats themselves. And because a human touched it, that conversation costs you $0 on Quidget. We priced it that way on purpose: an AI that gets paid for every answer has an incentive to answer everything. Ours doesn’t.

There’s a fifth piece for agents that do things – trigger refunds, update orders: approval gates, so a human holds the keys on real actions until you decide otherwise.

That’s the anatomy. Not a secret model. Four inspectable systems, each of which you can ask about, test, and watch working.

Refusing to answer is a feature

That fourth stage deserves a moment, because buyers consistently read it backwards.

A tool that answers 100% of questions is not better than one that answers 80% and escalates the rest. It’s worse – catastrophically worse – because the extra 20% is where the invented refund policies live. The most expensive sentence in AI support is the confident wrong one: it costs you the ticket, the customer’s trust, and the cleanup time, all at once.

So in your next demo, skip the softballs. Ask the agent something it can’t know – an account-specific edge case, a policy you never wrote down. The right behavior is a graceful stop and a handoff. If it improvises instead, you’ve learned what it will do to a real customer at 2am, unsupervised.

This is what “up to 80% auto-resolved” is made of, by the way – and why we hedge it with “up to.” The verified numbers under it: Softorino runs 60% of first-level responses through Quidget, deployed in under two days. JJESIM cut inbox volume 40%. Those aren’t model benchmarks. They’re harness results measured on real ticket queues.

Glass box, black box

Put the two stories side by side and the buying question gets simple.

A black-box pitch asks: trust our secret model. A harness pitch says: inspect the system. On Quidget you can see what the agent read before answering, what it said, and why it handed off when it did. Ask a mystery-model vendor to show you any of the three.

Transparency here isn’t a philosophical preference. It’s operational. When an answer goes wrong – and at some point one will, on any platform – the difference between “let me see what it retrieved” and “we’ve flagged it to the model team” is the difference between fixing your docs by lunch and waiting on a vendor’s roadmap.

The honest trade-offs

Fair is fair: fine-tuned models are a real advantage in narrow situations. If you’re processing millions of near-identical conversations in one domain, a tuned model can cut latency and per-token cost in ways that matter at that scale. If a vendor fine-tunes and shows you their evaluation results, that’s a serious operation. And someday we may fine-tune a small model on support conversations ourselves – if we do, we’ll say exactly what it is and what it’s trained on, because the whole point of this article is that you deserve to know.

But for an SMB doing hundreds or thousands of tickets a month? The model was never your bottleneck. The frontier models are already good enough to resolve most of your queue – what decides whether they actually do is the retrieval, the guardrails, the evals, and the handoff rule wrapped around them. A vendor who leads with a secret model is selling you the one part of the stack they didn’t build.

Three questions cut through any pitch: What model writes the answers? What happens when it’s wrong? Show me what it read before it answered. The vendors worth your money answer all three without flinching.

We built Quidget so the answers take one breath: the best available frontier models, a handoff to your team with full context at $0, and yes – we’ll show you.

See the harness work on your own content: create your bot, point it at your docs, and ask it the hardest question your customers ask you. Then ask it one it shouldn’t answer – that second test is the one that matters.

Share this article