Search quality

Search is what stands between a user and the right agent, so its quality decides whether the whole system feels reliable. This page gives the measured numbers, the conditions they hold under, and the method behind them.

The short answer

When your agents describe different things and the question is clear, Search2o finds the right agent more than 99% of the time. When two agents genuinely cover the same ground, search offers both rather than guessing. When nothing fits, search says so.

What decides accuracy is overlap, not catalogue size

Most people expect accuracy to fall as the catalogue grows. It barely does. What accuracy tracks is how distinct the agents are from one another.

Suppose one agent handles escalations and another handles incident triage. A question such as who signs off on this? belongs to both. No search engine can be certain, because the ambiguity is in the catalogue rather than in the search. The useful thing to do is offer both agents.

We measured this directly. Two catalogues of the same size, 50 agents each, scored 99.2% and 83.8% for the top match. The only difference was how far apart the agents' subjects sat. Fifteen points of accuracy came from catalogue design.

Catalogue design is in your hands, and Describing an agent covers what makes a description distinct.

A well-separated catalogue

50 agents, each from a different industry, and 500 questions.

Right agent is the single top match99.2%
Right agent is among the results shown99.6%
Searches answered with exactly one result95–98%
Confident but wrong single answers1 in 500 searches

Deliberately overlapping catalogues

Our hardest test material: six catalogues, 1,418 agents and 15,180 questions. Several are packed with near-synonym agent pairs on purpose, so that the difficult case is measured rather than the flattering one.

Right agent is the single top match86%
Right agent is among the results shown94.3%
Searches answered with exactly one result58%
Confident but wrong single answers0.77%
Nothing returned when something was available0.14%

Read the last lines with the first. Even on our hardest material, fewer than one search in a hundred gives a confidently wrong single answer.

These pooled figures describe our test set rather than a typical deployment. Per catalogue, the share of searches answered with a single result runs from 52% to 79%. The well-separated catalogues sit at the top of that range and the near-synonym ones at the bottom.

Scale: 40, 100 and 1,000 agents

Search2o handles a thousand agents in one catalogue. Accuracy at 1,000 agents is the same as at 100.

Catalogue sizeQuestionsTop match is correctRight agent is among the results
40 agents40094.0%98.2%
100 agents1,000 and 2,00085.8–86.5%94.3–95.4%
1,000 agents10,00085.8%93.6%

Going from 100 agents to 1,000 cost essentially nothing: 86% for the top match at both sizes.

The 1,000-agent test is the one worth describing in full, because it is the one people doubt.

  • Every one of the 1,000 agents was queried, with ten questions each. This is not a sample of the easy ones.
  • 367 agents answered all ten of their own questions correctly.
  • Not one agent failed all ten. There is no agent in a thousand that cannot be found.
  • The worst agents scored two out of ten, and every one of them was half of a near-synonym pair.

We also measured size on its own, holding subject separation roughly constant. At 50, 100 and 200 agents the top match was correct 83.8%, 83.3% and 79.5% of the time, and the right agent was among the top three 97.2%, 95.7% and 92.7% of the time. A fourfold increase in catalogue size costs a few points. Adding one near-duplicate agent costs more.

One, two or three results, and sometimes none

Search decides how many results to show from how close the candidates are to each other.

  • The best candidate is far ahead: one result, and the agent runs.
  • Two are close: both are offered.
  • Three are close: all three are offered.
  • Nothing is close enough to the question: no results, rather than a wrong answer.

Never more than three. This is what turns ambiguity into a choice instead of an error. On our hardest material the right agent is among the offered results 94% of the time, against 86% if search always insisted on a single answer. What the user sees for each of these cases is covered in How matching behaves.

The balance between the two is a real trade, and it is worth stating plainly. Answering with one result more often means being confidently wrong more often. Measured over 15,180 questions, the range runs from showing a single result for 88% of searches, where 7.2% of answers are confidently wrong, to showing one for 42% of searches, where 0.23% are. Search2o ships the setting that keeps confidently wrong answers below 1%. On a well-separated catalogue the choice stops mattering, because the best candidate is nearly always far ahead of the second.

Questions that nothing covers

A search engine that always answers is easy to embarrass, so we measured this separately. The accuracy tests above only ever asked questions that had an answer.

Against questions with no covering agent at all — general knowledge, other domains, unrelated subjects — 100% are refused in every language tested, and 99.6% across a larger English-only set. The cost is under 0.5% of genuine questions wrongly turned away.

Questions that sit beside the catalogue's territory without being covered by it are handled the other way on purpose. Between a third and two thirds of those are refused, and the rest are answered. Showing a plausible neighbouring agent helps a user; answering a general knowledge question from a payments catalogue does not.

Languages

Tested in Spanish, German, French, Mandarin and Japanese, at the 40-agent size, with 400 questions per language. The agent descriptions were translated, and the search code and settings were identical throughout.

Descriptions and questions in the same language:

LanguageTop match is correctRight agent is within the top three
English94.0%98.2%
Spanish92.8%98.2%
German92.5%98.0%
French92.0%98.0%
Mandarin91.8%98.2%
Japanese91.2%98.2%

The right agent is found just as reliably whatever the language: the top-three figure is flat at 98% in all six. The small differences are in how often search is confident enough to answer with a single result. Part of even that gap is method rather than capability, since the English text is the original and the others are translations of it.

Descriptions in one language and questions in another were tested across ten language pairs, 400 questions each. The top match was correct 88.2% to 90.8% of the time, and the right agent was within the top three 92.0% to 97.5% of the time.

Cross-language search works. A German description answers a French question, and a Mandarin description answers a Japanese one. The cost is about three points against staying in one language. The strongest pair of all was Mandarin and Japanese, at 97.5%: two different scripts with no shared vocabulary, which is understanding of meaning rather than word matching.

The refusal behaviour transfers as well. Unrelated questions were refused 100% of the time in every language, and no genuine question was wrongly refused in any of the five non-English ones.

Speed

Server-side time, catalogues up to about 120 agents70–130 ms
Server-side time, a 1,000-agent catalogueabout 400 ms
End to end from a browser at low load230–260 ms, under 600 ms at the 95th percentile
Throughputscales with provisioned capacity

Latency stays flat up to the capacity limit and then degrades sharply, with no gradual slowdown as a warning. Capacity is therefore provisioned to the expected rate rather than to the average.

Why these numbers are believable

Accuracy claims are cheap, so the method matters more than the figures.

  • Measured against the live system. Every figure comes from the live search service, running the production code with its shipped settings. We previously used a faster offline approximation. It was wrong by 14 points at the 1,000-agent size, so we deleted it and adopted a rule: no accuracy number that was not measured live.
  • More than 20,000 labelled questions, each with a known correct answer, across catalogues from 40 to 1,000 agents.
  • Settings were tuned on two catalogues and judged on four others they had never seen. The tuned settings were then applied unchanged to a catalogue ten times larger, and held up.
  • Ideas that did not survive were dropped. A cleverer rule for deciding when to show a single result looked seven points better on the data it was fitted to, and vanished on data it had not seen. It was not shipped, and several other candidate signals were rejected the same way.
  • The test catalogues are adversarially hard on purpose. The pooled figures are close to a worst case rather than a typical one.

What we do not claim

These are the limits of the numbers above, and a sceptical reader deserves them.

  • The test catalogues and their questions were generated rather than collected from live customers. They were built to be hard, and the questions for the largest catalogue were written without reference to the agent descriptions, so they do not flatter us. They are still not production traffic.
  • The better-than-99% case is a constructed catalogue of 50 agents from 50 different industries. It is a fair model of a focused deployment, and it is not an average of our test material.
  • The non-English results come from translated descriptions in one 40-agent catalogue, rather than from catalogues written natively at every size.
  • English is the strictest language on refusals. It turns away the most borderline questions, and it is the only language that wrongly refuses any genuine ones.
  • Every figure assumes the agent has been indexed and its description accepted. A description that says only how an agent works, rather than what it does for a user, is rejected at publish time and never reaches search.

Our commitment

We have lived and breathed search for the past twenty years. We make a simple commitment: Search2o's search quality will continue to improve over time.