All research

Before ranking a single product, we tested whether AI answers can be trusted.

We spent six weeks and a few hundred dollars of compute checking AI answers about software against an expert. Web search did not make the answers more correct, the AI changed its own answer on 7% of identical reruns, and memory almost never admits what it does not know.

Research6 min read

Clear laboratory jars, each holding a project management product tile, beside an evaluation tablet and test equipment. Illustration of products put under inspection.

At a glance

Key findings

  1. Web search did not make the answers more correct. Checked against an expert, answers with web search agreed 84% of the time and answers from memory alone 83% to 85%, while memory was tens of times cheaper and about twice as fast.
  2. The AI does not fully agree with itself. Asked the identical question again, it gave the same answer 93% of the time, so any accuracy figure for AI answers needs that ceiling beside it.
  3. Memory leans toward yes and almost never says it does not know. When memory disagreed with a web-checked answer it tended to credit a product with a feature, and on obscure vendors it did not once answer that it did not know.

The question we had to answer first

Juukbox tells a buyer which products fit their requirements. That only works if the answer to a simple question is right: does this product do this thing? Before building on AI answers, we needed to know how often they are right, when looking things up on the web helps, and where the answers go wrong.

So we tested it, at a scale that cost real money: more than eight thousand model calls and more than fourteen thousand live web searches in the main comparison alone, across different ways of asking, several models and many prompt designs. We logged the cost of every run.

Web search did not make the answers more correct

We asked the same capability questions with web search switched on and switched off, and checked every answer against an expert's labels. Answers that searched the web agreed with the expert 84% of the time. Answers from memory alone agreed 83% to 85% of the time, depending on how the questions were grouped.

Agreement with an expert

Share of answers that matched the expert's label, with web search and from memory alone.

Memory was tens of times cheaper per answer and about twice as fast. This does not mean search is useless: it found things memory could not, as the next sections show. It means search is not a truth machine. It retrieves pages; it does not judge them better.

The AI does not fully agree with itself

We ran the identical set of questions twice, with web search, and compared the two runs. They gave the same answer 93% of the time. Across our tests that ceiling sat between 91% and 95%.

That matters for anyone who reports on AI answers. If the assistant changes its own answer on 7% of identical reruns, an accuracy claim without a repeatability figure beside it is not a measurement. We now report both.

Memory leans toward yes

When the memory-only answer disagreed with the web-checked one, it tended to disagree in one direction: it credited the product with the feature. And even when the prompt told it to say so, it almost never answered that it did not know. On obscure vendors it did not once.

For a buyer this is the real risk. An assistant answering from memory is more likely to tell you a product has what you need than to admit it is unsure.

What memory cannot see, and what nobody judges well

Search earned its keep on a few kinds of question. It found third-party add-ons that extend a product, which memory missed, and it did better on features that move between plans and on certifications, the facts that change most often.

One kind of answer was hard for every method: "it does this, but only partly". With and without search, the assistant got 23% of those right. Finding the evidence is not the same as weighing it.

Rankings need enough requirements

Ranking products on a single requirement produced mostly ties: with one requirement, 67% of products tied for the top score, so the order was decided by the tie-break rather than by fit. A ranking built on one or two criteria is closer to alphabetical than to a recommendation.

What we changed because of it

We do not rely on one method. We check the web where memory is weakest, we measure how often answers repeat before we trust them, and we treat "not known" as an answer rather than a gap to fill with a guess. And we built Juukbox to ask buyers for enough requirements that a ranking reflects fit.

These tests also raised the questions we are now studying in public, starting with when an AI assistant searches the web at all.

Limitations

The expert labels are thin: 42 answers across a few products, labelled by one person, all in project management software. Most tests used one model family through its developer API, which is not the consumer chat app. Costs are list prices computed from what each run used, not invoices.

Method

Each test asked whether a named product meets a named requirement, taken from real buyer requirements, and recorded a verdict. Arms compared the Google Search tool switched on against memory alone, one question per call against questions grouped together, and several models, prompt designs and reasoning settings. The main comparison ran on gemini-3.7-flash.

Answers were scored against 42 answers labelled by an expert, with partial credit for a near miss, and against an identical rerun of the same run. Every call, verdict and cost was saved, and every figure here can be regenerated from those records.

Try it yourself

Start your own search.

Describe what you need and Juukbox will help you compare products, with the reasons behind every match.

by OpenDecision
Voice
No paid placements or sponsored results. Product visibility based solely on your priorities.