All research

When does AI search the web? When buyers say your name.

Ask an AI assistant about a software product by name and it usually checks the web first: 78% of the time in our study, against 7% for the same question asked in general. The rest of the time it answers from memory, and memory favors famous names.

Research7 min read

A stack of dark blocks on one side and a single teal block inside a ring on the other. Illustration of an assistant choosing between its memory and a web search.

At a glance

Key findings

  1. Naming a product makes the assistant check the web. It searched before answering 78% of the time when the question named a product, and 7% of the time when the same question was asked in general.
  2. Famous products are answered from memory, even by name. The assistant searched for 53% of named questions about well-known products and 99% about lesser-known ones.
  3. General questions run on memory, and memory favors famous names. Answering from memory, it brought up the product a question was built from in 64% of answers when that product was well known, and in 31% when it was lesser known.

Two ways to ask the same question

There are two ways to ask an AI assistant about software. You can name the product, or you can describe what you need. To a buyer they feel like the same question. We wanted to know whether the assistant treats them the same way: does it look things up on the web, or does it answer from what it learned in training?

The difference matters because the two sources can disagree. The web shows what a product does today. The assistant's memory reflects what was written about a product before the model was trained, which can be months or years old, and which covers famous products far more than small ones.

A well-known product, asked by name

Does Trello have visual boards for tracking project progress?

Answered from memory, no web search

Yes, absolutely. Visual boards are Trello's core feature. Trello is designed around the Kanban framework, making it simple to track project progress at a glance.

A lesser-known product, asked by name

Does Runn have capacity versus demand forecasting?

Searched the web before answering

Yes, Runn has capacity versus demand forecasting. It is one of the core features of the platform.

Naming a product makes the assistant check the web

When the question named a product, the assistant searched the web before answering 78% of the time. When the same question was asked in general, it searched 7% of the time.

How often the assistant searched the web

Share of answers where it ran at least one web search, with 95% intervals.

95% interval

The gap is 70 percentage points, and the 95% interval runs from 58% to 82%, so this is not a close call. In everyday terms: ask about a product by name and you will usually get an answer checked against the web. Ask in general and you will usually get an answer from memory.

The less famous the product, the more it checks

The assistant did not treat every name the same way. For well-known products it searched about half the time, 53%. It seemed confident it already knew them: asked by name about Trello, Asana, Wrike or Microsoft Project, it never searched in any of its tries. For lesser-known products it searched almost every time, 99%.

Search rate by how well known the product is

Share of answers where the assistant searched, for named and general questions, with 95% intervals.

95% interval

Naming a lesser-known product raised the search rate 48 percentage points more than naming a famous one, with a 95% interval from 23% to 71%. That is sensible behavior: when the assistant is unsure, it checks. The consequence is easy to miss. For a famous product, the answer often comes from memory even when you name it.

Ask in general, and memory decides who is named

General questions are where many buying journeys start, because early on a buyer does not know the names yet. The assistant answered more than nine in ten of them from memory.

Every general question was built from one real product whose own description states the capability asked about. So for each answer given from memory we checked one thing: did the assistant bring up the product the question came from?

Who the assistant remembered

General questions answered without searching: share of answers that brought up the product the question was built from.

It brought up the product in 64% of answers when the product was well known, and in 31% when it was lesser known. Other products it named may fit just as well. The point is who gets remembered, and the pattern was close to all or nothing: the assistant either brought a product up almost every time or never did.

Something we did not expect

A few general questions did make the assistant search, and most of them asked about capabilities that sound new: AI agents that sort customer feedback, test cases written automatically, creating tasks from WhatsApp. The assistant seems to check the web when a topic feels recent, whether or not a product is named. We did not set out to test this, so we treat it as a lead for the next round rather than a finding.

What it means for software vendors

When buyers already know your name, the assistant looks you up, and what it finds shapes the answer. Clear pages that say plainly what your product does are what it reads.

When buyers do not know your name yet, the assistant answers from memory, and a page you publish today does not change that answer directly. What counts is how often, and in how many places, your product has been written about, because that is what the model learned from. For a lesser-known product this second problem is the harder one, and it is where the buyer's search begins.

What it means for buyers

If you have a product in mind, ask about it by name: you are much more likely to get an answer checked against the web. Treat the list from a general question as a starting point drawn from memory. It leans toward the names the assistant already knows. Juukbox matches your requirements against its whole catalog, so a product can make your list without being famous.

Limitations

This is a pilot on one assistant. The results come from Google's Gemini and 50 question pairs; the full study repeats the test on other leading assistants with more questions. We used the developer version of the assistant, not the chat app people use every day, and the apps route questions their own way, so their search rates may differ.

All products come from project and portfolio management software, and each question used one wording. Popularity is measured by review and analyst presence, a stand-in for how much a model has read about a product. The recall measure checks for one product per question, so it describes who is remembered, not whether an answer was right.

Method

We wrote down our hypotheses before running any questions: that naming a product raises the search rate, and that the effect is larger for lesser-known products. The question set was fixed before the run and not changed after it.

We drew 50 products at random, with a fixed seed, from the 120 in the Juukbox project management catalog that have both a description and a popularity score, in three groups by score: 17 well known, 17 in between and 16 lesser known. For each product we took one capability its own catalog description states and asked about it two ways, with the same wording except for the name.

Every question went to gemini-3.8-flash through Google's developer API with the Google Search tool available and never forced. Each was asked 5 times in random order, named and general questions mixed together, for 500 answers. No call failed. An answer counts as searched when the assistant ran at least one web search for it. For 94% of questions, all five answers made the same choice about searching.

Intervals are 95% bootstrap intervals over the 50 question pairs, because the answers to one question are not independent of each other. The question set, every raw response and the analysis script are kept, and every figure here can be regenerated from them.

References

  1. Semrush (2026). ChatGPT used web search on 34.5% of queries in February 2026, down from 46% in late 2024.
  2. Nectiv (2025). About 31% of more than 8,500 prompts across 9 industries triggered a web search.
  3. Graphite (2026). "Do not force search in prompt tracking." In "best X" recommendation questions, forcing search shifted measured visibility by about 20 points.
  4. Kandpal, N. et al. (2023). Large Language Models Struggle to Learn Long-Tail Knowledge. ICML. arxiv.org/abs/2211.08411
  5. Mallen, A. et al. (2023). When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories. ACL. arxiv.org/abs/2212.10511
Try it yourself

Start your own search.

Describe what you need and Juukbox will help you compare products, with the reasons behind every match.

by OpenDecision
Voice
No paid placements or sponsored results. Product visibility based solely on your priorities.