Does Gemini search the web for every answer? What Google documents and a repeatable test shows

Search can be used and should not be assumed, the app and the API behave differently, and here is how to check one answer.

Published by AI Knows Us (Clyra Labs) · Updated 29 September 2026

No, you should not assume that Gemini searches the web for every answer. Google documents a grounding path that uses search, and that is a capability rather than a guarantee that it ran for the answer in front of you. The only evidence that a specific answer used live sources is what that answer displays: a visible search step, and source links that open and support the sentences they sit next to. As of 29 September 2026 we have not run the test below on Gemini, so this page publishes the method in full and no Gemini figure.

This matters commercially for one reason. If an answer about your category never retrieves anything, then publishing a better page cannot change that answer this month, and the work you should be doing is different.

Short answer: Search can be used, but do not assume it is used for every answer

Four statements, each separate, and mixing them is how people end up with the wrong plan.

  • The capability exists. Google's developer documentation describes grounding with search and what a grounded response returns about the sources used.
  • Use is decided per request, not per product and not per account. The same question can be answered from memory once and from fetched pages the next time.
  • Displayed citations are the evidence. Not the model's description of its own process, and not the confidence of the prose.
  • A confident answer with no links is a memory answer, in all likelihood, and that is worth recording rather than arguing with.

The strongest illustration of the last point we own is from a different assistant. Asked 78 buyer questions with no brand named, ChatGPT then confirmed it had run no live web search for any of them, so the whole reading described what the model remembered rather than what it could find. ChatGPT, 27 July 2026, on clawlaw.in. Those 78 answers read as informed market summaries, and the ranking inside them was a ranking of what the model had absorbed: ProVakil on 28 questions, CLAW on 21, Legistify on 18.

Separate the Gemini app from the Gemini API

These are two different products and a test that mixes them produces a number that means nothing. Five differences worth keeping in your notes.

  • Who decides on retrieval. In the consumer app the product decides. Through the API, the developer decides whether a search or grounding tool is available at all.
  • What you can see. The app shows a search step and links in the interface. The API returns structured data about grounding, which is far more auditable and is why a serious test uses it for the controlled half.
  • Personalisation. The app can be shaped by your account, your location, your history and any saved instructions. A plain API call is cleaner, which makes it the better instrument and a worse description of what your buyers actually experience.
  • Model identity. The app shows a product name. The API takes an exact model string. Record whichever you have, verbatim.
  • Cost and rate. App use is free or bundled. API use is billed, including any search tool. Read the current price off Google's own pricing page on the day you run the test and write the date next to it. We deliberately do not restate prices on this page, because an undated price copied into an article is the exact failure our own audit found on other sites.

Run both halves if you can. The API half tells you what is possible. The app half tells you what a buyer sees.

How to inspect a response

Seven checks on a single answer, in the order that takes the least time.

  • Watch for a retrieval step while the answer is being produced, and note whether it lists queries.
  • Count the displayed source links. Zero is a finding, not a missing value.
  • Open every link. A link that does not contain the claim it is attached to does not support it.
  • Look for a date in the answer. An answer that says "as of" and then names a period is telling you something about where it came from.
  • Ask a question that needs today's information. Something only a recently published page could answer. No link on that answer is strong evidence of memory.
  • Repeat in a clean session. A fresh chat with no memory and no earlier turns, because a previous turn can supply the sources for the next answer.
  • Do not ask the model whether it searched, or if you do, record the reply as a comment rather than as data. Asked afterwards how many of the questions it had actually searched for, one assistant withdrew its own earlier statement, saying it could not honestly substantiate the claim that it had run a live search for each question, and in another batch that its claim to have searched all five was not adequately supported. Perplexity, September 2026, on aiknowsus.com.

Repeatable test

The design

  • Twenty questions, ten that plainly need current information and ten that do not. The contrast is the whole point.
  • Three runs each, on three separate days, giving 60 observations per surface.
  • Two surfaces, counted separately: the Gemini app, and the Gemini API with grounding enabled.
  • Clean sessions throughout, no uploads, no pasted URLs, no follow-ups in the same thread.
  • Nothing named. Your brand appears in no prompt, so that the result is not a reaction to your own wording.

What counts as verifiable live-source evidence

One definition, applied the same way every time: the answer displays at least one source link, and at least one of those links, when opened, contains the claim it is attached to. A visible search step with no links does not count. Links that open to a home page do not count. This is stricter than most published tests and it is the only version worth reporting, because the loose version counts decoration.

What gets logged

Date, time, time zone, surface, model name or model string, prompt, full answer, every link, whether a retrieval step was visible, the strict evidence decision, and one line of reason. The figure produced is answers with verifiable live-source evidence over total runs, reported separately for the app and the API, and separately for the current-information questions and the others.

Results, dates, and limitations

Gemini app: not run by us as of 29 September 2026. Gemini API: not run by us as of 29 September 2026. No fraction, no split, no trend. When we run it, the counts, the denominators, the surface split and the dates will be published here.

What we hold, dated, on other assistants:

  • ChatGPT, 27 July 2026, on clawlaw.in. Zero live searches across 78 blind questions, confirmed by the assistant after the run.
  • ChatGPT, 6 August 2026, on clawlaw.in. A named URL on our own domain cited as a source for a vendor due diligence question, fourteen days after the page was published. That is a live-source answer, with a checkable URL in it.
  • Perplexity, September 2026, on aiknowsus.com. Across three separate batches the assistant withdrew its own claim to have searched, saying the claim was not adequately supported.
  • Claude, 17 September 2026, on clawlaw.in. Our pages ranked in the assistant's raw results and shaped what it wrote, and the site still landed as a name inside a list rather than as a linked recommendation. Retrieval happening is not the same as your domain being credited.

Five limitations that apply to anybody's version of this test.

  • It measures the test, not Gemini. Twenty questions on three days describe those 60 answers.
  • Interfaces change what they display, so a drop in visible links may be a product change rather than a behaviour change.
  • Self reports cannot be used, for the reason dated above.
  • Personalisation makes the app half yours alone. Report the API half if you want a number somebody else can reproduce.
  • Nothing here transfers to ChatGPT, Claude, Copilot or AI Overviews, each of which has to be run on its own terms.

Common questions

If Gemini did not search, does my website matter at all?

It still matters, for two reasons. The answers that do retrieve are the ones where a buyer is close to a decision, and those are the ones worth winning. And what the model absorbs over time comes from published material, so a site that says nothing specific gives the next model nothing specific to remember about you.

Why does the same question show links one day and not the next?

Because retrieval is decided per request, and product behaviour changes without announcements. This is why the method asks for three runs on three days and a reported range rather than one figure.

Is grounding through the API the same as what the app does?

Treat them as different until you have measured both. The API is a documented tool with a structured response. The app is a product making its own decisions. Numbers from the two should never be added together.

Can I force a search?

You can ask for current information, name a period, or ask for sources, and that makes retrieval more likely rather than certain. In a measurement you should not do it at all, because a prompt that demands sources measures your prompt rather than the assistant's default behaviour.

Does a search step guarantee my page was read?

No. A search step means queries were run. Your page appears in the answer only if it was retrieved, judged useful, and credited, and we have a dated case of the last step failing on its own: an assistant admitted using two arguments from our pages and dropping the attribution in both places. Claude, 17 September 2026, on clawlaw.in.

Primary sources and change log

Read Google's own material rather than summaries: the Gemini API documentation on grounding and on the search tool, including what a grounded response returns; Google's pricing page for the API, on the day you need a figure; the Gemini app help centre for what the consumer product does; and Google Search Central's documentation on AI features and site owner controls. Date every quotation.

Our observations come from the clawlaw.in programme from July 2026 and the aiknowsus.com audit of September 2026, 24 batches and 72 conversations. The 6 August 2026 citation is checkable from outside because it names a URL. Counts taken over our capture files are not, until we publish them.

Change log. First published 29 September 2026 with the Gemini results empty and marked empty. Disclosure: AI Knows Us sells AI visibility measurement.

What to do first

Ask Gemini two questions today in two clean chats: one that could only be answered from something published this month, and one general question about your category. Note for each whether a search step appeared and how many links came back. That five-minute contrast tells you whether your next month should go into publishing pages or into fixing what the model already believes about you.

See what AI says about you.

The first scan is free and takes about 20 seconds.

Free. No card. We ask 5 real buyer questions on 2 AI apps.