ChatGPT Built Its Own Search Index — What GEO Investigators Just Found

For about two years, one assumption sat quietly at the centre of most ChatGPT SEO advice: ChatGPT doesn't search the web itself. It retrieves through Bing and Google, so if you rank there, you're visible in ChatGPT. Optimise for Bing, the logic went, and you've covered the AI assistant's pipe.

An investigation published this week says that assumption is wrong — and the gap between "indexed in Bing" and "visible in ChatGPT" is wider than anyone thought. If you've been following our ChatGPT SEO guide or the wider how AI changed SEO picture, this is the finding you need to know about, because it quietly rewrites who you're actually optimising for.

Here's what investigators found, what's independently confirmed, and what it means for a real business.

The claim: ChatGPT owns a family of indexes, not one scraped result set

The report comes from Peec AI's GEO research team, the same people who build AI-visibility tracking tools. Their claim is direct: ChatGPT runs its own retrieval indexes — internally they've seen one label, "Labrador" — and it isn't a single index but a family of vertical ones covering general web, PDFs, YouTube, news, arXiv, Wikipedia, local listings, finance, legal, medical, shopping and images.

Their core evidence is a field called result_source that appeared in ChatGPT's server-sent events (the live data stream a browser receives while an answer generates) between May 21 and July 21, 2026. It carried only four values: Labrador (OpenAI's own), plus Bright, Oxylabs and SERP — three names that map to external scraping providers that pull Google/maps/news data. In other words, ChatGPT was telling attentive observers exactly where each result came from, and some of those results came from OpenAI's own index.

An important caveat before we go further: this is a fresh, single-investigator report, days old at the time of writing, and OpenAI has not confirmed the "Labrador" label or the full architecture. Treat the specific names and internal structures as a well-documented investigation rather than vendor-confirmed fact. The direction of travel, though, is backed by enough separate evidence that it's worth acting on.

The evidence that holds up independently

Not everything rests on one vendor's stream-sniffing. Three threads support the core idea that ChatGPT is building — and increasingly using — its own retrieval:

1. Microsoft's own grounding platform is a real, public thing. Microsoft Web IQ is Microsoft's "state-of-the-art grounding service for AI agents," built on twenty years of Bing infrastructure, returning "citation-ready context" with a claimed 164ms p95 latency. Its public page advertises enterprise customers — Nasdaq Boardvantage quotes its engineering director praising Web IQ's fast, secure external-data queries. Microsoft markets Web IQ precisely as the kind of retrieval layer an AI assistant like ChatGPT plugs into. So the "ChatGPT stitches together outside providers" picture isn't far-fetched — that plumbing is commercially real and right next door.

2. ChatGPT visibly fuses multiple retrievers with hybrid ranking. Separate from the Peec report, SEO researcher Metehan Yesilyurt documented finding Reciprocal Rank Fusion (RRF) code in ChatGPT's own dev console — parameters like rrf_alpha: 1 and handling of multiple result types (webpage, grouped_webpages, image_inline, and more). RRF is the standard method for merging results from multiple ranked lists into one (score = 1/(60+rank) summed across lists). You only need RRF if you're blending several retrieval channels together — a lexical/BM25 pass and a vector pass, exactly what an in-house index plus scraped SERPs would produce.

3. It's been a stated strategy since 2023. During the US antitrust trial of Google, OpenAI's head of ChatGPT, Nick Turley, testified that OpenAI began building its own search index in 2023, wanted access to Google's search data, and was refused. He also gave an honest horizon: even with full access to Google's index, OpenAI estimated it would take at least five years to know whether answering 100% of queries from its own index is even achievable. That's not a side project. It's a deliberate, years-long build — and every external-facing signal since (the job postings for "Foundations Search" and "Online Data Systems" teams running indexing at exabyte scale, the RRF code, the Microsoft grounding layer) says the build is continuing.

Why caching changes the game for your content

The piece also claims ChatGPT caches pages rather than always fetching them live — which would explain how it answers fast at scale. Their evidence is behavioural: test sites being crawled aggressively (a one-billion-page test domain had ~6 million pages crawled at roughly 35,000 requests/hour by early September 2026), and ChatGPT's "lockdown"/offline mode still returning cached versions of major SEO publishers' homepages without fetching them live. The scale logic checks out: with 58% of web pages taking longer than 0.8s to reach first byte, live-fetching everything on every query isn't feasible — so a cache (or index) must hold snapshots.

If ChatGPT is building its own index and cache of the open web rather than only mirroring Bing, the practical consequence is blunt: being indexed in Bing is no longer a reliable proxy for being visible in ChatGPT. The two may share providers (Microsoft Web IQ is often the actual routing, not vanilla Bing), but they're different products pulling from different layers. An under-crawled, slow, or poorly-structured page can fall out of ChatGPT's own coverage even while it ranks fine in Bing.

What actually matters if you run a business

Let's translate this into moves, and connect it to the rest of our SEO course:

  • Stop treating Bing visibility as ChatGPT visibility. They overlap less than the old assumption claimed. If your site is being crawled by GPTBot and indexed across your own content's freshness and structure, that feeds ChatGPT's index directly — reading the on-page fundamentals matters again.
  • Make your site crawlable and fast for its own crawler, not just Google's. The same Core Web Vitals that help you rank also decide whether ChatGPT's cache and index keep your pages fresh. Slow pages that never load into the cache simply don't exist in that system.
  • If you sell products, watch ChatGPT's shopping results specifically. The report found live A/B tests (like prefer-index-over-serp-v3, running on 8–20% of traffic) that decide whether ChatGPT uses its own product index or scraped results for shopping. Product structure, schema and fresh stock data feed that index directly.
  • Local? Fix Yelp and TripAdvisor regardless. In ChatGPT's provider mix, local SEO listings reach users partly through licensed partners like Yelp and TripAdvisor over which you have direct control — more than through Google Maps, which ChatGPT pulls without your say-so.
  • Test whether ChatGPT has cached you. Enable ChatGPT's lockdown/no-browse mode and ask it about your own latest page. If it answers from memory/cache even when it can't browse, you're in its stored index. If it can't, your freshness or crawl budget needs work.

This all folds into the GEO discipline we've been teaching in this course: ranking #1 on Google was never the whole game, and being the source an AI engine chooses to cite is becoming the real target. That's exactly why we wrote how to get cited in ChatGPT and laid out the raw ChatGPT-vs-Google traffic numbers — the fundamentals still hold, and this finding is one more reason to take them seriously.

The honest bottom line

This is a genuinely live story. A GEO vendor found technical evidence that ChatGPT runs an in-house index family, caches pages, and blends at least eight outside retrieval providers into its own ranking — and the picture is being tested and rebuilt week to week. Some labels and details will shift as OpenAI iterates or confirms. But the strategic conclusion is stable and actionable now: ChatGPT is not a thin client on Bing, your site needs to be discoverable to ChatGPT's own crawler and cache on its merits, and the businesses that start treating that seriously today will win the citations everyone else is chasing next quarter.

If you'd like a hand making sure your site is actually structured, crawled and cached the way ChatGPT (and Google) reward — our GEO and SEO services cover exactly that. Or browse the whole SEO course to keep building the picture yourself.

← Back to Blog

Comments (0)

No comments yet. Be the first to share your thoughts!

Leave a Comment

Your email address will not be published. Required fields are marked *