Segonzac: ChatGPT's Retrieval System analyzed - labrador, bright and more
- 4. Aug.
- 2 Min. Lesezeit

Key Takeaways:
Segonzac analyzed the retrieval system of ChatGPT and gained highly interesting insights:
Bright = web search engine but likely just Google
Labrador = an in-house index, topped up with press feeds, open scientific repositories and partner platforms (it is not Bing: as titles are too long)
The labrador snippet: 200 characters taken from the start of your page body (straight after your H1) and the page's title
What to do with this. Your grounding budget in instant mode is your full title plus roughly one hundred and fifty useful characters starting at your H1. Mark up an H1, clear the runway between it and the first paragraph, and keep an eye on the alt text of the first image, the one sitting right after the title.
Cache that is shared by users: a page that is accessed via ChatGPT User bot will be cached for other users
Some URLs are cited from memory
In instant mode, 85 % of citations point to a URL that appears nowhere in the retrieval net, against 31 % in thinking. Some of it comes from widgets — maps, product cards, entity cards — which carry their own links outside the search channel. The rest looks like brand homepages written from the model's own memory. We cannot yet tell the two apart.
The engines used for a certain prompt can not only switch over time but in 1/3 of the prompts switched instantly
5 Verticals in a nutshell:
Web Search: In-house index in free mode, scraped Google in thinking mode (Bright Search)
Images: Your images are re-hosted at OpenAI. They appear in answers without a single request reaching your infrastructure + in thinking mode, thumbnails transit through Bing's image delivery network
Local: Yelp, TripAdvisor, Google Maps. Each provider is recognisable by its id format.
Shopping: Ranking well in ChatGPT's links (web search) does nothing to get you into the product cards: these are two different jobs + Amazon is entirely absent from the offers, including on questions that name it. And the product card is an object cached for weeks, identical for everyone
News: Through the in-house index, ChatGPT receives a summary of about eleven hundred characters, six to eight times longer than the scraped snippets served by the other route
Really great insights that stress the importance of sophisticated tracking: just see a visibility score in answers drop (mentions or even citations) is not sufficient to get valid insights from the constant volatility
Prompt tracking tools need to incorporate a high level of information regarding which subsystems have been triggered how frequently and which modes were used. In Monitoring just looking at the raw answers, mentions and citations only provides close to useless volatility over time. - David Epding






Sources:


