top of page
Wissenswertes
Suche


Metehan & Gabe: How Common Crawl Rank might influence your Authority in AI visibility
Key Takeaways: Metehan analyzed Common Crawl (CC) data and found a correlation to citation frequency in LLM Chats such as ChatGPT, Perplexity and others: Most major LLMs were trained on CC data (64% of models studied, 80%+ of GPT-3 tokens) CC prioritizes high-authority domains in its crawling via Harmonic Centrality These same domains tend to be cited most frequently by LLMs Feel free to use his tool to analyze Common Crawl authority scores for your domain or industry:...
22. Jan.2 Min. Lesezeit


DEJAN: Google's grounding - chunk sizes from Vertex AI Search
Key Takeaways: Dan Petrovic analyzed 7,060 queries via Google's Vertex AI Search that likely powers Gemini Chunk size: The total words retrieved for grounding is surprisingly stable throughout queries at a median of 2.000 words (within 1.500 to 3.000 words) This budget is devided among sources based on their relevance ranking: First position with 28% to 5th position with 13% share. 77% of pages get 200-600 words selected. The typical page gets ~377 words. Longer, more extende
20. Dez. 20251 Min. Lesezeit
bottom of page