English
English

How Perplexity Finds and Cites Its Sources

·

Quint-Ia Vantage


Perplexity finds sources with its own crawler, PerplexityBot, and fetches pages live with a second agent, Perplexity-User, when someone asks a question. Every answer links back to the pages it used. That makes Perplexity the most transparent of the major answer engines (AI tools like ChatGPT, Gemini, and Perplexity that reply with a written answer instead of a list of links), and the one where your existing search work carries over most directly.

Each answer engine retrieves and cites content differently, so a single "AI visibility" number hides most of what matters. Perplexity behaves closer to a traditional search engine than any of the others.

This piece explains how Perplexity's retrieval works, using Perplexity's own crawler documentation and a 15,000-query citation study from Ahrefs, and what it means for your content.

Does Perplexity use Google's or Bing's index?

No. Perplexity runs its own crawler, PerplexityBot, which collects and indexes pages so it can surface and link them in Perplexity's answers [1]. Ahrefs also notes that Perplexity maintains its own search crawler instead of relying on Google or Bing [2].

Perplexity states that PerplexityBot is not used to collect content for training AI foundation models [1]. Its job is search: finding pages that can be shown and linked in an answer.

This has a practical consequence. If your robots.txt file (the file that tells crawlers which pages they may visit) blocks PerplexityBot, your pages are unlikely to appear in Perplexity's search results. Perplexity's documentation recommends allowing it and publishes the bot's IP ranges so you can verify real traffic from it [1].

How does Perplexity cite sources in real time?

When a user asks a question, Perplexity can send a second agent, Perplexity-User, to visit web pages on the spot and include links to them in the answer [1]. The answer draws on pages fetched at the moment the question is asked.

Perplexity-User acts on behalf of a person, so Perplexity says it generally ignores robots.txt rules [1]. Like PerplexityBot, it is not used to collect content for model training.

The two agents split the work:

  • PerplexityBot crawls and indexes pages so they can be surfaced and linked in answers. It follows robots.txt and is not used for model training.

  • Perplexity-User visits pages live when a user asks a question, then links them in the response. It generally ignores robots.txt and is not used for model training.

The result is an engine that cites by design. Ahrefs describes Perplexity as built to back nearly every statement with a source [2], and you can confirm it yourself: ask Perplexity any buyer question and the answer arrives with numbered source links inline.

For your content, that means a page Perplexity can read today can be cited today. You do not wait for the next model training cycle.

Bar chart of citation overlap with Google's top 10: Perplexity 28.6%, Gemini 8.6%, ChatGPT 6.1%

How closely do Perplexity's citations match Google rankings?

Closer than any other AI assistant Ahrefs measured. In a study of 15,000 long-tail queries collected in early July 2025, 28.6% of the URLs Perplexity cited also ranked in Google's top 10 for the same query [2].

The other assistants sat far lower. Across AI assistants, only 12% of citations also ranked in Google's top 10 on average [2].

  • Perplexity: 28.6%

  • Gemini: 8.6%

  • Copilot: 8.2%

  • ChatGPT (in-text citations): 8.0%

  • ChatGPT (reference list): 6.1%

The overlap is notable because Perplexity does not use Google's index. Its own crawler arrives at many of the same pages Google ranks highly. Ahrefs connects this to Perplexity's design: an engine built to cite consistently favors content that already ranks well [2].

The same study found the reverse for Bing. Only 3.3% of Perplexity's citations overlapped with Bing's top 10, the lowest of any assistant measured [2].

What should you do to get cited by Perplexity?

Start with access, then build coverage around the questions your buyers ask. Four steps follow directly from how Perplexity retrieves content.

1. Let PerplexityBot crawl your site. Check robots.txt and your CDN or firewall rules for blocks on PerplexityBot. A page the crawler cannot reach cannot be surfaced in Perplexity's search results [1].

2. Keep your Google rankings healthy. With 28.6% of Perplexity's citations also in Google's top 10 [2], solid SEO is the foundation for Perplexity visibility. If your pages do not rank for anything, fix that first.

3. Cover the variations of each question. Ahrefs explains that AI assistants expand one prompt into several related searches (a process called query fan-out) and merge the results. A page ranking sixth for three variations can beat a page ranking first for only the original wording [2]. Build topic clusters that answer the related questions alongside the main one.

4. Measure Perplexity on its own. Its citation pattern differs sharply from ChatGPT, Gemini, and Copilot, so a blended "AI visibility" score will hide where you stand. Track which of your pages Perplexity cites for your buyers' real questions, and which competitors it cites instead.

Standard SEO reporting does not cover step 4.

What this means for your AI search plan

Perplexity is the answer engine where traditional search work transfers most directly: it crawls independently, retrieves live, and cites openly. That makes it the clearest place to measure whether your content is citable, and a place where an improved page can be cited without waiting for a model to be retrained.

Frequently asked questions

Does Perplexity use Google search results? No. Perplexity runs its own crawler, PerplexityBot, and does not draw from Google's or Bing's index. Its citations still overlap with Google's top 10 results 28.6% of the time, according to a 15,000-query Ahrefs study from July 2025, the highest overlap of any AI assistant measured.

What is PerplexityBot? PerplexityBot is Perplexity's web crawler. It indexes pages so Perplexity can surface and link them in its answers. Perplexity states it is not used to collect content for training AI foundation models, and it follows robots.txt rules. Perplexity publishes its IP ranges so site owners can verify its traffic.

What is Perplexity-User? Perplexity-User is the agent Perplexity sends to visit web pages live when a user asks a question, so it can include links to those pages in the answer. Because a person triggers each visit, Perplexity says it generally ignores robots.txt. It is not used for model training.

Should I block PerplexityBot in robots.txt? Only if you do not want your pages surfaced in Perplexity. Blocking PerplexityBot keeps your pages out of its search index, which removes your main route into its cited answers. Perplexity states the crawler is not used for model training, so blocking it does not change whether your content trains AI models.

How is getting cited by Perplexity different from ChatGPT? Perplexity cites sources on nearly every answer and favors pages that rank well in Google. ChatGPT, Gemini, and Copilot cited Google top-10 pages only 6.1% to 8.6% of the time in the same Ahrefs study. For those three, 80% of citations did not rank in Google at all for the original prompt.

Does content need to be recent to be cited by Perplexity? Perplexity can fetch a page live at the moment of a question, so newly published or updated pages can be cited without waiting for a model to be retrained. The two sources used here do not measure how much Perplexity favors recency, so treat freshness as one factor among several.

See how your brand scores across 4 answer engines.

References

  1. Perplexity. "Perplexity Crawlers." Perplexity documentation. docs.perplexity.ai/docs/resources/perplexity-crawlers

  2. Linehan, L. and Guan, X. "AI Search Overlap." Ahrefs, published August 11, 2025 (data collected July 2025). ahrefs.com/blog/ai-search-overlap