Hypothesis

Does query type influence LLM citation sources? 

The starting assumption was straightforward: if you know what kind of question someone is asking, you can probably predict what kind of source the AI will pull from. A 'best of' question should pull from listicles and review sites. A 'who is' question should pull from Wikipedia and business databases. A 'latest news' question should pull from media and official sources.

If that pattern held up in the data, it would mean something concrete. Brands could reverse-engineer which source types they need to be present on, per query category, rather than trying to be everywhere at once.

Research Design

How we conducted this experiment

240 prompts were run through GPT-4o Mini in the UK, split evenly across 8 query types: Definitional, Comparison, Best-of, How-to, Opinion, Troubleshooting, Recency, and Entity. All prompts were within fintech topic space: open banking, Pay by Bank, instant payments, crypto on/off-ramp, and related fintech concepts.

Every URL citation returned in a response was extracted and classified by source type: Industry/Fintech, Wikipedia, News/Media, Official/Regulatory, Business Database, Product/Blog, or Other. The goal was to see whether citation patterns were predictable or effectively random.

Parameter Value
Total prompts 240
Query types 8 (30 prompts each)
Model GPT-4o Mini
Region United Kingdom
Total citations extracted 418
Query types with citations 4 of 8
Query types with zero citations 4 of 8

Outcome

The results

The first surprise: half the query types produced no citations at all

Definitional, Comparison, How-to, and Troubleshooting prompts returned zero URL citations across all 30 prompts each. The model answered entirely from its training data, no external sources referenced, no domains cited.

This matters for two reasons. First, it means there is no source competition happening in those categories, so there is nothing to optimise for through domain placement. Second, it immediately frames where the real citation game is being played: Best-of, Entity, Recency, and Opinion.

Query Type Citations AI cited external sources?
Recency 167 Yes
Entity 130 Yes
Best-of 90 Yes
Opinion 31 Partially, low volume
Definitional 0 No, answered from memory
Comparison 0 No, answered from memory
How-to 0 No, answered from memory
Troubleshooting 0 No, answered from memory

For the 4 query types that do cite: source type is predictable

Across Best-of, Entity, Recency, and Opinion, clear patterns emerged in which source types the model preferred. The hypothesis held.

Best-of queries: Fintech directories and niche blogs win

Industry/Fintech sites took 39% of citations, followed by Product/Blog at 24%. The standout top domain was yapily.com. Wikipedia made a minor appearance (6%) but was not the primary signal. Mainstream media and forums were absent entirely. For 'best X for Y' queries in fintech, the AI is not pulling from brand pages. It is pulling from aggregators, directories, and specialist comparison content.

Entity queries: Wikipedia and business databases dominate

'Who is X' and 'what is [company]' queries produced the clearest source pattern of all. Wikipedia appeared 21 times, tied with CB Insights. Business databases (Owler, Creditsafe, The Company Check) showed up almost exclusively in this category. If a brand is absent from those platforms, it likely does not exist in the AI's answer for entity queries.

Recency queries: EU regulation dominates the narrative

Recency had the largest citation pool (167), but the composition was striking. consilium.europa.eu and finance.ec.europa.eu were among the most cited domains. Official/regulatory sources accounted for 14% of all citations. When someone asks 'what is the latest in open banking,' the AI's answer is anchored in regulatory updates, not product launches or company news. The recency story is a policy story.

Opinion queries: small pool, strong fintech concentration

Opinion had the fewest citations (31), suggesting the model mostly answers opinion prompts from trained opinions rather than sourced ones. When it did cite, Industry/Fintech led at 48%. Wikipedia appeared at 13% here, higher than in Best-of. Official sources like openbanking.org.uk showed up here, which they did not in Best-of.

Evidence

Supporting evidence

Source Type Best-of Entity Recency Opinion
Industry / Fintech 39% 35% 11% 48%
Wikipedia 6% 16% 4% 13%
News / Media 0 8% 26% 23%
Official / Regulatory 3% 5% 14% 10%
Business Database 10% 12% 1% 0
Product / Blog 24% 2% 0 6%

Business databases appear almost exclusively for Entity. Product/blog content matters for Best-of but is nearly invisible elsewhere. Official/regulatory sources spike for Recency. These are consistent directional signals, not noise.

Domain Best-of Entity Recency Opinion Total
yapily.com 14x 6x 0 8x 28x
en.wikipedia.org 5x 21x 7x 4x 37x
cbinsights.com 0 21x 0 0 21x
openbankingexpo.com 0 4x 8x 0 12x
pymnts.com 0 0 6x 2x 8x
fca.org.uk 0 0 6x 0 6x
getivy.io 0 3x 0 2x 5x

Conclusion

Wrapping things up

The hypothesis was confirmed. Source citation is not random. It is strongly shaped by query intent. The AI draws from different source ecosystems depending on what kind of question it is answering, and those patterns are consistent enough to be useful.

The more counterintuitive finding is the zero-citation result. Four of the eight most common query types, including Definitional and How-to, which are high-intent and high-volume, produce no citations at all. That does not mean content in those categories is wasted. It means the goal there is not to be cited. It is to shape what the model was trained on.

For the four query types that do cite, the actionable read is this: figure out which source type wins each category in your space, and check whether your brand is present on those platforms.

  • Entity queries: Get on Wikipedia and CB Insights. Business database presence is disproportionately important for 'who is' recognition.
  • Best-of queries: The AI is citing aggregators and fintech directories, not brand pages. That is where to place content.
  • Recency queries: Regulatory coverage and policy commentary is what gets cited. Brand news alone will not move the needle here.

Zero-citation types: Shift the goal from citation to training influence. The model already has an answer; the question is whose framing it is drawing from.

Citations are not random and query type matters. The AI draws from different source ecosystems depending on what kind of question it is answering, and those patterns are consistent enough to be useful.
Study Source
Study Source
Study Source
Study Source
Study Source
Study Source
Study Source
Study Source
Study Source
Study Source
Study Source
Study Source
Study Source
Study Source
Study Source
Study Source
Study Source
Study Source
Study Source
Study Source
Study Source
Study Source
Study Source
Study Source
Study Source