Skip to content
AI Visibility Research

Relevance Beats Reach

What gets a YouTube video cited in Google AI Overviews? We rebuilt the ranking problem to find out.

Nick Clark

Nick Clark, Research Engineer, Gumshoe · September 2026

YouTube citations analyzed
1.1M
Videos in ranking tasks
212K
Cited video in top 10
71%

Key Findings

YouTube is one of the most cited domains in AI search, yet there is no rigorous account of why particular videos get picked. Most prior work compares metadata such as duration, engagement, and channel size across cited and uncited videos. Those descriptive comparisons cannot tell you what the AI system is actually responding to.

We took a different approach. Instead of describing cited videos, we tried to replicate the ranking decision behind Google AI Overviews source selection: given a prompt, pick the cited video out of a lineup of relevant alternatives that were not cited. Inspecting the model that does this well tells us which signals matter.

1

Views, likes, and comments barely separate cited from uncited videos. Against comparable alternatives, a cited video has only a 52% chance of having more views or likes. That is close to a coin flip.

2

Relevance to the prompt is the strongest single signal. A model using only semantic similarity between the prompt and the video's title and description outperforms a model using every piece of video and channel metadata.

3

Relevance plus a few metadata signals predicts citations well. Combined, they place the cited video in the top 10 of roughly 54 candidates 71% of the time on held-out data, versus 20% for random ordering. Longer videos, fuller descriptions, and established channels add signal on top of relevance.

The YouTube Citation Dataset

Gumshoe's citation data includes 1,105,198 YouTube citations across 395,775 unique videos, an average of 2.78 citations per video. This study focuses on Google AI Overviews.

Methodology

The goal is to identify which features of a YouTube video are associated with its inclusion in AI search results. To do that, we compare videos cited by AI Overviews against topically similar videos that were not cited for the same prompt. Two design choices shape that comparison.

Selecting cited videos

Our citation data is skewed towards the brands and verticals our users analyze, so a random sample of cited videos would be imbalanced. To correct for this, we kept only videos cited across at least three distinct prompts, assigned each video to the brand it appeared with most often, and sampled in seeded round-robin passes across brands until we reached 5,000 videos. The result spans 2,476 brands, with no brand contributing more than four videos (median of two).

Finding comparable videos that were not cited

For each cited video, we picked one of its AI Overviews prompts and generated two sets of four YouTube searches. The first set used only the prompt, to find other videos that could plausibly answer it. The second used only the cited video's title and description, to find videos similar to the one that was chosen.

For a prompt like "What are the best trail running shoes" with a cited review of a specific Nike trail shoe, the first set might search "trail running shoe reviews" and the second "Nike trail running shoes." Together they produce a diverse lineup of candidates the AI system could have cited but did not.

We excluded videos published on or after the prompt's run date, and videos AI Overviews had already cited for that exact prompt by that date. A video cited for the same prompt only later still counts as not cited, since using that future information would make the task artificially easy. Every video was then enriched with around 20 metadata fields, including views, likes, comments, publish date, duration, channel subscribers, and lifetime channel views.

Cited videos
5,000
Unique search queries
39,810
Uncited alternatives
207,494
Distinct channels
98,661
Median video age
1.74 years
Median duration
6.2 min

The Usual Metrics Barely Move the Needle

Before training anything, we compared the two groups on the characteristics most often cited in prior YouTube analyses: duration, popularity, engagement, and channel size.

Median Cited Not cited for the prompt
Duration (minutes) 9.4 6.2
Views 5,746 5,012
Likes 77 58
Comments 8 6
Channel subscribers 30,700 10,900
Histogram of video duration for cited videos and videos not cited for the prompt
Duration distribution. Dashed lines mark the medians: 9.4 minutes for cited videos, 6.2 for alternatives.
Histogram of channel subscriber counts for cited videos and videos not cited for the prompt
Channel subscribers. Cited videos skew toward mid-sized and larger channels, but the distributions overlap heavily.

With samples this large, almost any difference is statistically significant. To judge practical importance, we measured effect size with Cliff's delta and the probability that a randomly chosen cited video scores higher than a randomly chosen alternative.

Metric Cliff's delta Magnitude P(cited > not cited)
Duration +0.190 Small 59.5%
Views +0.033 Negligible 51.6%
Likes +0.053 Negligible 52.7%
Comments +0.043 Negligible 52.2%
Channel subscribers +0.193 Small 59.6%

Cited videos tend to be somewhat longer and come from channels with more subscribers, but the effects are small. Views, likes, and comments are negligible. None of these differences is large enough to explain citation behavior, which is why comparing cited videos against the general population can mislead: traits that look common among cited videos may simply reflect the kinds of videos that are relevant to those prompts in the first place.

Training a Ranking Model

Each ranking task contains one cited video and every deduplicated alternative returned by its searches, about 54 videos per task. The model's job is to rank the cited video above the rest. We split the 5,000 tasks 80/10/10 into training, validation, and test sets, keeping each task intact within one split. The test set was held back until all development was finished.

We used a CatBoost ranker, a gradient-boosting model that handles categorical features and grouped ranking objectives. The aim was not to find the best possible architecture but to test whether the features carry real signal, and to see how the model uses them.

Model one: metadata only

Using only video and channel metadata, the ranker put the cited video first in 7.6% of validation tasks and in the top 10 in 48.6%, against a random baseline of 2.0% and 19.7%. The median cited video landed at rank 11.

Bar chart of metadata feature importance: duration, description length, comments, lifetime channel views, tags, channel videos, title length, channel subscribers
Metadata feature importance, measured as the increase in ranking loss when the feature is removed.
SHAP dependence plots for eight metadata features
Direction of each feature's effect. Values above zero push a video up the ranking; values below zero push it down.

The model favors videos longer than about 10 minutes, longer descriptions, fewer comments, and channels with more uploads and subscribers. These relationships are not causal. Making a video longer or padding its description would not necessarily earn a citation. More likely, these features act as proxies for video types that tend to be relevant to the prompt, like explainers and reviews from established channels.

If metadata is mostly a stand-in for relevance, a model that measures relevance directly should do at least as well.

Model two: semantic similarity only

For each prompt and video, we computed four similarity scores: word overlap, phrase overlap, TF-IDF similarity, and cosine similarity between OpenAI text-embedding-3-large embeddings of the prompt and the video's title and description.

The semantic-only ranker beat the metadata-only ranker on every validation metric! It placed the cited video in the top 10 for 59.6% of tasks, versus 48.6%, and moved the median cited rank from 11 to 8. Embedding similarity did most of the work.

Bar chart of semantic feature importance: OpenAI embedding 0.061, bag of words 0.012, TF-IDF 0.002, word bigrams and trigrams 0.002
Semantic feature importance. Embedding similarity dominates; word-level overlap adds a little more.
SHAP dependence plot showing that higher prompt-video cosine similarity steadily pushes videos up the ranking
The closer a video's title and description are to the prompt, the higher it ranks. The effect turns positive around a cosine similarity of 0.48 and keeps climbing.

Model three: combined

Finally, we combined metadata and semantic features in a single ranker trained on the same splits. It substantially outperformed either feature family alone, ranking the cited video first in 20.4% of validation tasks and in the top 10 in 75.4%, with a median rank of 5.

Validation Hit@1 Hit@5 Hit@10 MRR NDCG@10 Median rank
Random baseline 2.0% 9.8% 19.7% 0.088 0.089 n/a
Metadata only 7.6% 31.2% 48.6% 0.205 0.251 11
Semantic only 9.6% 36.2% 59.6% 0.238 0.303 8
Combined 20.4% 53.8% 75.4% 0.363 0.444 5
Bar chart of combined model feature importance: OpenAI embedding 0.079, duration 0.020, channel subscribers 0.008, views 0.007, description length 0.006, tags, channel videos, likes
Combined model feature importance. Relevance leads by a wide margin, with duration and channel size as the next strongest signals.

Embedding similarity is the strongest feature by roughly four times. Duration, channel subscribers, and description length still add signal beyond relevance, which suggests that metadata and semantic fit capture different parts of how AI Overviews chooses what to cite.

Held-Out Test Results

With development complete, we scored the three frozen models once on the 500 held-out test tasks, which played no role in feature selection, training, or model comparison.

Test Hit@1 Hit@5 Hit@10 MRR NDCG@10 Median rank
Metadata only 8.8% 30.4% 46.0% 0.210 0.248 12
Semantic only 8.6% 37.4% 58.2% 0.236 0.299 8
Combined 18.0% 52.6% 71.0% 0.343 0.416 5

The ordering holds. The combined model ranks the cited video first in 18.0% of test tasks and in the top 10 in 71.0%, with a median rank of 5 out of about 54. Semantic features beat metadata on every metric except Hit@1, where metadata leads by 0.2 points. The combined model's NDCG@10 dips from 0.444 on validation to 0.416 on test, a modest and expected decline on unseen data.

It's worth reiterating the quality of our ranking proxy. Our combined model achieves Hit@1 of 18%, while you would expect less than 2% by chance!

What This Means for Brands

Chasing views and engagement is a weak strategy for AI visibility on YouTube. Among videos that could plausibly answer a prompt, the ones AI Overviews cites are the ones whose titles and descriptions most closely match what the buyer asked. Reach adds a little on top of that.

In practice, that points to a few priorities:

  • Build videos around the specific questions your buyers ask AI, and use that language in the title and description.
  • Write full descriptions that explain what the video covers, rather than a line and a list of links.
  • Favor substantive explainers and reviews over short clips. Cited videos run about 50% longer at the median.
  • Publish from an established brand channel where possible, since channel size adds signal beyond relevance.

These findings are correlational. The model shows what AI Overviews tends to favor, not what will cause a citation. The next step is to test the ranker directly by publishing videos that score highly on our own channel and tracking whether they get cited.

Find out which prompts your videos should answer. See how your brand appears across AI search engines, broken down by model, persona, and topic, along with the sources each model cites.

Ready to see your own data?

Run a persona-level AI visibility report across ChatGPT, Gemini, Claude, Perplexity, and more.

Free to start ยท No credit card required