Relevance Beats Reach
What gets a YouTube video cited in Google AI Overviews? We rebuilt the ranking problem to find out.
Nick Clark, Research Engineer, Gumshoe · September 2026
- YouTube citations analyzed
- 1.1M
- Videos in ranking tasks
- 212K
- Cited video in top 10
- 71%
Key Findings
YouTube is one of the most cited domains in AI search, yet there is no rigorous account of why particular videos get picked. Most prior work compares metadata such as duration, engagement, and channel size across cited and uncited videos. Those descriptive comparisons cannot tell you what the AI system is actually responding to.
We took a different approach. Instead of describing cited videos, we tried to replicate the ranking decision behind Google AI Overviews source selection: given a prompt, pick the cited video out of a lineup of relevant alternatives that were not cited. Inspecting the model that does this well tells us which signals matter.
Views, likes, and comments barely separate cited from uncited videos. Against comparable alternatives, a cited video has only a 52% chance of having more views or likes. That is close to a coin flip.
Relevance to the prompt is the strongest single signal. A model using only semantic similarity between the prompt and the video's title and description outperforms a model using every piece of video and channel metadata.
Relevance plus a few metadata signals predicts citations well. Combined, they place the cited video in the top 10 of roughly 54 candidates 71% of the time on held-out data, versus 20% for random ordering. Longer videos, fuller descriptions, and established channels add signal on top of relevance.
The YouTube Citation Dataset
Gumshoe's citation data includes 1,105,198 YouTube citations across 395,775 unique videos, an average of 2.78 citations per video. This study focuses on Google AI Overviews.
Methodology
The goal is to identify which features of a YouTube video are associated with its inclusion in AI search results. To do that, we compare videos cited by AI Overviews against topically similar videos that were not cited for the same prompt. Two design choices shape that comparison.
Selecting cited videos
Our citation data is skewed towards the brands and verticals our users analyze, so a random sample of cited videos would be imbalanced. To correct for this, we kept only videos cited across at least three distinct prompts, assigned each video to the brand it appeared with most often, and sampled in seeded round-robin passes across brands until we reached 5,000 videos. The result spans 2,476 brands, with no brand contributing more than four videos (median of two).
Finding comparable videos that were not cited
For each cited video, we picked one of its AI Overviews prompts and generated two sets of four YouTube searches. The first set used only the prompt, to find other videos that could plausibly answer it. The second used only the cited video's title and description, to find videos similar to the one that was chosen.
For a prompt like "What are the best trail running shoes" with a cited review of a specific Nike trail shoe, the first set might search "trail running shoe reviews" and the second "Nike trail running shoes." Together they produce a diverse lineup of candidates the AI system could have cited but did not.
We excluded videos published on or after the prompt's run date, and videos AI Overviews had already cited for that exact prompt by that date. A video cited for the same prompt only later still counts as not cited, since using that future information would make the task artificially easy. Every video was then enriched with around 20 metadata fields, including views, likes, comments, publish date, duration, channel subscribers, and lifetime channel views.
- Cited videos
- 5,000
- Unique search queries
- 39,810
- Uncited alternatives
- 207,494
- Distinct channels
- 98,661
- Median video age
- 1.74 years
- Median duration
- 6.2 min
The Usual Metrics Barely Move the Needle
Before training anything, we compared the two groups on the characteristics most often cited in prior YouTube analyses: duration, popularity, engagement, and channel size.
| Median | Cited | Not cited for the prompt |
|---|---|---|
| Duration (minutes) | 9.4 | 6.2 |
| Views | 5,746 | 5,012 |
| Likes | 77 | 58 |
| Comments | 8 | 6 |
| Channel subscribers | 30,700 | 10,900 |
With samples this large, almost any difference is statistically significant. To judge practical importance, we measured effect size with Cliff's delta and the probability that a randomly chosen cited video scores higher than a randomly chosen alternative.
| Metric | Cliff's delta | Magnitude | P(cited > not cited) |
|---|---|---|---|
| Duration | +0.190 | Small | 59.5% |
| Views | +0.033 | Negligible | 51.6% |
| Likes | +0.053 | Negligible | 52.7% |
| Comments | +0.043 | Negligible | 52.2% |
| Channel subscribers | +0.193 | Small | 59.6% |
Cited videos tend to be somewhat longer and come from channels with more subscribers, but the effects are small. Views, likes, and comments are negligible. None of these differences is large enough to explain citation behavior, which is why comparing cited videos against the general population can mislead: traits that look common among cited videos may simply reflect the kinds of videos that are relevant to those prompts in the first place.
Training a Ranking Model
Each ranking task contains one cited video and every deduplicated alternative returned by its searches, about 54 videos per task. The model's job is to rank the cited video above the rest. We split the 5,000 tasks 80/10/10 into training, validation, and test sets, keeping each task intact within one split. The test set was held back until all development was finished.
We used a CatBoost ranker, a gradient-boosting model that handles categorical features and grouped ranking objectives. The aim was not to find the best possible architecture but to test whether the features carry real signal, and to see how the model uses them.
Model one: metadata only
Using only video and channel metadata, the ranker put the cited video first in 7.6% of validation tasks and in the top 10 in 48.6%, against a random baseline of 2.0% and 19.7%. The median cited video landed at rank 11.
The model favors videos longer than about 10 minutes, longer descriptions, fewer comments, and channels with more uploads and subscribers. These relationships are not causal. Making a video longer or padding its description would not necessarily earn a citation. More likely, these features act as proxies for video types that tend to be relevant to the prompt, like explainers and reviews from established channels.
If metadata is mostly a stand-in for relevance, a model that measures relevance directly should do at least as well.
Model two: semantic similarity only
For each prompt and video, we computed four similarity scores: word overlap, phrase overlap, TF-IDF similarity, and cosine similarity between OpenAI text-embedding-3-large embeddings of the prompt and the video's title and description.
The semantic-only ranker beat the metadata-only ranker on every validation metric! It placed the cited video in the top 10 for 59.6% of tasks, versus 48.6%, and moved the median cited rank from 11 to 8. Embedding similarity did most of the work.
Model three: combined
Finally, we combined metadata and semantic features in a single ranker trained on the same splits. It substantially outperformed either feature family alone, ranking the cited video first in 20.4% of validation tasks and in the top 10 in 75.4%, with a median rank of 5.
| Validation | Hit@1 | Hit@5 | Hit@10 | MRR | NDCG@10 | Median rank |
|---|---|---|---|---|---|---|
| Random baseline | 2.0% | 9.8% | 19.7% | 0.088 | 0.089 | n/a |
| Metadata only | 7.6% | 31.2% | 48.6% | 0.205 | 0.251 | 11 |
| Semantic only | 9.6% | 36.2% | 59.6% | 0.238 | 0.303 | 8 |
| Combined | 20.4% | 53.8% | 75.4% | 0.363 | 0.444 | 5 |
Embedding similarity is the strongest feature by roughly four times. Duration, channel subscribers, and description length still add signal beyond relevance, which suggests that metadata and semantic fit capture different parts of how AI Overviews chooses what to cite.
Held-Out Test Results
With development complete, we scored the three frozen models once on the 500 held-out test tasks, which played no role in feature selection, training, or model comparison.
| Test | Hit@1 | Hit@5 | Hit@10 | MRR | NDCG@10 | Median rank |
|---|---|---|---|---|---|---|
| Metadata only | 8.8% | 30.4% | 46.0% | 0.210 | 0.248 | 12 |
| Semantic only | 8.6% | 37.4% | 58.2% | 0.236 | 0.299 | 8 |
| Combined | 18.0% | 52.6% | 71.0% | 0.343 | 0.416 | 5 |
The ordering holds. The combined model ranks the cited video first in 18.0% of test tasks and in the top 10 in 71.0%, with a median rank of 5 out of about 54. Semantic features beat metadata on every metric except Hit@1, where metadata leads by 0.2 points. The combined model's NDCG@10 dips from 0.444 on validation to 0.416 on test, a modest and expected decline on unseen data.
It's worth reiterating the quality of our ranking proxy. Our combined model achieves Hit@1 of 18%, while you would expect less than 2% by chance!
What This Means for Brands
Chasing views and engagement is a weak strategy for AI visibility on YouTube. Among videos that could plausibly answer a prompt, the ones AI Overviews cites are the ones whose titles and descriptions most closely match what the buyer asked. Reach adds a little on top of that.
In practice, that points to a few priorities:
- Build videos around the specific questions your buyers ask AI, and use that language in the title and description.
- Write full descriptions that explain what the video covers, rather than a line and a list of links.
- Favor substantive explainers and reviews over short clips. Cited videos run about 50% longer at the median.
- Publish from an established brand channel where possible, since channel size adds signal beyond relevance.
These findings are correlational. The model shows what AI Overviews tends to favor, not what will cause a citation. The next step is to test the ranker directly by publishing videos that score highly on our own channel and tracking whether they get cited.
Find out which prompts your videos should answer. See how your brand appears across AI search engines, broken down by model, persona, and topic, along with the sources each model cites.
Ready to see your own data?
Run a persona-level AI visibility report across ChatGPT, Gemini, Claude, Perplexity, and more.
Free to start ยท No credit card required