# Gumshoe: Platform Reference

Gumshoe is an AI search visibility platform. It measures how large language models describe a brand when buyers ask for recommendations, and gives teams the diagnostics and content tooling to change those answers.

This file is a factual reference about the platform. Last reviewed September 2026. Figures reflect the product as shipped on that date.

---

## Measurement and metrics

### How does Gumshoe calculate Brand Visibility, and what does the number mean?

> **Summary:** Brand Visibility is the percentage of a report's AI-generated answers that mention the brand. It is a share-of-answers measure, not a search-volume or traffic measure.

A Gumshoe report runs a set of prompts against a set of models, once each, as a set of buyer personas. Every prompt-model pair is one conversation. Brand Visibility is the share of those conversations whose answer mentions the brand.

The same calculation runs at every level of the report, which is what makes the number decomposable: overall Brand Visibility, then Topic Visibility, Persona Visibility, and Model Visibility, each the same percentage restricted to one slice. A brand at 62% overall that sits at 90% on Gemini Flash and 20% on Google AI Overviews has a channel problem, not a content problem, and the model breakdown is what shows it.

Alongside visibility, Gumshoe records mentions, which is the raw count of times a brand is named, and rank, which is the position a model gives a brand inside a ranked answer. Visibility and rank move independently. A brand can appear in 90% of answers and sit sixth in most of them, which means the models know it exists but do not lead with it.

### What is the Leaderboard and how does competitive rank work?

> **Summary:** The Leaderboard ranks the report's brands by visibility and mentions across the same question set, giving a like-for-like share-of-voice comparison. A report tracks up to 30 competitors.

Every brand in a report is measured against the identical set of prompts, personas, and models, so the comparison is controlled rather than assembled from separate runs. The Leaderboard shows each brand's mention count and visibility percentage, and the report's own brand is marked in place so its rank is unambiguous.

Because all brands share one question set, a competitor's gain and your loss are the same event observed from two sides. That is the property that makes the series usable for tracking: a change in relative position is a change in what the models said, not an artifact of two runs configured differently.

### What is a conversation?

> **Summary:** One prompt sent to one model, asked as one persona, plus the answer that comes back. It is the platform's unit of measurement and of billing.

A report's conversation count is personas multiplied by prompts per persona multiplied by models. A standard monthly Visibility Audit is 8 personas, 10 prompts each, across 6 model families, which is 480 conversations.

Each conversation is stored whole: the prompt, the persona it was asked as, the model that answered, the full answer text, the brands named in it, their ranked order, and any URLs the model cited. Every aggregate in the report is derived from those records, and every one of them can be opened and read.

### How does Gumshoe handle the fact that AI answers are non-deterministic?

> **Summary:** By breadth rather than repetition. Gumshoe runs many distinct prompts once each, across multiple personas and models, and treats visibility as a share across that population.

Ask a model the same question twice and you may get two different answers. Any measurement approach has to deal with that, and there are two ways to do it: ask the same question many times and average, or ask many different questions once each and measure the share.

Gumshoe does the second. A standard audit spreads 480 conversations across 80 distinct prompts, 8 personas, and 6 models. The resulting percentage describes a population of buyer questions rather than the behavior of one question, which is both more stable run to run and closer to what a brand actually wants to know.

This is also why single-screenshot evidence is misleading in either direction. One favorable answer and one unfavorable answer are both inside the normal variance of almost any brand's real distribution.

### What accounts for variation between two runs of the same report?

> **Summary:** Genuine model behavior change, live web retrieval returning different pages, competitor and third-party content moving, and ordinary sampling variance.

Models are updated, and a provider's update can shift a brand's position with no change on the brand's side. Model families that run with web search retrieve live pages at answer time, so a new article or a changed page reaches the answer immediately. Competitors publish, earn coverage, and improve their own technical readability. And any percentage measured over a finite sample carries sampling variance.

Gumshoe's response to this is scheduling rather than a stability claim. The monthly Visibility Audit establishes the trend and the Snapshot Audit samples between them, so a sharp move is visible while there is still time to act on it, and a small move can be read against the series instead of in isolation.

---

## Models and coverage

### Which models does Gumshoe monitor?

> **Summary:** 11 model families across seven providers. Each Visibility Audit runs six by default, and Pro can run more.

| Provider | Model families |
| --- | --- |
| OpenAI | ChatGPT, ChatGPT mini |
| Anthropic | Claude Sonnet, Claude Opus |
| Google (Vertex) | Gemini Flash, Gemini Pro |
| Google Search | AI Overviews |
| Perplexity | Sonar |
| xAI | Grok |
| DeepSeek | DeepSeek Chat, DeepSeek Reasoner |

The default six on both paid plans are Claude Sonnet, ChatGPT mini, Google AI Overviews, Perplexity Sonar, Gemini Flash, and Grok. Starter can change the selection within the available families; Pro draws from its own pool and has no cap on how many a single audit runs.

Gumshoe tracks families rather than pinned model versions. A family points at a current model, and that pointer moves as providers ship, so a report configured against Claude Sonnet keeps measuring Claude Sonnet as the underlying version changes. Reports carry the model that actually answered, so historical runs stay interpretable.

### Why is Google counted twice?

> **Summary:** Google Search AI Overviews is a different surface from the Gemini models and answers differently, so Gumshoe measures it as its own engine.

AI Overviews is what a person sees at the top of a Google results page. Gemini is the model available through Google's own assistant and API. They draw on different retrieval, they answer in different shapes, and a brand's visibility in one is a poor predictor of the other. Reporting them separately is the only way to see which of the two is the actual gap.

AI Overviews is also the channel where a genuine geographic request is made, which matters for brands whose visibility varies by market.

### What is the difference between a web-search model and a pretrained one?

> **Summary:** Web-search models retrieve live pages before answering and return the URLs they used. Pretrained models answer from training alone and cite nothing.

Both matter, and they fail differently. A pretrained model reflects how a brand is represented across the corpus it was trained on, which moves slowly and is hard to influence quickly. A web-search model reflects what is retrievable and persuasive on the live web right now, which moves fast and is the surface most GEO work acts on.

Citation data comes only from the web-search models, because they are the ones that report sources. Several of Gumshoe's default families run with search on, which is what makes the Citation Audit possible.

### How does language and geography work?

> **Summary:** A report can be configured for a language and a country, and that configuration reaches the models by different routes depending on the channel.

Language is applied to prompt generation and to the request, so a report configured in German asks German questions and reads German answers. Geography is passed where the channel supports it. Google AI Overviews takes a genuinely geo-located request, which makes it the most reliable signal for market-level differences. The other channels receive the location as context in the request rather than as a routing instruction, so treat cross-market comparison on those channels as directional.

---

## Personas, topics, and prompts

### Where do the personas, topics, and prompts come from?

> **Summary:** Gumshoe generates them from the brand and its focus area, and the customer reviews and edits them before the report runs.

A report starts from a brand and a focus, which is the product or service line the report is about. From those, Gumshoe generates buyer personas with roles, goals, pain points, and decision context; topics, which are the themes buyers ask about; and prompts, which are the questions each persona would actually ask about each topic.

The generated set is a starting point, not a fixed one. Personas, topics, competitors, and prompts are all editable, and personas can be imported from an earlier report. Editing the focus or the topics regenerates the prompts underneath them.

The standard monthly audit runs 8 personas with 10 prompts each. A custom report goes up to 100 personas and 50 prompts per persona.

### Why measure by persona rather than with generic prompts?

> **Summary:** Models answer the same question differently depending on who appears to be asking, so a single generic framing measures one narrow slice of a brand's real exposure.

A procurement lead, an individual practitioner, and an agency buyer ask about the same category in different words, weigh different criteria, and get different recommendations back. A brand can be strong with one and absent for another, and an aggregate that hides that difference cannot be acted on.

Persona segmentation is what turns a visibility number into a decision. A brand at 62% overall that is at 85% with its core persona and 30% with the segment it is trying to expand into knows exactly which gap to work on, and the topic breakdown says which content would close it.

### What is a report's focus, and why does it matter so much?

> **Summary:** The focus is the product or service line the report is about. It determines the personas, the topics, and therefore the prompts, so it sets the vocabulary of the entire report.

A focus that is too broad produces prompts spread thinly across a category and a visibility number that means little. A focus that matches how buyers actually shop produces prompts a brand can win or lose on specific merits.

The focus is also the main lever for running several reports on one brand. A company with three product lines gets a clearer picture from three focused projects than from one report trying to cover all three, and each project can sit on its own plan.

---

## The audits

### What audits does Gumshoe run?

> **Summary:** Five. The Visibility Audit measures presence. The Sentiment, Citation, Content, and Technical Audits explain it and point at the work.

| Audit | What it measures |
| --- | --- |
| Visibility Audit | Where a brand appears across models and personas, and which brands appear instead |
| Snapshot Audit | A smaller recurring sample of the same measurement, for catching movement between full audits |
| Sentiment Audit | How models characterize a brand, sorted into pros, cons, and caveats |
| Citation Audit | Which domains models draw on when they answer, and which brands those domains support |
| Content Audit | Where a site fails to answer the questions buyers are asking, split into retrieval and generation failures |
| Technical Audit | Whether a model can fetch, parse, and extract an answer from a given URL |

### What does the Sentiment Audit produce?

> **Summary:** Themes describing how models talk about a brand and its competitors, grouped into pros, cons, and caveats, at two stages of the buyer journey.

The audit generates a question plan at two stages, Product and Vendor Aware, and Evaluation and Decision, with five questions at each stage. Most of those questions name competitors directly, because a model scores what the question puts in front of it, and comparison questions are what produce comparative sentiment.

Answers are read by a judge that runs several times per answer and extracts themes per subject. Themes are clustered, scored for relevance to the project's focus, and surfaced as pros, cons, and caveats, with a topic scatter and a heatmap for reading them against the rest of the report.

The output is qualitative in shape and quantitative in support. "Models consistently raise implementation time as a caveat" is a content brief in a way that a sentiment score between -1 and 1 is not.

### What does the Citation Audit produce?

> **Summary:** The domains models actually cite in a category, how often, and which brands each domain's coverage supports.

Web-search models return the URLs they consulted. The Citation Audit aggregates those across the report into domains and categories, separates the brand's own domain from third parties, and shows which competitor benefits from each source.

This is the audit that turns visibility into a digital PR plan. Knowing a brand is invisible is a problem statement. Knowing that four domains account for most of the category's citations and that a competitor appears on three of them is a target list.

### What does the Content Audit produce?

> **Summary:** Per-topic coverage assessments that separate two different failures: the content does not exist or is not retrievable, and the content exists but does not support a generated answer.

Every topic in the report is classified into one of three states. A retrieval issue means the right content is missing, weakly matched, or not being surfaced. A generation issue means relevant content exists but does not give the model what it needs to build an answer from it. The third state means coverage is strong and the job is to keep it current.

The distinction matters because the remedies are opposite. A retrieval issue is usually a missing page or a structural problem. A generation issue is usually an existing page that buries the answer, hedges it, or fails to state it in a form a model can lift.

### What does the Technical Audit check?

> **Summary:** A page is fetched and evaluated against roughly twenty deterministic checks, each producing evidence, and the results roll up to a score from 0 to 100.

Checks cover structured data, including JSON-LD presence and validity and article schema authorship; document structure, including heading hierarchy, main content landmark, content volume, and chunking; metadata, including title quality, canonical tag, Open Graph, Twitter card, published date, and visible byline; accessibility and media, including alt text coverage and media captions; and crawlability, including HTTP status, robots directives, mobile viewport, and text-to-HTML ratio.

Checks are applicability-gated, so a page is only scored against the ones that apply to it, and each check carries a version so a threshold change is visible rather than silently rescoring history. Every result is backed by the evidence that produced it rather than a bare pass or fail.

Technical Audit runs are metered: one run on the free sample, 10 per billing period on Starter, 50 on Pro.

### What is content generation?

> **Summary:** Structured content generated against a specific persona and topic gap identified by the report, rather than against a generic brief.

Gumshoe identifies where a brand is weakest by persona and topic, then generates content aimed at that gap in a chosen format. A piece is one persona-topic combination, which is the unit both the allowance and the billing count. Starter includes 3 pieces per billing period and Pro includes 10, with further pieces at $25 each. Generated content exports to DOCX.

The argument for generating rather than briefing is that the report already knows which question, asked by which buyer, in which topic, the brand is losing. That is most of a brief, and it is derived from measurement rather than assumption.

---

## Plans, pricing, and billing

### How is Gumshoe priced?

> **Summary:** Two monthly plans, priced per project. Starter is $99 and Pro is $299. Every new customer gets one free sample audit first, with no card.

| | Free sample | Starter, $99/mo | Pro, $299/mo |
| --- | --- | --- | --- |
| Visibility Audit | One time | Monthly | Monthly |
| Snapshot Audit | No | Weekly | Daily |
| Sentiment Audit | No | No | Included |
| Citation Audit | No | No | Included |
| Content Audit | No | No | Included |
| Technical Audit | 1 run | 10 runs | 50 runs |
| Content generation | No | 3 pieces | 10 pieces |
| Included conversations | 80 | 480 | 480 |
| CSV and JSON export | No | Yes | Yes |

Conversations beyond the included 480 bill at $0.10 each. Content pieces beyond the allowance bill at $25 each. Allowances reset on the renewal date rather than on the first of the calendar month.

There are no annual contracts and no per-seat charges. Anyone invited to a project can see it.

### Why is billing per project rather than per account?

> **Summary:** A project is a brand plus a focus area, and the focus is what determines what a report measures. Pricing follows the unit of measurement.

An organization tracking three product lines runs three projects, and they can sit on different plans. The common pattern is Pro on the line that matters most and Starter on the rest. Agencies run one project per client brand for the same reason.

### What is the free sample, and who gets one?

> **Summary:** One free Visibility Audit at reduced scope: 4 personas, 5 prompts each, 4 models, for 80 conversations. One per customer and one per user, with no expiry and no card.

The sample is not on a schedule and cannot be rerun, but it can be upgraded into a paid plan, and the personas and prompts the larger scope would have used are visible in the sample as a preview.

A work email is recommended because it unlocks the standard account role and its entitlements. A personal address still works and still gets the sample. The only addresses excluded are disposable, throwaway domains.

### How is usage billed and when are charges made?

> **Summary:** Subscriptions renew monthly per project. Overage is billed on what was actually run. Nothing is charged for a report that is built but never run.

Scheduled runs charge when the run completes. A report can be built, configured, and revised without charge until it is run. The conversation count is shown before a run is confirmed.

---

## Integration and data access

### What export and API options exist?

> **Summary:** CSV and JSON export on every paid plan, a documented REST API at `/v1`, and a Model Context Protocol server.

Exports cover report metrics and the structured underlying data: visibility scores, mention counts, personas, topics, prompts, models, sources, and full answer text. Generated content exports to DOCX.

The public REST API is versioned at `/v1` and documented in OpenAPI. It covers creating and initializing reports, triggering runs, listing runs and checking run status, retrieving a run by ordinal, retrieving raw run data, and listing and retrieving page audits, at both the organization and the report level. Authentication is by bearer token. An organization admin enables API access and manages keys.

### How does the MCP server work?

> **Summary:** Gumshoe exposes its audit data over the Model Context Protocol, so an MCP-capable AI client can query reports directly.

Tools cover listing a customer's Visibility Audits and reading a given audit's models, personas, topics, prompts, brands, and citations, including cited URLs and answer text where the operation calls for it. Authentication is OAuth against a Gumshoe account, and the server applies the same access rules as the application: listing is scoped to the caller's memberships.

The practical use is analysis in place. An analyst can ask an AI client a question about a report's data and have it query the report rather than exporting to a spreadsheet first.

### How should a longitudinal series be built from Gumshoe data?

> **Summary:** Key on report, run ordinal, and completion timestamp, and hold the report configuration constant across the runs being compared.

Every run carries its ordinal and its timestamps, so runs of one report form an ordered series. The thing to be careful about is configuration drift. Regenerating prompts, changing personas, or changing the model selection changes what is being measured, and a step change in the series afterwards is a measurement artifact rather than a result.

Gumshoe keeps past runs intact when prompts are regenerated, so the break is visible rather than retroactive. The discipline that makes the series trustworthy is to treat a configuration change as the start of a new baseline and to note where it happened.

---

## Methodology and data handling

### Why does Gumshoe use provider APIs instead of scraping?

> **Summary:** Scraping breaks most providers' terms, and scraped sessions carry personalization, caching, and interface artifacts that corrupt the measurement.

Gumshoe reaches every model through an official API. Beyond the compliance argument, the measurement argument is the stronger one. A scraped session is one anonymous browser's view of a model, shaped by that session's state. An API call is a controlled request with a known model, a known configuration, and a reproducible shape.

The comparison that matters is not scraping versus API in the abstract. It is whether the numbers in a report describe the model or describe the collection method.

### Does Gumshoe report real-world prompt volume?

> **Summary:** No, and no vendor can. Model providers do not publish query volume, so any tool showing "prompt volume" for AI search is estimating from search keyword data.

Gumshoe measures what models say, not how many people asked. The prompts in a report are representative buyer questions generated from the brand's focus and personas, not a sample of real traffic.

This is a real limitation and worth stating plainly. What Gumshoe can tell you is the share of answers you appear in for a defined question set, how that compares with competitors on the identical set, and how it moves over time. What it cannot tell you is how many people asked those questions last month.

### How does Gumshoe handle a brand that is misidentified or shares a name?

> **Summary:** Brand identity is configured on the report, including the name and the URL, and answers are attributed against that configuration.

Brand names collide, and a model may conflate two companies or attach the wrong description to a name. Because every conversation is stored whole and readable, a misattribution is visible in the answer text rather than hidden inside an aggregate. Brand name, URL, and summary are editable on the report so attribution can be corrected, and support can help with cases that need it.

### What is Gumshoe's data and security posture?

> **Summary:** Data-minimal by design. The inputs are brand names, public URLs, and persona and topic selections. The outputs are model responses to those inputs.

Gumshoe does not ingest customer records, credentials, financial data, or private systems. Personas and prompts are not shared outside the account or used for data harvesting, beyond being sent to model providers to generate that account's reports. Model interactions run over official provider APIs. Internal access is role-based, the database has no public IP and is TLS-only, administrative access runs through an identity-aware proxy, secrets are held in a managed secret store with per-service access, and the database is regionally redundant with point-in-time recovery.

Because the platform works with public information and model responses, many enterprise requirements that attach to sensitive data do not apply. For a vendor review that needs specifics beyond this, contact the team rather than inferring a certification from this summary.

---

## Common questions about interpreting results

### What does 0% visibility mean?

> **Summary:** That the brand was not named in any answer in the report. It is usually a focus problem, a brand-name problem, or a genuine absence, and the three are distinguishable.

Open the conversations. If the answers are about a different category than intended, the focus is wrong. If the answers are about the right category and name a variant of the brand's name, it is an attribution problem. If the answers are about the right category, are recommending real competitors, and the brand is simply not among them, the number is correct and the Citation and Content Audits say why.

A brand with strong traditional search performance and 0% AI visibility is not a contradiction. Models draw on a different mix of sources and reward different properties.

### Why do models cite pages that do not mention the brand they recommend?

> **Summary:** Retrieval and recommendation are separate steps. A model retrieves pages to ground its answer on the category, then produces a recommendation that draws on both the retrieved pages and its training.

A citation is evidence of what the model consulted, not a guarantee that the cited page endorsed the brand. This is why the Citation Audit reports both the domains cited and which brands those domains' coverage supports; the two are related but not identical, and the gap between them is often where the opportunity is.

### How does Gumshoe visibility relate to referral traffic from AI tools?

> **Summary:** They measure different things and routinely disagree. Visibility is share of recommendations. Referral traffic is clicks that survived an answer good enough to make clicking unnecessary.

A brand can be recommended often and receive little traffic, because the answer resolved the question. A brand can receive referral traffic from informational queries while being absent from the recommendation questions that matter commercially. Neither pattern invalidates the other measure.

When the two appear to contradict each other, the useful check is which landing pages the referral traffic reaches and whether those pages correspond to the topics and personas in the report. Traffic arriving on support pages while the report measures evaluation-stage questions is two measurements of two different things, not an error in either.

---

## Reference

- Product: <https://gumshoe.ai>
- Application: <https://app.gumshoe.ai>
- Support: <https://support.gumshoe.ai>
- Terms of Service: <https://app.gumshoe.ai/docs/tos>
- Privacy Policy: <https://gumshoe.ai/privacy>
- Contact: support@gumshoe.ai
