A hypothetical tracker reports that your brand appears in 55% of AI answers, which sounds healthy. Then you split the prompts by what buyers ask, and the "best tools in my category" group sits at 13%. One average hid the number that mattered.
This guide shows small-site owners how to build a tagged set of 30-50 short prompts, score it by tag, and label each gap with a problem type and a next action. You can run it by hand in a spreadsheet or load it into any tracker.
Step 1: Start with business questions, not prompts
The Ahrefs guide on AI visibility workflows argues that AI visibility is a collection of different problems that look alike from a distance. Your brand may be absent, losing recommendations to competitors, tied to the wrong topics, described with outdated facts, or hard for assistants to reach at all. Each problem needs a different fix. (Ahrefs)
Decide what you want to learn first, then give every prompt one tag for the question it answers. The template below covers five groups that suit most small businesses.
| Group | Question it answers | Example prompts (placeholders) | Tag | |---|---|---|---| | Comparisons | When AI compares us with a named competitor, do we hold up? | [us] vs [competitor], [us] or [competitor] | cmp | | Recommendations | When someone asks for the best option in our category, are we named? | best [category] tools, [category] for small teams | rec | | Use cases | Are we tied to the specific jobs and audiences we sell to? | [category] for [audience], [job] software | use | | Product facts | Does AI state our pricing, plans and features correctly? | [us] pricing, [us] free plan | fact | | Reputation | What criticism repeats, and is it accurate? | is [us] worth it, [us] reviews | rep |
The Ahrefs guide uses a similar structure and adds groups for how-to jobs and expansion categories. Add those only if you have a real reason to track them.
Step 2: Collect 30-50 short prompts
Thirty to fifty prompts is a practical suggestion for a small site, not a standard. That range allows six to ten prompts per group and stays small enough to run by hand.
Keep each prompt short. The Ahrefs author aims for about six words or fewer, built around one real need. The goal is a stable set of probes, not a replica of every possible conversation.
Good candidates come from sources you already have:
- Your own organic keywords. They turn existing SEO demand into prompts.
- Competitors' organic and paid keywords. They show which searches rivals rank for or pay for.
- Sales and support questions. The Ahrefs guide calls these especially valuable because they come from real customers and use their wording.
- Reddit, Quora and niche forums. People there describe problems before they translate them into marketing language.
- Keyword research. Search volume hints at which topics matter.
- "Best of" lists and comparison pages in your category.
- AI brainstorming. It helps fill gaps, but treat the output as candidates, not evidence of demand.
Cut duplicates. If two prompts would get the same answer, keep one. Then count prompts per tag, because a group with two prompts cannot tell you much.
Step 3: Define your brand match to avoid false positives
Before you score anything, define what counts as your brand. The Ahrefs guide warns that names that are also common words or have other meanings, such as Square, Stripe or Asana, produce false positives when you match the text alone.
If your name is generic, write down what a correct mention looks like: the right company, the right product, the right domain. Do the same for each competitor you track. This pairs well with does schema markup still help ai search citations? explained.
Step 4: Run the set and record a scoring sheet
Use one row per prompt per run. Run every prompt at least three times in fresh sessions, logged out where possible. Note the date and the assistant. These columns earn their place:
| Column | What to enter | |---|---| | Prompt and tag | The text and one tag | | Run number | 1, 2, 3 | | Brand mentioned | Y or N, after reading the sentence | | Competitors named | List them | | Claim check | OK, wrong, outdated, or n/a | | Topic association | Does the answer place you in the right category? | | Cited sources | URLs, and whether any is your own domain |
Your core metric is mention rate per tag: mentions divided by runs. Calculate it per tag first, and calculate an overall rate last, if at all.
Step 5: Read a worked audit (hypothetical)
This example is invented. "Tallybee" is a fictional invoicing tool for freelance designers, and every number below is illustrative, not a benchmark. Assume 40 prompts, three runs each, 120 answers.
| Tag | Prompts | Runs | Mentions | Mention rate | |---|---|---|---|---| | cmp | 6 | 18 | 12 | 67% | | rec | 8 | 24 | 3 | 13% | | use | 10 | 30 | 12 | 40% | | fact | 8 | 24 | 21 | 88% | | rep | 8 | 24 | 18 | 75% | | All | 40 | 120 | 66 | 55% |
The 55% overall looks fine, but it hides several separate problems. Breaking the tags apart shows four stories:
- rec: Tallybee is nearly absent while competitors fill the answers.
- use: Prompts about freelance designers produce 11 mentions in 12 runs. Prompts about agency retainers produce 1 in 18.
- fact: The mention rate is high, yet 9 of the 21 mentions repeat a free plan that no longer exists.
- rep: Several answers cite one two-year-old forum thread.
The Ahrefs guide describes the same masking effect: a brand can look strong overall while nearly invisible for an important use case. In its own data, mention rates run from 96.2% for the head category "best seo tools" down to 0.0% for one niche. Those are Ahrefs's numbers for its own brand and do not transfer to your business. (Ahrefs)
Label each gap and pick an action
| What you see | Problem type | Next action | |---|---|---| | Not named; competitors are | Absence | Check whether you have a page that directly answers the prompt. Look at which pages get cited instead. | | Named, but a rival is preferred head to head | Competitor win | Read the winning answer for its stated reason, then compare your page on that point. | | Named under the wrong category or audience | Wrong association | Clarify positioning on your homepage, product and about pages. | | Named with a stale price, plan or feature | Outdated claim | Update your own source pages and the third-party pages the answers cite. | | Your pages are never cited, even for facts only you publish | Possible access problem | Check robots.txt, bot blocking, server errors and rendering. |
The access row is an inference. Uncited pages are a symptom, not proof. The Ahrefs guide lists technical checks as a separate workflow step, and the list above is a starting point, not that guide's checklist. This pairs well with go deeper on turn off shopify agentic storefronts? audit settings first.
In the Tallybee audit, rec maps to absence. Agency use cases map to absence plus wrong association. The fact tag maps to an outdated claim, and rep maps to an outdated claim with a source problem. Four tags call for four different jobs.
If wrong facts reach buyers directly, fix the outdated pricing claim first. That order is a judgment call, not a rule.
Also read: context: claude opus 5.5 migration checklist for developers
Step 6: Verify the findings are not noise
AI answers change between runs, and 40 prompts is a small sample. Work through this checklist before acting:
- Repeat every prompt at least three times. Mark each result stable (same outcome every run) or unstable, and act on stable gaps first.
- Require a pattern. One missed prompt is a lead. Several misses in one tag, or the same wrong claim across prompts, is a finding.
- Spot-check sources. Open the cited URLs and confirm each page says what the answer claims. A wrong claim can trace back to a single stale page.
- Read every mention for false positives. A same-name company, a generic word, or a competitor's comparison page that name-drops you does not count.
- Keep conditions constant. Use the same assistant, a similar time window and the same logged-out state, and note any change.
- Log the date. Rerun the identical set later instead of editing prompts, or you cannot compare results.
Limitations to state in every report
A prompt set is a probe, not a measurement of real user behavior. Results vary by assistant, location and run. Small samples exaggerate swings, so a move from 3 to 5 mentions is not a trend.
Nothing here predicts rankings, citations or traffic, and a better mention rate does not guarantee sales. The Ahrefs figures belong to Ahrefs, and the Tallybee figures are invented for teaching.
What to expect next
Your first run is a baseline, nothing more. Expect some prompts to flip between runs and some tags to look worse than the average suggested.
Fix one problem type at a time, rerun the same set after a few weeks, and compare tag by tag. If a tag stays unstable, add prompts to it before you draw conclusions.



