4-step manual AI visibility workflow with steps for goal setting, data collection, assessment, and reporting

Most founders discover their AI visibility problem by accident – someone on the team asks ChatGPT to recommend a tool in the category, the company does not show up and a slightly uncomfortable Slack thread follows. The instinct after that moment is usually to look for a platform that will explain what is happening.

But, before any dashboard makes sense, it helps to have done the work by hand at least once because manual tracking teaches you what the numbers actually mean when a platform eventually hands them to you.

Running a manual AI visibility audit means picking a fixed set of buyer-phrased prompts, running them across three or four AI engines on a regular cadence and logging brand mentions, position and framing in a spreadsheet you can compare over time. It is not a substitute for a tracking platform but it teaches you what a platform’s numbers actually mean once you have one.

Gartner projects that by 2027, 95% of sellers research workflows will begin with AI compared with less than 20% in 2024. That shift alone is reason enough to know where your brand stands before you invest in tracking infrastructure. The problem is that most teams either guess based on a handful of casual prompts or they wait until a platform is in place before looking at all. Both approaches misses the step where you actually learn how to track your ai visibility across models.

A manual audit forces you to sit inside the same interfaces your buyers use and ask the same questions they would ask. It is slower and messier than any dashboard and that friction is exactly what makes it useful the first time around.

Four-step editorial flow diagram showing how to write buyer-focused prompts, run them across ChatGPT, Gemini, Perplexity, and Claude, log the answers in a spreadsheet, and tag each response as positive, neutral, or negative.
A four-step workflow for tracking brand visibility across AI search engines

What a Manual Audit Actually Requires

There is a reason most teams never do this properly – it requires deciding on a fixed set of prompts, running them consistently across engines and recording what comes back in a way that lets you compare results over time. None of that is technically hard but all of it is tedious enough that people either skip it or do it once and never repeat it.

The value of doing it anyway is that it exposes assumptions you did not know you had. Teams often assume their competitors dominate every AI answer in their category and sometimes that is true. But more often, the actual pattern is quite uneven – strong presence in one engine, complete absence in another and a framing in a third that nobody on the team would have predicted. You only find that by looking directly engine by engine, prompt by prompt.

Building the Prompt Set

Start with three or four engines – ChatGPT, Perplexity and Gemini cover most of the buyer behavior worth tracking and Claude is worth adding if your buyers skew toward technical or analytical roles. Running the same prompt across all of them matters more than running many different prompts on one engine because the differences between engines are often the most useful signal in the whole exercise.

Write between 8 and 12 prompts that map to how a real buyer would phrase a question not how a marketer would phrase a keyword. A buyer does not type “best AI visibility platform.” They type something closer to “what tools help track how our brand shows up in ChatGPT answers” or “alternatives to manually checking AI search visibility.” The gap between keyword phrasing and buyer phrasing is where most brands lose visibility without realizing it because they have optimized content for the query nobody actually types.

Cover four categories of intent –  category questions, direct comparison questions, problem-first questions where your category is one of several possible answers and questions specific to your positioning or niche. This mix matters because BrightEdge’s analysis of prompts across ChatGPT, Perplexity and Google’s AI systems found that the engines disagreed on which brands to recommend 62% of the time. If you only test one type of prompt, you are measuring a narrow and potentially misleading slice of your actual visibility.

Read — For a deeper breakdown of how each engine weighs sources differently, the Multi-Engine AI Search Optimization Guide covers the mechanics behind why ChatGPT, Gemini, and Perplexity rarely converge on the same answer.

Logging Mentions in a Spreadsheet

Once the prompts are set, the spreadsheet does the real work.

Keep the structure simple enough that you will actually maintain it. A workable version includes the prompt itself, the engine, the date, whether your brand appeared at all, the position or prominence within the answer, which competitors appeared alongside you, and, a column for the exact language the engine used to describe your brand.

That last column is where most of the useful insight lives.

Position matters less in AI answers than it does in traditional search because these are synthesized responses rather than ranked lists. What matters more is framing. An engine that describes you as a niche tool for a narrow use case is telling you something different than one that describes you as a category leader even if both technically “mention” your brand. Recording the actual phrasing, not just a yes or no on visibility is what turns a spreadsheet from a checklist into something you can actually learn from.

Run this same prompt set on a fixed schedule, ideally every two to four weeks, and keep the historical rows rather than overwriting them. A single snapshot tells you almost nothing because AI answers shift as models update and as the underlying source material online changes. The pattern across several snapshots is where the real signal appears.

Tagging Sentiment and Framing by Hand

This is the step most people skip entirely and it is the one that separates a real audit from a vanity check on whether you got mentioned. For each row where your brand appears, tag the sentiment as positive, neutral or negative and separately tag the framing as accurate, outdated, or misleading.

These two tags are not the same thing. An engine can describe your product accurately but with lukewarm enthusiasm. It can also describe you enthusiastically using outdated positioning from an old landing page or a review site that has not been updated since your last pricing change. Neither of those shows up if you are only tracking whether you appeared. Both directly affect whether a buyer reading that answer would actually consider you.

Doing this by hand also teaches you something a platform’s automated sentiment scoring often misses in early iterations – context.

An answer that says your tool “is more limited than alternatives for enterprise use cases” is not simply negative – it might be accurate for your current stage, in which case the fix is roadmap and positioning, not content. Or, it might be based on an outdated feature set in which case the fix is closing a content gap.

Two-column comparison table outlining five benefits and five limitations of manually tracking AI search visibility, including low starting cost, direct engine comparisons, subjective tagging, limited scalability, and growing time requirements.
Pros and cons of manually tracking brand visibility across AI search engines.

Where the Manual Method Hits Its Wall

The honest part of this piece is that the method above works and it also breaks down faster than most people expect.

Running 8 to 12 prompts across four engines produces roughly forty data points per audit cycle. That is manageable once but running it every two weeks across a growing prompt list while also tracking three or four competitors on the same prompts to have any comparative baseline, turns into well over a 100 rows to log, read, and tag by hand every cycle.

The real constraint is not effort, it is consistency. AI answers are not static and the same prompt run twice in the same week can return a meaningfully different response even without any change on your end.

A 2026 analysis of citation data by BrightEdge found only 11% domain overlap in the sources ChatGPT and Perplexity draw from for comparable queries which means the variance you are seeing across engines is not a measurement error – it is the underlying system behaving differently by design. Capturing that variance manually means running more prompts more often which is exactly where a spreadsheet stops being sustainable.

There is also a ceiling on what manual tagging can reasonably capture.

Sentiment and framing are judgment calls and judgment calls made by one person on a Tuesday afternoon are not always the same judgment calls that same person would make a month later under time pressure. That inconsistency is fine for a first pass but it becomes a real problem the moment you try to use the data to make a defensible claim about whether your visibility is actually improving.

None of this means the manual method was a waste of time – it means it has done its job which was never to become your permanent measurement system. Its job was to teach you what to look for so that when you do reach for something built to track this at scale, you already know exactly what a useful answer looks like.

If you want a faster first read before building out a full manual audit, GeoRankers built a free AI Visibility Score tool that runs a version of this check for you. It will not replace the judgment you build by doing the manual version once but it is a reasonable place to start if you want a number to react to before committing to the spreadsheet.

Read — For a broader framework on how AI search visibility fits into a category-level strategy, the GEO Guide walks through the mental models this audit method is built on.

Frequently Asked Questions

How many prompts do I need for a useful first audit?

Eight to twelve well-constructed prompts run across three or four engines is enough to surface real patterns. More prompts help once you have the process running smoothly, but a smaller, carefully written set beats a long list of near-duplicate questions.

Should I track competitors in the same spreadsheet?

Yes, at least two or three direct competitors on the identical prompt set. Without a comparative baseline, it is hard to tell whether your visibility is actually weak or whether the entire category is underrepresented in AI answers right now.

How often should the audit be repeated?

Every two to four weeks is a reasonable starting cadence. AI answers change as models update and as source material online shifts, so a single audit is a snapshot, not a trend.

Is manual tracking accurate enough to make decisions from?

It is accurate enough to spot direction and obvious gaps, such as complete absence from an engine or consistently outdated framing. It is not precise enough to make fine-grained comparative claims, since the same prompt can return different results run to run and different people will tag sentiment slightly differently.

What should I do with the framing data once I have it?

Route it to whoever owns content and positioning. Outdated framing usually points to a content gap that can be closed directly. Accurate but unfavorable framing usually points to a positioning or product gap that content alone will not fix.

Leave a Reply

Designed with WordPress

Discover more from GeoRankers Blogs

Subscribe now to keep reading and get access to the full archive.

Continue reading