<!-- Generated from content/blog. -->
# Which answer-engine signals should be compared over time?

Learn which AI visibility signals to track across engines, prompts, citations, competitors, traffic, and qualified actions, and how to compare them consistently.

Canonical URL: https://orathis.ai/blog/which-answer-engine-signals-should-be-compared-over-time/
Author: Quinn Bean
Published: 2026-09-14T22:02:17.331Z
Last verified: 2026-09-14T22:02:17.331Z

**A rising mention count can hide weaker citations, stronger competitors, and no business impact.**

Answer engines change their wording, sources, and brand choices from one response to the next. A useful report must separate those changes, preserve the conditions of each test, and connect visibility to outcomes without pretending that one caused the other.

## What should you compare at a glance?

**Compare engine and prompt conditions first; then mentions, prominence, sentiment, citations, source coverage, competitors, traffic, and qualified actions.** Keep each signal as its own history over time, also called a time series. Join them using stable fields such as engine, product surface—the specific app or mode tested—prompt ID, locale, date, and repeat run.

The practical takeaway is simple: do not reduce answer-engine visibility to one score. A single score can tell you that something moved, but not what moved or why.

| Signal group | What to record | What the comparison can reveal |
|---|---|---|
| Test conditions | Engine, model or specific app or mode, prompt ID, exact wording, intent, locale, whether the user was signed in or personalized settings applied, run time, repeat number | Whether two observations are genuinely comparable |
| Brand presence | Mention present, brand named, product named | Changes in recurring visibility |
| Prominence | Original engine-specific position, answer section, order of appearance | Whether the brand became more or less noticeable |
| Sentiment | Positive, neutral, negative, mixed, plus the relevant text | How the answer describes the brand |
| Citations | Cited URL, domain, citation location, accessible or not | Which pages and domains support the answer |
| Source coverage | Source type, unique domains, owned versus third-party sources | Whether visibility rests on a narrow or broad source base |
| Competitors | Competitor mentions, prominence, citations, substitutions | Who appears when the brand does not |
| Demand and traffic | AI referrals, search clicks and impressions, branded search, landing page | Whether visibility is followed by measurable visits or demand |
| Business outcomes | Qualified actions, conversions that AI visibility may have helped even when it was not the final source, pipeline, revenue | Whether visible prompts overlap with valuable journeys |

This comparison set follows the measurement scope Orathis describes on its [service site](https://orathis.ai): more than 1,000 prompts organized around categories, problems, comparisons, use cases, and buyer journeys, with monitoring for mentions, sentiment, citations, source coverage, and answer rankings. That is a first-party description of the service, not independent proof that tracking these fields improves business results.

## Start with a stable unit of comparison

A visibility change means little if the underlying test also changed. “ChatGPT visibility rose” is not a clean comparison when one sample used a product-comparison prompt in the United States and the next used a broad category prompt elsewhere.

Give every prompt a permanent ID. For each run, retain:

- The exact prompt wording
- Prompt intent, such as discovery, comparison, or purchase
- Engine and product surface, meaning the specific app or mode used
- Model or version when it is visible
- Locale and language
- Account state, such as signed in, signed out, or using personalized settings
- Search or browsing availability
- Timestamp and repeat-run number
- Any known change to the test method

Do not overwrite old wording when a prompt changes. Create a new version and keep its relationship to the original. Otherwise, the history will combine two different questions under one label.

This matters because answer variation has several possible sources. The engine may have changed, its search index may have changed, or the phrasing may have triggered a different interpretation. Stable identifiers will not remove that variation. They make it visible.

## Compare engines and surfaces separately

**A citation in one product is not automatically equivalent to a citation in another.** Each system presents answers, links, and publisher data differently, so the original observation must survive any later normalization—the process of translating different systems’ observations into shared reporting categories.

Official documentation shows why:

- [OpenAI](https://help.openai.com/en/articles/12627856-publishers-and-developers-faq) says public web content may be cited and linked in ChatGPT search. Referral URLs include `utm_source=chatgpt.com`, which publishers can measure in analytics.
- [Anthropic](https://support.anthropic.com/en/articles/10684626-enabling-and-using-web-search) says Claude web-search answers include citations that users can follow. The cited documentation does not identify a dedicated publisher referral parameter.
- [Google](https://support.google.com/webmasters/answer/7042828?hl=en) defines clicks, impressions, and position for AI Mode within its established Search measurement rules. A follow-up question counts as a new query.
- [Microsoft](https://learn.microsoft.com/en-us/microsoft-365/copilot/extensibility/plugin-citations) documents automatic citations in synthesized Microsoft 365 Copilot responses, including URL citations for public-web grounding. It also notes that cited information may not always be accessible to the user.
- [Perplexity](https://docs.perplexity.ai/docs/agent-api/tools/web-search) provides structured source records in its Agent API. These are consistent data fields for each source, including a citation index, canonical URL, title, snippet, and dates. Those API records should not be assumed to match its consumer interface or publisher referral data.

Record the literal engine event first: a linked citation, an unlinked source label, a Search impression, a structured API result, or another observable event. A shared field such as “citation present” can support a dashboard, but it should point back to that original event.

There is no documented cross-engine standard for answer rank or citation position in the sources above. If you create a common prominence scale, label it as an internal convention rather than a platform metric.

## Measure mentions and prominence together

A brand can be present yet practically invisible. It might appear in the opening recommendation, sit last in a long list, or show up only in a warning.

Track at least two separate facts:

1. Was the brand or product mentioned?
2. Where and how did it appear?

Keep the engine’s native presentation, then add a simple reporting category if needed. For example:

| Original observation | Internal prominence label |
|---|---|
| Brand is the direct recommendation in the opening answer | Primary |
| Brand appears in a short comparison list | Included |
| Brand appears only in supporting text | Supporting |
| Brand is absent | None |

This label is useful for analysis, but it is not a universal ranking. Save the answer text or an allowed snapshot so a reviewer can inspect the classification later.

Recurring presence matters more than a lucky appearance. Big Human’s [AEO guidance](https://www.bighuman.com/blog/answer-engine-optimization) recommends examining repeated visibility rather than isolated mentions. That is vendor guidance, not a validated measurement standard, but the underlying practice is sound: repeated samples expose volatility that one run conceals.

## Preserve sentiment with the evidence behind it

**A sentiment label without the relevant sentence is hard to audit.** “Negative” could mean a factual limitation, a direct criticism, or a classifier mistake.

Use a small set of labels, such as positive, neutral, negative, and mixed. Store the text that justified the label and, where possible, what the sentiment refers to. Product quality, pricing, suitability, safety, and customer support are different subjects even when they share the same overall label.

Break sentiment down by source type too. Siteimprove warns in its discussion of [third-party signals](https://www.siteimprove.com/blog/third-party-signals-answer-engine-representation) that reviews and user-generated content may reinforce or distort a brand’s representation. This is commercial guidance rather than proof of how every engine weighs sources. It still supports a useful reporting choice: do not merge owned pages, editorial coverage, reviews, forums, and social posts into one source total.

A negative answer backed by one forum thread calls for a different investigation from a negative answer repeated across several independent publications.

## Track citations as URLs, domains, and roles

Citation count alone throws away the part that helps a team act. Keep the cited URL, its domain, the page type, and the claim or answer passage it appears to support.

Useful citation measures include:

- Citation incidence: the share of repeated, comparable tests that cite the brand’s domain
- Unique cited URLs and domains
- Citation persistence across repeat runs
- Citation location within the answer
- Whether the link is accessible
- Owned, independent editorial, review, forum, directory, or other source type
- New and lost citation targets
- Citations supporting competitors but not the brand

Use one counting rule for every reporting period. A comparable run is a completed test that uses the approved engine, surface, prompt version, locale, account state, and sampling method for that comparison. Count each run once:

- Mention rate = comparable runs containing at least one brand or product mention ÷ all comparable runs
- Primary-prominence rate = comparable runs classified as Primary ÷ all comparable runs
- Brand-domain citation incidence = comparable runs containing at least one citation to the brand’s domain ÷ all comparable runs

If an answer mentions the brand several times or cites several pages on the brand’s domain, it still adds only one run to the relevant numerator. Keep those individual mentions, URLs, and citation events for separate volume and source analysis. Report each rate with its numerator and denominator, such as 35 of 100 runs, rather than showing the percentage alone.

A citation is not proof that the answer is correct or favorable. OpenAI explicitly warns that ChatGPT search results and citations may be incomplete, outdated, or incorrect in its [search documentation](https://help.openai.com/en/articles/9237897-searching-the-web-with-chatgpt). Review the answer’s claim and the cited page together.

For Perplexity API measurement, use the consistent source fields returned by the API instead of extracting URLs from written answers. Perplexity’s [prompt guide](https://docs.perplexity.ai/docs/agent-api/prompt-guide) directs API users to collect source details from `search_results` across search steps. Its [changelog](https://docs.perplexity.ai/docs/resources/changelog) also records the replacement of the old `citations` field with `search_results`. A measurement pipeline that ignores this change in data format could report a false drop caused by data collection rather than visibility.

## Compare source coverage, not just citation volume

Ten citations from one domain represent a different source base from ten citations across eight domains. The first pattern may depend on a single publisher. The second may show wider coverage, though it says nothing by itself about source quality.

Compare:

- Number of unique domains
- Number of unique cited pages
- Source categories
- Coverage by prompt topic and intent
- Owned versus third-party representation
- Domains shared with competitors
- Domains that cite only a competitor
- Concentration in the top one, three, or five domains

This is where citation monitoring becomes an editorial tool. A lost URL suggests inspecting that page. A competitor appearing across sources where the brand is absent suggests a coverage gap outside the company site. Neither observation proves what caused the answer engine to choose its sources, but both narrow the next investigation.

The same logic informs technical work. [Schema for AEO](/blog/how-should-schema-be-used-for-aeo) can clarify entities and page meaning, but structured data should be evaluated alongside the content and third-party sources that support a claim.

## Treat competitor substitution as its own event

A missing brand mention is only half the finding. Record which competitor took its place, whether that competitor gained prominence, and which sources appeared with it.

A delta log, as recommended in Click Laboratory’s [monitoring workflow](https://www.clicklaboratory.com/aeo/monitor-ai-answer-engine-changes), can flag structural drops and competitor substitutions. A delta log is a record of what changed between two reporting periods. That advice comes from a commercial agency, not an independent study. Its operational value lies in preserving the joins among prompt, engine, page, citation, and competitor.

For each material change, log:

- Previous and current brand presence
- Previous and current competitor presence
- Lost and gained citations
- Affected prompt IDs
- Known site, content, or tracking changes
- Engine or interface changes visible to the team
- Traffic and conversion movement in the same period

Avoid assigning a cause in the change log unless the evidence supports it. “Competitor replaced us after our page edit” establishes sequence, not causation.

## Keep traffic separate from visibility

**Traffic is a downstream observation, not a synonym for answer-engine presence.** An answer can mention a brand without offering a link, and a linked citation may answer the user’s question so well that no click follows.

Use the platform-specific evidence available:

- For ChatGPT, track referrals carrying `utm_source=chatgpt.com`, as documented by [OpenAI](https://help.openai.com/en/articles/12627856-publishers-and-developers-faq).
- For Google AI Mode, use Google’s documented click, impression, and position definitions. Do not assume every generative Search surface can be isolated unless Google provides the required filter or dimension.
- For other engines, use observable referral data, but do not invent or infer a dedicated tracking parameter.

Compare landing page, session quality, and prompt theme where the data permits it. Also watch branded-search movement and assisted conversions—valuable actions that AI visibility may have helped even when another source received final credit. Big Human argues that these may reveal influence missed by last-click reporting, which gives all credit to the final recorded source before an action. This remains vendor guidance rather than a formal method for deciding how credit is shared among sources.

A useful chain might read:

`Prompt run → brand mention → citation URL → referral session → qualified action`

Most records will not complete that chain. Showing the gaps is more honest and more useful than forcing them into a single visibility-to-revenue number.

## Define qualified actions before reporting them

A qualified action needs a stable business rule. It might be a demo request from an eligible company, a completed consultation form, or another agreed action that signals real buying interest. A page view or unqualified form submission should not quietly enter the same total.

For every action, retain:

- Action type
- Qualification rule and rule version
- Landing page
- First-touch source and any source that helped before the action, when available
- Timestamp
- Associated campaign or content
- Pipeline stage and revenue only when the underlying system supports them

Orathis says its reporting connects AI visibility with qualified actions, pipeline, and revenue through landing experiences and conversion workflows. That describes the company’s approach. It does not establish that a mention or citation caused a later outcome.

Good reporting shows the sequence and states the attribution model—the rule used to decide which sources receive credit for an action. It does not turn correlation into credit.

## Choose a cadence that exposes variation

One run per prompt is a snapshot. It is not a trend.

The right cadence, or testing schedule, depends on the decision. Frequent sampling helps detect answer and citation volatility. Longer comparison windows help teams judge whether a change persisted. Repeat important prompts within each period so that a single unusual response does not control the result.

Vested Marketing recommends observing featured-result and click-through movement over a consistent 90-day window before a major rewrite. It also suggests that initial changes may appear within 30 to 90 days for some high-intent queries. These are [practitioner recommendations](https://www.vested.marketing/blog/what-is-answer-engine-optimization), not universal benchmarks.

A practical schedule could combine:

- Frequent repeat runs for priority prompts
- Weekly review of large gains, losses, and competitor substitutions
- Monthly comparisons using a fixed prompt set
- Longer-period review before major strategic changes
- Immediate annotations for site releases, content changes, tracking changes, and known engine updates

This is a hypothetical operating model, not a claimed performance schedule. The important rule is consistency: use the same sampling method across the periods being compared.

## Build the report as a chain of evidence

Once the fields are stable, the report should move from observation to business relevance without collapsing the steps.

1. Confirm that engine, prompt, locale, and sampling conditions match.
2. Compare brand presence and prominence.
3. Inspect sentiment and the text behind its label.
4. Review gained, lost, and persistent citation URLs.
5. Check source diversity and competitor substitutions.
6. Compare traffic and demand signals.
7. Examine qualified actions under a stated rule for assigning credit.
8. Annotate changes and possible explanations without presenting guesses as causes.

This sequence prevents a common error: starting with a traffic change and searching backward for an answer-engine story that seems to explain it.

For a broader measurement framework, see [Measuring AI visibility](/blog/measuring-ai-visibility). Teams that need to turn these fields into a repeatable operating process can also use the principles in [designing AI workflows that scale](/blog/from-chaos-to-clarity-designing-ai-workflows-that-scale).

## What should never be merged into one metric?

Mentions, prominence, sentiment, citations, source coverage, traffic, and qualified actions answer different questions. Combining them into one index may help with executive scanning, but it can also hide opposing movements.

Consider this explicitly hypothetical month-to-month result. Each period contains 100 completed, comparable runs:

| Signal | Earlier period | Later period |
|---|---:|---:|
| Mention rate | 35 of 100 runs (35%) | 50 of 100 runs (50%) |
| Primary-prominence rate | 20 of 100 runs (20%) | 10 of 100 runs (10%) |
| Brand-domain citation incidence | 18 of 100 runs (18%) | 8 of 100 runs (8%) |
| Competitor substitutions | 12 events | 21 events |
| Qualified actions | 7 actions | 6 actions |

The three rates use the same run population and count a run no more than once for each measure. For example, three brand-domain citations in one answer contribute one run—not three—to citation incidence.

A composite score might rise because mentions increased. The underlying picture is less comfortable: the brand appeared more often but held weaker positions, earned fewer citations, lost more answers to competitors, and recorded no growth in qualified actions.

That is the central reason to preserve separate histories over time. The disagreement between signals is often the finding.

## Frequently asked questions

### How many prompts should a company monitor?

Use enough prompts to represent the decisions customers make, including category, problem, comparison, use-case, and buyer-journey questions. The right number depends on the breadth of the market and the team’s ability to sample consistently.

Orathis states that it builds and tracks more than 1,000 prompts for clients. This establishes the scope of its service, not a minimum that every company must adopt. A smaller stable set is more useful than a large set that changes without version control.

### Should answer position be compared across engines?

Only with care. Preserve the native observation from each engine and surface. If the reporting team creates labels such as primary, included, supporting, and absent, document the rules and treat them as an internal convention.

No official source cited here defines a universal answer-position scale.

### Does a citation mean the source influenced the answer?

It shows that the engine presented the source in connection with the answer. It does not by itself prove that the source caused the wording, supported every claim, or deserved the citation. Check the answer passage and source page together.

### Can AI referral traffic be isolated in analytics?

Sometimes. OpenAI documents `utm_source=chatgpt.com` for ChatGPT referrals. Google documents clicks, impressions, and position for AI Mode, subject to Search Console’s available reporting. The cited official materials do not establish equivalent dedicated referral markers for Claude, Microsoft Copilot, or Perplexity.

### When is a trend reliable?

Confidence improves when the same prompt is run repeatedly under recorded conditions and the movement persists across comparison periods. There is no universal sample count in the cited evidence. Report the number of runs and the observed variability so readers can judge the trend.

## The useful signal is often the disagreement

Answer-engine monitoring becomes valuable when it preserves the path from a prompt to an observable answer, from that answer to its sources, and from those sources to measurable reader behavior.

The goal is not to make every chart point upward. It is to see when the charts disagree and know where to look next. More mentions paired with weaker citations call for source analysis. Stable visibility with falling traffic calls for link and landing-page review. More referrals without qualified actions shifts attention to intent, qualification, and the conversion path.

That is a better decision system than one visibility score because it keeps the evidence intact.

Orathis works across answer-engine strategy, technical implementation, content, third-party distribution, prompt tracking, and reporting. To discuss a comparison framework for your prompt set and business goals, [contact Orathis](/contact).

## About the author

[Quinn Bean](https://www.linkedin.com/in/quinn-bean-0b38282b8) is Director of Orathis, focused on answer-engine strategy, AI visibility, governed content systems, technical implementation, and connecting AI discovery to measurable business outcomes.

## Citation guidance

Use the canonical HTML URL when citing this page. Verify time-sensitive or third-party platform claims against linked primary sources before repeating them.
