A confident answer can hide both a factual error and a weak trail of evidence.
An AI answer says your company has discontinued a product, changed its pricing, or operates in the wrong market. The natural response is to correct the first page you find and run the prompt again. That may produce a different answer, but it does not tell you where the original error came from or whether the change will last.
What should the team do first?
Save the complete answer and the conditions that produced it before changing anything. Then split the answer into claims, verify the disputed fact against the current authoritative record, inspect every visible citation, correct any source you control, report the problem through an appropriate provider channel where available, and repeat the test under recorded conditions.
The practical sequence is:
- Preserve the answer and its test conditions.
- Define the exact claim under dispute.
- Check the claim against the source responsible for that fact.
- Inspect the answer’s citations one by one.
- Classify the error by what the visible evidence shows.
- Correct inaccurate information on controlled pages.
- Submit careful provider feedback where the product supports it.
- Retest the original prompt and useful variants several times.
- Record what changed without claiming an unproven cause.
This process establishes what the engine displayed, whether the claim was accurate, and whether later tests differ. It usually cannot reveal the engine’s hidden reasoning or prove that one corrective action caused a changed answer.
1. Preserve the original answer
Capture the evidence before a page edit, new conversation, model update, or interface change makes the first result hard to reproduce.
A useful investigation record includes:
| Field | What to save |
|---|---|
| Answer | Full text, not only the inaccurate sentence |
| Prompt | Exact wording, spelling, and follow-up questions |
| Engine | ChatGPT, Claude, Gemini, Copilot, Perplexity, or another product |
| Model or mode | Any visible model name, search mode, research mode, or browsing setting |
| Time | Date, time, and time zone |
| Market | Country, locale, and language |
| Account context | Signed in or signed out, plan type, and relevant organization settings |
| Conversation context | Earlier messages that may have shaped the answer |
| Sources | Citation labels, links, quoted passages, and access dates |
| Visual record | Screenshots or an exported conversation where permitted |
Treat these fields as a practical reproducibility checklist, not a universal list required by every provider.
Prompt wording matters. MIT Sloan’s guidance on AI hallucinations and bias notes that prompt specificity affects output quality. Recording the exact prompt makes a later comparison more meaningful. It does not mean the prompt caused the error.
Before sharing screenshots or exporting a conversation, remove personal data, credentials, private pricing, customer information, and other secrets.
2. Turn the answer into claims you can test
A single response may mix correct details with one invented, outdated, or distorted statement. The University of Maryland Libraries explains that AI output can combine accurate information with fabricated material. Its information-literacy guidance supports checking the answer claim by claim.
Suppose an answer says:
Hypothetical example: “Northstar offers a free plan, serves healthcare companies, and closed its European operations in 2025.”
That sentence contains at least three separate claims:
- Northstar offers a free plan.
- Northstar serves healthcare companies.
- Northstar closed its European operations in 2025.
Write the disputed claim as a short factual statement. Do not investigate a vague label such as “the answer is wrong.” A clear claim lets the team identify the right owner and record.
Different facts belong to different authoritative sources:
| Claim | Strong starting record |
|---|---|
| Current price or plan | Live pricing page and approved billing record |
| Product availability | Product documentation or release record |
| Office status | Current company location page or corporate record |
| Executive role | Official leadership page or relevant filing |
| Certification | Issuing organization’s current registry |
| Policy term | Current policy page and its effective date |
The source responsible for a fact should take priority over a reseller page, an old news story, or a search snippet. This follows the general principle of primary-source verification: confirm information with the organization or record that issues it. The Joint Commission’s definition applies specifically to credentials, so its use here is an analogy rather than an AI-product rule.
3. Check whether the company’s own record is clear
The investigation may expose a company information problem before it exposes an engine problem.
Check the live page, page title, structured data, downloadable documents, help articles, press releases, partner listings, and regional versions. Look for conflicting dates, plan names, legal names, addresses, or product descriptions. Search the site for the old fact as well as the new one.
Record:
- The current correct value
- The page or internal record that establishes it
- The owner of that record
- Its effective or last-updated date, if known
- Any controlled page that still shows the old value
- Any ambiguity between regions, plans, products, or dates
A current pricing page does not necessarily prove what the price was last year. Likewise, an archived announcement may have been accurate when published but wrong as a description of the company today. Attach time and scope to the fact.
This is also where a team should resolve its own contradictions. Updating one page while an old PDF and regional landing page remain live leaves answer engines and human readers with competing evidence.
4. Open every citation
A citation is a lead to inspect, not proof that the sentence is supported.
For each cited page, ask:
- Does the page actually contain the claimed fact?
- Is the page current enough for the question?
- Does it refer to the same company, product, market, and time period?
- Is the engine summarizing it accurately?
- Is the source authoritative for this fact?
- Does the page cite another source that should be checked?
This distinction matters. In the Tow Center’s test of eight generative search tools, ChatGPT incorrectly identified 134 of 200 requested news articles and expressed uncertainty only 15 times. The Tow Center report supports close inspection of citations and confident answers. Those figures describe a specific news-retrieval experiment, not a general accuracy rate for company facts.
Microsoft likewise says Copilot may make mistakes or misrepresent third-party material even when it provides web citations. Its Copilot privacy FAQ therefore supports comparing the answer with the cited text rather than accepting the citation at face value.
5. Classify what the visible evidence establishes
The next action depends on the error class.
| What you observe | What you can conclude | Appropriate response |
|---|---|---|
| A cited controlled page contains the old fact | The visible company source is outdated | Correct that page and related controlled records |
| A cited third-party page contains the old fact | The answer has surfaced an outdated external source | Document it and contact the publisher if correction is appropriate |
| The cited page says something different | The answer misrepresents its visible citation | Save the relevant passage and include it in provider feedback |
| The citation supports only part of the sentence | The answer extends beyond its cited support | Separate the supported and unsupported claims |
| No source is shown | The claim lacks visible support | Verify the fact, but do not guess where the statement originated |
| Sources conflict by date or market | The answer may have collapsed distinct contexts | Clarify the current scope and dates on authoritative pages |
Do not label an uncited statement “training-data contamination” or a “retrieval error” without evidence. It might reflect training data, hidden retrieval, earlier conversation context, personalization, or generated text. The displayed answer alone often cannot distinguish among them.
That limit is useful. It keeps the team focused on evidence it can inspect instead of inventing a technical diagnosis.
For ongoing monitoring, separate brand mentions from the pages used to support them. Orathis explains this distinction in why teams should track sources as well as brand mentions.
6. Correct the public record you control
Once the authoritative value is settled, fix incorrect or unclear controlled pages. Give each important fact one clear, current statement, then align other relevant pages with it.
Depending on the claim, that work may include:
- Updating the main product, pricing, location, leadership, or policy page
- Removing or marking obsolete documents
- Adding an effective date where time changes the meaning
- Clarifying differences between regions, plans, and product editions
- Updating structured data so it agrees with visible page content
- Fixing internal links that still point readers toward obsolete material
- Requesting corrections from third-party publishers with the supporting record
Structured data can help machines read explicit page information, but it should match what a person can see. It is not a hidden correction channel. The practical boundaries are covered in how schema should be used for AEO.
A page edit improves the public record. It does not guarantee that an answer engine will crawl the page, select it, interpret it correctly, or revise an answer on a set schedule.
7. Use provider feedback with care
Provider controls differ across products, account types, regions, and organization settings. Check the current interface and official help page for the exact product you tested.
| Provider | Documented option | Important limit |
|---|---|---|
| ChatGPT | OpenAI documents thumbs-up and thumbs-down response feedback. Its reporting flow also covers specified safety or legal concerns. | The documented reporting route is not a universal factual-correction request. Feedback may include the associated conversation. |
| Claude | Anthropic documents response feedback in Claude Console and reporting for shared content. | An organization administrator may disable Console feedback. Shared-content reporting has a narrower purpose than ordinary fact correction. |
| Gemini Apps | Google documents feedback on inaccurate responses and a general problem-reporting option. | Response feedback includes the associated conversation, including prompts and responses. |
| Microsoft Copilot | Microsoft documents thumbs-up or thumbs-down feedback with an explanation field. | An administrator may disable feedback for work or school users. |
| Perplexity | Perplexity provides a dedicated process for reporting an incorrect or inaccurate answer. | The help page does not promise correction of a particular future response. |
Relevant primary instructions include:
- OpenAI’s ChatGPT reporting guidance
- Anthropic’s Claude Console feedback settings
- Google’s Gemini Apps feedback guidance
- Microsoft’s Copilot feedback guidance
- Perplexity’s inaccurate-answer reporting guidance
The cited documentation does not establish a current general fact-correction procedure for Grok or for ordinary factual errors in Google AI Overviews. Use a relevant visible control if one is available, but do not assume an undocumented route.
A useful report is brief and testable:
On 15 September 2026 at 14:10 UTC, the answer to “[exact prompt]” stated “[exact disputed claim].” The current authoritative record at [URL] states “[correct fact]” and is dated [date, if available]. The answer cited [URL], but that page [contains an outdated value / does not support the statement]. Screenshot attached.
Submit only the context needed to explain the error. OpenAI says the entire associated conversation may be used when response feedback is submitted. Google also says Gemini response feedback includes the associated conversation. Review and redact sensitive material before sending feedback where the product permits it.
8. Retest without moving the goalposts
Run the exact original prompt again under conditions as close as practical to the recorded test. Then test a small set of reasonable variants, such as:
- The same prompt in a new conversation
- A direct question about the disputed fact
- A version that asks for sources
- A version that specifies the relevant country or date
- The same prompt while signed out, if appropriate and permitted
- Another engine that is relevant to the team’s audience
Repeat important tests. Washington State University reports that its researchers found inaccurate and inconsistent behavior across more than 700 AI-generated responses. The university’s summary of that research supports repeated observation rather than reliance on one run, though the cited release does not provide the complete underlying methodology.
Record results by run:
| Run | Prompt | Conditions | Claim correct? | Source shown? | Source supports claim? |
|---|---|---|---|---|---|
| 1 | Original prompt | Original conditions | No | Yes | No |
| 2 | Original prompt | Matched retest | Yes | Yes | Yes |
| 3 | Original prompt | Matched retest | No | No | Not assessable |
| 4 | Date-specific variant | Same engine and account | Yes | Yes | Yes |
Count runs, not just distinct answer wordings. If five runs produce the same wrong claim, that is five observed responses containing the claim. It is not five different errors or five affected users.
Teams that monitor a larger prompt set also need stable comparison fields across dates, engines, and modes. See which answer-engine signals should be compared over time for a broader measurement framework.
What does a changed answer prove?
A changed result proves only that the observed output changed under the recorded test.
It does not by itself prove:
- The page edit caused the change
- Provider feedback caused the change
- The engine removed an old fact from every system
- Other accounts, markets, or modes will receive the same answer
- The correction will persist
- The answer is now accurate in every respect
Search indexes, model versions, source selection, account context, location, language, and normal response variation may all differ between tests. Avoid attributing cause unless the evidence isolates it.
A practical outcome label is more honest:
- Not reproduced: The wrong claim did not appear in the recorded retests.
- Intermittent: Both correct and incorrect versions appeared.
- Persisting: The wrong claim continued to appear.
- Changed but still unsupported: The wording changed, but the visible evidence still did not support it.
- Correct with supporting source: The tested answer matched the authoritative fact and its cited source.
This turns a frustrating screenshot into a measured issue without claiming more than the tests show.
How much investigation is enough?
Match the effort to the harm the error could cause.
A wrong office opening time may need a quick record check and page update. A false statement about regulatory status, product safety, executive identity, eligibility, or contractual terms deserves legal or compliance review and a fuller evidence trail.
Ask three questions:
- Could someone make a material decision based on this claim?
- Is the claim appearing repeatedly or across several engines?
- Is the underlying authoritative record clear and current?
Escalate when the possible harm or uncertainty is high. The principle is proportional verification: spend more effort where a false fact could cause greater damage. Research on source-data verification in clinical trials describes such verification as resource-intensive quality assurance. Its PLOS ONE study is not an AI correction guide, but the proportionality principle is useful here.
Frequently asked questions
Should we ask the AI where it got the wrong fact?
You may ask, but treat the reply as another claim to verify. Open the cited pages and compare their actual text with the answer. A model’s explanation of its own source or reasoning is not a substitute for visible evidence.
Should we delete old company pages?
Not automatically. An old page may still be useful as a historical record. If it remains public, label its date and status clearly, link to the current information, and remove wording that presents an obsolete fact as current.
How soon should we retest after a correction?
The available evidence does not establish a universal update time. Save a baseline, retest on a schedule appropriate to the risk, and keep the conditions and intervals in the log. Avoid promising that any provider will process a page edit or feedback report by a particular date.
What if the answer has no citations?
Verify the disputed claim against the authoritative record and save the uncited response. You can report the visible error where an appropriate control exists, but you cannot infer whether it came from training, retrieval, conversation context, or generation.
When is one correct retest enough?
One correct result shows that the engine produced a correct answer once. It does not establish consistency. Claims that affect purchases, compliance, reputation, or customer action deserve repeated tests across the relevant prompts and conditions.
Build a record that survives the next answer
The investigation starts with one wrong sentence, but its real output should be a durable chain of evidence: what the engine displayed, which fact was disputed, which record established the truth, what each citation supported, what the team changed, and what later tests showed.
That chain matters more than a single corrected screenshot. It lets marketing, operations, legal, and technical teams work from the same facts while keeping uncertainty visible.
Orathis runs month-to-month answer-engine optimization across strategy, technical GEO, AI-ready content, third-party distribution, prompt tracking, and reporting. Its monitoring covers more than 1,000 prompts across ChatGPT, Claude, Gemini, Grok, Perplexity, Copilot, and Google AI Overviews, with measurement of mentions, sentiment, citations, source coverage, and answer rankings. Teams that need a structured investigation and ongoing visibility program can contact Orathis.
About the author
Quinn Bean is Director of Orathis, focused on answer-engine strategy, AI visibility, governed content systems, technical implementation, and connecting AI discovery to measurable business outcomes.