Answerworthy.ai

Learn

Updated August 3, 2026

Our methodology

Every AI-visibility tool will show you a dashboard. The question worth asking any of them, including us, is: where did this number come from, and would it survive scrutiny? This page is our full answer. It is the page we most want you to read before trusting us, because in a category where a screenshot passes for evidence, methodology is the product.


Replicates, or it didn’t happen

AI engines are probabilistic. The same prompt, on the same engine, minutes apart, can produce answers that differ in whether your brand appears at all. Any tool that reports a single run as settled fact, with no stated sample size and no stored raw answer, is reporting a coin flip as your visibility.

So Answerworthy treats every engine answer as a sample, never a verdict. Today, each prompt-engine cell in a scheduled panel run executes as a single measurement, and every raw answer is stored verbatim so the evidence behind each number stays inspectable. The replicate count is a configuration value, not a promise: it can increase as measurement budgets allow, and when a cell runs with multiple replicates, a brand counts as cited only when the majority of replicates cite it. Whatever the count, every run and every report states the number of replicates actually executed. We would rather show you an honest sample size than a dressed-up one.

One measurement is a point. Two are a line you should not trust. We do not draw trend conclusions from less than eight weeks of data, and alerts fire only when a metric moves more than two standard deviations from its own eight-week history, per metric, per engine. Not on single-prompt changes, not on day-to-day jitter.

The methodology itself is versioned and locked: once set, metric definitions do not quietly change underneath your trend lines. If a definition ever must change, that is a new baseline, flagged visibly, never a silent re-cut of history.

OBSERVED vs INFERRED, enforced structurally

Every finding Answerworthy produces carries one of two tags, at the database level, not as a style choice:

  • OBSERVED: backed by a reproducible measurement, with a required reference to its source: a stored engine response, a verified log line, a validator output, a Search Console export. You can click through to the evidence.
  • INFERRED: a professional interpretation of the evidence, labeled as exactly that.

The two never share a sentence unmarked, in dashboards or in PDF reports. Where evidence does not exist, the report says what could not be verified rather than papering over it. We built it this way because agency reporting lives or dies on defensibility, and “the tool said so” is not a defense.

Raw responses, stored forever

Every engine answer behind every metric is stored verbatim. When a report says your citation rate on Perplexity fell, you can open the actual responses that produced the number. Reproducibility is what turns a claim into a measurement.

Native APIs, and an honest caveat

Answerworthy measures engines through native APIs with live search grounding wherever the engine offers one: OpenAI’s API for the ChatGPT surface, Anthropic’s for Claude, Google’s Gemini grounding, and Perplexity’s Sonar, with Google AI Overviews and AI Mode tracked as additional, separately labeled surfaces. The methodology is documented in full and every raw answer is stored, so APIs give clean, reproducible measurement at statistical volume, which browser-scraping approaches struggle to guarantee.

Other vendors bury this caveat; we lead with it: an API answer is a rigorous proxy for the consumer surface, not a screenshot of it. Consumer apps layer personalization and memory on top of the same engine, so individual sessions can diverge from API results. Every engine surface in Answerworthy carries a provenance note saying precisely what was measured. We will not claim to show you “what ChatGPT says,” because nobody measuring at scale honestly can. What we show you is a statistically stable, reproducible measurement of the engine that powers it.

Crawler data is verified, not assumed

Layer 2 numbers come from real server and CDN logs, and every visit is verified: user-agent matched against a maintained bot registry, then checked against the vendor’s published IP ranges (with reverse-DNS verification where vendors support it). Traffic that fails verification is set aside as spoofed, visible but excluded from analytics. Bots are classified by behavior (training, search, user-triggered), the classes are never summed into one “AI traffic” number, and crawler hits are never conflated with human visits: crawl-to-referral ratios run from tens to thousands to one, and pretending otherwise flatters everyone’s charts but informs no one.

Audits measure what engines experience

The on-site audit fetches your pages the way AI bots do, compares raw HTML against the JavaScript-rendered page, and gates its own findings honestly: if a bot cannot reach the page, downstream content checks report as blocked rather than failing misleadingly. Schema is validated, not pattern-matched, and the generator refuses to emit thin markup because sparse schema measurably underperforms none.

What we do not do

  • We do not report single-run results as trends.
  • We do not invent a proprietary composite score we cannot explain.
  • We do not conflate crawler hits with traffic, or mentions with citations.
  • We do not claim attribution precision that does not exist. AI referral data undercounts systematically; we report the observed floor and the inferred range, labeled.
  • We do not mark a fix “done” because it shipped. Verified means re-measured.

Every metric in the product carries an explanation panel describing the principle behind it, so your team and your clients learn the methodology while using it.

Run a free AI-visibility scan · Create your free account