> The Answerworthy measurement methodology: stated replicate counts, majority-cited scoring, OBSERVED vs INFERRED evidence, verified crawler data, honest caveats.

Published by Answerworthy · Canonical: https://answerworthy.ai/learn/methodology · Updated 2026-08-03

# Our methodology

Every AI-visibility tool will show you a dashboard. The question worth asking any of them,
including us, is: **where did this number come from, and would it survive scrutiny?** This page
is our full answer. It is the page we most want you to read before trusting us, because in a
category where a screenshot passes for evidence, methodology is the product.

---

## Replicates, or it didn't happen

AI engines are probabilistic. The same prompt, on the same engine, minutes apart, can produce
answers that differ in whether your brand appears at all. Any tool that reports a single
run as settled fact, with no stated sample size and no stored raw answer, is reporting a coin
flip as your visibility.

So Answerworthy treats every engine answer as a sample, never a verdict. Today, each
prompt-engine cell in a scheduled panel run executes as a **single measurement**, and every
raw answer is stored verbatim so the evidence behind each number stays inspectable. The
replicate count is a configuration value, not a promise: it can increase as measurement
budgets allow, and when a cell runs with multiple replicates, a brand counts as cited only
when **the majority of replicates cite it**. Whatever the count, every run and every report
states the number of replicates actually executed. We would rather show you an honest sample
size than a dressed-up one.

## Trends need history, alerts need statistics

One measurement is a point. Two are a line you should not trust. We do not draw trend
conclusions from less than eight weeks of data, and alerts fire only when a metric moves more
than two standard deviations from its own eight-week history, per metric, per engine. Not on
single-prompt changes, not on day-to-day jitter.

The methodology itself is versioned and locked: once set, metric definitions do not quietly
change underneath your trend lines. If a definition ever must change, that is a new baseline,
flagged visibly, never a silent re-cut of history.

## OBSERVED vs INFERRED, enforced structurally

Every finding Answerworthy produces carries one of two tags, at the database level, not as a style
choice:

- **OBSERVED**: backed by a reproducible measurement, with a required reference to its source: a
  stored engine response, a verified log line, a validator output, a Search Console export. You
  can click through to the evidence.
- **INFERRED**: a professional interpretation of the evidence, labeled as exactly that.

The two never share a sentence unmarked, in dashboards or in PDF reports. Where evidence does
not exist, the report says what could not be verified rather than papering over it. We built it
this way because agency reporting lives or dies on defensibility, and "the tool said so" is not
a defense.

## Raw responses, stored forever

Every engine answer behind every metric is stored verbatim. When a report says your citation
rate on Perplexity fell, you can open the actual responses that produced the number. Reproducibility
is what turns a claim into a measurement.

## Native APIs, and an honest caveat

Answerworthy measures engines through native APIs with live search grounding wherever the
engine offers one: OpenAI's API for the ChatGPT surface, Anthropic's for Claude, Google's
Gemini grounding, and Perplexity's Sonar, with Google AI Overviews and AI Mode tracked as
additional, separately labeled surfaces. The methodology is documented in full and every raw
answer is stored, so APIs give clean, reproducible measurement at statistical volume, which
browser-scraping approaches struggle to guarantee.

Other vendors bury this caveat; we lead with it: **an API answer is a rigorous proxy for the consumer
surface, not a screenshot of it.** Consumer apps layer personalization and memory on top of the
same engine, so individual sessions can diverge from API results. Every engine surface in
Answerworthy carries a provenance note saying precisely what was measured. We will not claim to
show you "what ChatGPT says," because nobody measuring at scale honestly can. What we show you
is a statistically stable, reproducible measurement of the engine that powers it.

## Crawler data is verified, not assumed

Layer 2 numbers come from real server and CDN logs, and every visit is verified: user-agent
matched against a maintained bot registry, then checked against the vendor's published IP
ranges (with reverse-DNS verification where vendors support it). Traffic that fails
verification is set aside as spoofed, visible but excluded from analytics. Bots are classified
by behavior (training, search, user-triggered), the classes are never summed into one "AI
traffic" number, and crawler hits are never conflated with human visits: crawl-to-referral
ratios run from tens to thousands to one, and pretending otherwise flatters everyone's charts
but informs no one.

## Audits measure what engines experience

The on-site audit fetches your pages the way AI bots do, compares raw HTML against the
JavaScript-rendered page, and gates its own findings honestly: if a bot cannot reach the page,
downstream content checks report as blocked rather than failing misleadingly. Schema is
validated, not pattern-matched, and the generator refuses to emit thin markup because sparse
schema measurably underperforms none.

## What we do not do

- We do not report single-run results as trends.
- We do not invent a proprietary composite score we cannot explain.
- We do not conflate crawler hits with traffic, or mentions with citations.
- We do not claim attribution precision that does not exist. AI referral data undercounts
  systematically; we report the observed floor and the inferred range, labeled.
- We do not mark a fix "done" because it shipped. Verified means re-measured.

Every metric in the product carries an explanation panel describing the principle behind it, so
your team and your clients learn the methodology while using it.

**[Run a free AI-visibility scan](https://app.aeo-test.ai)** · [Create your free account](https://app.aeo-test.ai)
