Sample data. yoursite.com is a fictional brand and every figure below is generated, not measured.

Skip to main content

Sample workspace · yoursite.com · fictional

From a recommendation to a shipped change to a measured difference

A score you cannot act on is a vanity metric. This is the queue of fixes, one of them measured 14 days after it shipped, and the two kinds of trouble the product watches for: an answer that changed, and a claim that is simply untrue.

The work loop

Four columns, and nothing leaves the last one unmeasured

Mirrors/work

Recommended

3
  • Publish per-seat pricing as a table, not an image

    /pricing

    Effort S+3 to +6 on commercial prompts
  • Add a comparison page against Northlake

    /compare/northlake

    Effort M+2 to +5 on comparison prompts
  • Mark up the FAQ with schema.org/FAQPage

    /help

    Effort SCitation share, not score

In progress

2
  • Rewrite the integrations page to name each tool

    /integrations

    Effort M+2 to +4 on integration prompts
  • Ask three customers for a G2 review

    Off-site

    Effort LReview-source citations

In review

1
  • Correct the 'enterprise only' line in the About page

    /about

    Effort SRetires one false claim

Shipped

2
  • Public pricing page with a plain-text plan table

    /pricing

    Effort SMeasured below
  • Per-project margin explainer with worked numbers

    /guides/margin

    Effort MMeasured next window

Every recommendation carries the surface it touches, the effort it costs and the effect it is expected to have — stated before it ships, so the measurement afterwards can disagree.

Shipped, then measured

What the pricing page was actually worth

Mirrors/work

Shipped 12 Aug 2026

Public pricing page with a plain-text plan table

14 days before

48

Visibility on the affected prompts only, not the whole workspace.

14 days after

54

Same prompts, same engines, same cadence.

Difference

+6

95% CI +1.4 to +10.6

Readings compared

336

168 before, 168 after.

Verdict

Distinguishable from no change at 95%, but only just. The interval still contains a +1 effect, so treat this as evidence, not proof.

Fourteen days of readings before the change against fourteen days after, on the same prompts and the same engines. The interval decides whether it counted.

Answer change

The sentence that changed, word for word

Mirrors/answer-changes
ChatGPT19 Aug 202626 Aug 2026

best project cost tracking software for agencies

Added
  • yoursite.com, which publishes per-project margin in real time
  • priced per seat with no minimum
Removed
  • which is mainly used by enterprise finance teams
  • though its pricing is not public

Both removals track the pricing page you shipped on 12 Aug. The addition quotes it almost verbatim.

The product keeps the previous answer, so when an engine starts saying something new you get the diff rather than a number that moved for unexplained reasons.

Brand truth

A claim that is wrong, and how long it has been wrong for

Mirrors/brand-truth
Severity: highOpen 74 daysFirst seen 19 Jun 2026 · last seen 30 Aug 2026

yoursite.com is enterprise-only and does not sell to teams under 50 people.

There is a 3-seat plan. The claim traces to a 2023 press quote that is still the top news citation.

Where the false claim appears
EngineOccurrencesOf readingsRate
Google Gemini192407.9%
Claude122405.0%
ChatGPT72402.9%
Total387205.3%

Likely source: theregister.example — 2023 interview, still cited by two engines

A wrong answer repeated across three engines is a sales problem with an age. The row carries where it shows up, how often, and the source it most likely came from.