---
title: The Vendor Claim Audit Checklist
description: "Fifteen checks in five groups for auditing any vendor engagement or AI claim: what to ask, the red flag, and what a pass looks like."
url: https://preview.artificialpoets.com/resources/vendor-claim-audit-checklist/
site: Artificial Poets
type: a13s_content
date: 2026-08-09T20:37:16+00:00
modified: 2026-08-18T03:09:42+00:00
image: https://preview.artificialpoets.com/wp-content/uploads/a13s-cards/1535-social-336d4804.png
---
# The Vendor Claim Audit Checklist

Fifteen checks in five groups — baseline, comparison, windows, mechanism, reversibility — for auditing any vendor’s engagement claim. Including ours.

Fifteen checks in five groups. Each has what to ask, the red flag, and what a pass looks like. A vendor who survives all fifteen has measured their product; a vendor who volunteers the material before you ask is the one to keep. We publish our own numbers to this standard — which is why we can hand you the list.

**Free checklist**

## Get the full checklist

Enter your work email and the fifteen checks unlock instantly — readable in the browser, printable for the meeting.

## Group A · Baseline

- **1 · Is there a pre-period at all?** *Ask:* "What did these metrics look like before enablement, over how many weeks?" *Red flag:* results with no baseline, or a baseline shorter than one seasonal cycle. *Pass:* a named baseline window — ideally 12+ weeks — collected before anything changed.
- **2 · Same instrumentation across the series?** *Ask:* "Did tracking, tagging or consent setup change between baseline and result?" *Red flag:* "we improved tracking as part of onboarding" — the metric changed, not the readers. *Pass:* identical event definitions across the whole series, stated in writing.
- **3 · Metrics frozen before enablement?** *Ask:* "When was the success metric chosen — before the deployment or after the results?" *Red flag:* the headline metric differs from deal to deal — the tell of post-hoc selection. *Pass:* metric and windows fixed in the pilot agreement, before the switch was flipped.

## Group B · Comparison

- **4 · What is the counterfactual?** *Ask:* "What would these sites have done without you, and how do you know?" *Red flag:* before/after only, or "industry benchmarks" standing in for a comparison. *Pass:* comparison properties measured over identical calendar weeks.
- **5 · How was the comparison group selected?** *Ask:* "Is this every eligible property, or a selection?" *Red flag:* comparison sites chosen after results existed. *Pass:* a stated selection rule — ideally the entire eligible set.
- **6 · Was comparison tracking intact throughout?** *Ask:* "Show me last-seen dates per comparison site across the window." *Red flag:* any comparison property whose data thins or dies inside the window — a dying denominator inflates the lift mechanically. *Pass:* a balanced panel — every site present every week, or the week dropped.

## Group C · Windows

- **7 · What window is the headline from?** *Ask:* "Exact dates, and why those dates?" *Red flag:* a window ending suspiciously close to launch, or a reason that is the result itself. *Pass:* window bounded by data-integrity events, not by what flatters.
- **8 · What happens under adjacent windows?** *Ask:* "Shift the baseline ±2 months and the endpoint month by month — what happens?" *Red flag:* the lift swings by multiples or flips sign. We have watched identical data produce +32pp and −4pp this way. *Pass:* movement of a few points at most, or movement with a stated explanation.
- **9 · Where is the steady state?** *Ask:* "What is it running at now, versus the launch quarter?" *Red flag:* only the launch window exists. *Pass:* peak and run rate side by side — and distrust anyone whose run rate isn't a fraction of launch.

## Group D · Mechanism

- **10 · Is the metric recomputable?** *Ask:* "Define it precisely enough that my analyst can reproduce it from our own data." *Red flag:* proprietary composite scores; "engagement index". *Pass:* a definition with its known failure modes attached.
- **11 · Can anything inflate it that isn't a reader?** *Ask:* "What non-reader events move this number — auto-loads, refreshes, bots, retags?" *Red flag:* a raw pageview or event count presented as engagement with no de-duplication story. *Pass:* counting rules that exclude machine and mechanical inflation — e.g. distinct articles rather than raw loads.
- **12 · Does the segment cut hold up?** *Ask:* "Break it by device and by channel. All segments, including the ones that fell." *Red flag:* only the blended number exists — or every single segment improved. *Pass:* the full cut with at least one honest negative. Real interventions have boundaries.

## Group E · Reversibility

- **13 · What happens when it's switched off?** *Ask:* "Have you measured the product not running on a live site?" *Red flag:* the question has never come up. *Pass:* an off-period where the metric returned toward baseline and recovered on resume — the strongest single piece of evidence a vendor can hold.
- **14 · What will they not claim?** *Ask:* "What did the product not do?" *Red flag:* it improved everything it touched. *Pass:* stated non-results, unprompted — traffic, search, revenue, whatever the honest boundary is.
- **15 · Can you leave?** *Ask:* "What does removal look like — technically, contractually, and in the data?" *Red flag:* removal is undefined, or the data leaves with the vendor. *Pass:* a clean uninstall path, and your measurement history stays yours.

## Scoring

**13–15 passes** — a vendor whose numbers you can reuse internally. **9–12** — proceed, with the failed checks written into the pilot as conditions. **Under 9** — the deck is marketing, not measurement.

## You just estimated. Want it measured?

Every tool here approximates. In twenty minutes on your analytics we compute the real distribution: your single-page rate, your depth by device and channel. No deck.

[Run my real numbers](/request-a-demo/)
