---
title: We Ran the Same Analysis Five Ways and Got Answers From +32% to −4%
description: Same sites, same metric, same data — five defensible measurement windows produced engagement lifts from +32% to −4%. How window-shopping manufactures results, and the three-check protocol that catches it.
url: https://preview.artificialpoets.com/blog/engagement-lift-window-shopping/
site: Artificial Poets
type: post
date: 2026-08-10T03:32:09+00:00
modified: 2026-08-14T18:30:48+00:00
categories:
  - Measurement
---
# We Ran the Same Analysis Five Ways and Got Answers From +32% to −4%

## What is window-shopping in analytics?

A pre/post analysis needs four dates: where the baseline starts, where it ends, where the post-period starts, and where it ends. Every one of those is a choice, every choice moves the answer, and almost nobody discloses which they made.

The insidious part is that no single choice looks wrong. Starting the baseline in August rather than September is defensible. Ending the post-period in April rather than May is defensible. But run the grid of defensible choices and you get a *distribution* of answers — and whoever presents one number from that distribution, without telling you the rest, has made a decision on your behalf.

## How we got five answers from one dataset

We measured pageview change on two publisher titles that had enabled our Platform, against comparison titles in the same network, across a fixed intervention date. Then we varied only the windows.

| Baseline starts | Post-period ends | Treatment-vs-comparison gap |
| --- | --- | --- |
| September | late January | +32.2pp |
| October | late January | +29.8pp |
| August | late January | +24.7pp |
| September | late March | +18.7pp |
| September | late April | +11.4pp |
| September | mid May | +4.4pp |
| August | mid May | −3.7pp |

Twenty combinations in the full grid. The gap ranged from **+32.2 points to −3.7 points, including a sign flip.** An analyst with a target could have honestly produced almost any story from this data — "strong lift", "modest lift", "no effect", even "slightly negative" — by moving nothing but the calendar.

Three mechanics drive the swing.

**Seasonality asymmetry.** If the treatment and comparison groups have different seasonal shapes — one spikes in December, the other in spring — then which months sit inside your window decides who looks better. That is a property of the sites, not of the intervention.

**Endpoint sensitivity.** A single unusual week near a window boundary carries enormous leverage over a period average. Move the boundary one week and the number jumps.

**Panel composition.** If any site's tracking degrades inside the window — a migration, an outage, a retag — the group containing it craters and the other group "wins" mechanically. This is worth its own article, and it has one: [the comparison group that went dark](#).

## The number we rejected

The +32-to-−4 metric was raw pageview lift, and we rejected it entirely — not averaged, not "triangulated", rejected. A number that moves 36 points under defensible window choices is not measuring the intervention. It is measuring the calendar.

That rejection had a cost. Pageview lift was the biggest, most saleable number available, and early window choices produced headline figures we had already drafted around. Killing it meant publishing smaller numbers. It was still the right call, because the alternative was publishing a number any competent analyst on the buyer's side could reverse with a one-week window shift.

## What a robust number looks like

The same sweep that killed pageviews validated a different metric. Pages per session, measured as the treatment-versus-comparison gap, behaved completely differently under the same torture test:

- **Across every baseline-start choice, it moved less than a third of a point** — 19.6 to 20.1pp on the full post-window. August, September, October, November: same answer.
- **Across post-period lengths it declined smoothly** — from ~34pp measured over the first weeks to ~20pp over the full period. That is not window artifact; that is the effect itself attenuating after launch, which is real, expected, and honest to report as a decay curve rather than hide by picking the early window.

That is the distinction that matters. Window-dependence you cannot explain means the metric is broken. Window-dependence you *can* explain — a launch effect settling toward run rate — is a finding, and the honest report shows both ends of it. Ours does: the launch-quarter figure and the steady-state figure, side by side, in everything we publish.

## The three-check protocol

Before believing any pre/post number — a vendor's, an agency's, your own team's — run these:

1. **Sweep the baseline start.** Recompute with the baseline starting one and two months earlier and later. If the answer moves more than a couple of points, the metric has a seasonality problem the presenter has not handled.
2. **Sweep the post-period end.** Recompute ending the post-period at each available month. A smooth decline is attenuation — fine, report it. A jump or sign change at a specific endpoint means something happened at that endpoint; find it before quoting anything.
3. **Demand the grid, or the reason there isn't one.** Anyone presenting a lift should be able to show what the number does under adjacent windows. "We chose this window because…" is an acceptable sentence only when it ends with a reason that isn't the result.

If a presenter cannot survive these three, the conversation is over — and if they volunteer the sweep before you ask, you have found the rare vendor whose numbers you can reuse internally.

- Same data, five defensible windows: treatment gap from **+32.2pp to −3.7pp**, including a sign flip
- Full window grid: **20 combinations**, one metric rejected outright
- The surviving metric moved **<0.3pp** across every baseline choice
- Its only window-dependence was a smooth **~34pp → ~20pp** decay — the launch effect settling, reported openly

## FAQ

No. Most window-shopping is unconscious — an analyst tries a few ranges, the best one "feels representative", and confirmation does the rest. The output is identical to deliberate manipulation, which is why the protocol tests the number, not the intent.

Because launch quarters are real and larger. That is legitimate *if* the steady-state figure sits next to it. Our own launch-quarter figure is roughly four times our run rate on one title — both are published, and the run rate is the one to plan against.

Long enough to cover both groups' seasonal shape (ideally the same months pre and post), bounded by data-integrity events — migrations, retags, outages — not by which endpoint flatters the result.

## Evaluating a vendor's lift right now?

The Vendor Claim Audit Checklist is the fifteen-point version of this article — what to ask, what a red flag looks like, and what a pass looks like, for every claim in the deck.

## See what your session depth looks like

We read your analytics with you for twenty minutes and tell you what share of your sessions stop at the first page. You leave with the annotated read, whether or not we ever talk again. 20 minutes, your analytics, no deck.

[Book my session-depth read](/request-a-demo/)
