← Back to blog

PostHog Case Study: Reading About 400 MILA Session Replays With AI to Find Where the Funnel Broke

PostHog Case Study: Reading About 400 MILA Session Replays With AI to Find Where the Funnel Broke

Watching session replays one by one is a good way to understand a single user, but it only scales to a handful of sessions. When a funnel leaks in several places, what you need is to see many sessions together and know which screen each problem keeps showing up on. That is what we did for MILA Stories, running PostHog's AI agent over the recordings of their purchase funnel. This note covers what we found, what MILA's team changed because of it, which numbers moved, and what the recordings could not see.

MILA makes collaborative books and videos to give as gifts for an occasion. One person organizes and invites family, friends or coworkers over WhatsApp, and each of them sends audio, photos and text that get assembled into a gift, in print or digital format. Their site sums it up as "many voices, one book." The underlying question for MILA was whether growth was held back by the product or by distribution, and answering it meant looking closely at what people actually did inside the funnel.

About 400 recordings in a single pass

We ran PostHog's AI agent over about 400 session recordings from the printed-book funnel and cross-checked them against the events from the previous 30 days. The difference from watching replays by hand is scale: instead of picking a few sessions and drawing conclusions from what they showed, the AI read all of them in one pass, and the events told us how much each pattern weighed in the total.

What came up was not a single drop-off point but four bottlenecks in series:

  • Half of visitors left on the first screen, before touching a field.
  • The date screen jammed badly enough that one person spent more than 6 minutes trying to enter a date.
  • The "why?" step asked for something creative written from a cold start, and people typed, deleted and left.
  • At checkout there were frustration clicks and failed payments.

Each of these is the kind of finding a recording shows and an event funnel doesn't. The funnel tells you which step people drop at. The recording shows someone typing and deleting an answer, or spending more than 6 minutes on a date. Read together, the sessions stopped being anecdotes and described the whole funnel.

From diagnosis to redesign, and what can be attributed to it

MILA's team redesigned the funnel around the steps the analysis flagged and landed on three screens: date and format (digital or print) together, with the "why?" as an optional step. Our part was the diagnosis and the before-and-after measurement.

Comparing 12 days before to 12 days after, with 6% less Meta spend, the share of people who saw the price and then clicked "Pay" went from 34% to 59%. That is the number that can be attributed to the redesign. Over the same period sales grew 2.6x (+157%) and cost per sale fell 64%, but those two numbers need careful reading: new landing pages and ads went live in those same days, and part of the improvement at the top of the funnel comes from that.

Checkout friction, read through recordings and events

The fourth bottleneck, checkout, deserved a closer look. The recordings and events showed something counterintuitive: the people who fought hardest with the payment screen were the ones who ended up buying, at 1 rage click per session versus 0.1 among those who reached "Pay" and didn't buy. PostHog counts a rage click as three rapid clicks in the same spot, and its own documentation notes that this is a heuristic based on proximity and timing that is worth confirming in recordings, which is exactly what we did. Errors, on the other hand, were low across the whole funnel, so the problem was not site failures but a wall of friction at the payment step.

One concrete fix came out of that. Stripe's one-click payment button, Link, accounted for almost 3 in 4 declines. It was removed, and sales were not hurt.

A bigger change followed: in one of its markets, MILA moved from Stripe to a local processor. PostHog did not trigger that decision. What it did was point to payment as the friction zone and, once the change was made, measure the result. The first week against the one before, sales doubled and cost per sale fell 47%, with 3% less spend on that campaign and no change in traffic. The barrier was mechanical, not about the product, and that was the first strong evidence that MILA's problem was distribution.

What recordings can't see

The most useful lesson of the case about the method came from a mistake the AI made. When we asked PostHog to summarize abandoned-checkout sessions, 98% of the summaries said the form had not loaded. That was false.

The payment form lives in an iframe from another company, and PostHog can't record iframes from domains you don't control, with third-party embeds like Stripe Checkout among the examples in its documentation. In the recording that space shows up as an empty element. The AI described what the recording showed, a gap where the form should be, and drew the wrong conclusion from it.

This doesn't take away from the value of reading recordings with AI, but it sets the boundaries. The AI summarizes what the recording sees, and the analyst's job is to know what it can't see. Clicks and navigation on your own pages are recorded; what happens inside a third-party form is not. For that stretch, PostHog's documentation suggests using the provider's own analytics or sending its results to PostHog as events.

Marketing questions answered with a query

The speed of this case was not limited to the first read. We connected PostHog to our AI agent through MCP, so a new question gets answered with a query instead of a week of spreadsheet exports. Two examples from this project:

  • How much WhatsApp sells: in one of the weeks we analyzed, more than half of buyers had chatted with the seller before paying, versus 1 in 7 the week before.
  • Why one market buys digital and another buys print: by crossing the occasion date with the payment date, we saw that most digital buyers had the occasion 2 weeks out, and the printed book couldn't arrive in that window.

What made the measurement trustworthy

Reading the funnel is of little use if the numbers that measure the outcome can't be trusted, so we also cleaned up the measurement. In the PostHog data warehouse we brought together site and app events with Stripe, the local processor, Meta Ads and Google Ads, and added WhatsApp data, which came in through a different route. The rule that came out of cross-checking everything was to count a sale as money collected, not as a browser "purchase event."

That rule surfaced two problems. Meta saw 1 in 20 sales, and its ROAS looked 7 times lower than the one derived from collections. And the dashboards the client was looking at counted a single processor, to the point that one week they showed only 3% of what had been collected.

We also record each sale from the server and send it to Meta through the Conversions API using a PostHog destination. In the first days measured, Meta went from crediting 1 in 20 sales to 3 in 5. The next step, an expected-value signal for each funnel step, is running in an A/B test against the classic campaign and doesn't have a result we can share yet.

What to check when you analyze your own product with recordings

For any product with a multi-step funnel, we recommend starting with these questions:

  • Are you reading recordings at scale or a handful? A few hand-picked sessions show cases; many read together show patterns and the screen where they repeat.
  • Are you cross-checking recordings against events? The recording explains what happens to one person, and the events tell you how many run into the same thing.
  • What part of your flow can't the recording see? If checkout or a form lives in a third-party iframe, neither the replay nor an AI summary will show what happens inside, and that stretch is better measured with events.
  • Does your before-and-after comparison isolate the change? It helps to use windows of the same length, note the spend in each period, and list what else shipped in those days, so you know how much of the improvement belongs to what.

And one more, on the measurement side: it is worth counting sales as money collected rather than a browser event, because that is the number any redesign gets judged against.

For a broader map of what PostHog offers, see our note on what to look at first in an ecommerce store. MILA's case has its own page in our case studies.

This content was developed with AI assistance and reviewed by the Zenda team. The bad ideas are 100% ours.