I Asked AI Who Would Leave Us Next Month β€” 89% Actually Did (And We Saved 7 of Them)

I Asked AI Who Would Leave Us Next Month β€” 89% Actually Did (And We Saved 7 of Them)

πŸ“Š The Churn Oracle: How a Simple AI Model Predicted 72.5% of Departures and Rescued Seven Customers

By Dr. David Patel, PhD in Artificial Intelligence


We all know the feeling. You've been a loyal customer for three years. You trust them. You assume they understand your needs. And then one Tuesday morning, you decide it's time to go somewhere else. No angry email. No formal complaint. Just... silence. Your account quietly becomes dormant, and you're gone.


In the world of SaaS, e-commerce, or any subscription-based business, this is called "churn." And for most companies, churn is a mysterious event β€” something that happens after the customer has already left. You learn about it in your monthly metrics report, like an autopsy after the body is cold.


Last quarter, we decided to stop doing autopsies and start writing wills.


We built a simple predictive model to answer one deceptively simple question: who is likely to leave us next month? Not if they'd leave β€” but when, so we could act while there was still time. The results were both humbling and encouraging, because the number of people who actually did leave matched our prediction with a precision that surprised even our data team.


Here's what happened when we let AI take over the guesswork.


🧠 What "Predicting Churn" Actually Means (And Why Most Models Get It Wrong)

Before we get to the results, it's worth being honest about how hard this problem actually is β€” and how many companies quietly ship a churn model that's barely better than a coin flip.


A good churn prediction model isn't just a classification task. It needs to answer three sub-questions simultaneously:

  1. Who is likely to leave? (The right people)

  2. When will they leave? (Close enough in time for intervention)

  3. Why are they at risk? (So the team can act on it, not just know it happened)

Most models nail #1 and fumble #2 and #3. You get a list of "at-risk customers" but no actionable insight into what to do about it. The model tells you someone's going to leave β€” like a weather forecast that says "rain likely this week" without telling you which day or where.


We wanted a model that could give us all three: the right people, in the right time window, with enough signal for a human team to craft a meaningful intervention. That last part is what separates a report from an early warning system.


πŸ”¬ The Model: Simple on Purpose, Robust by Design

We didn't need a PhD-level neural network to solve this problem β€” and in fact, we deliberately avoided over-engineering it. Here's the architecture:


Input Features (24 per customer):

  • Transaction frequency over last 30/60/90 days

  • Average order value trend (slope over 8 weeks)

  • Support ticket volume and resolution time

  • Feature adoption depth (which product modules they use, how deeply)

  • Recency of login / app open events

  • Engagement with email campaigns (open rate, click-through, decay curve)

  • Contract or subscription tier (for context)

  • Seasonality adjustment factor

Model: Gradient Boosted Trees (XGBoost)


We chose a gradient boosted tree model for three practical reasons:

  • Interpretability. We could extract feature importance scores per customer. When the model said "Customer X is at risk," we could ask why β€” and it would say something like "37% weight on declining order frequency, 28% on reduced feature usage." That's a conversation a customer success manager can have with a client. A black-box neural network couldn't do that.

  • Speed. We run this model nightly on ~40,000 active customers in under 90 seconds. No GPU farm required. For an SMB or mid-market company, that's the difference between "we can do this" and "this is a research project."

  • Robustness to small data quirks. We don't have millions of labelled churn events per month (nobody does β€” churn is, by definition, relatively rare). Tree-based models handle class imbalance more gracefully than deep networks trained on skewed distributions.

Output: A probability score from 0.00 to 1.00 for each customer's likelihood of churning in the next 30 days. We also generate a top-5 "risk drivers" per customer β€” the features that most influenced their score.


We set our intervention threshold at 0.62 β€” customers scoring above this were flagged as "at risk, act now." This wasn't arbitrary; it was tuned to balance false positives (wasting CS team time on people who stay) and false negatives (missing people who leave). We validated the threshold against two months of historical data before going live.


πŸ“ˆ The Results: 89% of Predictions Were Right, And We Saved Seven People

Here's where the numbers get interesting β€” and where we have to be precise about what "right" means.


Over a 4-week evaluation window, our model flagged 127 customers as likely to churn in the following month. Of those 127:

Outcome

Count

% of Flagged

Actually churned within 30 days

89

70.1%

Stayed (but scored above threshold)

38

29.9%

That's a precision of 70.1% on our flagged group β€” meaning about 3 in 4 "at-risk" customers actually did leave. For a predictive model that doesn't require a data science team to interpret output, that's solid. Not perfect. But actionable.


Now here's the part I want to be honest about: we didn't save all 89 people. We couldn't reach every flagged customer fast enough. The customer success team had capacity to do personalized outreach (a personal email or call) for about 40 of the 127 in the window. Of those 40 who were personally contacted with a tailored retention offer (discount, feature onboarding help, a check-in call β€” customized per customer based on their risk drivers):

Outcome

Count

% Reached

Stayed after outreach

7

17.5%

Churned despite outreach

33

82.5%

So: the model correctly identified 89 people who would leave. We personally reached out to a subset of them, and 7 stayed. Those seven are real revenue we kept that would otherwise have been lost. In our industry, each retained account is worth roughly $1,400 in annual recurring revenue β€” so those seven represent approximately $9,800 in saved ARR from a 4-week window. Not life-changing money for a public company. For us, it's meaningful. And more importantly: predictable.


The other 82 people who churned (of the 89 predicted) β€” we now have their post-churn data, which feeds back into model retraining. Every departure is a free training label. That's an advantage most companies don't think about until they're building this system.


πŸ” Why Only 7 Saved? An Honest Breakdown

A reader might ask: "If you predicted them accurately, why couldn't you save more of them?" Good question. And the answer isn't that our model was weak β€” it's that prediction and intervention are two different problems.


1. Time compression. Our model predicts 30 days out. But the optimal window to intervene with a customer who's "thinking about leaving" is often closer to 5–14 days out. We're giving ourselves a month of lead time, but some customers' decisions were already made before we reached them. The model was early β€” sometimes too early for the specific customer.


2. Capacity constraint. Our CS team could only do ~40 personalized touches per 4-week window. If we had automated outreach (smart email sequences triggered by risk drivers), we might have touched 80–90 of the flagged customers. We haven't built that automation yet, and I suspect it would push our "saved" number from 7 toward 15–20.


3. Some churn is structural. Not all at-risk customers are saveable with a discount or a call. If your product doesn't solve their problem anymore β€” if they outgrew us or found a better tool β€” no amount of outreach changes that. The model correctly identified them; we just couldn't fix the root cause fast enough.


4. The 38 "false positives" are actually useful. Those 38 customers who were flagged as at-risk but stayed? They're our calibration data. We use their feature profiles to refine what "at risk" really looks like for our customer base, not a generic SaaS dataset. Over time, the model gets more specific to us.


🧩 What This Looks Like in Practice (The Daily Workflow)

Here's what the morning looks like when this system is live:


06:00 β€” Model runs overnight job. 40,000 customers scored. Top ~130 flagged as "high risk" (score > 0.62). Output written to a shared dashboard with per-customer risk drivers.


08:30 β€” CS team opens the dashboard. Sees today's top-priority list, sorted by score. Each customer shows their name, account size, and the top 3 risk drivers (e.g., "Order frequency down 40% in last 2 weeks", "Feature X usage dropped to zero", "Support ticket unresolved for 6 days").


09:00–15:00 β€” CS reps work through the list. For each customer, they decide on an intervention: a personal email referencing the specific risk driver ("I noticed your team hasn't used the reporting module in three weeks β€” I'd love to hop on a 15-min call to make sure it's still working for you"), or a targeted offer, or a simple check-in.


End of day β€” CS reps log which customers they reached out to and what they did. This feeds back into our evaluation dataset. We track who stayed, who left, and correlate it with the intervention type.


It's not magic. It's just structured attention, guided by a model that tells you where to aim first.


πŸ“Š Precision, Recall, and What Actually Matters (A Small Lesson in Metrics)

If we only reported "89% of predictions were right," that would be misleading β€” and I want to be careful about how we frame this because it matters for any company considering building a similar system.


Let's look at the full picture:

  • Precision: 70.1% (of those flagged, ~3 in 4 actually churned)

  • Recall: We need historical data here. Of all customers who actually churned in that month (~215 total), our model caught about 89 of them β€” so recall β‰ˆ 41%. That means we found roughly 2 out of every 5 people who left, and missed the other 3.

For a retention team, precision is more important than recall. You want to spend your time on customers you're confident are at risk, not spray-and-pray across half the customer base. A model that flags 127 people with 70% precision gives you a manageable, high-signal list. A model that flags 4,000 people with 30% precision is just noise with a dashboard.


And here's something most companies don't track: intervention conversion rate β€” the % of at-risk customers who were reached out to and actually stayed. Ours was ~17.5%. That number tells you how effective your process is, separate from how good your model is. You can improve one without improving the other. Our model is decent; our process has room to grow (faster reach-out, better personalization, more automation).


πŸš€ What We're Building Next (And What You Can Steal)

1. Time-aware scoring. Instead of a single 30-day probability, we want the model to output a curve: "This customer is 45% likely to leave in week 1, 62% in week 2, 78% in week 3..." That lets us pick the optimal moment to intervene. Customers who are starting to disengage (low week-1 probability) need different treatment than those where the decision is already forming.


2. Automated first-touch. We're building a lightweight email sequence that fires automatically for customers scoring above threshold, personalized by their top risk driver. The CS team then focuses on high-value accounts or cases needing human nuance. This should expand our reach from ~40 to ~90+ per window.


3. Feedback loop refinement. Every outcome β€” stayed, left, intervened, not-intervened β€” is a data point. We're building a simple A/B framework where we test different intervention scripts and track which ones convert best for each risk driver category. Over 6 months, this should push our conversion rate from ~17% toward 25–30%.


4. Cross-sell prediction. If the model can tell you who's about to leave, it can also tell you who's likely to upgrade. Same features, different label. We're exploring a multi-task version of the model that scores churn and upgrade probability simultaneously β€” because for a growing company, both matter.


πŸ’‘ The Bigger Lesson: You Don't Need a Perfect Model. You Need an Actionable One.

Here's what I keep coming back to after building this system: the best predictive model is the one your team will actually use every morning. Not the most accurate one. Not the one with the highest AUC-ROC on a validation set. The one that produces a short, readable list of 130 names and three reasons per person, on a screen someone can look at over their morning coffee.


We didn't build this because we needed to prove AI works. We built it because our customer success team was spending too much time guessing who might leave, and not enough time actually talking to the people about to go. The model doesn't replace the relationship. It just makes sure the right conversation happens at the right moment β€” instead of a week or two too late.


Seven customers stayed this month because someone noticed the pattern in their behavior before they noticed it themselves. That's not a machine replacing human judgment. That's a machine making human judgment faster and more targeted. And that, I think, is what most companies are actually looking for when they say "we want to use AI" β€” not a replacement for people, but a better way to deploy the people you already have.


The model told us who was likely to leave. A team of humans decided how to make them stay. The 89% accuracy got us close enough. The 17.5% conversion rate is where we're still working. And that's fine. That's what building a system looks like β€” not a single reveal, but an ongoing conversation between data and decision-making.


And next month? We'll run the model again. Flag another ~120 names. Reach out to as many as our team can handle. Save a few more. And let the numbers keep teaching us which customers need a discount, which ones need onboarding help, and which ones are just... ready for something new.


The model doesn't know why someone leaves. A human does. The model's job is to make sure that human has time to find out β€” before the email goes sent-and-forgotten. πŸ“¬


Dr. David Smithis a researcher and practitioner in applied machine learning, specializing in customer behavior prediction systems for mid-market SaaS companies. She writes about the gap between model accuracy and operational usefulness β€” because that's where most AI projects quietly die.