ITSMAugust 4, 2026· 7 min read

The Watermelon Effect: When Every SLA Is Green and Everyone Is Unhappy

Green on the outside, red on the inside. The watermelon effect is what happens when service level agreements measure process compliance instead of whether anyone was actually helped, and it is a design flaw, not a reporting error.

The Watermelon Effect: When Every SLA Is Green and Everyone Is Unhappy

Key takeaways

  • The watermelon effect describes SLAs reporting green while user satisfaction is red
  • It is caused by measuring the handling of work rather than the outcome of it
  • Ticket-splitting, premature resolution and clock-stopping are rational responses to the measurement, not misconduct
  • The fix is to measure outcomes users recognize: resolution that stuck, effort expended, requests that never needed to exist

Green on the outside. Red on the inside.

The watermelon effect is the state of a service desk whose SLA dashboard is comprehensively green while the people it serves are comprehensively unimpressed. Every target met, every report clean, and an executive who keeps hearing complaints that the data insists are not happening.

It is tempting to treat this as a reporting problem: the wrong things are being counted, so count better things. That is part of it. But the deeper issue is that the discrepancy is not an accident of measurement. It is the predictable output of measuring process compliance in an environment where people are rewarded for the measurement.


What the SLA is actually measuring

A typical service level agreement commits to things like: respond within 30 minutes, resolve priority-3 incidents within two business days, maintain 95% compliance across the month.

Every one of those is a statement about the handling of a work item. None is a statement about whether a person got what they needed.

Those two things correlate, loosely, when the desk is healthy. Under pressure they decouple, and they decouple in a specific direction, because the work item is what is measured and the person is not.


The behaviours it produces

None of these require anyone to act in bad faith. Each is a locally rational response to what is being counted.

Ticket splitting. One user problem with three components becomes three tickets. Each closes inside target. The user's actual issue took eleven days and three separate conversations.

Premature resolution. Resolve the ticket, tell the user to reopen if the problem persists. The clock stops. The user, reasonably, opens a new ticket instead: which starts a fresh clock and reports as a new, promptly-handled request.

Clock management. "Pending customer" pauses the SLA. A question asked at 4:45pm on Friday buys three days of stopped clock. The status is accurate. The effect is to make waiting invisible.

Priority deflation. If P1 targets are tight and missing them is career-relevant, marginal incidents become P2. The P1 compliance figure improves and nothing about the service does.

Optimizing the measured segment. Attention flows to tickets approaching breach, which means work is prioritized by clock position rather than by impact. A trivial ticket at 90% of target outranks a serious one logged this morning.

The uncomfortable summary: your compliance figure improves as the measurement decouples from reality. A perfectly green dashboard, sustained over a long period, in an organization that is nonetheless unhappy, is evidence of the decoupling rather than against it.


Why the usual fixes do not work

Tightening the targets increases the pressure that produced the gaming. Compliance will hold. The behaviours will intensify.

Adding a satisfaction survey helps only if anyone acts on it. Attached to a green dashboard, a low CSAT score reads as a survey problem: response bias, unrepresentative sample, people only respond when annoyed.

Auditing for gaming treats a design flaw as a conduct issue. The people splitting tickets are responding correctly to the incentive you installed. Punishing them produces more careful gaming.


What to measure instead

The principle: measure things a user would recognize as describing their own experience.

Time to actual resolution, per user problem, not per ticket. Link split tickets and measure the span from first contact to the point the person stopped having the problem. This one change removes most of the value of ticket splitting.

Repeat contact rate. Did they come back about the same thing within a fortnight? A resolution that did not stick was not a resolution. This is the single most effective counterweight to premature closure.

Customer Effort Score. How much work did the user have to do to get helped? Effort correlates with service desk quality more tightly than satisfaction does, and it is harder to game.

Requests that never needed to exist. The strongest signal available, and invisible in any SLA framework: because the best outcome for a user is not a fast ticket, it is no ticket. Across our platform roughly 70% of requests are resolved before they are ever tickets; those never appear in an SLA report at all.

Ageing distribution rather than compliance percentage. "97% compliant" conceals the 3%. If those are the same twelve users every month, you have a systemic failure reported as rounding.


Keep the SLA

None of this argues for abolishing service level agreements. They serve a real purpose: they set expectations, they create accountability, and in commercial relationships they are contractually necessary.

The argument is narrower. An SLA is a floor, not a definition of good. It describes the worst acceptable performance, and a dashboard confirming you are above the floor tells you nothing about whether you are anywhere near the ceiling.

Report both. Put the compliance figure next to repeat contact rate and effort score, and watch what happens to the conversation when one is green and the others are not. That gap is the actual state of your service, and it is far more useful management information than either number alone.

Last updated on August 27, 2026

See the agentic service desk in action

Watch Rezolve.ai autonomously resolve real IT and HR tickets: governed, auditable, glass-box.

Book a demo

Frequently asked questions

What is the watermelon effect in ITSM?

A service desk reporting green across every SLA while users are dissatisfied: green outside, red inside. It arises when SLAs measure how work items are handled rather than whether people got what they needed, and the two decouple under pressure.

What causes the watermelon effect?

Measuring process compliance in an environment that rewards the measurement. Ticket splitting, premature resolution, clock-stopping via 'pending customer', priority deflation and prioritizing by clock position are all locally rational responses to the metric, not misconduct.

How do you fix the watermelon effect?

Measure outcomes a user would recognize: time to resolution per user problem rather than per ticket, repeat contact rate, Customer Effort Score, and requests that never needed to exist. Keep the SLA as a floor, but report it alongside those measures rather than alone.

Shano K. Sam
LinkedIn ↗

Get service-desk AI insights in your inbox

Practical guidance on agentic AI for IT and HR support: one email, no spam.

By submitting, you agree we may use the details you’ve provided to contact you. See our Privacy Policy.