The message that becomes a one-star review almost always arrived first as a slightly terse email nobody flagged.
TL;DR Use it to flag for human attention, never to respond. Tune it to over-flag. Track the trend across your whole book, not individual scores. It misses sarcasm and politeness-masking-anger, so it supplements judgement rather than replacing it.
What it is for
Flagging, not responding.
Legitimate uses
- Surfacing an upset customer in a busy inbox.
- Prioritising the response queue by tone as well as by age.
- Spotting a trend across many conversations that no individual would notice.
- Reviewing a technician’s written communications for coaching.
Illegitimate uses
- Automatically replying based on detected sentiment.
- Scoring individual staff for performance management.
- Any decision affecting a customer taken without a human reading the message.
Tune it to over-flag
The cost asymmetry is stark.
False positive: a person reads a message that was fine. Cost: thirty seconds.
False negative: an upset customer is not flagged, does not get called, and leaves a review. Cost: substantial and public.
Set the sensitivity high. Ten unnecessary flags a week is a fair price for catching one that mattered.
Review the misses specifically. When a complaint arrives that was not flagged, go back and look at the earlier messages. That is how you improve the configuration.
Early frustration signals
Some of these are detectable by rules alone, without AI, and rules are more reliable.
Rule-based flags worth building
| Signal | Why it matters |
|---|---|
| Second message before a reply | They are chasing. Always a flag |
| “Still waiting”, “as I said”, “again” | Explicit frustration |
| All caps or multiple question marks | Obvious |
| Mentions of a refund, a review, or “elsewhere” | Immediate escalation |
| Legal or regulatory words | Immediate escalation |
| Message after 10pm | Often indicates distress or urgency |
| Formality increasing suddenly | A customer who was casual and is now formal is preparing a complaint |
The formality shift is the most underrated signal and it is easy to miss without a system watching for it.
Rules first, AI on top. Rules are transparent, explainable and free. Use AI for the ambiguity the rules cannot catch.
The honest limitations
Be clear about what it gets wrong.
- Sarcasm. Reads as positive, routinely. “Great, another day off work waiting” scores well.
- Politeness masking anger. Very common in some cultures and dialects. Extremely polite complaint messages score neutral and are often the most serious.
- Brevity. Short messages carry little signal. A one-word “Fine.” is deeply ambiguous.
- Cultural and regional variation. Directness that reads as hostile in one region is ordinary in another.
- Technical language. Fault descriptions contain negative words and get scored negative.
The last one produces most false positives in a trade business. Tune for it or you will flag every message containing “broken”, “leak” and “failed”.
Escalation flow
When something flags
- Alert to a named person, within minutes.
- Human reads the whole thread, not the flagged message alone.
- Human decides: real, or a false positive.
- If real, ring them. Do not email.
- Log the outcome, which improves the configuration.
Never let the system send anything. An automated “we noticed you seem frustrated” message to someone who is frustrated is worse than silence.
Trend analysis, which is the real value
Individual flags are useful. Aggregate trends are more so.
Track monthly
- Proportion of conversations flagged negative.
- Flags by service type.
- Flags by technician, for coaching only.
- Flags by stage. Booking, during, after, invoicing.
- Common themes in flagged messages.
The stage breakdown is the most actionable. If flags cluster after invoicing, you have a pricing communication problem. If they cluster during the job, you have a communication cadence problem. The aggregate tells you something no individual message does.
Staff coaching, done carefully
Reviewing written communication for tone is legitimate and useful.
Rules that keep it constructive
- Coaching only. Never linked to pay or discipline on sentiment scores alone.
- Tell the team it is happening, and what it is used for. Covert monitoring of staff communications is both a trust problem and, in many jurisdictions, a legal one.
- Show examples, good and bad, rather than reporting a score.
- Consider the context. A technician dealing with an abusive customer will score badly and should not be penalised for it.
Sentiment scores measure the conversation, not the person.
Data protection
Customer communications contain personal data.
- Know where the analysis happens. If messages are sent to a third-party API, that is a processor relationship with obligations attached.
- Check your privacy notice covers it.
- Set a retention period and honour it.
- Do not feed customer messages into a general-purpose tool that may use them for training, unless the terms explicitly prevent it.
Check the terms of any tool you use for exactly this point. It is the most commonly overlooked risk in adopting AI on customer data.
Measure it
- Complaints caught before escalation, versus after.
- Time from first negative signal to human contact.
- False positive rate, tracked to tune sensitivity.
- Public reviews from customers who were flagged and contacted, versus flagged and missed.
- Negative flag rate over time, which is your real quality trend.
Build the rule-based flags this week before touching AI at all. Second message before a reply, plus the refund and review keywords, will catch most of what matters and requires no tooling beyond your inbox filters.
Need a pro to set it up? [BOOK A CALL]
