0a3f71c2-8e04-4b1d-9f27-5c60a884e113
Adviser A · 04 Nov, 09:12
Not reviewedNews: Avesta Labs is exhibiting at The AI Summit Australia 2026, 7–9 September at the MCEC in Melbourne. Meet us there →
AVESTA LABS
CAPABILITY
INSIDE LABS
Case study · Life insurance · Australia
Life Insurance Direct has been comparing life cover for Australians since 2006, across nine insurers and six product types, and almost all of it happens on the phone. We built the system that reads every one of those calls, scores both sides, and hands the team a review queue in risk order instead of a sample.
Change the product. The job stays the same.
If the phone call is where your business gets sold, gets serviced and gets audited, then the small sample you review by hand is the part of your business you have quietly decided not to look at.
Before
5steps
A capable, careful process, run the only way it could be run by hand. Pull the reports, work out which calls look worth hearing, pick a few per adviser, listen end to end, write it up. Nothing wrong with any single step. The limit is what they add up to.
The call that most needed a listen was almost never the one that looked worth picking.
After
2steps
Every call is read and scored first. Only then does a person decide what to open. The high risk ones first, then the middle band, then the ones that are simply an opportunity to do something better.
No longer done by hand
Customer experiences they would not have picked up before are now the ones getting picked up.
Before the build
The cheapest option was the one they were already running, and running properly. So the honest question was never whether hand picking works. It was what happens to the calls it never reaches.
Around five hundred calls a week, so roughly two thousand a month. A careful manual process reaches somewhere between two and five of every hundred. Reading the rest by hand is not one more reviewer. It is a department.
Feedback reached an adviser three to four weeks after the call it was about. By then they had taken hundreds more, the same way, with the same habit.
Risk based selection is still selection. It sorts the calls a person thought to look at, which means the one nobody thought to look at stays exactly where it was.
The things worth checking here are boundaries inside a conversation, not phrases. Whether an explanation was complete. Whether the customer was led. Keyword spotting cannot see any of that.
The two things they asked for
And underneath both, a change of direction most vendors in this market never make. The brief moved off whether an adviser might cross the compliance line, and onto whether they stop dead at it, say what they cannot do, and lose a customer who only wanted to understand what they were buying.
So the brief was never compliance policing. It was to see the whole floor clearly enough to coach it, which is a different product and a much harder one.
Why the bar was set where it was
Life Insurance Direct holds an Australian financial services licence, and the phone call is where the whole business happens. It is the sale, it is the service, and it is the thing a regulator would read if it ever came to that. So a system that reads every call is not a reporting tool. It is something with an opinion about a licensed business, running across a hundred per cent of its most sensitive material. That set the bar, and the bar was mostly about restraint.
Every score had to be:
Every score points at the exact place in the transcript that produced it, with the audio beside it.
A flag is a candidate issue with evidence attached. A person confirms it. The system reports nothing to anyone on its own.
It runs in their own account and their own region, with personal information stripped before any transcript reaches a model.
Nothing they send trains a model shared with anyone else, and no client's calls go anywhere near another client's.
Concepts to production
The client called this order himself and he was right. The universal checks go first, because they run on every call from day one. Call type scripts layer on top of something already working, rather than being the thing everything waits for.
Their own scripts and standard procedures came across marked up three ways. Said word for word, covered in your own words, or left to the adviser. Each checkable item became one plain English instruction the system runs on every call.
Their QA lead reviewed calls beside the system and told us where it was right and where it was not, with the reason. That loop is the whole method. We tuned against real disagreements rather than imagined ones.
Reviewing calls turns up something worth watching. That becomes a new concept, written as one sentence in plain English rather than a retrain, and from then on it runs on every call. The loop closes in days, not quarters.
What the first round changed
An adviser who stops at "I cannot give you advice" is compliant, and has also just ended a conversation the customer wanted to have. Blocking still has to happen. What we changed was what gets measured: not whether the adviser stopped, but whether they followed the stop with what they can do, and whether the customer left understanding the product.
That is the difference between compliance policing and compliant selling, and it is worth being clear that we did not think of it. The client did. Our job was to notice that it changed the product, and rebuild the checks around it.
The hard part
The hardest check they asked for was a comparison. Did the specialist read the mandatory sections of the insurer's application as written, did they skip a question, did they lead the customer. The obvious way to build that is to hand a model the application and the transcript and ask it to compare them.
Why that fails
Give a model a whole application form and attention lands on the front, the second page and the end. The middle gets skimmed. On an insurance application the middle is where the medical questions sit, which is exactly the part that has to be checked word for word. So the risk of a confident wrong answer is highest precisely where being right matters most.
What we did instead
The check never sees the document. It sees one marked line range, with the personal information already removed, and a single purpose built check runs against that section and nothing else. The client's QA lead put it better than we did: the less we give it, the less it makes up.
On a compliance check a confident wrong answer costs more than no answer. A QA team that opens ten flags and finds eight were fine stops opening flags, and at that point the whole thing is an expensive dashboard.
“The thing I have enjoyed the most is the responsiveness and the willingness to challenge yourselves and how you might have deployed a solution. Open to challenging us as well, and giving us different options to consider, and options that we never even thought of.”
In production
Every call is scored on the customer side and the adviser side, so a call can be flagged for how it was handled, not only for how it ended.
The queue is ranked. The highest risk calls get a person first, then the middle band, then the ones that are simply worth a look.
Checks are written in plain English. No tagging exercise, no retraining. Deploy one and it runs across every call from then on.
Open a score and it lands on the exact point in the transcript that produced it, with the audio playing alongside.
A composite score out of a hundred, built from accuracy, satisfaction, resolution, compliance and communication. The spread inside one person matters more than the number on the front, because that is what tells you which axis to coach.
The whole thing runs inside their own account, in Australia, with personal information stripped before any transcript reaches a model.
Not just how someone is performing, but which areas they are weaker in than the rest of the floor. Onboarding for a new starter and refresher training for an experienced one both get aimed at the specific thing, rather than the whole team sitting through the same session.
A flag is a candidate issue with evidence attached, not a finding. A person confirms it, and the team kept their own scorecards and their own judgement. Metricsense reports nothing to a regulator, and nothing to anyone else, on its own.
Every call already read and already sorted. The only decision left is where to start.
0a3f71c2-8e04-4b1d-9f27-5c60a884e113
Adviser A · 04 Nov, 09:12
Not reviewed6b18d40f-2ca7-4e59-8d13-b0947fe2a5c6
Adviser D · 04 Nov, 11:47
Not reviewedc92e5a77-31bd-4a08-9c64-72f1e0d3b845
Adviser B · 03 Nov, 15:26
Not reviewedThe columns fill left to right as you scroll, which is the order a reviewer works in. Not Applicable is the fourth column for a reason: a call where nothing was triggered still got read.
The tree on the left is the real library of checks, at its real depth and with its real counts. The names are held back at the request of the client, so read the shape rather than the labels. Several of these checks exist only because this client asked for them, including the reframe from earlier on this page, which sits in the product as a check of its own. Nothing off the shelf ships with that.
Agent Scorecard
| Concept ↓ · Adviser → | Adviser A | Adviser B | Adviser C | Adviser D | Adviser E | Adviser F | Adviser G | Total |
|---|---|---|---|---|---|---|---|---|
| Topic 1(3) | ||||||||
| Topic 1.1 | clean | |||||||
| Topic 1.2 | 1 | 1 | 2 med | |||||
| Topic 1.3 | ||||||||
| Topic 2(7) | 2 | 1 | 1 | 2 | 1 | |||
| Topic 2.1 | ||||||||
| Topic 2.2 | 2 med | |||||||
| Topic 2.3 | 2 med | |||||||
| Topic 2.4 | 2 med | |||||||
| Topic 2.5(3) | 1 | 1 | ||||||
| Topic 3(3) | ||||||||
| Topic 3.1 | clean | |||||||
| Topic 3.2 | ||||||||
| Topic 3.3 | ||||||||
| Topic 4(1) | ||||||||
| Topic 4.1 | clean | |||||||
| Call | 15 | 8 | 8 | 6 | 5 | 4 | 4 | 50 |
Read a row and you are looking at one concept across the whole floor. Read a column and you are looking at one person's training plan.
Scope
The version of the application check that runs automatically on every application call is not built. It stayed a pilot on two marked up forms. Everyone assumed there was a clean structured record to check the call against, and there was not: their customer record holds around ten fields where the application itself collects a hundred. So the shortcut did not exist, and the honest options were a narrow pilot or a wide guess.
A half connected version would have looked like a feature and behaved like a guess, on the one check where a guess does the most damage. It stays a pilot until there is a way to do it that we would defend in front of their regulator, not just in front of them.
9,218
calls read, both sides
Every call between 16 April and 10 August 2026, scored on the customer side and the adviser side. Not a sample of them. All of them.
2 to 5%
what hand picking reached
Their own coverage before this, running a careful risk based process. The gap between that and the line above is the entire argument.
26
checks on every call
Four topic groups, running from the advice boundary through to how the call itself was handled, each check written as one plain English instruction. The library grows whenever a review turns up something worth watching.
0
findings reported on its own
Every flag is a candidate issue with the evidence attached, and a person decides. The system is a triage instrument. It is not an adjudicator.

“Before Metricsense, checking calls was a manual process: we selected files for review. Now the platform analyses every call, and we review by risk. It has identified customer experiences we would not have picked up before, and it lets us scale our quality assurance program without increasing headcount, with our team still the human in the loop. It helps us protect our customers, our business and our agents.”
Six months from now
Nobody notices a QA system going wrong. A check gets tightened, a model gets swapped, and nothing looks broken. What happens instead is that a reviewer opens three flags, finds all three were fine, and slowly stops opening them. So before any check changes, it runs against a set of calls the client has already reviewed and marked themselves.
Without that marked set, every change means re-reviewing the whole library by hand. Which in practice means it does not get re-reviewed, and confidence goes quietly rather than loudly.
What it exists to catch
Early on, positive findings were turning up inside the high risk box. Tightening the rule cleaned the queue and cost recall somewhere else, which is a real trade and not a bug fix. The marked set is how that trade gets made deliberately and measured, rather than discovered later by a reviewer who has quietly stopped trusting the queue.
He would put it less finally than we just did
Asked how it was going, he said the platform continues to need training, and that they are in a good position to keep training it and improving it as it matures further into the life cycle of the business. That is the right description and we are not going to write a tidier one over the top of it. This is a system being taught a business, and the loop above is what makes that a plan rather than a hope.
A short discovery session, then a pilot on a sample of your real calls with your own checks configured. You score us against your own QA before you commit to anything. If the precision is not there, you have lost a fortnight and learned something useful about your sample.
Book an AI KickoffOr email hello@avestalabs.ai