News: Avesta Labs is exhibiting at The AI Summit Australia 2026, 7–9 September at the MCEC in Melbourne. Meet us there

Life Insurance Direct

Case study · Life insurance · Australia

They stopped choosing which calls to listen to.

Life Insurance Direct has been comparing life cover for Australians since 2006, across nine insurers and six product types, and almost all of it happens on the phone. We built the system that reads every one of those calls, scores both sides, and hands the team a review queue in risk order instead of a sample.

Client
Life Insurance Direct
Sector
Life insurance distribution, Sydney
Status
Live in production
What we built
Metricsense, call quality and compliance QA
Book an AI Kickoff

Change the product. The job stays the same.

  • Mortgage brokers
  • Intermediaries
  • Life insurance distributors
  • Advice licensees
  • Any regulated contact centre

If the phone call is where your business gets sold, gets serviced and gets audited, then the small sample you review by hand is the part of your business you have quietly decided not to look at.

Read first. Choose second.

Before

5steps

A capable, careful process, run the only way it could be run by hand. Pull the reports, work out which calls look worth hearing, pick a few per adviser, listen end to end, write it up. Nothing wrong with any single step. The limit is what they add up to.

  1. 1Pull the call reports
  2. 2Work out which calls look risky
  3. 3Pick a handful per adviser
  4. 4Listen to each one end to end
  5. 5Write up the scorecard

The call that most needed a listen was almost never the one that looked worth picking.

After

2steps

Every call is read and scored first. Only then does a person decide what to open. The high risk ones first, then the middle band, then the ones that are simply an opportunity to do something better.

  1. 1Every call read and scored
  2. 2Review in risk order

No longer done by hand

  • Pull the call reports
  • Work out which calls look risky
  • Pick a handful per adviser
  • Listen to each one end to end

Customer experiences they would not have picked up before are now the ones getting picked up.

Before the build

Why not just listen to more calls?

The cheapest option was the one they were already running, and running properly. So the honest question was never whether hand picking works. It was what happens to the calls it never reaches.

Do the arithmetic

Around five hundred calls a week, so roughly two thousand a month. A careful manual process reaches somewhere between two and five of every hundred. Reading the rest by hand is not one more reviewer. It is a department.

The lag is the other half

Feedback reached an adviser three to four weeks after the call it was about. By then they had taken hundreds more, the same way, with the same habit.

A sample can only rank what you already suspect

Risk based selection is still selection. It sorts the calls a person thought to look at, which means the one nobody thought to look at stays exactly where it was.

Listening for words is not reading a call

The things worth checking here are boundaries inside a conversation, not phrases. Whether an explanation was complete. Whether the customer was led. Keyword spotting cannot see any of that.

The two things they asked for

  • Cover more calls without growing the QA team.
  • Keep a person deciding. Always.

And underneath both, a change of direction most vendors in this market never make. The brief moved off whether an adviser might cross the compliance line, and onto whether they stop dead at it, say what they cannot do, and lose a customer who only wanted to understand what they were buying.

So the brief was never compliance policing. It was to see the whole floor clearly enough to coach it, which is a different product and a much harder one.

Why the bar was set where it was

What was at stake.

Life Insurance Direct holds an Australian financial services licence, and the phone call is where the whole business happens. It is the sale, it is the service, and it is the thing a regulator would read if it ever came to that. So a system that reads every call is not a reporting tool. It is something with an opinion about a licensed business, running across a hundred per cent of its most sensitive material. That set the bar, and the bar was mostly about restraint.

Every score had to be:

Pinned to the moment

Every score points at the exact place in the transcript that produced it, with the audio beside it.

A candidate, not a verdict

A flag is a candidate issue with evidence attached. A person confirms it. The system reports nothing to anyone on its own.

Inside their own cloud

It runs in their own account and their own region, with personal information stripped before any transcript reaches a model.

Theirs alone

Nothing they send trains a model shared with anyone else, and no client's calls go anywhere near another client's.

Concepts to production

How we worked.

  1. 01First

    Concepts, then scripts

    The client called this order himself and he was right. The universal checks go first, because they run on every call from day one. Call type scripts layer on top of something already working, rather than being the thing everything waits for.

  2. 02A few days each

    Their QA became the checks

    Their own scripts and standard procedures came across marked up three ways. Said word for word, covered in your own words, or left to the adviser. Each checkable item became one plain English instruction the system runs on every call.

  3. 03Weekly, still running

    They marked it

    Their QA lead reviewed calls beside the system and told us where it was right and where it was not, with the reason. That loop is the whole method. We tuned against real disagreements rather than imagined ones.

  4. 04Ongoing

    The library grows

    Reviewing calls turns up something worth watching. That becomes a new concept, written as one sentence in plain English rather than a retrain, and from then on it runs on every call. The loop closes in days, not quarters.

What the first round changed

It was scoring the block. They wanted it to score what came after.

An adviser who stops at "I cannot give you advice" is compliant, and has also just ended a conversation the customer wanted to have. Blocking still has to happen. What we changed was what gets measured: not whether the adviser stopped, but whether they followed the stop with what they can do, and whether the customer left understanding the product.

That is the difference between compliance policing and compliant selling, and it is worth being clear that we did not think of it. The client did. Our job was to notice that it changed the product, and rebuild the checks around it.

The hard part

The obvious build was the wrong one.

The hardest check they asked for was a comparison. Did the specialist read the mandatory sections of the insurer's application as written, did they skip a question, did they lead the customer. The obvious way to build that is to hand a model the application and the transcript and ask it to compare them.

Why that fails

A long document does not get read evenly

Give a model a whole application form and attention lands on the front, the second page and the end. The middle gets skimmed. On an insurance application the middle is where the medical questions sit, which is exactly the part that has to be checked word for word. So the risk of a confident wrong answer is highest precisely where being right matters most.

What we did instead

Give it less, and give it the right part

The check never sees the document. It sees one marked line range, with the personal information already removed, and a single purpose built check runs against that section and nothing else. The client's QA lead put it better than we did: the less we give it, the less it makes up.

On a compliance check a confident wrong answer costs more than no answer. A QA team that opens ten flags and finds eight were fine stops opening flags, and at that point the whole thing is an expensive dashboard.

The thing I have enjoyed the most is the responsiveness and the willingness to challenge yourselves and how you might have deployed a solution. Open to challenging us as well, and giving us different options to consider, and options that we never even thought of.
Russell Cain, Chief Executive Officer

In production

What it does now.

Reads both sides

Every call is scored on the customer side and the adviser side, so a call can be flagged for how it was handled, not only for how it ended.

Risk order, not date order

The queue is ranked. The highest risk calls get a person first, then the middle band, then the ones that are simply worth a look.

A new check is one sentence

Checks are written in plain English. No tagging exercise, no retraining. Deploy one and it runs across every call from then on.

Every score links to the moment

Open a score and it lands on the exact point in the transcript that produced it, with the audio playing alongside.

A quality index per adviser

A composite score out of a hundred, built from accuracy, satisfaction, resolution, compliance and communication. The spread inside one person matters more than the number on the front, because that is what tells you which axis to coach.

In their cloud, in their region

The whole thing runs inside their own account, in Australia, with personal information stripped before any transcript reaches a model.

Where each person is on their training journey

Not just how someone is performing, but which areas they are weaker in than the rest of the floor. Onboarding for a new starter and refresher training for an experienced one both get aimed at the specific thing, rather than the whole team sitting through the same session.

What it deliberately does not do.
A flag is a candidate issue with evidence attached, not a finding. A person confirms it, and the team kept their own scorecards and their own judgement. Metricsense reports nothing to a regulator, and nothing to anyone else, on its own.

The screen they open in the morning.

Every call already read and already sorted. The only decision left is where to start.

DashboardsRisk Review Queue

Sample data
Last 90 Days
Apply
High Risk3

0a3f71c2-8e04-4b1d-9f27-5c60a884e113

Adviser A · 04 Nov, 09:12

Not reviewed
T13/3
T31/3
T41/1
T24/7T2.53
Topic 2.3Topic 2.5

6b18d40f-2ca7-4e59-8d13-b0947fe2a5c6

Adviser D · 04 Nov, 11:47

Not reviewed
T13/3
T32/3
T41/1
T25/7T2.53
Topic 2.4

c92e5a77-31bd-4a08-9c64-72f1e0d3b845

Adviser B · 03 Nov, 15:26

Not reviewed
T13/3
T31/3
T41/1
T24/7T2.53
Topic 3.3
Medium Risk9
Low Risk34
Not Applicable4

A recreation of the real screen. Every call, adviser and count in it is constructed.

Keep scrolling

The columns fill left to right as you scroll, which is the order a reviewer works in. Not Applicable is the fourth column for a reason: a call where nothing was triggered still got read.

And the one that decides who gets coached on what.

The tree on the left is the real library of checks, at its real depth and with its real counts. The names are held back at the request of the client, so read the shape rather than the labels. Several of these checks exist only because this client asked for them, including the reframe from earlier on this page, which sits in the product as a check of its own. Nothing off the shelf ships with that.

DashboardsAgent Compliance Scorecard

Sample data

Agent Scorecard

Concept ↓ · Adviser →Adviser AAdviser BAdviser CAdviser DAdviser EAdviser FAdviser GTotal
Topic 1(3)
Topic 1.1clean
Topic 1.2112 med
Topic 1.3
Topic 2(7)21121
Topic 2.1
Topic 2.22 med
Topic 2.32 med
Topic 2.42 med
Topic 2.5(3)11
Topic 3(3)
Topic 3.1clean
Topic 3.2
Topic 3.3
Topic 4(1)
Topic 4.1clean
Call1588654450

The concept tree is the real library. The advisers are masked and every state in the grid is constructed.

Read a row and you are looking at one concept across the whole floor. Read a column and you are looking at one person's training plan.

Scope

What we left out on purpose.

The version of the application check that runs automatically on every application call is not built. It stayed a pilot on two marked up forms. Everyone assumed there was a clean structured record to check the call against, and there was not: their customer record holds around ten fields where the application itself collects a hundred. So the shortcut did not exist, and the honest options were a narrow pilot or a wide guess.

A half connected version would have looked like a feature and behaved like a guess, on the one check where a guess does the most damage. It stays a pilot until there is a way to do it that we would defend in front of their regulator, not just in front of them.

9,218

calls read, both sides

Every call between 16 April and 10 August 2026, scored on the customer side and the adviser side. Not a sample of them. All of them.

2 to 5%

what hand picking reached

Their own coverage before this, running a careful risk based process. The gap between that and the line above is the entire argument.

26

checks on every call

Four topic groups, running from the advice boundary through to how the call itself was handled, each check written as one plain English instruction. The library grows whenever a review turns up something worth watching.

0

findings reported on its own

Every flag is a candidate issue with the evidence attached, and a person decides. The system is a triage instrument. It is not an adjudicator.

Life Insurance Direct
Russell Cain, Chief Executive Officer, Life Insurance Direct
Before Metricsense, checking calls was a manual process: we selected files for review. Now the platform analyses every call, and we review by risk. It has identified customer experiences we would not have picked up before, and it lets us scale our quality assurance program without increasing headcount, with our team still the human in the loop. It helps us protect our customers, our business and our agents.

Six months from now

The failure here is quiet.

Nobody notices a QA system going wrong. A check gets tightened, a model gets swapped, and nothing looks broken. What happens instead is that a reviewer opens three flags, finds all three were fine, and slowly stops opening them. So before any check changes, it runs against a set of calls the client has already reviewed and marked themselves.

  1. 1Do the checks still fire where the reviewer said they should?
  2. 2Has anything new fired on a call a reviewer already cleared?
  3. 3Does every score still open on the moment that produced it?
  4. 4Is the flag still a candidate, with a person deciding?
  5. Goes live

Without that marked set, every change means re-reviewing the whole library by hand. Which in practice means it does not get re-reviewed, and confidence goes quietly rather than loudly.

What it exists to catch

Early on, positive findings were turning up inside the high risk box. Tightening the rule cleaned the queue and cost recall somewhere else, which is a real trade and not a bug fix. The marked set is how that trade gets made deliberately and measured, rather than discovered later by a reviewer who has quietly stopped trusting the queue.

He would put it less finally than we just did

Asked how it was going, he said the platform continues to need training, and that they are in a good position to keep training it and improving it as it matures further into the life cycle of the business. That is the right description and we are not going to write a tidier one over the top of it. This is a system being taught a business, and the loop above is what makes that a plan rather than a hope.

Start with fifty of your own calls.

A short discovery session, then a pilot on a sample of your real calls with your own checks configured. You score us against your own QA before you commit to anything. If the precision is not there, you have lost a fortnight and learned something useful about your sample.

Book an AI Kickoff

Or email hello@avestalabs.ai