Skip to main content
Engagement

Answering insurance enquiries from the wording, and escalating everything else

The design settled what would never be automated before it settled what would. Everything after that was a question of evidence: what the answer was built from, and whether it could be shown.

Industry
Insurance
Delivery
Voice and digital channels, delivered as subcontracted scope inside a partner programme.
Results

Key Outcomes

GroundedJourney control

Reference design target: answers stay within approved policy and knowledge.

ContextualEscalation

Reference design target: handover carries the conversation record forward.

ContinuousQuality review

Reference operating target across automated and human review.

Baseline-ledOutcome measure

Reference design target: containment is balanced with correctness and effort.

The Challenge

The operating challenge

A general insurance contact centre in Australia was taking on more claims and policy contact than its roster could grow to meet, and most of that contact never needed a person. The operation was not short of technology.

It ran a mature contact centre platform, a maintained knowledge base and a workforce management tool that could forecast demand to the half hour, but none of it helped in the moment a customer was on the line. Agents opened every conversation cold, pulled context together from three systems by hand, and searched the knowledge base whilst the customer waited.

Waiting times had passed the point where customers were abandoning calls, and the calls being abandoned were mostly the straightforward ones: policy details, claim status, payment arrangements, document requests. Hiring was the obvious answer and the wrong one.

Recruitment and accreditation take longer than a seasonal spike lasts, so the operation was either carrying capacity it did not need or short of it exactly when correctness mattered most. In regulated financial services an overloaded queue is not only a service problem: it is the condition under which rushed and inconsistent answers get given about excesses, exclusions and claim outcomes, and those answers carry conduct consequences long after the queue has cleared.

What we found

  • A small group of enquiry types carried the bulk of contact. They were repetitive, fully documented in the policy wordings, and low regulatory risk when answered from the wording rather than from memory.
  • Demand was volatile in a patterned way rather than a random one. Weather events and end of financial year renewals produced sharp spikes the roster could not follow, whilst the mix of enquiries inside those spikes barely changed.
  • Agents were not slow. They were starting from nothing on every interaction, because the customer's history, the claim record and the applicable wording lived in three places and had to be assembled by hand before anyone could answer.
  • The floor for correctness sat higher than the ceiling for coverage. Quoting the wrong excess, misstating an exclusion or implying a claim outcome is a conduct issue, not a defect to fix in a later release, so any automated answer had to be traceable to the document it came from.
  • Risk and compliance held a veto over the design and had seen enough opaque machine learning to be ready to use it. An answer that could not be explained after the event was, for this operation, the same as a wrong answer.

Where the bottleneck sat

The constraint was not agent capability or headcount. It was that every interaction, routine or complex, began from an empty screen and had to be assembled by hand before it could be answered.

Automation aimed only at deflection would have relieved the queue and left that untouched for the calls that still reached a person.

Design rationale

The first decision on the engagement was not a technical one. Operations, risk and compliance sat in the same room and agreed which interactions would never be automated at all, ahead of any platform or model choice.

Hardship, complaints and anything carrying a vulnerability indicator came off the table permanently rather than sitting behind a confidence threshold that could be tuned later. That agreement shaped the build more than any subsequent piece of engineering.

It turned an open-ended automation question into a bounded one, and it meant the escalation path was designed first, rather than retrofitted around whatever the system turned out to be unable to handle.

What was at stake

Waiting times had passed the point where customers were abandoning calls, and the calls being abandoned were mostly the straightforward ones: policy details, claim status, payment arrangements, document requests.

Hiring was the obvious answer and the wrong one. Recruitment and accreditation take longer than a seasonal spike lasts, so the operation was either carrying capacity it did not need or short of it exactly when correctness mattered most.

In regulated financial services an overloaded queue is not only a service problem. It is the condition under which rushed and inconsistent answers get given about excesses, exclusions and claim outcomes, and those answers carry conduct consequences long after the queue has cleared.

Build

Our Solution

The first decision on the engagement was not a technical one. Operations, risk and compliance sat in the same room and agreed which interactions would never be automated at all, ahead of any platform or model choice.

Hardship, complaints and anything carrying a vulnerability indicator came off the table permanently rather than sitting behind a confidence threshold that could be tuned later. That agreement shaped the build more than any subsequent piece of engineering: it turned an open-ended automation question into a bounded one, and it meant the escalation path was designed first, rather than retrofitted around whatever the system turned out to be unable to handle.

What was built is an assist and containment layer across the existing voice and digital channels, shipped as three capabilities in order: a copilot for human agents first, then grounded containment for the narrow set of enquiries agreed in discovery, then a supervision layer sitting over both. The copilot went first deliberately, putting the frontline team in the position of critic before anything faced a customer, and producing weeks of labelled interaction data on real calls that made containment far more accurate once it was switched on.

Everything runs inside the insurer’s own cloud tenancy, alongside the data rather than moving the data out to reach it. The hardest part was grounding, not generation: policy wordings and product disclosure statements are written to be read by a person holding the whole document in their head, and a retrieved passage that is locally true and globally wrong does not look like a failure in a demo.

The fix was structural rather than a prompt change, restructuring the corpus so a clause and its qualifiers are retrieved together, and giving the evaluation model the same document set so it can mark an answer down for being incomplete rather than only for being unsupported.

Intent classifier

Sits in front of every voice and chat interaction and routes against the taxonomy built in discovery. In scope and low risk goes to the containment agent. Everything else goes straight to a person, with no attempt to have a go first.

Grounded containment agent

Answers only from retrieved evidence: policy wordings, product disclosure statements, internal procedures, and the systems of record for anything specific to a customer. Where the evidence does not support a confident answer it escalates rather than composing one.

Agent copilot

Assembles the customer's history, drafts a first response and surfaces the exact clause an agent needs, so the calls that do reach a person start from something rather than from nothing.

Supervision layer

A separate evaluation model scores every drafted response for groundedness, tone and policy compliance before a customer sees it. Low confidence and sensitive interactions are removed from automation rather than downgraded.

Escalation with context

A handover carries the transcript, the retrieved evidence, the citations and the reason for escalating into the agent’s screen. The customer does not repeat themselves and the agent does not begin again.

Governed model gateway

Model calls, retrieval and evaluation all route through one gateway inside the tenancy. One place to enforce policy, one place to log, one place to change a model without touching the application.

Decision trail

Every regulated interaction records what was asked, what evidence the answer was built from, how the evaluation scored it, and why it was contained or escalated. Audit reconstructs an interaction from the record, not from a reconstruction.

The line the system does not cross

It never settles a claim, never varies a policy and never makes a decision that binds the insurer. It resolves enquiries and prepares work. Anything with a financial or contractual consequence stays with an accountable person.

Grounding, not generation. Policy wordings and product disclosure statements are written to be read by a person holding the whole document in their head. Exclusions are qualified by clauses several pages away, and endorsements change the meaning of wording they never sit next to. The failure that mattered was a retrieved passage that was locally true and globally wrong, and it does not look like a failure in a demo. The fix was structural rather than a prompt change: the corpus was restructured so a clause and its qualifiers are retrieved together, and the evaluation model was given the same document set, so it can mark an answer down for being incomplete rather than only for being unsupported.

Outcome

What changed

  • Routine enquiries are answered when they arrive rather than when the queue reaches them, and each answer carries the wording it was built from.
  • Agents open a conversation with the history, the applicable clause and a drafted response already in front of them, instead of assembling context whilst a customer waits.
  • Escalations arrive with the transcript, the retrieved evidence and the reason for the handover attached, so customers stop repeating themselves at the point of transfer.
  • Hardship, complaints and vulnerability interactions reach a specialist team directly and with context prepared, rather than being triaged by whoever happens to answer.
  • Seasonal peaks are absorbed by the containment layer rather than by the roster, so the operation no longer chooses between carrying unused capacity and making customers wait when the weather turns.
  • New starters work from a default suggestion that encodes what experienced agents already do, instead of learning it by sitting next to someone who does.
  • Internal audit can reconstruct any automated interaction end to end from the record alone, which moved the risk function from gatekeeping the programme to sponsoring its extension.
Contact centreGeneral insuranceRetrieval groundingEscalation designAI governance

Ready to shipEnterprise AI?

Get the Executive Guide