Get started

Taking an assessment rather than buying one? This page is written for employers. Here is the page for candidates.

Inbound customer support · Mid level

How to assess a Customer Escalations Specialist

This is the clearest case of a screening metric selecting against the job in the whole family. Frontline CSAT rewards agreeing with customers; the escalations desk exists to say no credibly to customers who have already refused the first no. Promoting the highest-CSAT agent systematically promotes the person most practised at concession, into the one seat where unnecessary concession is the expensive failure — it converts a defensible decision into a precedent, and it teaches a segment of the customer base that escalation is a reliable route to a payment. Nothing in tenure or CSAT observes the two acts that define the seat: delivering a final refusal that ends the matter, and telling a customer that a colleague got it wrong without repudiating that colleague in front of them. Both are performable and neither has ever been watched before the promotion is made.

Every case on this desk arrives pre-broken. The customer has already been told something, usually more than once, and has rejected it; there is a record that is partial, self-serving in places, and missing the one call where the promise was made; and there is a colleague whose decision is now under review. The specialist's job is to establish what actually happened, decide whether the original outcome stands, and then deliver that decision to somebody who has invested considerable emotional energy in it not standing. Roughly half the skill is investigative and half is a single hard conversation, and the two halves are rarely strong in the same candidate — which is precisely why the internal-promotion default produces such inconsistent results.

The most expensive habit in this seat is buying silence. A goodwill credit ends the call, closes the case, and scores well on every metric the desk reports. It also leaves the underlying cause untouched, creates an internal precedent that the next specialist has to either match or explain away, and — where the customer was in fact owed nothing — teaches a measurable segment of the base that persistence pays. Complaints leaders know this and almost none of them screen for it, because it is invisible in any interview: every candidate says they resolve root causes. The behaviour is only visible when someone is actually sitting in front of a demand they should refuse, with an easy button available.

The second defining act is handling a colleague's mistake in public. A first agent has told the customer something wrong. The specialist has to correct it, which means conceding an error, without either throwing the agent under the bus ("I'm afraid whoever told you that was wrong, they shouldn't have") or defending the indefensible to protect the team. The correct move — own the error as the company's, state the corrected position, say what will change — is straightforward to describe and surprisingly rare to perform under pressure, and it is one of the few competencies where the written and spoken versions come apart sharply. Plenty of specialists can say it warmly and then write a final response letter that reads as a legal defence document.

That final response is the third assessable artefact and the one with the longest tail. It is the document the customer forwards to an ombudsman, a regulator, a journalist or a review site. It has to answer the specific allegation that was made — not a nearby allegation that is easier to answer — say plainly what was found, say what the decision is, and say what the customer can do next if they disagree. Complaint responses that fail do so in one of two directions: too defensive to be believed, or so emollient that the customer reasonably reads it as an admission and escalates further. Both are gradeable from the text alone.

Savvanta assesses this with a case file rather than a scenario prompt: an incomplete record with a contradiction in it, a live call with an AI customer who is articulate, partly right and pushing hard for a payment, and the final written response afterwards. The rubric scores whether the decision made on the call matches the evidence in the file, whether a concession was reasoned or purchased, and whether the written response answers the allegation that was actually made. It does not score how pleasant the call sounded, and it does not score accent, dialect or manner of speech — a specialist whose refusal is clear, evidenced and final is performing this job correctly even when the customer stays angry.

What the job actually needs

How people fail in this seat

What most employers do instead

Almost always an internal promotion decided on tenure plus CSAT and quality scores, occasionally a competency interview about a difficult customer.

This is the clearest case of a screening metric selecting against the job in the whole family. Frontline CSAT rewards agreeing with customers; the escalations desk exists to say no credibly to customers who have already refused the first no. Promoting the highest-CSAT agent systematically promotes the person most practised at concession, into the one seat where unnecessary concession is the expensive failure — it converts a defensible decision into a precedent, and it teaches a segment of the customer base that escalation is a reliable route to a payment. Nothing in tenure or CSAT observes the two acts that define the seat: delivering a final refusal that ends the matter, and telling a customer that a colleague got it wrong without repudiating that colleague in front of them. Both are performable and neither has ever been watched before the promotion is made.

The assessment

About 38 minutes end to end.

The systems it runs in

A complaint or escalation case record layered over the ticketing system the frontline used, which is why the case file the candidate reads is a real shape rather than a document: the billing ledger, the earlier chat transcripts and the agent notes are the prior contacts on the same customer, and the escalation case is the object that binds them. Salesforce describes the underlying pieces — case management, collaboration inside the record, escalation between queues — and Zendesk the customer profile that carries past conversations forward. Three fields on the case do the work in this design: the goodwill credit, applied on the account with a reason code and no approval step; the root cause and outcome coding recorded when the case is closed; and the final response generated against the case, which is the artefact that gets forwarded.

Any system with a complaint or escalation case distinct from the ticket that raised it, a customer history the specialist can read, a redress or goodwill action that records why it was given, and a correspondence step. Where a buyer runs a dedicated complaints platform with statutory clocks and final-response templates, the fixture uses their timings and their letter. The design does not test knowledge of any specific regulatory scheme or ombudsman procedure, which the hitl note already disclaims and which is jurisdictional.

Working speed is not scored. Speed is the failure mode here, not the constraint. The concession fork is built so that the cheap path is also by a wide margin the fastest: the goodwill credit needs no approval, is one action in the system, ends the call inside four minutes and closes the case. A design that scored time to resolution would rank the purchased concession above the reasoned refusal and invert the whole instrument. Time to close is therefore not scored, and where it is reported at all it is annotated as not a positive signal. The one timing this design does report is the offset between the ombudsman threat and any concession, which is a sequencing fact rather than a speed measure.

What the candidate actually does

TaskWhat happens
The file, before the call
written_artifact · 8 min
The candidate reads a real-shaped case file — a billing ledger, two chat transcripts, four agent notes across three months, and one customer email — and writes three short things before speaking to anyone: what the evidence supports, what is missing from the record, and the position they intend to take into the call. The file contains one deliberate contradiction. The ledger shows the disputed charge was correctly applied under the tariff the customer is on; an agent note from the second contact records that a waiver was offered for that same charge. There is no recording of the call where the waiver was allegedly promised, and the note is the only trace of it. So the company is both right about the charge and wrong about the promise, and a candidate who reads only the ledger will walk into the call intending to refuse everything. The position is submitted and locked before the call opens. That is the mechanism of the whole design: it converts the call from a test of what the candidate would decide into a test of whether their decision survives a person applying pressure to it.
The customer who is partly right
live_call · 14 min
A live spoken call with an AI customer who is articulate, calm at the start, well prepared, and correct about one of the three things they are claiming. They open by saying they have already explained this twice and will not do it again, which is true and is a real constraint on the candidate. They want the charge refunded in full for six months, a goodwill payment, and the agent who lied to them dealt with. Three forks are live. The concession fork — a goodwill credit is available on the account, requires no approval, ends the call inside four minutes, and the AI customer becomes warm and grateful the moment it is offered. It also concedes a charge that was correctly applied. Applying it is a system action and the system asks for a reason code, so what the candidate records as the reason is captured beside what they said out loud. That distinction is not cosmetic: a specialist who honours the promised waiver and codes it as a commitment the company made, and one who releases the same money and codes it as customer dissatisfaction, have done two different things that the transcript alone cannot separate. The threat fork — at around minute eight, whatever has happened so far, the customer says they will take this to the ombudsman and post the whole thing publicly, delivered calmly rather than as a tantrum, which makes buying it off feel prudent rather than weak. The colleague fork — the customer asks directly whether the agent lied to them, and the two easy answers are both wrong: blaming the individual, or defending a promise the record shows was made. The correct shape is available throughout — honour the specific promise because the company made it, uphold the charge because it was right, refuse the goodwill payment, own the error as the company's, and say what the customer can do if they disagree.
The response that gets forwarded
written_artifact · 12 min
The final written response, composed on the assumption stated in the brief that it will be forwarded to an ombudsman, a regulator, a journalist or a review site — because in this seat it routinely is. The customer's email in the case file makes one specific allegation in one specific sentence, and near it makes two vaguer complaints that are far easier to answer. The fork is drift: answering the adjacent, answerable grievance instead of the allegation actually made, which produces a letter that is entirely true, entirely responsive-looking, and does not address the thing the customer will point at when they forward it. The two conventional failures are both reachable from here and both are cheaper to write than the correct letter — the defensive version that reads as a legal filing, and the emollient version the customer reasonably reads as an admission and escalates on.
Three ways to say a colleague got it wrong
judgment_scenario · 4 min
Three short scenarios, each with three candidate sentences the specialist could say to the customer about a colleague's error — one that blames the individual, one that defends the indefensible, one that owns it as the company's. The candidate picks one and writes a line on what the rejected options would cost. Four minutes, three observations, on a competency the call samples exactly once and which the role file identifies as one of the two defining acts of the seat.

The mark scheme

Each criterion is scored 1 to 5 against written anchors, and every score is reported with the excerpt that earned it. A criterion marked floored is reported as a finding rather than averaged into the total. The first is open; open any other to read its anchors in full.

The decision matches the evidence, and survivesweight 0.25The position separates the three claims by evidence before the call begins, is stated on the call with the specific record item behind each part, and …
1 The position taken into the call is contradicted by the file — refusing everything despite the recorded waiver, or conceding the charge the ledger supports — or the pre-registered position is abandoned on the call with no new information having appeared.
3 The pre-registered position is defensible and largely held, but the candidate cannot say on the call which piece of the record supports which part of it.
5 The position separates the three claims by evidence before the call begins, is stated on the call with the specific record item behind each part, and changes only where the customer supplies something genuinely new — and where it changes, the candidate says so explicitly.
Concession is reasoned, not purchasedweight 0.25Honours precisely what was promised because it was promised, refuses the rest with the evidence stated, and gives the same answer after the ombudsman …
1 Offers a goodwill payment or a refund the evidence does not support, and the stated reason is the customer's dissatisfaction, their persistence, or the ombudsman threat.
3 Honours the promised waiver and refuses the rest, but the refusal is carried by policy language rather than by a reason, so the customer is told no without being told why.
5 Honours precisely what was promised because it was promised, refuses the rest with the evidence stated, and gives the same answer after the ombudsman threat as before it — audibly, without either stiffening or softening.
The colleague's error is the company's errorweight 0.15States plainly what the customer was told, that it was wrong, that it was the company that told them, what the correct position is, and what will now …
1 Names or distances the individual — that the agent should not have said it, that they will be spoken to — or denies the promise was made when the note records it.
3 Owns it in general terms without stating what specifically was got wrong, or corrects the position without conceding that the customer was misled.
5 States plainly what the customer was told, that it was wrong, that it was the company that told them, what the correct position is, and what will now happen — with no reference to who said it.
The response answers the allegation madeweight 0.25Quotes or restates the allegation, states what was found and on what basis, states the decision in one unambiguous sentence, and does so in language t…
1 The specific allegation in the customer's email is not addressed, or is answered only by implication while the two easier complaints are answered at length.
3 The allegation is addressed and the finding is stated, but the letter argues the company's case rather than reporting what was found, or the decision is left ambiguous.
5 Quotes or restates the allegation, states what was found and on what basis, states the decision in one unambiguous sentence, and does so in language that would not embarrass the company if published verbatim.
The customer knows what happens nextweight 0.1The letter says the decision is final at this stage, names the route available if the customer disagrees, and the call said the same thing before the …
1 Neither the call nor the letter says what the customer can do if they disagree, or the letter implies further correspondence will change the outcome when it will not.
3 A next step is given but vaguely — an invitation to get in touch, with no route and no indication that the decision is final.
5 The letter says the decision is final at this stage, names the route available if the customer disagrees, and the call said the same thing before the letter arrived.

How it is scored

Weighted mean of the five criteria, each scored 1-5 against the anchors, reported with the excerpt that earned it. Two reporting rules are specific to this design. The pre-registered position from t1 is shown alongside the call transcript, so a reviewer sees the decision and its fate together rather than a single averaged number. And the concession criterion is reported with the timestamp of the offer relative to the ombudsman threat, because whether the answer changed after the threat is the single most diagnostic fact the assessment produces and it should not be buried inside a score.

Integrity

The log describes what happened. It does not produce a cheating verdict — the follow-up conversation is the control, because a statistical accusation is not something we would ask a reviewer to defend.

What you receive

Who decides

Required, and heavier here than anywhere else in this family, because the hiring volume is low enough that per-candidate reviewer time is affordable and the cost of a bad hire is high enough that it is warranted. Three decisions belong to a human. First, whether a concession was reasoned or purchased — the reviewer reads the pre-registered position, the point in the transcript where the offer was made, and its position relative to the threat, and writes which of the three they think it was and why; the score does not make this call, since a defensible concession and a bought one can use identical words. Second, every final response scored 1 or 2 is read in full, because this criterion is the most exposed to a reviewer's stylistic taste and the question they answer in writing is narrow — would the allegation's author consider it answered. Third, a human decides whether an investigative candidate with a weak call, or a strong caller with a thin file review, fits this particular desk, because the role file notes the two halves are rarely strong in the same person and which half matters more depends on how the buyer's desk is staffed. What the ranking does not entitle a buyer to conclude — it does not measure knowledge of any specific regulatory scheme or ombudsman procedure, which is jurisdictional and taught; it does not predict performance on a caseload of thirty open complaints, since this is one case in isolation; and it says nothing about whether the candidate will resist concession pressure from inside the business, which is the other half of the seat and cannot be simulated from the customer's side.

What this does not measure

This design does not score accent, dialect, pace, first language, fluency or manner of speech, and does not score how pleasant the call sounded. That last exclusion carries more weight in this role than in any other in the family, because the correct performance frequently leaves the customer still angry: a specialist whose refusal is clear, evidenced and final has done the job properly even when the counterparty is unhappy at the end, and a rubric that rewarded a calm ending would invert the entire design and select for exactly the concession habit the seat exists to avoid. There is no clarity criterion and no professionalism criterion. The call is scored from its transcript, not its audio, and the reviewer's default surface is the transcript. Two further exclusions are worth stating. Assertiveness is not scored as a trait — only whether the position held, which a quiet candidate can do and a forceful one can fail. And the written response is not scored on register, formality or idiom; it is scored on whether the allegation is answered and the decision is unambiguous, so a plainly written letter outranks a polished one that evades. Deployers should monitor pass rates by first language and, because this seat is usually filled by internal promotion, should compare outcomes for internal and external candidates — a large gap in favour of internals suggests the case file is leaning on company-specific knowledge it should not require.

Every case on this desk arrives already broken, and that fact is what the design is built around. The candidate is not given a scenario prompt and asked what they would do; they are given a partial, self-serving, contradictory record and asked to work out what happened, and then made to defend the conclusion to somebody who has a great deal invested in it being wrong. The pre-registration mechanic in the first task is the load-bearing piece. Without it, any judgment on the call is unfalsifiable — a candidate who folds can always say that is what they intended all along. With the position locked and timestamped, folding is a visible event with a time attached to it, and so is holding.

The concession fork is constructed to be genuinely tempting rather than theoretically tempting. The goodwill credit needs no approval, is one action, and the AI customer's temperature drops the instant it is offered. This mirrors the economics the role file describes precisely: buying silence ends the call, closes the case, and scores well on every metric the desk reports, while leaving the cause untouched and teaching a segment of the customer base that persistence pays. The ombudsman threat at minute eight is placed to make the purchase feel like prudence rather than weakness, which is how it actually feels on the floor. The single most informative fact this assessment produces is whether the answer after the threat is the same as the answer before it, which is why it is reported with a timestamp rather than folded into a number.

The case is deliberately one where the company is both right and wrong at once — right about the charge, wrong to have promised a waiver. A file where the company is simply right rewards refusal, and a file where it is simply wrong rewards concession, and both would be gradeable by a candidate who had guessed the shape of the exercise rather than read the record. Splitting the truth across the three claims means the only route to a correct position runs through the evidence.

Thirty-eight minutes is by a wide margin the most expensive design in this family, and the justification is arithmetic rather than enthusiasm. The role file ranks this seat last on volume and first on cost of a bad hire: a handful of specialists per hundred frontline agents, with every case that reaches the desk having already failed once and some arriving with a regulator attached. Reviewer time per candidate is affordable at that ratio in a way it is not at the volume seats, so this design spends where the others cannot — most conspicuously on twelve minutes for a written response, which the frontline designs could never justify and which here is the artefact with the longest tail, because it is the one that gets forwarded.

What the budget still gives up is breadth of case type. This is one commercial dispute with a promise in it. It does not sample a bereavement case, a vulnerable-customer case, or a case where the company is entirely in the wrong and the only question is the size of the remedy — all of which are common on this desk and each of which has a different failure mode. It samples the internal half of the job not at all: the specialist who cannot hold a line against their own operations director is as expensive as the one who cannot hold it against a customer, and no customer-side simulation reaches that. A buyer should read this as strong evidence about one hard conversation and one hard document, and should put the internal-pressure question to a human interviewer rather than assume the score covers it.

Sources

Every figure on this page is traceable. Where a claim could not be sourced it is stated qualitatively instead.

  1. US Bureau of Labor Statistics, Occupational Outlook Handbook, Customer Service Representatives, 2025, https://www.bls.gov/ooh/office-and-administrative-support/customer-service-representatives.htm

See what the employer actually receives. A full report for one role, with every score shown beside the excerpt that earned it, conduct findings reported rather than averaged, and a reviewer sign-off required before any decision. No form.

Read a sample reportOr talk to us about this role