Get started

Taking an assessment rather than buying one? This page is written for employers. Here is the page for candidates.

Inbound customer support · Entry level

How to assess a Voice Customer Service Representative

The phone screen is the only conversation in the entire process, and it is the one conversation where the other party is being deliberately agreeable. At no point does anybody watch this candidate be told they are wrong by someone who is angry and partly right. The typing test measures throughput, which workforce management measures directly and accurately from week one, so it buys information the employer already gets for free. A personality inventory predicts a disposition, not a decision: the failure that costs money here is not a temperament, it is a specific verbal act — saying "I'll have that sorted by Friday" with no basis, at minute six, because the silence was unbearable. That act is trivially observable on a recorded call and invisible everywhere else.

The unit of work is one stranger, one problem, one conversation. A queue item arrives, the person on the other end has already spent time being annoyed before the call connected, and the representative has somewhere between four and twelve minutes to work out what actually went wrong, decide what the business will commit to, say so in a way the customer accepts, and leave a record the next person can pick up cold. Everything surrounding that — the product, the CRM, the disposition codes, the tone-of-voice guide — is trainable inside a fortnight. The part that is not trainable in a fortnight is commitment discipline: what you are willing to promise to somebody who is upset with you right now.

Two structural facts should shape how this seat is screened, and both cut the same way. First, the hiring is replacement hiring at enormous scale in a shrinking occupation, so cost and speed per candidate are the binding design constraints. A ninety-minute assessment will not be run at 289,500 openings a year; it will be quietly skipped by the recruiter under pressure to fill. Second, the easy contacts are the ones being deflected to self-service and automation first, which means the residue arriving at a human is systematically harder than the average contact these teams were hiring against five years ago. The role is getting narrower and more demanding while the headcount falls. A screen built around the old job description — polite, fast, follows the script — selects for the part of the work that is disappearing.

What separates the top quartile is legible on a transcript and almost nowhere else. The median performer answers the question the customer asked. The top quartile notices when the question the customer asked is not the problem the customer described, and says so before spending eight minutes solving the wrong thing. The median performer fills silence with reassurance; the top quartile says the smaller true thing — "I can see the refund was raised, I can't see when it will land, and here is who can tell you by Thursday" — and is willing to sit through the customer's displeasure at that being the answer. The median performer escalates when the customer demands it; the top quartile escalates when the call stops being winnable, which is usually two or three turns earlier and is a judgment nobody has ever asked them to demonstrate before their first week live.

The written half matters more than most job descriptions admit. The NICE survey of 400 contact centre leaders found 79 percent say their agents handle multiple channels concurrently on the same shift most of the time or every shift, so the case note is not paperwork after the real work — it is half the real work, and it is the artefact that determines whether a repeat contact starts from where the last one ended or from zero. A note that records "customer frustrated, advised of process" is a note that costs the business a second call.

A live simulated call reproduces this loop rather than proxying it: an AI customer with a specific grievance who escalates if over-promised and calms if given a bounded, honest answer, followed by the written note the customer and the next agent actually receive. Two things this deliberately does not measure, and they should be stated to any buyer in writing. It does not score accent, dialect, pace, first language or manner of speech — these are protected-characteristic proxies with no established relationship to resolution quality, and in an industry with this much offshore delivery the risk of building an accent filter under a competency label is real and specific. And it does not score warmth as a personality trait; it scores whether the commitment made on the call survives into the note, which is a behaviour and not a temperament.

What the job actually needs

How people fail in this seat

What most employers do instead

CV sift for prior contact-centre tenure, a friendly twenty-minute phone screen, sometimes a typing or grammar test, sometimes an off-the-shelf personality inventory.

The phone screen is the only conversation in the entire process, and it is the one conversation where the other party is being deliberately agreeable. At no point does anybody watch this candidate be told they are wrong by someone who is angry and partly right. The typing test measures throughput, which workforce management measures directly and accurately from week one, so it buys information the employer already gets for free. A personality inventory predicts a disposition, not a decision: the failure that costs money here is not a temperament, it is a specific verbal act — saying "I'll have that sorted by Friday" with no basis, at minute six, because the silence was unbearable. That act is trivially observable on a recorded call and invisible everywhere else.

The assessment

About 22 minutes end to end.

The systems it runs in

A ticketing agent workspace with a softphone attached: the call arrives on the ticket, the customer record sits in the same window, and the ticket carries both an internal note the next agent reads and a customer-facing contact summary that goes out. Zendesk describes exactly this shape — a unified agent workspace, macros as pre-defined responses to common issues, customer profiles capturing past conversations, and voice among the channels brought into one place. Salesforce calls the same object a case and the workspace a Service Console. The fixture is built on the generic shape rather than on any one vendor's menus: ticket, record, macro, internal note, public summary.

Any ticketing or case system with a queue, a customer record, macros, and a separation between internal notes and customer-facing replies. Nothing in the rubric scores navigation of a particular menu, and the design says so explicitly under adverse_impact — product and system familiarity are a proxy for prior access, not for ability. Rebuilt against the buyer's own instance where they provide a sandbox.

Working speed is scored. Handle time and queue depth are real constraints in this seat, so the session reports time to first substantive answer and the gap between call end and note submission, beside the quality scores and never alone. A fast candidate who invented a Friday ranks below a slow one who did not, and the report states that on its face. Two exclusions hold regardless. Speed at navigating the employer's own systems is not measured — that is the familiarity effect the hitl note disclaims and it is training. Typing speed is not measured either, because workforce management measures throughput directly and accurately from week one, so a pre-hire proxy for it sells the employer information they already receive free.

What the candidate actually does

TaskWhat happens
The refund that was already promised
live_call · 10 min
A live spoken call with an AI customer who was told by a previous agent that a refund would land last Friday. It did not. The candidate has a read-only account record open, and the record does not support any confident answer: the refund was raised, it was returned by the bank because the card it was issued against has since been closed, and it is now sitting unallocated with no processing date attached to it. There is genuinely no honest date available. Two forks are built in. The first is the commitment fork — around minute four the customer asks, twice, for a date, and the cheap path is to invent one, because a date ends the discomfort immediately and the AI customer visibly softens the moment it is given. The counterparty then presses on it ("so I have your word it is Friday?" and then "and if it is not?"), which is where an invented date becomes an escalation and a bounded answer becomes an acceptable one. The second is the buried-problem fork — in passing, while complaining, the customer mentions that the closed card is the one their monthly subscription bills to. Nobody asks about it. The stated request is the refund; the more expensive problem is next month's failed payment and a service interruption, and the record contains the subscription line if the candidate goes looking.
The note and the line
written_artifact · 6 min
Immediately after the call, in the same monitored window, the candidate writes two short pieces from memory and the record, both into the same ticket: the internal note the next agent will pick up cold, and the single customer-facing sentence that goes out as the contact summary. Because they are two fields on one record rather than two documents, the reviewer reads them side by side as a colleague would, which is what makes a commitment that is firm out loud and soft in writing visible rather than inferred. The fork here is that the easy note is the one that discharges the writer's own feelings — it records that the customer was upset and that the advisor apologised and explained the process — and it is faster to write than a note that states what was committed to, by whom, with what dependency. Both notes close the ticket. Only one of them prevents a second call.
Three calls, cut mid-turn
judgment_scenario · 6 min
Three short call transcripts, each stopped at a different point, each already going wrong in a different way — one where the customer has repeated the same sentence three times, one where the advisor has just been asked something the advisor's authority does not cover, one where the customer has become abusive but the underlying complaint is valid. For each, the candidate marks the turn at which the call stopped being winnable at their level and writes one line saying what they would do at that turn. This exists because escalation timing is the competency the role file identifies as invisible before week one, and a single ten minute call can only ever produce one data point on it. Three cut transcripts produce three, for six minutes.

The mark scheme

Each criterion is scored 1 to 5 against written anchors, and every score is reported with the excerpt that earned it. A criterion marked floored is reported as a finding rather than averaged into the total. The first is open; open any other to read its anchors in full.

Commitment disciplineweight 0.35Separates the three states out loud — what the record confirms, what is not yet known, and what the customer will receive next and when — and holds th…
1 States a date, time or outcome that nothing in the record supports, and repeats or firms it up when the customer presses.
3 Gives a range rather than a fixed date and names at least one thing the timing depends on, but does not say what will happen if the range is missed.
5 Separates the three states out loud — what the record confirms, what is not yet known, and what the customer will receive next and when — and holds that position through the second and third push without either hardening into a policy sentence or drifting into a date.
Finds the problem behind the requestweight 0.2Raises the subscription billing consequence unprompted, before the customer has finished with the refund, and states it as a thing that will happen ra…
1 Works only the refund. The closed card is never connected to anything beyond the failed refund, and the call ends with the subscription problem intact.
3 Notices the closed card matters beyond this refund, but raises it only after the customer has already been told the call is finishing, or flags it without acting on it.
5 Raises the subscription billing consequence unprompted, before the customer has finished with the refund, and states it as a thing that will happen rather than a thing that might.
Record fidelityweight 0.25The note states what was promised, by when, conditional on what, and the customer-facing sentence says the same thing in the same terms — no commitmen…
1 The case note describes the customer's mood and the advisor's conduct. A colleague reading it cannot tell what the customer was told.
3 The note records the outcome and the main commitment, but omits either the dependency the commitment rests on or the second issue raised.
5 The note states what was promised, by when, conditional on what, and the customer-facing sentence says the same thing in the same terms — no commitment appears in one that is absent or softer in the other.
Escalation timingweight 0.2Marks the turn where the advisor's authority ran out or the same exchange began repeating — earlier than the demand in at least two cases — and in the…
1 Marks the escalation point at or after the customer demands it in every transcript, and gives no reason beyond the demand.
3 Marks a plausible point in at least two transcripts and names a reason tied to the content of the call rather than to the customer's volume.
5 Marks the turn where the advisor's authority ran out or the same exchange began repeating — earlier than the demand in at least two cases — and in the abusive-contact item separates the conduct from the complaint, escalating the conduct while keeping the valid issue alive.

How it is scored

Weighted mean of the four criteria, each scored 1-5 against the anchors above. Every score is reported with the excerpt that earned it: a quoted turn from the call transcript, a quoted line from the note, or the candidate's marked turn and reason. A score with no excerpt attached is not reported. The call is scored from its transcript rather than its audio, and the reviewer sees the transcript first — this is a deliberate control against the accent and manner effects named under adverse_impact, not a convenience.

Integrity

The log describes what happened. It does not produce a cheating verdict — the follow-up conversation is the control, because a statistical accusation is not something we would ask a reviewer to defend.

What you receive

Who decides

Required, not optional, on two decisions. First, any candidate scoring 1 on commitment discipline is a reject on the strength of a single verbal act, and a human must read the actual exchange and confirm that a date was invented rather than mis-transcribed. Second, all rejects within half a point of the threshold are read rather than filtered. The reviewer is looking at the transcript and the note side by side, checking one thing above all: does the commitment made out loud appear in the note in the same words. Overrides are recorded with a written reason. What the ranking does not entitle a buyer to conclude — it does not measure product knowledge, system speed, attendance, schedule adherence or how the candidate will perform in month six, and a rank order across candidates within a band this narrow is not a reliable ordering. Use it to decide who to interview, not who to hire.

What this does not measure

This design does not score accent, dialect, pace, first language, fluency or manner of speech, and the rubric is built so it cannot smuggle them back in. There is no clarity criterion and no professionalism criterion — the two places where an accent filter conventionally hides. Clarity is only ever scored as an attribute of the record, from text. The call is scored from the transcript, and the reviewer's default surface is the transcript, so the scoring act itself is one step removed from voice. In a family with this much offshore delivery, an accent filter wearing a competency label is a real and specific risk, and the control has to be structural rather than an instruction to reviewers to be fair. It also does not score warmth or agreeableness as traits. A candidate who is flat, brisk and correct outscores one who is delightful and invents a Friday. Deployers should monitor pass rates by candidate location and first language and treat any material gap as a defect in this design rather than a property of the applicant pool.

The whole argument for this design is that the expensive failure in this seat is a single verbal act, performed at a predictable moment, and that nothing in the conventional funnel puts a candidate in front of that moment. The role file names it: saying "I'll have that sorted by Friday" with no basis, at minute six, because the silence was unbearable. So the call is built to manufacture minute six on purpose. The record is constructed so that no honest date exists, the AI customer asks for one twice, and the counterparty rewards the invented date immediately — it thanks the candidate and relaxes — before testing it. That sequencing matters. If the cheap path were punished on the spot the fork would not be a fork; the candidate has to actually feel the invented date working before they find out what it costs.

Twenty-two minutes is the design constraint, not an accident of scoping. This is the highest-volume seat in the corpus and the hiring is replacement hiring in a shrinking occupation, which means the assessment competes against a recruiter's incentive to skip it. Anything that a busy contact centre will quietly stop running has a real predictive validity of zero. So every minute here is spent against a named failure and nothing else: ten on the only thing that cannot be observed any other way, six on the artefact that determines whether a repeat contact starts from zero, six on the one competency that a single call cannot sample more than once.

What that budget gives up is worth stating to a buyer plainly. There is one customer, one temperament and one grievance type, so this predicts handling of an angry-and-partly-right billing contact and infers the rest. It does not test the same candidate against a confused elderly caller, a fluent aggressive one, and a silent one, which a forty-minute design could and which would produce a more stable picture. It samples escalation judgment from transcripts rather than live, which measures whether the candidate can recognise the moment but not whether they can act at it while a person is shouting. And it produces no signal at all on stamina — the fortieth call of a shift is a different act from the first, and no pre-hire assessment of any length reaches it.

The design refuses two measurements that buyers frequently ask for. Typing speed is absent because workforce management measures throughput directly and accurately from week one, so a typing test sells an employer information they already receive free. Personality inventories are absent because they predict a disposition and the failure that costs money here is a decision. The distinction is the whole product argument: an agreeable person who over-promises is a worse hire than a blunt one who does not, and only one of these two instruments can tell them apart.

Sources

Every figure on this page is traceable. Where a claim could not be sourced it is stated qualitatively instead.

  1. US Bureau of Labor Statistics, Occupational Outlook Handbook, Customer Service Representatives, 2025, https://www.bls.gov/ooh/office-and-administrative-support/customer-service-representatives.htm
  2. IT and Business Process Association of the Philippines (IBPAP), The Philippine IT-BPM Industry Overview 2026, 2026, https://admin.ibpap.org/storage/hub-resources/7Hv9U2uoLJVx2kxyRMObNm6W6L0MR57l9dZfrez7.pdf
  3. NICE, 2025 Workforce Management Trends for Contact Center Leadership, survey of 400 contact centre leaders in North America and EMEA, 2025, https://resources.nice.com/wp-content/uploads/2025/04/Managing-the-Modern-Contact-Center-Current-Employer-Trends-2025-Survey.pdf

See what the employer actually receives. A full report for one role, with every score shown beside the excerpt that earned it, conduct findings reported rather than averaged, and a reviewer sign-off required before any decision. No form.

Read a sample reportOr talk to us about this role