Taking an assessment rather than buying one? This page is written for employers. Here is the page for candidates.
Inbound customer support · Entry level
How to assess a Chat and Messaging Support Advisor
Every step in this funnel is a serial test of a concurrent job. The defining constraint of the chat seat is that three or four conversations are live at once and attention is the scarce resource — and nothing in a phone screen, a grammar test or an untimed writing sample observes attention allocation at all. A grammar test scores an unhurried sentence; the sentence that damages this business is the hurried one, written while two other customers are waiting, where the advisor reaches for the nearest macro because composing a real answer costs forty seconds they feel they do not have. Nor does anything in the current process show whether the candidate can hold a firm line in text, which is materially harder than in speech: there is no tone of voice to soften a no, so an advisor who has only ever been screened by phone has never been observed doing the actual difficult act of this job.
Chat looks like a cheaper version of the voice seat and is treated as one by most hiring processes. It is a different job. The voice advisor holds one conversation and controls its pace; the chat advisor holds several and controls none of them, because every customer sets their own tempo and expects a reply inside the window they consider reasonable. The work is therefore not primarily about writing well. It is about deciding, forty times an hour, which of the live threads gets the next ninety seconds of real thought and which gets a holding line — and then writing the holding line so it actually holds rather than reading as a brush-off.
The two failure modes that dominate quality reviews in this channel are both products of that constraint. The first is macro misuse: a template exists that is adjacent to the customer's question, the advisor is three threads deep, and the template goes out. It is grammatically perfect, on-brand, and does not answer what was asked, which guarantees a second contact and usually an angrier one. The second is the silent thread — a conversation that receives no message for five or six minutes while the advisor works another, with no signal to the customer that they have not been abandoned. Neither failure is a writing deficiency and neither is detectable in a writing sample produced with unlimited time and one thing to think about.
Text also changes the emotional mechanics of the job in a way that matters for selection. On a call, a difficult message can be carried by tone: a warm voice makes "no, we cannot refund that" survivable. In chat there is no carrier signal, so the same content lands harder, and the advisor's instinct is to soften it into vagueness — hedging until the customer no longer knows whether they have been refused. The top quartile does the opposite: they state the outcome plainly in the first clause, then give the reason, then give the one thing that can still be done. That ordering is a learnable, observable, gradeable behaviour, and it is the single strongest separator in this seat. The median advisor buries the refusal in the third paragraph and hopes the customer stops reading.
There is a second, newer separator worth naming to any buyer: judging when a canned or model-generated draft is wrong for the situation in front of them. Most chat operations now run suggested-reply tooling, which means the advisor's real job has shifted from composition toward editorial judgment — accepting, amending or discarding a draft that is fluent and sometimes subtly wrong. That shift is very recent and almost no screening process has caught up with it. A candidate who accepts a plausible draft that contradicts the account record is a specific, expensive risk, and it is directly assessable by giving them a draft and a record that disagree.
Savvanta assesses this with a concurrent written exercise rather than a call: several simultaneous inbound conversations with different tempers and different information needs, a macro library that is nearly-but-not-quite right for one of them, and a system record that contradicts one of the customers. The score comes from whether every thread was closed with the customer knowing what happens next, whether the macro was amended or sent blind, and whether the refusal was stated before it was explained. What it does not measure is typing speed, which workforce management already measures live, or dialect and idiom in the candidate's writing — the criterion is whether the customer would understand and accept the message, not whether it matches a house style the candidate has not yet been trained on.
What the job actually needs
- written clarity under concurrency
- holding context across parallel conversations
- judging when a template answers the question and when it does not
- tone calibration in text
- closing a thread rather than abandoning it
How people fail in this seat
- pastes a macro that does not address what was asked
- lets one of three concurrent threads go silent for six minutes
- writes a sentence that is technically accurate and reads as dismissive
- mirrors the customer's escalating register instead of steadying it
- ends a chat with the customer unsure whether anything will now happen
What most employers do instead
A CV sift, a voice-shaped phone screen, and sometimes a grammar or typing test. Occasionally a single untimed written sample.
The assessment
About 24 minutes end to end.
The systems it runs in
A messaging agent workspace holding several conversations at once, with a macro library, a customer record in the same window, and a suggested-reply tool proposing drafts. Zendesk describes this shape directly — a unified agent workspace, macros as pre-defined responses to common issues, customer profiles that capture past conversations, and live chat and messaging apps among the channels brought into one place. The fixture reproduces the objects rather than any vendor's chrome: the conversation, its own history, the macro, the record, the generated draft, and per-message send timestamps that make silence a measurement rather than an impression.
- Zendesk
- Intercom
- Salesforce Service Cloud
- Freshdesk
Any console that presents concurrent conversations with canned responses, a customer record and per-message timestamps, whether the channel is web chat, WhatsApp, SMS or in-app messaging. Where the buyer's console proposes AI drafts, the fixture uses theirs; where it does not, the draft trap in thread C is supplied by ours, because editorial judgment over a generated draft is now part of the seat regardless of which vendor supplies the generator. Rebuilt against the buyer's own instance where a sandbox exists.
What the candidate actually does
| Task | What happens |
|---|---|
| Three threads at once written_artifact · 14 min | Three AI customers open conversations in a single chat console within the first ninety seconds, each on their own clock and none of them waiting politely. A macro library and a read-only account record sit in the same window. The threads are built to compete rather than to queue. Thread A is a slow-burn shipping query that will accept a holding line but goes cold and then hostile if it receives nothing for four minutes — the counterparty sends "hello?" at three minutes and closes the chat with a complaint at six. Thread B asks a question that is adjacent to a macro but not answered by it: the macro explains how to change a delivery address, the customer is asking how to change it after dispatch, which the macro does not cover and which has a different answer. Sending the macro is one click, is grammatically perfect, is on-brand, and buys back forty seconds for the other two threads. Thread C is where a suggested-reply tool proposes a fluent, confident draft stating the customer's plan includes a feature that the account record, two clicks away, says was removed at their last downgrade. Accepting it is the fastest possible resolution and the customer thanks them warmly for it. The console records send times per thread, so silence is measurable rather than inferred. |
| The refusal, alone written_artifact · 5 min | One thread, nothing else live, no time pressure worth speaking of. A customer wants a return processed eleven days outside the window, is polite, is a long-standing account, and has a reason that is sympathetic and irrelevant. The answer is no and there is one genuine partial remedy available. The fork is structural rather than temporal — with the concurrency removed, the only remaining pressure is the discomfort of writing a refusal into text with no tone of voice to carry it, and the cheap path is to hedge: to open with the sympathy, describe the policy at length, mention that it is being looked into, and let the customer leave the conversation not actually knowing they have been refused. That message is easier to write, feels kinder, and produces a second contact. Isolating this from t1 is deliberate — it separates a candidate who cannot write a refusal from one who could not get to it while three threads were live, which are different diagnoses with different fixes. |
| Accept, amend, discard judgment_scenario · 5 min | Five suggested-reply drafts, each paired with the account record snippet it was generated against. Two are correct and should go as written. One is fluent and contradicts the record. One is correct but answers a narrower question than the customer asked. One is correct, accurate and badly timed — it delivers a policy refusal to a customer whose previous message was an apology. For each the candidate marks accept, amend or discard and writes one line of reason; where they amend, they write the amended sentence. Five items in five minutes is a deliberate rate: it makes deliberation expensive, which is the point, because on the floor this decision is made in seconds. |
The mark scheme
Each criterion is scored 1 to 5 against written anchors, and every score is reported with the excerpt that earned it. A criterion marked floored is reported as a finding rather than averaged into the total. The first is open; open any other to read its anchors in full.
Thread stewardshipweight 0.3Every thread receives contact inside its own tolerance, every holding line names a specific action and a time, and all three conversations end with th…
Editorial judgment over macros and draftsweight 0.25Opens the record before answering thread C, states the corrected position rather than the drafted one, and either rewrites the address macro or explic…
The refusal is legibleweight 0.25The outcome is in the first sentence, the reason follows it, and the one thing that can still be done is named specifically enough for the customer to…
Context integrityweight 0.2Each thread's own history is carried correctly through every reply, and where the candidate returns to a thread after a gap they resume from where it …
How it is scored
Weighted mean of the four criteria, each scored 1-5 against the anchors above, reported with the excerpt that earned it. Thread stewardship is scored partly from console timestamps rather than judgement — the gap lengths are a measurement, not an impression — but the quality of a holding line is scored from the text. The judgment items in t3 contribute to the editorial judgment criterion alongside the live threads, so a candidate who caught the trap under no pressure but missed it under three-thread load is visibly that candidate rather than an average of the two.
Integrity
- single monitored console with per-thread send timestamps and paste detection
- the account record contradiction in thread C is randomised per session, so knowing that a trap exists does not tell a candidate where it is
- a 90-second written follow-up asking the candidate to say why they amended the specific macro sentence they amended, answerable only by someone who read it
- t3 items rotated from a larger pool, with order and correct-answer distribution varied per session
- unusually uniform inter-message intervals across concurrent threads flagged for review as a possible automation artefact rather than scored
The log describes what happened. It does not produce a cheating verdict — the follow-up conversation is the control, because a statistical accusation is not something we would ask a reviewer to defend.
What you receive
- full three-thread transcript with per-message timestamps and gap lengths
- macro library usage log showing what was sent as-is and what was amended
- whether and when the account record was opened in thread C
- the isolated refusal message as submitted
- five accept/amend/discard decisions with reasons and any amended sentences
- per-criterion score with quoted excerpt
Who decides
Required on three decisions. A reviewer must confirm any thread stewardship score of 1 by reading the gap in context, because a four-minute silence spent correctly resolving thread C is a different act from a four-minute silence spent stuck. A reviewer must read every refusal scored 1 or 2, since this is the criterion most exposed to a reviewer's own stylistic preference, and the question they answer in writing is narrow — would a reader finish this message knowing the answer was no. And a reviewer must decide, not the score, whether a candidate who is strong on concurrency and weak on refusals is a hire for this particular operation, because that trade is a function of the queue mix and the assessment does not know it. What the ranking does not entitle a buyer to conclude — it says nothing about how the candidate handles five or eight concurrent threads rather than three, nothing about sustained performance across a shift, and nothing about voice, which is a different seat with a different design.
What this does not measure
Typing speed is not a criterion and is not reported. This is worth being precise about because the temptation to reintroduce it is strong. Throughput is measured directly and continuously by workforce management from the first week, so a pre-hire proxy for it buys the employer nothing they will not have free in seven days; and speed correlates with keyboard familiarity, input method, language of the layout and physical condition, none of which is the competency. What t1 does measure is whether a candidate allocates attention well enough that no customer is abandoned at a three-thread load. That is not a typing measure and it does not tell a buyer what the candidate would do at eight threads, which is a staffing question the buyer must answer themselves. The design also does not score dialect, idiom, register or non-native phrasing. Spelling and grammar are scored only where an error changes what the customer would understand the answer to be — a misplaced comma is not a finding, a missing "not" is. House style is explicitly out of scope, since it is taught in onboarding and a candidate cannot know it. Deployers should monitor pass rates by first language and by assistive technology use, and should treat any material gap on the refusal criterion in particular as a defect in the anchors rather than a property of the pool.
Chat is assessed here as a concurrency problem, not a writing problem, and that single decision is what keeps this design from being the voice assessment with a keyboard attached. The role file is explicit that attention is the scarce resource in this seat and that every conventional screening step — phone screen, grammar test, untimed writing sample — is a serial instrument pointed at a parallel job. So the central task is built to make attention genuinely scarce: three counterparties on independent clocks, one of which punishes neglect on a timer, and two traps that are only visible to a candidate willing to spend the forty seconds they feel they cannot spare.
Both traps are drawn from the same underlying economics. The macro in thread B and the suggested draft in thread C are each the cheapest available action, each produce something fluent and on-brand, and each make a customer happy in the moment — the thread C customer actively thanks the candidate for the wrong answer. That is what makes them forks rather than tests. A trap that looks like a trap measures only whether the candidate is being watched. The draft trap in particular is the newest thing in this design and the one least covered by any existing screening step, because suggested-reply tooling has moved the advisor's job from composition toward editorial judgment faster than hiring processes have adjusted, and accepting a plausible draft that contradicts the account record is now a specific and recurring way to lose a customer.
The refusal is deliberately pulled out of the concurrent task and given its own five quiet minutes. Under load, a hedged refusal and an unwritten refusal look identical, and they are not the same problem: one candidate is uncomfortable saying no in text, the other simply ran out of time. Separating them costs five minutes and converts an ambiguous score into an actionable one, and it isolates the behaviour the role file identifies as the strongest single separator in this seat — outcome first, reason second, remaining remedy third.
Twenty-four minutes is two more than the voice design and every one of them is spent on concurrency, which cannot be sampled in a short window: a three-thread exercise needs long enough for the threads to actually collide, and collisions do not begin until the second or third turn on each. What the budget gives up is scale of load. Real chat seats run four to eight simultaneous conversations and this design runs three, because at eight the exercise stops discriminating between advisors and starts measuring reading speed. A buyer who staffs at eight should read a strong score here as evidence that the candidate allocates attention deliberately, not as evidence that they will hold eight. It also gives up channel breadth — this is synchronous chat, and asynchronous email and social care have different tolerances and different failure modes that go unsampled.
Sources
Every figure on this page is traceable. Where a claim could not be sourced it is stated qualitatively instead.
- US Bureau of Labor Statistics, Occupational Outlook Handbook, Customer Service Representatives, 2025, https://www.bls.gov/ooh/office-and-administrative-support/customer-service-representatives.htm
- NICE, 2025 Workforce Management Trends for Contact Center Leadership, survey of 400 contact centre leaders in North America and EMEA, 2025, https://resources.nice.com/wp-content/uploads/2025/04/Managing-the-Modern-Contact-Center-Current-Employer-Trends-2025-Survey.pdf
See what the employer actually receives. A full report for one role, with every score shown beside the excerpt that earned it, conduct findings reported rather than averaged, and a reviewer sign-off required before any decision. No form.
Read a sample reportOr talk to us about this role