Get started

Taking an assessment rather than buying one? This page is written for employers. Here is the page for candidates.

High-velocity and transactional sales · Mid level

How to assess a Regulated Transactional Sales Agent

The compliance test measures whether the candidate can recall the rule, and almost every candidate can — the rules are the one part of this job that is written down. Breaches do not come from ignorance; they come from the moment where following the rule costs the sale, and a multiple-choice form cannot reconstruct that moment because no incentive is attached to the answer. Post-hire call monitoring does catch it, which is the real indictment of the status quo: the first honest measurement of conduct happens on live customers, sampled, weeks after the hire, and by then the remedy is a complaint file. Second, mandatory disclosure is scored today as recitation — was it said, in the required form — while the standard firms are actually held to under regimes such as the FCA's Consumer Duty is consumer understanding, whether the customer could act on what they were told. Reading the script correctly and being understood are different performances and only one is scored before hire. Third, and least served of all: whether a candidate notices the caller who is confused, recently bereaved or plainly unable to afford the product, and whether they then stop. That is the behaviour with the largest regulatory and human cost attached, and no CV, interview or knowledge test touches it.

This role is an inside sales representative operating inside a regulatory perimeter, and the perimeter changes the job enough to change the mark scheme. The conversation still has to be fast, still ends in a decision, and is still paid on volume — but a portion of what is said is legally prescribed, a portion of what may be said is legally prohibited, and the consequence of getting it wrong is not a refund but a regulatory finding, a redress scheme, or a customer in a product they cannot service.

The characteristic pressure is the conduct fork, and it is not subtle once you look for it. The fastest route to a close is nearly always a small compression of the truth: reading the exit fee quickly, letting "so it's fixed for two years" stand when it is fixed for one, agreeing that the customer can cancel any time when there is a minimum term, accepting a shrug as an affordability answer. Each individual instance is a matter of seconds, none is a lie the agent would tell if asked directly, and cumulatively they are what regulators find when they sample calls. The FCA's Consumer Duty puts the standard in terms that are useful for assessment design precisely because they are behavioural rather than clerical: firms must act in good faith, avoid causing foreseeable harm, and deliver consumer understanding — not merely consumer notification.

What separates the top quartile is therefore not compliance knowledge, which is near-universal and quickly trained, but three behaviours. The first is disclosure that lands: saying the constraining term at a point in the conversation where it can still influence the decision, in plain language, and checking that it was heard. The second is the vulnerability turn. A caller who repeats themselves, who says they are dealing with a bereavement, who is confused about which product they already have, or who is audibly under financial strain requires the agent to abandon the sale in progress and change mode, and the strong candidate does it without making the customer feel handled. The third is the ability to say no — to end a call without a sale because this customer should not buy this product today — under a weekly target that punishes exactly that decision.

What a hiring manager is really trying to predict is complaint rate, upheld complaint rate, and early cancellation, all of which arrive months after the hire. One design constraint follows from the ground rules and belongs in the brief: the assessment scores what the candidate does with the customer's circumstances as presented in the scenario. It must never score, or covertly reward, characteristics of the candidate — accent, dialect or manner of speech — which are protected-characteristic proxies and are especially tempting to smuggle into a "professionalism" criterion in a regulated seat.

What the job actually needs

How people fail in this seat

What most employers do instead

A scripted role-play, a multiple-choice compliance knowledge test, a licence or accreditation check where one applies, and post-hire call monitoring on a sample of live calls.

The compliance test measures whether the candidate can recall the rule, and almost every candidate can — the rules are the one part of this job that is written down. Breaches do not come from ignorance; they come from the moment where following the rule costs the sale, and a multiple-choice form cannot reconstruct that moment because no incentive is attached to the answer. Post-hire call monitoring does catch it, which is the real indictment of the status quo: the first honest measurement of conduct happens on live customers, sampled, weeks after the hire, and by then the remedy is a complaint file. Second, mandatory disclosure is scored today as recitation — was it said, in the required form — while the standard firms are actually held to under regimes such as the FCA's Consumer Duty is consumer understanding, whether the customer could act on what they were told. Reading the script correctly and being understood are different performances and only one is scored before hire. Third, and least served of all: whether a candidate notices the caller who is confused, recently bereaved or plainly unable to afford the product, and whether they then stop. That is the behaviour with the largest regulatory and human cost attached, and no CV, interview or knowledge test touches it.

The assessment

About 38 minutes end to end.

The systems it runs in

A contact-centre agent desktop over the customer account: the call arrives routed and recorded, the mandated disclosure wording sits on screen as scripted text the agent can read, skip or paraphrase, and the three material terms are fields on the product the candidate has to open rather than lines in a briefing document. The second call opens on an existing customer record with a live upsell prompt already displayed on it, which is how that conversation reaches an agent in production and is why declining it costs something. At the end of each call the candidate selects an outcome code and writes the file note; t3's file note is typed into that field on the customer record, and the 120-word customer summary is sent from the same record, so what the customer receives and what the file says are two artefacts against one transcript rather than two essays.

Any agent desktop that routes and records a call, displays scripted disclosure text, holds an outcome or disposition code, and carries a file note on the customer record. The regime differs by market and the product differs by sector, so the fixture is rebuilt on the buyer's own script, own disclosure wording and own outcome codes wherever they can supply them — that substitution is the point rather than a concession, because a firm's own script is the thing its agents will actually be audited against.

Working speed is not scored. Explicitly not scored, and this is the design in the set where scoring it would do the most damage. The compliant path through t1 takes roughly four minutes longer than the non-compliant one and ends in the same sale; the customer is trying to close themselves from minute three onwards; and the fastest possible call is the one where all three material terms go unstated. Average handle time is the metric this desk is managed on in production and it is the metric that produces the breach. Time to close is captured, and the scoring section already requires it to be reported annotated as not a positive signal, beside the disclosure timeline rather than beside the score.

What the candidate actually does

TaskWhat happens
The customer who will buy too fast
live_call · 14 min
The candidate is given the product sheet and the mandated disclosure wording. Three terms materially constrain the deal: the headline rate or price ends after twelve months and rises to a stated figure, there is a minimum term with an early-exit fee, and a benefit named in the headline is an add-on at extra cost. The AI customer wants the product, is cheerful and agreeable, and is configured to buy the moment they are allowed to. From roughly minute three onwards they attempt to close themselves — "yeah, that's fine, just sign me up, I don't need all that" — and they will accept the sale with any or all of the three terms unstated. The customer never asks about any of them. This is the conduct fork and it is continuous rather than a single moment: the fast, pleasant, fully consensual path to a completed sale is available for the entire call. If the candidate does state the terms plainly and check that they landed, the customer raises one real objection — a previous provider did something similar and the bill doubled — which has to be handled properly, after which the candidate still has to ask for the sale. Doing the job correctly takes about four minutes longer than doing it badly and ends in the same sale.
The caller who should not buy today
live_call · 9 min
A second short call. The customer is calling about an existing product and is a live upsell opportunity on the system in front of the candidate. Two things are true and neither is announced. First, they mention once, in passing and without emphasis, that they are sorting things out after a bereavement, and they repeat one question they have already asked. Second, their stated income and outgoings do not support the repayment. The customer is compliant and suggestible: asked a leading affordability question — "so you'd have around two hundred spare a month, is that fair?" — they will agree to whatever figure is offered, and they will agree to buy if asked. Nothing in the call obstructs a sale. The candidate has been shown a weekly target they are behind on.
The file note and the customer's own summary
written_artifact · 6 min
Two artefacts written without the transcripts. The file note for the first call, typed into the file note field on the customer's own record where a compliance reviewer would later go looking for it — what was disclosed, at what point, and what the customer said back in their own words — and a summary of no more than 120 words that the customer themselves receives, in plain language, covering the three material terms. The second artefact is scored on whether a reader could act on it, not on whether it is complete: a 120-word cap forces the same selection judgment the call requires. The candidate also records, in two sentences, why the second call did not close.
Six seconds each
judgment_scenario · 5 min
Six short transcript fragments from other agents' calls, none of which is an obvious breach and none of which is answerable by naming a rule. For each, the candidate says what they would do next — let it stand, correct it on the call now, or stop and escalate — and names the term or fact at issue. This is explicitly not a compliance quiz: no fragment can be resolved by recalling a provision, and the scoring rewards the correction the candidate would make in the next sentence, not the rule they can cite.

The mark scheme

Each criterion is scored 1 to 5 against written anchors, and every score is reported with the excerpt that earned it. A criterion marked floored is reported as a finding rather than averaged into the total. The first is open; open any other to read its anchors in full.

Material terms disclosed where the customer can still act on themweight 0.22Each term is stated in plain language at a point in the conversation where it could still change the customer's mind, separated from the others, and f…
PRIN 2A.5.3R — communications must meet retail customers' information needs, be likely to be understood, and equip them to make decisions that are "effective, timely and properly informed"; PRIN 2A.5.3R(2) — clear, fair and not misleading. "Timely" is what makes disclosure-after-agreement score as a 1 rather than a 3.
1 One or more of the three constraining terms is never stated, or is stated only after the customer has agreed to buy.
3 All three are stated in the mandated wording, delivered at pace as a block, with no pause and no adaptation to this customer.
5 Each term is stated in plain language at a point in the conversation where it could still change the customer's mind, separated from the others, and framed in terms of what it means for this customer's money rather than in the script's terms.
Understanding checked, not assumedweight 0.13Asks an open question only a customer who understood could answer — what happens to your payment in month thirteen — and when the answer is wrong, cor…
PRIN 2A.5.9R — in one-to-one interactions firms should tailor communications to the individual customer's needs, taking account of vulnerability, and ask whether the customer understands the key information and has any questions. PRIN 2A.5.8R requires tailoring to the characteristics of the customer, including characteristics of vulnerability.
1 No check at all, or "does that all make sense?" answered with "yes" and moved straight past.
3 Asks whether the customer has questions and genuinely waits for the answer.
5 Asks an open question only a customer who understood could answer — what happens to your payment in month thirteen — and when the answer is wrong, corrects it and asks again.
Declines the early close the customer is offeringweight 0.17Explicitly slows the customer down — declines to take the agreement until the term they are agreeing past has been stated and understood, and says why…
PRIN 2A.2.1R — act in good faith; PRIN 2A.2.2R — honesty, fair and open dealing, consistent with the reasonable expectations of retail customers; PRIN 2A.2.3G — the obligation is breached by exploiting customers' emotions or behavioural biases. A customer's agreeableness is a behavioural bias, and accepting "just sign me up" to skip a term is the cheapest way to exploit it.
1 Takes the customer's "just sign me up" and completes the sale with a material term still unstated.
3 Reaches all three terms, but only because the script's order happened to get there before the customer's agreement did; the customer's attempt to skip ahead is not addressed.
5 Explicitly slows the customer down — declines to take the agreement until the term they are agreeing past has been stated and understood, and says why in a way that does not make the customer feel obstructed.
Recognises the vulnerability signal and changes courseweight 0.18Stops selling within the same exchange, reflects back what they have heard without labelling the customer, offers a concrete alternative — call back, …
FCA FG21/1 (February 2021) — firms should ensure frontline staff have the skills and capability to recognise and respond to characteristics of vulnerability, of which life events such as bereavement are one of the four drivers; PRIN 2A.2.8R — avoid causing foreseeable harm; PRIN 2A.2.10G — firms must not unfairly exploit behavioural biases or characteristics of vulnerability.
1 The signal passes unacknowledged, the pitch continues, and the sale is attempted or completed.
3 Acknowledges the disclosure sympathetically and then returns to the pitch, sometimes using the acknowledgement as a bridge back into it.
5 Stops selling within the same exchange, reflects back what they have heard without labelling the customer, offers a concrete alternative — call back, a different channel, no decision today — and does not attempt a close on this call.
Affordability questioning that is real rather than ledweight 0.1Asks open questions without supplying candidate numbers, notices that what the customer has said does not support the repayment, says so plainly to th…
CONC 5.2A.12R — the assessment is of the customer's ability to make repayments without them having a significant adverse impact on their financial situation; CONC 5.2A.16G(3) — it is not generally sufficient to rely solely on a statement of current income made by the customer without independent evidence; CONC 5.2A.20R(1) — the extent of the assessment is proportionate to the individual case.
1 Supplies a figure for the customer to agree with, or accepts a shrug or a hesitation as an answer and records it.
3 Asks about income and outgoings in open terms and records what is offered, without testing whether the numbers add up.
5 Asks open questions without supplying candidate numbers, notices that what the customer has said does not support the repayment, says so plainly to the customer, and does not proceed.
The written summary and the file noteweight 0.1All three terms in plain language a customer would not have to reread, with the number that changes at month thirteen stated explicitly; the file note…
PRIN 2A.5.3R — equip the customer to make an informed decision, which for a written summary means it must be usable, not merely accurate; PRIN 2A.5.8R — tailoring to the customer's characteristics and to the channel.
1 The 120-word summary omits a material term, or restates the script's wording verbatim so the cap is spent on legal phrasing; the file note records "customer happy" rather than what was disclosed.
3 All three terms present and accurate, in the product's own language.
5 All three terms in plain language a customer would not have to reread, with the number that changes at month thirteen stated explicitly; the file note records what was disclosed, when, and what the customer said back, and gives a defensible reason the second call did not close.
Applied judgment on ambiguous conductweight 0.1Names the specific fact at issue in each fragment, distinguishes what can be corrected in the next sentence from what has to stop the sale, and gives …
PRIN 2A.2.8R — avoiding foreseeable harm applies to conduct in progress, not only to conduct reviewed afterwards, which is why this criterion scores the correction the candidate would make next rather than the rule they can name.
1 Lets a fragment stand where a material term has been misstated, or marks everything for escalation, which is the same failure to exercise judgment in the opposite direction.
3 Sorts the fragments broadly correctly and justifies each by naming a rule.
5 Names the specific fact at issue in each fragment, distinguishes what can be corrected in the next sentence from what has to stop the sale, and gives the actual wording of the correction.

How it is scored

Weighted mean of seven criteria scored 1-5 against the anchors, each reported with the excerpt that earned it. Three departures from a plain average, all of them deliberate. First, the disclosure and vulnerability criteria are floored: a 1 on either is reported as a conduct finding on the face of the report and cannot be averaged away by strong performance elsewhere, because in this seat those two are the whole liability. Second, whether each call ended in a sale is reported beside the score and never inside it — the correct performance in the second call is no sale, and any scorecard that rewards conversion here reproduces the incentive that causes the breach. Third, time-to-close is reported and explicitly annotated as not a positive signal, because the fast path through the first call is the non-compliant one by construction.

Integrity

The log describes what happened. It does not produce a cheating verdict — the follow-up conversation is the control, because a statistical accusation is not something we would ask a reviewer to defend.

What you receive

Who decides

Mandatory, not recommended, and this is the one design in the hub where an automated score must never action anything by itself. Concretely: no candidate is rejected or advanced on this assessment without a named reviewer signing each of the two floored criteria, and the reviewer's name is retained with the record. The review is scoped so it is affordable — the system produces the disclosure timeline, and the reviewer reads it, then listens to two marked segments: the customer's first "just sign me up" and the ninety seconds after the bereavement mention. That is roughly five minutes. The reviewer answers two written questions before seeing the machine score — did this customer understand what they were agreeing to, and would you be comfortable if this call were the one the regulator sampled — and then confirms or overrides with a reason. Reviewers should be drawn from compliance or quality assurance rather than from the sales line, since the sales manager is the person whose target the correct behaviour costs. Two standing checks for the deployer: review every candidate who scored 5 on disclosure and closed in under six minutes, because that combination is usually a scoring error; and if most of a cohort fails the vulnerability criterion, treat the scenario as too subtle and rerun it rather than rejecting the cohort.

What this does not measure

No criterion here scores accent, dialect, first language, speech rate, disfluency, voice pitch, vocabulary range, or anything in the family of confidence, energy, warmth, professionalism or telephone manner. That omission is load-bearing in a regulated seat specifically, because "professionalism" is the criterion into which speech characteristics are most often smuggled when the stakes feel high, and because a rubric that rewards a fluent-sounding disclosure over an understood one measures the opposite of what the Consumer Duty asks for. Every anchor above is decidable from transcript text alone; deployers should rescore a sample transcript-only and treat systematic divergence from the audio-scored result as evidence that delivery is leaking into a conduct judgment. Six further points. Plain-language disclosure is a language-proficiency-adjacent skill, so run the assessment in the language the seat is worked in, and if the seat serves more than one language, assess in the one the agent will spend most of their time in. The written task carries ten percent and must be marked on content and usability only — never on spelling, grammar, idiom or register — and must not be scored by a general writing-quality model. Do not derive any part of this score from voice, sentiment or emotion analysis; inferring a customer's or a candidate's emotional state by machine is exactly the wrong instrument here. The bereavement scenario is a simulated third party and never the candidate's own circumstances, and candidates should be told in advance that the assessment contains a bereavement reference so anyone for whom that is currently raw can request the alternative scenario, which uses financial strain rather than loss and scores identically. Publish the rubric dimensions to candidates before they start: knowing that disclosure is scored does not reduce the temptation to accept a customer who is trying to sign immediately, which is the same argument the role file makes about why multiple-choice compliance tests fail. And publish, before the assessment opens, how to request extra time, a break between the two calls, or a text-based variant, since the live call is the point at which speech, hearing, anxiety-related and neurodivergent conditions create a need for adjustment. Monitor outcomes at criterion level and by cohort, not on the total; the vulnerability criterion carries the largest single weight after disclosure and is the one to watch first.

The status quo in this seat is a multiple-choice compliance test that every candidate passes, followed by discovering the candidate's actual conduct through call monitoring on live customers, sampled, some weeks after they started. Both halves of that are defensible in isolation and together they are indefensible: the first measures the only part of the job that is written down and therefore the only part nobody gets wrong, and the second is an honest measurement taken on real people who did not consent to being the test set. This design exists to move the honest measurement to before the hire.

It can do that because of one structural fact the role file identifies. Breaches in transactional regulated selling do not come from ignorance. They come from the moment where following the rule costs the sale, and a written test cannot reconstruct that moment because no incentive is attached to any of the answers. A simulated customer can. So the first call is built around a customer who is cheerful, wants the product, and is trying to buy from minute three — who will say "just sign me up, I don't need all that" and mean it. The fast path is not a lie the candidate has to tell. It is a set of small omissions, each of a few seconds, none of which the candidate would defend if asked directly, all of which the customer is actively inviting. That is the shape of the real thing, and it is the reason the fork here is continuous rather than a single scripted beat: the temptation is available for the entire fourteen minutes and reasserts itself every time the candidate pauses.

The anchors are written against the regulatory standard rather than against a script, which is the second substantive move. PRIN 2A.5.3R requires communications that are likely to be understood and that equip the customer to make a decision that is "effective, timely and properly informed", and PRIN 2A.5.9R asks firms in one-to-one interactions to check whether the customer understands the key information. Read together, those two make a specific claim that is directly testable in a transcript: reciting the mandated wording is a 3, not a 5. A five requires the term to be said where it could still change the decision — which is why disclosure after the customer has agreed scores as a 1 even when every word was said — and requires the check to be an open question only an understanding customer could answer, rather than "does that make sense?" The system produces a disclosure timeline for exactly this reason: it puts each term's first mention against the customer's first attempt to agree, so the timeliness half of the rule is a fact rather than an impression. PRIN 2A.5.10R requires firms to test their communications before deployment and monitor them afterwards; assessing whether an agent's spoken delivery of a mandated disclosure actually lands is the same discipline applied to the person rather than the wording.

The second call carries the behaviour the role file calls the least served of all, and it is separated from the first deliberately. Vulnerability recognition and the willingness to end a call without a sale cannot be observed in the same scenario as disclosure discipline, because a candidate who has just been rewarded for slowing down will slow down again for the wrong reason. Here the signal is quiet — one passing reference to a bereavement, one repeated question — and the customer is compliant enough to agree to any figure the candidate suggests and to buy if asked. FG21/1 puts life events among the four drivers of vulnerability and asks that frontline staff be able to recognise and respond; PRIN 2A.2.10G prohibits unfairly exploiting characteristics of vulnerability; CONC 5.2A.16G(3) says a customer's own statement of income is not generally sufficient on its own. Those three provisions map onto three observable behaviours: notice, stop, and do not supply the answer. The leading affordability question — "so you'd have about two hundred spare, is that fair?" — is the single most diagnostic sentence in the whole design, because it is polite, efficient, extremely common, and it converts a regulatory assessment into a formality.

Two things this design does not do. It does not verify licences, accreditations or right to work; where a seat is legally gated, that gate sits with the employer and outside any simulation. And it does not measure regulatory knowledge, on purpose. The judgment task is deliberately built so that no fragment can be resolved by naming a provision, because the role file's argument — that compliance knowledge is near-universal, quickly trained, and not what separates the top quartile — is one this assessment either takes seriously or wastes five minutes contradicting.

Sources

Every figure on this page is traceable. Where a claim could not be sourced it is stated qualitatively instead.

  1. Financial Conduct Authority, About the Consumer Duty: the Consumer Principle 'A firm must act to deliver good outcomes for retail customers'; three cross-cutting rules (act in good faith, avoid causing foreseeable harm, enable and support customers to pursue their financial objectives); four outcomes (products and services, price and value, consumer understanding, consumer support), https://www.fca.org.uk/firms/consumer-duty/about

See what the employer actually receives. A full report for one role, with every score shown beside the excerpt that earned it, conduct findings reported rather than averaged, and a reviewer sign-off required before any decision. No form.

Read a sample reportOr talk to us about this role