Taking an assessment rather than buying one? This page is written for employers. Here is the page for candidates.
Collections and retention · Entry level
How to assess a Early Arrears Collections Agent
A multiple-choice compliance test establishes that the candidate can recognise a rule on a page in week one, when nothing is at stake. It says nothing about whether they apply it at minute nine of a live call with someone who is crying, or lying, or shouting, when applying it costs them the outcome their bonus is calculated on. That gap is not a training problem, it is a selection problem, and it is the only place in this corpus where the regulator has effectively published the answer key: the CFPB's 2025 FDCPA annual report records that of roughly 207,800 debt collection complaints in 2024, the most common issue was attempts to collect a debt not owed, and among communication tactics complaints 34 percent concerned continued contact after the consumer asked it to stop. Those are transcript-level behaviours. The scripted role-play cannot surface them because the team leader running it never disputes the debt and never asks the agent to stop calling.
Early arrears is first-party collections: the agent works for the original creditor, calls in that creditor's name, and is usually talking to somebody who is one, two or three payments behind rather than someone who has been in default for years. This distinction has real legal weight and is worth stating on the page, because it is routinely muddled in job adverts. Under 15 U.S.C. 1692a(6) the FDCPA's definition of "debt collector" excludes an officer or employee of a creditor collecting in the creditor's own name, so a first-party arrears agent is generally outside the FDCPA proper. They are not outside conduct regulation: the CFPB supervises first-party collections under the Dodd-Frank prohibitions on unfair and abusive practices at 12 U.S.C. 5531 — where "abusive" expressly includes taking unreasonable advantage of a consumer's lack of understanding of material risks, costs or conditions — and in the UK a lender's arrears handling sits squarely under FCA CONC 7, which at CONC 7.3.4R requires firms to treat customers in or approaching arrears with forbearance and due consideration. A candidate who has learned collections compliance as "the FDCPA does not apply to us" has learned the most dangerous possible half-truth, and that belief is detectable in a simulation within about four minutes.
The commercial skill and the compliance skill are separable, and a good design scores both, because the seat fails in two independent directions. The compliance failure is well documented. The commercial failure is quieter: an agent who extracts an unrealistic commitment books a promise-to-pay this week and manufactures a broken arrangement, another contact cycle, and a worse recovery next month. Every collections operation has this pattern in its data and very few hire against it, because promise-to-pay is the metric on the wallboard and kept arrangements are the metric on the quarterly review. The behaviour that predicts kept arrangements is unglamorous and observable: the agent who asks what the customer's income and essential outgoings actually are before proposing a figure, and who talks the customer *down* from an over-optimistic offer, produces arrangements that hold. The agent who accepts whatever number ends the call produces the opposite.
The third competency is vulnerability, and it is the one that is trained everywhere and screened nowhere. The FCA's FG21/1 sets out four drivers of vulnerability — health, life events, low resilience and low capability — and expects frontline staff to have the skills to recognise and respond to them in a conversation. Recognition here does not mean inferring anything about the customer; it means noticing what the customer has actually told you. The CFPB's own supervisory findings describe agents responding to consumers who explained a recent hospital stay by taking an aggressive tone. That is not a subtle case. In a simulation it is a single fork: the AI customer mentions, in passing and without emphasis, that they have been signed off work after a diagnosis, and the transcript then shows whether the next turn changed. Most candidates who fail this do not fail it by being unkind; they fail it by carrying on with the scripted next question as though nothing had been said, because the disclosure did not arrive in the format the training deck used.
Two design constraints should be written into any deployment of this assessment and repeated to the buyer. Score the behaviours — was the stop request honoured, did pressing continue after a dispute, was the arrangement affordable, did the approach change after a disclosure — and never the candidate's accent, dialect, manner of speech or apparent warmth. And never build a scenario that rewards a candidate for inferring anything about a customer's circumstances beyond what the customer says on the call; a design that pays for good guesses about who is struggling is a design that pays for stereotyping.
What the job actually needs
- establishing affordability rather than extracting a promise
- recognising and responding to disclosed vulnerability
- honouring a contact-preference or stop request the moment it is made
- explaining consequences accurately without threatening
- holding the ask after the first refusal
How people fail in this seat
- books an arrangement the customer cannot keep to hit a weekly promise-to-pay target
- continues to press after the customer says the amount is wrong
- keeps talking after a stop-calling request
- misstates the consequence of non-payment to create urgency
- treats a disclosed illness or bereavement as an objection to overcome
What most employers do instead
CV screen for prior collections or sales tenure, a phone screen, an induction-week compliance module with a multiple-choice test, and a scripted role-play with a team leader who is reading the same script.
The assessment
About 30 minutes end to end.
The systems it runs in
A first-party collections and recovery platform sitting over the lender's own account: the balance and the arrears stage, a promise-to-pay the agent books with an amount and a date, a payment arrangement or instalment plan, a place to record income and essential outgoings, contact-preference and channel flags, an account query or dispute marker, and a support or vulnerability marker on the customer. Finvi describes Katabat as collections and recovery software for first-party lenders such as banks and fintechs; C&R Software describes Debt Manager as an all-in-one collections solution used by banks and other creditors on past-due accounts. Every conduct behaviour this design scores is a field in that shape, not a sentence in a free-text box, which is the point of naming it.
- Katabat
- C&R Software Debt Manager
Any collections platform, or the collections module of a core banking or loan servicing system, that carries an account in arrears, a promise-to-pay with an amount and a date, an affordability or income-and-expenditure record, contact-preference flags and a support marker. The fixture is rebuilt against the buyer's own instance and their own arrears stages where they provide a sandbox; where they do not, the generic shape above is used and the buyer is told which fields it assumes.
What the candidate actually does
| Task | What happens |
|---|---|
| The customer who offers too much live_call · 14 min | An AI customer two payments behind on a personal loan taken with the candidate's own employer. The brief given to the candidate states a weekly promise-to-pay target and says they are behind it, because the cheap path has to be genuinely attractive for the task to measure anything. The customer opens cooperative and embarrassed and, within ninety seconds, volunteers to clear the whole arrears on Friday. That offer is not affordable and the customer does not yet know it; if asked, they describe an income and a set of essential outgoings that leave roughly a third of the offered sum. Three forks are planted. The customer mentions in passing, inside a sentence explaining why the payments slipped, that they have been signed off work since a diagnosis, and never returns to it. The customer asks whether the car will be taken, which the scenario brief does not support and which the candidate can leave standing, correct, or amplify. And the customer, if talked down from the Friday figure, resists once because they want the debt gone. Booking the Friday promise ends the call in six minutes and hits the target. |
| The arrangement the customer receives written_artifact · 7 min | The confirmation the customer gets after that call, plus what the candidate leaves behind in the collections system: the promise-to-pay or arrangement booked with its amount and date, the contact-preference flag if one was requested, the query marker if the balance was disputed, and the account note. It is scored against the candidate's own transcript, not against a model answer, so a candidate who agreed a poor arrangement cannot recover marks by writing up a better one. The temptation is to book the figure the system was expecting rather than the figure actually agreed, and to leave the disclosure out of the note because recording it feels intrusive. Requiring the fields rather than only the prose is what separates an agent who said the right thing from one who did it: a contact preference honoured in conversation and never flagged binds nobody, and the next agent calls the workplace again on Monday. |
| Four next turns judgment_scenario · 8 min | Four short call extracts, each stopping mid-conversation. The candidate states what they say next and gives a one-sentence reason. A customer says not to call them at work. A customer's partner answers and asks what the call is about. A customer says they will put the arrears on a credit card to clear it today. A customer says the balance is wrong because a payment in June is missing. In each item the answer that keeps the collection moving is the wrong one, and in the credit-card item the wrong answer is the one that books a payment inside the hour. |
The mark scheme
Each criterion is scored 1 to 5 against written anchors, and every score is reported with the excerpt that earned it. A criterion marked floored is reported as a finding rather than averaged into the total. The first is open; open any other to read its anchors in full.
Establishes affordability before accepting a figure (FCA CONC 7.3.4R, forbearance and due consideration; CONC 7.3.10R, no pressure to pay unreasonably large amounts in an unreasonably short time)weight 0.25Establishes stated income and essential outgoings, says the arithmetic out loud, and proposes or accepts a figure lower than the one offered, explaini…
Changes something after the customer discloses a change of circumstance (FCA FG21/1, chapter 3, frontline staff must have the skills to recognise and respond to characteristics of vulnerability; four drivers include health and life events)weight 0.2Acknowledges it, asks one question about how it affects what the customer can pay or when, and changes at least one of the amount, the date or the con…
States consequences accurately, neither softened nor inflated (12 U.S.C. 5531(d)(2)(A), taking unreasonable advantage of a consumer's lack of understanding of the material risks, costs or conditions of the product)weight 0.2States what will happen, at what point, and which parts are automatic and which are decisions somebody makes, using only what the brief supports, and …
Treats a contact-preference request, a third-party pickup and a disputed amount as instructions rather than obstacles (scored against the scenario brief's stated contact policy; the workplace-contact and third-party items follow the pattern set by 15 U.S.C. 1692c(a)(3) and 1692c(b), which most first-party policies copy but which do not themselves bind an employee of the creditor collecting in the creditor's own name under 15 U.S.C. 1692a(6)(A))weight 0.15Complies in the turn the request is made, records the constraint where the next agent will see it, and for the disputed payment stops asserting the co…
Written arrangement matches what was actually agreed on the callweight 0.2Amount, date, missed-payment consequence and the customer's query are all present and consistent with the transcript, and the disclosure is recorded i…
How it is scored
Weighted mean of the five criteria, each scored 1 to 5 against the anchors above, reported with the transcript excerpt or document line that earned each score. Two criteria are also reported separately as an unweighted conduct pair, because a buyer usually wants to see the affordability and consequence scores without them being averaged away by strong performance elsewhere. No criterion is a hard gate; that decision belongs to the deployer, who is told what a 1 on the consequence criterion means before they set one.
Integrity
- monitored session with the call and the written task in one unbroken sitting
- the written task is scored against the candidate's own transcript, so a prepared answer cannot fit it
- one live follow-up question asking the candidate to say how they arrived at the figure they agreed
- timing anomaly flags between the end of the call and the start of the note
- scenario variants rotated so the planted disclosure differs between sittings
The log describes what happened. It does not produce a cheating verdict — the follow-up conversation is the control, because a statistical accusation is not something we would ask a reviewer to defend.
What you receive
- full call transcript with the three planted forks time-marked
- the confirmation and account note as submitted
- the account as the candidate left it, being the arrangement booked with its amount and date, and any contact-preference, query or support flag set
- per-criterion score with the excerpt that earned it
- the four judgment answers with their stated reasons
Who decides
Required, not optional, for the consequence criterion and the disclosure criterion. A reviewer reads the two time-marked excerpts and confirms or overrides the score with a written reason, because both turn on whether a sentence was accurate given the brief, and that is a content judgment a human makes in seconds and an automated scorer makes badly. The reviewer is looking at the turn immediately after the disclosure and at every sentence in which a consequence is asserted. The ranking entitles the buyer to conclude that this candidate did or did not do these specific things in this specific conversation. It does not entitle them to conclude that the candidate is compliant, empathetic, or suited to the role in general, and it is not a substitute for the employer's own conduct training and monitoring after hire.
What this does not measure
This assessment does not read accent, dialect, fluency, vocabulary range, vocal warmth or manner of speech, and no rubric criterion is satisfiable by sounding a particular way. Every anchor above is a thing the candidate did or did not do, and the two anchors most likely to smuggle manner back in, the disclosure response and the affordability discussion, are written as changes to the figure, the date or the contact plan rather than as tone. It does not score call duration, talk-time ratio or promise-to-pay conversion, and a deployer who adds those metrics back on top has rebuilt the selection problem this design exists to solve. It does not ask the candidate to infer anything about the customer's circumstances: the disclosure is scripted, explicit and spoken aloud, and a design that rewarded good guesses about who is struggling would be rewarding stereotyping. The written task grades content and not surface form, and spelling, punctuation and idiom carry no weight, which matters because this seat is staffed heavily by second-language speakers. Deployers should monitor score distributions by hiring channel, region and language background, and should watch specifically for a gap on the written task, which is where a construct-irrelevant literacy effect would show up first. Candidates who need longer for the written task or a screen reader should get it; the time limits are there to control cost, not to measure speed.
The design starts from an unusual position: the regulator has effectively published the answer key for this seat, and the industry screens against none of it. Early arrears is first-party work, so the agent sits outside the FDCPA's definition of debt collector under 15 U.S.C. 1692a(6)(A), and a great many induction decks stop there. What actually governs the call is the Dodd-Frank prohibition on abusive practices at 12 U.S.C. 5531(d), which expressly covers taking unreasonable advantage of a consumer's lack of understanding of the material risks, costs or conditions of a product, and in the UK CONC 7.3.4R, which requires firms to treat customers in or approaching arrears with forbearance and due consideration. Both of those are about what happens in the middle of a specific conversation, which is why a multiple-choice module cannot reach them and a fifteen-minute simulation can.
The central fork is deliberately not the one buyers expect. Most collections screening imagines the failure mode as an agent who is too aggressive, so it tests whether the candidate can avoid shouting. The expensive failure is the opposite shape and much quieter: a customer offers more than they can afford, the agent takes it, the promise-to-pay is booked, the arrangement breaks in three weeks, and the account re-enters the queue in worse condition than it left. Every collections operation has this pattern in its own data and almost none of them hire against it, because promise-to-pay is the number on the wallboard. So the AI customer opens by offering the full arrears, the candidate is told they are behind target, and the correct move is to talk the customer down. That is a fork where the cheap path is not merely tempting but actively rewarded by the metric the candidate has spent their career being measured on.
The vulnerability fork is built to fail the way it fails in production. Most candidates who miss it are not unkind. They miss it because the disclosure does not arrive in the format the training deck used: it is one clause inside a sentence about something else, said without emphasis, never repeated. FG21/1 asks firms to ensure frontline staff have the skills to recognise and respond to characteristics of vulnerability across four drivers including health and life events, and recognition here means noticing what was said, not inferring anything about the person. The rubric therefore scores a change in the ask, the date or the contact plan, and explicitly not sympathy, because sympathy is cheap, unfalsifiable, and shades directly into scoring manner.
The third fork is the sentence that manufactures urgency. Agents rarely threaten anything; they let a fear stand. The customer asks whether the car will be taken, and saying nothing is free, accurate-sounding and enormously effective. The anchors treat leaving that fear uncorrected as a 1, at the same level as asserting the consequence outright, because the effect on the customer is the same and 5531(d)(2)(A) is about the consumer's understanding, not the agent's literal words.
What this design refuses to measure is anything about how the candidate sounds. That refusal is not decorative. A collections floor staffed by second-language speakers, hired at volume, in a seat where warmth is the folk theory of good performance, is exactly the setting in which an assessment built on impressions would do real damage while looking rigorous. Every criterion here resolves to a turn a reviewer can point at, and the written task is scored against the candidate's own call, which is the cheapest available guard against both preparation and impression.
Sources
Every figure on this page is traceable. Where a claim could not be sourced it is stated qualitatively instead.
- US Bureau of Labor Statistics, Occupational Outlook Handbook, Bill and Account Collectors, 2025, https://www.bls.gov/ooh/office-and-administrative-support/bill-and-account-collectors.htm
- Consumer Financial Protection Bureau, Fair Debt Collection Practices Act: CFPB Annual Report 2025, November 2025, https://files.consumerfinance.gov/f/documents/cfpb_fdcpa-2025-annual-report_2025-11.pdf
- 15 U.S.C. 1692a(6), definition of debt collector and the exclusion for an officer or employee of a creditor collecting in the creditor's own name, https://www.law.cornell.edu/uscode/text/15/1692a
- 12 U.S.C. 5531(c)-(d), Dodd-Frank standards for unfair and abusive acts or practices, https://www.law.cornell.edu/uscode/text/12/5531
- FCA Handbook, CONC 7.3.4R (forbearance and due consideration), CONC 7.3.10R (pressurising a customer to repay in one payment or by further borrowing) and CONC 7.3.14R (disproportionate action), https://www.handbook.fca.org.uk/handbook/CONC/7/3.html
- Financial Conduct Authority, FG21/1 Guidance for firms on the fair treatment of vulnerable customers, 23 February 2021, https://www.fca.org.uk/publications/finalised-guidance/guidance-firms-fair-treatment-vulnerable-customers
See what the employer actually receives. A full report for one role, with every score shown beside the excerpt that earned it, conduct findings reported rather than averaged, and a reviewer sign-off required before any decision. No form.
Read a sample reportOr talk to us about this role