Taking an assessment rather than buying one? This page is written for employers. Here is the page for candidates.
Collections and retention · Mid level
How to assess a B2B Credit Controller
The reconciliation test screens for the half of the job that finance systems are steadily automating, and it screens for it in isolation from the half that is actually hard. The hard half is a negotiation with a professional counterparty who has their own cash-flow problem, their own boss, and a standing relationship with your sales team that somebody senior wants preserved. Nothing in a spreadsheet exercise reveals whether this candidate can sit on a call where the customer's accounts payable manager says the invoice is disputed, decide in real time whether that is true or convenient, and choose between holding the account, accepting a part-payment, and handing it to the account owner. That decision is the job, it is made dozens of times a quarter, and the current process observes it exactly never.
This role is filed inside a family marked `regulated: true` and is itself marked false, deliberately, because the distinction is load-bearing and is constantly got wrong in job descriptions that treat credit control and consumer collections as the same discipline. The FDCPA defines "debt" at 15 U.S.C. 1692a(5) as an obligation of a *consumer* arising out of a transaction in which the money, property, insurance or services are primarily for personal, family or household purposes. Commercial trade debt is not that. Neither the FDCPA's communication rules nor the FCA's CONC 7 arrears regime apply to a credit controller chasing an unpaid invoice from another business, and a hiring process that imports consumer collections compliance wholesale into this seat is screening for the wrong thing. Two caveats belong on the page rather than in a footnote: many US states have their own commercial collection statutes and licensing regimes, and a controller whose portfolio includes sole traders or unincorporated customers may be handling consumer debt without realising it. Neither changes the shape of the assessment; both belong in the employer's own compliance gate.
What replaces the compliance fork is a commercial one, and it is genuinely harder to screen for. The counterparty is not distressed and not embarrassed; they are a professional whose own incentive is to hold cash, and who has an extensive repertoire for doing so politely. The invoice is queried. The purchase order number does not match. The approver is on leave. The goods were short. Some of these are true. The controller's skill is not scepticism — a controller who disbelieves everything burns the relationship and gets escalated over — it is the ability to test a stated dispute cheaply and quickly, on the call, by asking for the one specific thing a genuine dispute can produce and an invented one cannot. Which line is queried. What the delivery note says. Who signed for it.
The second competency is knowing whose decision it is. A credit controller who unilaterally puts a top-ten account on credit hold on a Friday afternoon has usually created a bigger problem than the receivable, and one who never puts anything on hold has no leverage at all. The judgment is about escalation routing: when to make the call themselves, when to loop the account owner before acting, when to route a payment-plan request to whoever owns the credit limit. This is the seat's equivalent of the escalation-timing judgment in inbound support, and like that one it is invisible in an interview because every candidate describes themselves as commercially minded.
The third is documentary discipline, and it is where the data half of the assessment earns its place. A promise extracted on a call is worth nothing unless the confirmation email states the amount, the date, what happens if it is missed, and what remains disputed and therefore unpaid. Controllers who write that email loosely create ledgers that nobody can age accurately, and the aged-debt report becomes fiction six weeks later. This is directly assessable: give the candidate a messy ledger with a duplicate credit note and an unallocated remittance, have them work out what is genuinely owed, then have them defend that number on a call against a customer who has a different figure and is confident about it.
Savvanta's design therefore pairs a data task with a live call, in that order, and scores the consistency between them: whether the number the candidate defended on the call is the number their own reconciliation supports, whether the dispute was tested or accepted, and whether the follow-up email would let a colleague pick up the account cold. What it does not assess is anything about the candidate's manner or accent; the criterion for the call is what was agreed and what was written down afterwards.
What the job actually needs
- distinguishing a genuine invoice dispute from a delay tactic
- negotiating a payment plan against a counterparty with its own cash-flow problem
- protecting the commercial relationship while holding the ledger
- escalating to the account owner at the right moment
- reconciling a remittance against a messy ledger
How people fail in this seat
- treats a legitimate dispute as an excuse and chases the full balance for six weeks
- agrees a plan without escalating the credit limit decision
- puts a strategic account on stop without warning the account owner
- accepts a partial payment that quietly restarts a limitation clock
- chases the wrong contact for two months because nobody validated who approves payment
What most employers do instead
A CV screen for ledger size and ERP names, an accounts-receivable reconciliation or Excel test, and an interview about aged debt reduction.
The assessment
About 35 minutes end to end.
The systems it runs in
The accounts receivable sub-ledger in an ERP — open invoices, credit notes, unallocated cash, aging buckets, the credit limit and the credit hold flag — worked through a receivables platform layered over it that supplies the collections worklist, the dunning correspondence, the dispute or deduction code and the promise-to-pay. Versapay describes that layer as accounts receivable automation across the whole invoice-to-cash cycle, covering collections management and cash application, and names Oracle NetSuite, Sage Intacct and Microsoft Dynamics 365 among the ERPs it connects to. That two-layer shape is the point of the first task: the aged-debt total the candidate is tempted to chase is what the ledger shows, and disagreeing with it is a deliberate act against the system of record.
- Versapay
- Oracle NetSuite
- Sage Intacct
- Microsoft Dynamics 365
Any ERP or accounting system with an AR sub-ledger and an aging report, with or without a collections layer over it. The reconciliation task is presented as a plain table precisely so that no spreadsheet or platform skill is scored, which the adverse_impact section states; the hitl note is equally explicit that a score here does not establish knowledge of this buyer's ERP, credit policy or escalation thresholds, all of which are training. Rebuilt against the buyer's own ledger extract where they supply a redacted one.
What the candidate actually does
| Task | What happens |
|---|---|
| What is actually owed data_task · 12 min | A single customer's ledger, deliberately messy in the ordinary way real ledgers are messy. Nine open invoices, a credit note that has been posted twice, an unallocated remittance for a round sum that matches no single invoice but matches three of them together less a settlement discount, and one invoice raised against a purchase order number that appears on no delivery note. The candidate states the figure they will chase, the figure that is properly in question, and their working. The trap runs in both directions, which is what makes it a fork rather than an arithmetic test. The cheap path is to chase the aged-debt report total, because that is the number the system shows and the number the collections target is set against. The over-corrected path concedes the duplicate credit note twice and hands back money the customer never claimed. |
| The accounts payable manager live_call · 13 min | An AI counterpart who is a professional, not a distressed consumer. Polite, well briefed, and holding cash on purpose. They raise three obstacles in sequence. One is a genuine dispute, a short delivery on a specific invoice, and it will survive any test the candidate applies. One is the approver being on leave. One is a purchase order mismatch on an invoice that in fact matches, which the candidate's own reconciliation should already have established. Midway through, they offer a part-payment as a gesture in exchange for the account coming off credit hold today, and they mention their own cash-flow problem and the long relationship with the candidate's sales team. The fork is that the part-payment closes the call with a result, and taking it means conceding a hold decision that is not the controller's to make. |
| Two documents that have to agree written_artifact · 8 min | The confirmation email to the accounts payable manager and the internal note to the account owner, written back to back, plus what the candidate records on the account itself: the dispute or query marker against the one invoice that is genuinely in question, the promise-to-pay for the undisputed balance, and whether the credit hold is left as it stands with the decision raised, or changed. That last field is the whole escalation criterion made observable — a controller who says on the call that the hold is not theirs to lift and then lifts it, and one who says it and routes it, produce identical transcripts. The temptation is asymmetry: a warm, vague email to the customer and a franker note internally, or the reverse. The correct pair states the same amount, the same date, the same list of what remains disputed and therefore unpaid, the same evidence requested with a deadline, and flags the hold decision to the account owner as pending rather than made. |
The mark scheme
Each criterion is scored 1 to 5 against written anchors, and every score is reported with the excerpt that earned it. A criterion marked floored is reported as a finding rather than averaged into the total. The first is open; open any other to read its anchors in full.
The number defended on the call is the number the candidate's own reconciliation supportsweight 0.25Names the figure, explains on the call which invoices make it up, which credit and which remittance have been applied, and identifies the specific lin…
Tests a stated dispute instead of believing or disbelieving itweight 0.2For each obstacle names the one artefact that would resolve it, the delivery note and who signed for it, the queried line, the purchase order the invo…
Routes the decisions that are not the controller's to make, rather than taking or refusing them aloneweight 0.15Declines to decide it on the call, names who does decide it and by when, commits to a specific next contact, and alerts the account owner in the inter…
Holds the ledger without spending the relationshipweight 0.15Separates the genuine dispute from the convenient ones explicitly on the call, offers a route for the genuine one, and holds the rest, so that the cou…
The customer email and the internal note say the same thing and would let a colleague actweight 0.25Both state the amount, the date, the specific evidence requested and its deadline, and what remains disputed and therefore unpaid, and the internal no…
How it is scored
Weighted mean of the five criteria, each scored 1 to 5 against the anchors, reported with the excerpt or ledger line that earned each score. The reconciliation task is scored on the stated figure and the working, not on the method or the tool used to reach it. Because the first and last criteria both test consistency across artefacts, the report shows the three figures side by side, the one in the working, the one spoken on the call and the one in the email, which is usually the single most informative line in the whole output.
Integrity
- monitored session with all three tasks in one unbroken sitting
- the call and the written pair are scored against the candidate's own reconciliation, so a prepared answer cannot fit
- one live follow-up question asking the candidate to walk through how the unallocated remittance was applied
- ledger variants rotated so the position of the duplicate credit note and the genuine dispute differ between sittings
The log describes what happened. It does not produce a cheating verdict — the follow-up conversation is the control, because a statistical accusation is not something we would ask a reviewer to defend.
What you receive
- the reconciliation with the candidate's stated figure and working
- full call transcript with the three obstacles time-marked
- both written documents as submitted
- the account as the candidate left it, being the dispute marker, the promise-to-pay booked against the undisputed balance, and the state of the credit hold
- per-criterion score with the excerpt or ledger line that earned it
Who decides
Recommended rather than required, and the reviewer's job here is narrower than in the consumer collections designs. One decision needs a person: whether the candidate's treatment of the genuine dispute was commercially right for this buyer, because a controller in a business with a strategic customer base and a controller in a business selling to a long tail of small accounts should behave differently on that call, and the rubric cannot encode which one the buyer is. The reviewer reads the time-marked exchange on the short-delivery invoice and confirms or overrides the relationship criterion with a written reason. The ranking entitles the buyer to conclude that this candidate can reconcile a messy ledger and defend the resulting number against a prepared counterparty. It does not establish that they know this buyer's ERP, this buyer's credit policy, or this buyer's escalation thresholds, all of which are training, and it says nothing about whether the candidate holds any state commercial collection licence the employer may require.
What this does not measure
This is the one role in this family where the assessment carries no consumer credit compliance criteria at all, and that is a deliberate design decision rather than an omission. The FDCPA defines debt at 15 U.S.C. 1692a(5) as an obligation of a consumer arising from a transaction primarily for personal, family or household purposes; commercial trade debt is not that, and neither the FDCPA's communication rules nor the FCA's CONC 7 arrears regime reach a controller chasing an unpaid invoice from another business. Importing consumer collections criteria here would look rigorous and would be measuring a rule that does not apply, penalising candidates for not performing compliance theatre. Two caveats belong in the deployer's own gate rather than in this rubric: several US states have their own commercial collection statutes and licensing regimes, and a controller whose portfolio includes sole traders or unincorporated customers may be handling consumer debt without realising it. The design does not read accent, dialect, fluency or manner of speech, and no anchor is satisfiable by sounding commercially confident. The most likely source of construct-irrelevant disadvantage is the data task: it is presented as a plain table, any arithmetic method is permitted including a calculator or paper, no spreadsheet formula use is scored, and speed is not scored within the time box. Deployers should still monitor the data task separately by group, because a gap that appears there and not on the call points at numeracy-test format familiarity rather than at the competency. Candidates who need extended time or a screen reader should have it; the time limits exist to control cost.
Every other assessment in this family is built around a fork where the profitable move and the permitted move come apart. This one is not, because for this seat there is no such fork, and pretending otherwise would be the single most common error in how credit control is hired for. The FDCPA's definition of debt is confined to consumer obligations, so the communication rules that dominate the rest of this family are simply absent here, and a screen that tests them is testing something the job does not contain. What replaces the compliance fork is a commercial one, and it is genuinely harder to observe.
The counterparty in this design is not distressed or embarrassed. They are a professional whose own incentive is to hold cash and who has a large, polite repertoire for doing it: the invoice is queried, the purchase order does not match, the approver is on leave, the goods were short. Some of these are true. The controller's skill is not scepticism, because a controller who disbelieves everything gets escalated over and loses the account, and it is not credulity either. It is the ability to test a stated dispute cheaply and on the call, by asking for the one specific thing a genuine dispute can produce and an invented one cannot. That is why the 5 anchor on the dispute criterion asks for a named artefact and a date rather than for a general challenge, and why the scenario includes one obstacle that will survive testing. A design where every obstacle was a delay tactic would reward blanket aggression and would teach the wrong thing to anyone who took the report seriously.
Sequencing the data task first is the load-bearing choice. Reconciliation tests are common in this market and they are testing the half of the job that finance systems are steadily automating. Run in isolation they measure arithmetic. Run immediately before a call in which the candidate has to defend their own number against a confident counterparty with a different figure, the same exercise measures something else entirely: whether the candidate believes their own working under pressure, and whether the number they defend is the one they computed. The report shows those two figures next to the one in the confirmation email, and the gap between them is frequently the most diagnostic thing in the output. Candidates who reconcile well and then quietly chase the system total on the call are common, and no existing screen surfaces them.
The escalation criterion covers the failure that costs the most and is complained about the least. A controller who unilaterally puts a top-ten account on hold on a Friday afternoon has created a bigger problem than the receivable. A controller who never holds anything has no leverage and the ledger ages. The judgment is about routing rather than courage, and it is invisible in an interview because every candidate describes themselves as commercially minded. Here it is a specific moment: the counterpart offers a part-payment in exchange for the hold coming off today, and the candidate either takes a decision that is not theirs, defers it into a void, or routes it with an owner and a date.
The written pair closes the loop, because a promise extracted on a call is worth nothing unless the confirmation states the amount, the date, what happens if it is missed, and what remains disputed and therefore unpaid. Controllers who write that email loosely produce ledgers nobody can age accurately, and the aged-debt report becomes fiction six weeks later. Asking for the internal note alongside it exposes a specific and common asymmetry, where the customer is told something softer than the account owner is, and both readers later act on their own version.
Sources
Every figure on this page is traceable. Where a claim could not be sourced it is stated qualitatively instead.
- US Bureau of Labor Statistics, Occupational Outlook Handbook, Bill and Account Collectors, 2025, https://www.bls.gov/ooh/office-and-administrative-support/bill-and-account-collectors.htm
- 15 U.S.C. 1692a(5), definition of debt as an obligation arising out of a transaction primarily for personal, family or household purposes, https://www.law.cornell.edu/uscode/text/15/1692a
See what the employer actually receives. A full report for one role, with every score shown beside the excerpt that earned it, conduct findings reported rather than averaged, and a reviewer sign-off required before any decision. No form.
Read a sample reportOr talk to us about this role