Get started

Taking an assessment rather than buying one? This page is written for employers. Here is the page for candidates.

Accounting and finance operations · Entry level

How to assess a Accounts Payable Specialist

Matched invoices now post themselves, so a keying-speed test measures the part of the job the software already does and the CV screen measures which vendor's buttons the candidate has pressed. Neither observes the only moment in the role that moves money: what the candidate does when the invoice, the purchase order and the goods receipt disagree, the requisitioner says pay it, and the supplier is threatening to stop shipping. Nor does anything in the process test the single most expensive judgment an AP clerk makes — whether to act on a change of supplier bank details — which is not a knowledge question that a quiz can ask, because every candidate knows the right answer in the abstract and a meaningful number of them still process the payment.

An accounts payable specialist spends the working day on the invoices that did not go through. The ones that did — the great majority, where the invoice matches a purchase order and a goods receipt within tolerance — are posted by the system before anyone looks at them, and that share rises every year. What lands in a human queue is the residue: an invoice with no PO because a manager ordered directly from a supplier they like, a price three percent above the agreed rate, a delivery receipted for eight units against an invoice for ten, a credit note that references an invoice number nobody can find, and a supplier emailing twice a day about a payment that was already made against a different entity in the group.

The job is to decide, for each of those, whether to pay, hold, or query — and then to say so to somebody in writing. That last part is what most job descriptions leave out and most hires are weakest at. Every exception in AP has a person attached: a supplier who wants their money, and a budget holder who wants the goods and who did not read the purchase order process. The specialist sits between them with no authority over either. A median performer resolves this by deferring to whoever pushed hardest, which reliably means the price increase gets paid and the discrepancy is never recorded. A top-quartile performer resolves it by writing three sentences — this is the discrepancy, this is what I checked, this is what I recommend — and sending them to the person who can actually decide. The difference is not diligence. Both candidates checked. The difference is whether they wrote it down and put it in front of someone, which is what turns a personal query into a decision the business owns.

The second thing that separates the top quartile is a specific kind of suspicion. Duplicate payments do not arrive labelled; they arrive as the same work re-billed under a new invoice number, or the same invoice submitted once by the supplier's portal and once by email from a different address. And the most consequential exception of all looks entirely ordinary: a supplier writes to say their bank details have changed, please update the account before the next run. The FBI's IC3 reported $3,046,598,558 in business email compromise losses in 2025, the second-largest cyber-enabled fraud loss category behind investment fraud, and payables is the function these attacks are aimed at. A strong AP hire treats that email as an out-of-band verification task rather than an administrative update, and does so even when the request is well written, plausibly urgent, and apparently from a supplier they recognise.

What a hiring manager is actually trying to predict, then, is not throughput. Throughput is a function of the system and of tolerances the controller sets. What they are trying to predict is where this person's threshold sits: how large a discrepancy has to be before they stop, how much pressure it takes to move them off that threshold, and whether they can write the exception up in a form that a non-finance budget holder will act on rather than ignore. That is observable in a single well-built invoice batch, and it is observable in essentially nothing that a normal AP recruitment process currently does.

What the job actually needs

How people fail in this seat

What most employers do instead

CV screen for ERP and P2P tool names (SAP, Oracle, NetSuite, Coupa, Sage), a competency interview about attention to detail, and occasionally a data-entry speed and accuracy test.

Matched invoices now post themselves, so a keying-speed test measures the part of the job the software already does and the CV screen measures which vendor's buttons the candidate has pressed. Neither observes the only moment in the role that moves money: what the candidate does when the invoice, the purchase order and the goods receipt disagree, the requisitioner says pay it, and the supplier is threatening to stop shipping. Nor does anything in the process test the single most expensive judgment an AP clerk makes — whether to act on a change of supplier bank details — which is not a knowledge question that a quiz can ask, because every candidate knows the right answer in the abstract and a meaningful number of them still process the payment.

The assessment

About 32 minutes end to end.

The systems it runs in

A purchase-to-pay queue in an ERP, where the invoice, the purchase order and the goods receipt sit against one vendor record, the bill carries a state the candidate sets, and a tolerance is configured on the account rather than held in the candidate's head. NetSuite's vendor bill approval workflow is the reference build for the fixture: published Oracle documentation describes installing the approvals SuiteApp, enabling approval routing on vendor bills, and setting tolerance and difference levels, which is precisely the shape this batch needs. The disposition is therefore a record in the system — the bill is approved, held or queried — and not a note about a decision taken elsewhere. Two further surfaces matter. The vendor master is where the bank-detail request has to be worked, because the correct move is to check the record the business already holds against a route it already held, and that is only a real action if the record exists. And a spreadsheet sits beside the queue for the batch working: the three-way comparison across invoice, order and receipt is a reconciliation across three lists that do not share a key column, and what is scored is whether the twelve lines came out right, never which function got them there.

Any system that holds order, receipt and invoice against one vendor, carries an approve/hold/query state, exposes a configurable tolerance, and keeps verified bank details on the vendor record separately from correspondence. The fixture is rebuilt inside the buyer's own instance where they provide a sandbox; where they do not, the twelve lines ship as a plain table with the supporting documents attached and the queue states are recorded on the face of it. No criterion is satisfied by knowing where a menu item lives in any named product: the platform decides what the artefact looks like, not what earns a score.

Working speed is scored. Throughput against a treasury cut-off is a real constraint in this seat and the brief states it, so how much of the batch was dispositioned inside the window is scored — but only in combination with the calibration and duplicate criteria, never alone. A candidate who cleared all twelve fastest and processed the bank-detail change has produced the worst result in the set, and the report says so on its face. What is not measured anywhere is keying speed: there is no typing component, arithmetic may be done by any method, and the clock runs on decisions rather than on data entry.

What the candidate actually does

TaskWhat happens
The brief
judgment_scenario · 2 min
Read once, identical for every candidate, and scored on nothing. You are the only AP specialist covering this ledger. The payment run closes at 16:00 and the treasury cut-off cannot be moved. Your team is measured on invoices cleared per day and there are 140 items behind the twelve in front of you. Two of these suppliers have already put other customers on stop this month. Tolerance is stated as a number, so that over-stopping is a scoreable error and not a matter of taste. Nothing in the brief asks the candidate to flag anything, query anything, or look for anything unusual; it asks them to process the run.
The fork. This segment exists to make the fork in t1 expensive. Without a stated cost to stopping, every candidate stops everything and the batch measures caution rather than calibration.
The payment run
data_task · 14 min
Twelve invoices with whatever purchase order and goods receipt exists for each, presented as a plain table with the supporting documents attached. Two match cleanly and must be paid. One sits inside the stated tolerance and must also be paid. One is priced above the purchase order, with a covering note from the supplier saying the increase was agreed verbally with the buyer, and no evidence anywhere in the pack that it was. One invoices ten units against a goods receipt for eight. One arrives with no purchase order at all from a supplier a manager uses directly. One is the same work, same amounts, same description, re-billed six weeks later under a new invoice number. One credit note references an invoice number that appears nowhere. And attached to one supplier's account is a courteous, well-written email from the contact the candidate has been corresponding with all week, giving new bank details and asking that they be used for this run — under which the requisitioner has already replied, visibly, "yes that's them, go ahead." The candidate records a disposition against every line as the bill's own state in the payables queue rather than as a comment beside it — pay, hold or query — with the reason and what they checked. Dispositions are committed before t3 opens and cannot be revised.
The fork. Paying the price variance clears the queue, the supplier is grateful, and nobody ever finds out; querying it costs the run and makes the candidate the clerk who blocks everything. The bank-detail change is the sharper fork: it is polite, it is from a real contact, the budget holder has already endorsed it in writing, and processing it is the only response that produces no friction whatsoever today.
Three sentences to the budget holder
written_artifact · 7 min
The note to the non-finance manager who raised the purchase order behind the price variance. Short by design, and graded against the candidate's own dispositions rather than a model answer, so a prepared paragraph cannot fit. The recipient does not work in finance, does not know what a three-way match is, and will act on this or ignore it depending entirely on how it is written.
The fork. The fastest defensible-looking note asks a question — "can you confirm this is approved?" — which moves the paperwork without moving the decision, and reliably returns "yes, fine" from a manager who wants their goods. Writing the discrepancy, the check and a recommendation with a deadline attached is slower, and is the only version that leaves a record the business owns.
Six more exceptions
judgment_scenario · 9 min
Six short cards, three or four lines each, answered with a disposition and one line of reasoning. A supplier threatening to stop supply over a disputed line. A requisitioner saying pay it now and the purchase order will follow. A credit note that would conveniently clear an aged debit nobody can explain. A supplier statement that disagrees with the ledger by one invoice. A first invoice from a new supplier with no bank record on file. A same-day payment request from a director who is travelling and will sign when they land. The cards are answered after the batch is locked, so they resample the same threshold in situations the candidate cannot align to their earlier answers.
The fork. Each card offers a cheap path that is socially correct in the moment: keep the supplier shipping, help the manager who is under pressure, clear an old debit, do what the director asked. Two of the six are cards where the compliant answer costs the candidate something in front of someone more senior than them.

The mark scheme

Each criterion is scored 1 to 5 against written anchors, and every score is reported with the excerpt that earned it. A criterion marked floored is reported as a finding rather than averaged into the total. The first is open; open any other to read its anchors in full.

Stop-threshold calibration across the batchweight 0.25Pays the three that should be paid including the tolerance item, stops the four that should be stopped, and for each stop names which of the three doc…
1 Pays every discrepancy in the batch, or holds nine or more of the twelve including the two clean invoices and the within-tolerance item. Both directions score 1; a candidate who stops everything has not demonstrated judgment, they have demonstrated that they are unwilling to make a call.
3 Stops the large price variance and the short delivery, but either misses the duplicate or the missing purchase order, or stops the within-tolerance item, so the queue is held for something the buyer has already decided not to care about.
5 Pays the three that should be paid including the tolerance item, stops the four that should be stopped, and for each stop names which of the three documents disagrees and by how much. Stated reasons are specific to the line rather than a general statement about needing approval.
Finds the re-billed work without being told there is a duplicateweight 0.15Identifies the earlier invoice it repeats, states what matches — the amounts, the period, the description — and holds it as a suspected duplicate rath…
1 Pays it. The invoice number is new, the date is current, and nothing on the face of the document indicates it has been seen before.
3 Notices that the amount or the description looks familiar and queries it as an unclear item, without establishing what it duplicates.
5 Identifies the earlier invoice it repeats, states what matches — the amounts, the period, the description — and holds it as a suspected duplicate rather than a general query, so the supplier conversation starts from evidence.
Verification before a payment destination changes flooredweight 0.2Treats the change as unverified regardless of who endorsed it, contacts the supplier through a route the business already held rather than one supplie…
1 Updates the bank details, or schedules the payment to them, on the strength of the email and the requisitioner's endorsement. This is reported on the face of the report as a finding and is never averaged into the total.
3 Hesitates and replies to the email asking for confirmation, which is verification through the same channel the request arrived on and would be answered by the same party. Or holds the payment without saying what would release it, leaving a supplier unpaid and no process started.
5 Treats the change as unverified regardless of who endorsed it, contacts the supplier through a route the business already held rather than one supplied in the request, holds only the payment and not the relationship, and says so plainly to the requisitioner without implying the supplier or the manager has done anything wrong.
The note converts a private query into a decision somebody ownsweight 0.25States the discrepancy against the purchase order in money, what was checked and what was not found, a recommendation, and what will happen by when if…
1 Asks the budget holder to approve, or states that the invoice is on hold pending approval, without saying what the discrepancy is or what was checked. The manager cannot tell from the note whether they are being asked to confirm a price they agreed or to authorise an overcharge.
3 States the discrepancy and the amount clearly but stops at a question, so the decision is transferred rather than informed, and there is no date by which silence becomes an outcome.
5 States the discrepancy against the purchase order in money, what was checked and what was not found, a recommendation, and what will happen by when if there is no reply. Uses no finance vocabulary the recipient would have to look up, and is consistent with the disposition the candidate actually recorded in t1.
Pressure from a supplier or an internal senior does not move a dispositionweight 0.15Holds the disposition and moves the problem, naming who can decide it, what would resolve it, and by when, so that the counterparty leaves with a next…
1 Reverses or waives on the cards where the pressure is applied — pays the disputed line to keep the supplier shipping, or releases the director's same-day payment on the strength of who asked.
3 Holds the line but has no route through it, so the answer is a refusal with nothing offered: the supplier is told no and the director is told the process does not allow it, with no alternative and nobody else brought in.
5 Holds the disposition and moves the problem, naming who can decide it, what would resolve it, and by when, so that the counterparty leaves with a next step rather than a grievance. Distinguishes the card where the pressure is legitimate commercial urgency from the card where it is a substitute for a control.

How it is scored

Weighted mean of the five criteria, each scored 1 to 5 against the anchors and reported with the invoice line or the sentence that earned it. Two raw numbers are reported beside the score and explicitly annotated as not positive signals: how many of the twelve lines the candidate paid, and how many they stopped. Buyers will ask for a throughput figure, and left unannotated they will read a high paid count as efficiency. In this batch both a high and a low stop count are failures, and the annotation says so in the report rather than in a footnote nobody opens. Where the disposition recorded in t1 disagrees with what the t2 note asserts, both are shown side by side; that gap is frequently the most informative line in the output.

Integrity

The log describes what happened. It does not produce a cheating verdict — the follow-up conversation is the control, because a statistical accusation is not something we would ask a reviewer to defend.

What you receive

Who decides

Recommended, with one decision that must be a person's. The tolerance in the brief is a policy the buyer sets, and the honest position is that a candidate who stops a variance this buyer would wave through is wrong for this buyer and right for a different one. The reviewer confirms the tolerance line before the batch is scored and overrides the calibration criterion in writing where their own policy differs. The ranking entitles a buyer to conclude that this candidate can work an exception queue, will stop for the right reasons, and can write an escalation a non-finance manager will act on. It does not establish that they know this buyer's ERP, this buyer's approval matrix, or this buyer's supplier base, all of which are training.

What this does not measure

This design deliberately does not measure keying speed, and no criterion rewards data entry. Where throughput is scored it is the count of lines dispositioned inside the run window, read only in combination with whether those dispositions were right. That is the point of the role page, and it is also the main adverse-impact control: a timed accuracy test at volume is the instrument most likely to disadvantage candidates for reasons unrelated to the competency, and it measures the part of the job the software already does. The batch is presented as a plain table, arithmetic may be done by any method including a calculator, and no spreadsheet formula use is scored. No criterion reads fluency, register or idiom in the written note; the note is scored on whether the discrepancy, the check, the recommendation and the deadline are present and consistent with the candidate's own decisions, and a note written in plain, unpolished English that contains all four scores 5. What the design does not reach is stamina: a real AP specialist works this judgment across a hundred and forty items a day, and twelve items in fourteen minutes samples the judgment without sampling the fatigue. Nor does it observe the candidate after six months, when the same supplier's third price increase arrives and the escalation has become socially expensive. Deployers should monitor the data task's score distribution separately by group, because a gap that appears there and not in the written note points at test-format familiarity rather than at the competency. Candidates who need extended time or assistive technology should have it; the time limits exist to control cost, not to create difficulty.

Everything in this design exists to locate one number that no CV carries: where this candidate's stop-threshold sits. An accounts payable specialist who stops nothing pays price increases nobody agreed to and duplicates that arrive with new invoice numbers. An accounts payable specialist who stops everything is a different and equally expensive problem, because the queue backs up, the suppliers escalate to the budget holders, and within a quarter the business has raised the tolerance to get the run out. The distance between those two failures is the whole hire, and it is a calibration question rather than a diligence question — which is why asking a candidate whether they are detail-oriented has never once produced a useful answer.

The batch is built so that both errors are available and both are punished. Three of the twelve lines are correct and must be paid, including one that differs from its purchase order by an amount inside the tolerance the brief states. That item is the control on over-stopping, and it is the reason the tolerance is a number in the brief rather than a matter of judgment: without it, a candidate who holds all twelve looks careful, and there is no scoreable difference between caution and paralysis.

The brief carries the commercial pressure explicitly, and every candidate reads the same sentences: the run closes at 16:00, throughput is the team's measure, there are a hundred and forty items behind these twelve, and two of these suppliers have put other customers on stop this month. This follows the pattern that a scenario with nothing at stake tests knowledge rather than resistance. Every candidate knows that a price variance should be checked. The question the buyer is paying to answer is what happens to that knowledge when checking costs the candidate their day's numbers, and if the design does not put anything on the table it cannot observe the answer. Equally, the brief does not tell the candidate to flag anything or to look for anything unusual. It tells them to process the run. Instructing them to raise exceptions would convert the whole exercise into a test of compliance with an instruction; leaving it out and planting the silent cases is what makes the raising itself the observation.

The bank-detail change is the most carefully built object in the file, because it is the one exception in accounts payable where a single decision moves an irrecoverable sum. Business email compromise is among the largest cyber-enabled fraud loss categories the FBI's IC3 reports, and payables is the function it aims at. Every candidate knows the rule in the abstract, which is exactly why a quiz on it is worthless. So the fixture is built to win socially before it fails: the request is courteous and well written, it comes from the contact the candidate has been dealing with all week, it references a real open invoice, and the requisitioner has already replied underneath it endorsing the change. Processing it is the only action that generates no friction with anyone today. The correct response — verifying through a route the business already held, rather than one supplied in the request — costs the candidate a small embarrassment with a manager who has already said it is fine. That criterion is floored: a score of 1 on it is reported on the face of the report as a finding and never averaged away, because a candidate who redirects a payment on an email must not surface as a strong hire on the strength of a tidy batch.

The three-sentence note is the second half of the role and the half most processes skip entirely. Both a median and a strong candidate check the price variance; the difference is whether anything leaves their head. The commonest weak artefact is a request for approval, which looks responsible, takes fifteen seconds, and reliably returns "yes, fine" from a manager who wants their goods and has not been told what they are approving. It transfers the decision without informing it, and it leaves a record that says the business agreed to the increase. Grading the note against the candidate's own dispositions rather than a model answer does two things at once: it makes a memorised template unfittable, and it surfaces the candidates whose note quietly asserts something different from what they actually decided fourteen minutes earlier.

The six cards resample the same threshold rather than lengthening the batch, because a stop-threshold observed once is close to luck, and because pressure applied once tells you very little. Two of the cards put the candidate in front of someone senior to them. Answering them after the batch is locked prevents the candidate from tidying their earlier answers into consistency, which is itself worth knowing about.

Thirty-two minutes is the budget, and the arithmetic is against a real alternative rather than an ideal one. This is the highest-volume seat in its family, hired by a controller or a shared-services manager who is filling several at once and whose realistic alternative is a CV screen and a conversation. Anything over about half an hour and the assessment loses to skipping the assessment. So there is no live call in this design, even though the role certainly involves talking to suppliers, and that is a real omission rather than a scoping convenience: how this person sounds when a supplier is angry with them on the telephone is not observed here, and a buyer who cares about it should put it in front of a human. What thirty-two minutes buys instead is the thing that actually moves money — twelve locked dispositions, one written escalation, and six more chances to see whether the threshold holds when somebody more senior leans on it.

Sources

Every figure on this page is traceable. Where a claim could not be sourced it is stated qualitatively instead.

  1. US Bureau of Labor Statistics, Occupational Outlook Handbook, Bookkeeping, Accounting, and Auditing Clerks, 2025, https://www.bls.gov/ooh/office-and-administrative-support/bookkeeping-accounting-and-auditing-clerks.htm
  2. FBI Internet Crime Complaint Center, 2025 IC3 Annual Report, business email compromise losses of $3,046,598,558 in 2025, the second-largest cyber-enabled fraud loss category, https://www.ic3.gov/AnnualReport/Reports/2025_IC3Report.pdf

See what the employer actually receives. A full report for one role, with every score shown beside the excerpt that earned it, conduct findings reported rather than averaged, and a reviewer sign-off required before any decision. No form.

Read a sample reportOr talk to us about this role