Taking an assessment rather than buying one? This page is written for employers. Here is the page for candidates.
Accounting and finance operations · Mid level
How to assess a Bookkeeper
Software-name screening is weaker here than anywhere else in the family, because the software is the easy part and the practice or business doing the hiring will retrain the person on their own stack within a fortnight regardless. A certification badge certifies familiarity with a menu. What no part of this process observes is the condition that actually defines the job: the bookkeeper is alone with a messy file and nobody reviews the work before it becomes the basis of a tax return. Given a transaction they cannot categorise and an owner who says "just put it through", the candidate either books it, queries it, or refuses it — and which of the three they do is the entire hire. It is also the one thing a friendly interview with the owner is structurally unable to surface, because the owner is the counterparty in that scenario.
A bookkeeper is the whole finance function for an organisation too small to have one, or the delivery layer of an accounting practice serving dozens of such organisations. Either way the defining structural fact is the absence of a reviewer. A staff accountant's journals are checked by a financial controller before they reach a statutory account. A bookkeeper's coding decisions go straight into the records that a tax return is built from, and the first person to look at them critically may be an accountant nine months later, at which point the cost of a bad convention is nine months of rework.
The work is the full cycle at small scale: sales and purchase invoices, expenses, bank feeds, reconciliation, payroll journals, sales tax or VAT returns, and a set of numbers handed to an accountant or an owner. In practice most of the hours go to two things. The first is coding — deciding what each transaction actually is, from a bank feed line that says "SQ *THE COFFEE HOUSE 14.20" or a card payment to a marketplace that might be stock, might be equipment, and might be the owner's spouse's laptop. The second is reconciliation, which in this context is less a control procedure than a detection method: for most small businesses, the bank reconciliation is the only mechanism that will ever catch a missing invoice, a duplicated payment or a subscription nobody remembers signing up to.
The failure mode that costs the most is quiet and cumulative. A weak bookkeeper faced with an ambiguous transaction posts it to a miscellaneous or suspense code and moves on, because stopping costs time and asking the owner feels like admitting ignorance. Do that forty times a month and the accounts stop describing the business: the profit figure is wrong in a way nobody can decompose, the VAT treatment on a chunk of spend is unexamined, and the year-end accountant's first job is a forensic exercise billed at three times the bookkeeper's rate. A strong bookkeeper does the opposite thing and does it in a way that is easy to observe — they maintain a short, live query list, they push it to the owner in a form that can be answered in two minutes, and they do not let it exceed a page.
The relationship with the owner is the part that makes this a judgment role rather than a processing one. Small business owners routinely run personal spend through the business, sometimes out of confusion about what is allowable and sometimes not. The bookkeeper is the person who sees it, is usually junior to the person doing it, is often paid by them directly, and has to say something. There is a wide band of legitimate answers here — a lot of expenditure is genuinely borderline and the right response is to flag it and route it to the accountant — but there is also a floor, and candidates differ enormously in where they put it. Some will book anything they are told to book. That is the single most important thing to know about a bookkeeper before hiring them, and it is the thing least likely to come up in the interview.
The other thing that separates the top quartile is legibility for the next person. Practice bookkeepers are moved between clients, go on holiday, and leave. A file where the coding is consistent, the queries are documented and the reconciliations show what was checked can be picked up by a colleague in an hour. A file that only made sense to the person who built it is a liability that the practice discovers at the worst possible moment.
What the job actually needs
- bank reconciliation
- transaction coding judgment
- sales tax and VAT treatment
- working without a reviewer
- explaining finance to a non-finance owner
How people fail in this seat
- codes anything unfamiliar to a catch-all account
- reconciles the bank by posting an adjustment rather than investigating
- treats an owner's personal spending as a business cost because they were told to
- accumulates unresolved queries and surfaces them at year end
- keeps the file tidy for themselves and illegible to the next person
What most employers do instead
CV screen for software names (QuickBooks, Xero, Sage, FreeAgent), platform certification badges, and a conversational interview, often with the business owner rather than a finance professional.
The assessment
About 38 minutes end to end.
The systems it runs in
A small-business cloud ledger with four surfaces the candidate has to move between: a bank feed carrying the thirty-five lines, a chart of accounts on which the catch-all code genuinely exists and is unblocked, a document inbox holding the purchase invoices and receipts, and the ledger's own bank reconciliation screen. QuickBooks Online and Xero are the reference builds, with Dext or Hubdoc as the document source; both publish the pattern this fixture depends on, in which a captured document becomes a transaction with the source image attached to it, and both expose bank rules that pre-code a line before a human looks at it. Naming the reconciliation screen is what changes the task: an unreconciled difference stops being a claim in a query list and becomes a state of the file that the next person opens into. The same is true of the coding decisions — the six awkward items leave this task as postings against named accounts with or without a document attached, which is the artefact an accountant will actually inherit. The bank rules matter for the fork rather than for fluency: the recurring subscription nobody has mentioned is exactly the kind of line a rule codes silently and correctly every month, and the question is whether the candidate notices that nothing behind it was ever checked. A spreadsheet is the fifth surface and it carries the reconciliation working: the feed on one sheet and last month's reconciliation on another, where the difference has to be decomposed into the uncashed payment and the missing invoice rather than plugged with one figure.
- QuickBooks Online
- Xero
- Sage Business Cloud Accounting
- Dext
- Hubdoc
- Microsoft Excel
- Google Sheets
Any ledger with a bank feed, an editable chart of accounts, a place to attach a source document to a transaction, and a reconciliation that can be left unfinished. The fixture is rebuilt inside the buyer's own instance where they provide a sandbox, which is worth doing for a practice, because a bookkeeper is hired into a specific file and the chart of accounts is the part that is genuinely theirs. Where no sandbox exists, the month ships as plain tables and document images and the codes are recorded against the lines. No criterion mentions a platform and none is satisfied by knowing where a menu item lives; naming the ledger fixes what the artefact looks like, not what earns a score.
What the candidate actually does
| Task | What happens |
|---|---|
| A month of somebody else's file data_task · 16 min | One month of a small trading company: a bank feed of roughly thirty-five lines, a folder of purchase invoices and receipts, a short sales listing, and last month's reconciliation. The pressure is stated in the brief in identical words for every candidate: this is one of fourteen client files, it is on a fixed fee, the hours budgeted for it are already spent, the owner answers queries about once a week, and the year-end pack goes to the accountant on Friday. Most lines code themselves. Planted among them are a card payment to a marketplace that could be stock, equipment or personal; a hotel booking during a week the owner was at a trade show and also on holiday; a recurring subscription nobody has mentioned and no invoice supports; the same supplier payment made twice eleven days apart; a supplier invoice where the sales tax charged does not match the rate the ledger has always applied to that category; and a payment to a named individual with no invoice, no payroll record and no description. The bank will not reconcile, and the difference is made up of one uncashed payment and one genuinely missing purchase invoice. The candidate codes the month in the ledger, reconciles the bank on the ledger's own reconciliation screen so that any difference remains an unfinished state of the file rather than a remark in a note, and records what they did with each item they could not resolve. The fork. A miscellaneous or suspense code is available and nothing prevents its use. Coding the six awkward items there takes two minutes, produces a file that looks complete, reconciles the bank with a small adjustment, and protects a fixed fee that is already overspent. Nobody will look at it until an accountant does, nine months from now, at three times the rate. |
| The query list and the file note written_artifact · 10 min | Two artefacts written back to back and graded against the candidate's own coding rather than a model answer. The first is the query list that goes to the owner — a person with no accounting vocabulary, no interest in the subject, and a strong preference for being asked as little as possible. The second is the note left in the file for whoever picks this client up next: the colleague covering a holiday, or the accountant preparing the year end. Neither artefact is specified in the brief beyond who reads it. The fork. A twenty-item query list is thorough, defensible, and will not be answered. A list with nothing on it protects the fee and the relationship. The artefact that works is neither: a short list, ordered by what actually changes the accounts, each item answerable in a sentence by somebody who does not know what a nominal code is. |
| What the owner says next judgment_scenario · 12 min | Five short exchanges with the owner, delivered in sequence and answered with an action and one line of reasoning each. The equipment purchase that is for a family member who "does some work for us". The trade-show trip that also contained a holiday. An instruction to date a purchase invoice into the earlier period. An instruction to leave the cash sales out and "sort them at the year end". And, when the candidate has pushed back on anything at all, the closing line — that they are the bookkeeper, not the auditor, and the accountant will pick up anything that matters. Two of the five are genuinely borderline and have a legitimate answer that is neither refusal nor compliance. One is not borderline at all. The candidate is not told which is which. The fork. The owner is not a villain and is not being tested on. They are pleasant, they are the person who pays the practice, they have a reason for each request that makes sense to them, and every request is easier to accept than to question. Refusing all five is available too, and it is wrong: it would mean a bookkeeper who cannot tell a genuinely borderline expense from a document that is being backdated, and who escalates the practice's client relationship over a hotel bill. |
The mark scheme
Each criterion is scored 1 to 5 against written anchors, and every score is reported with the excerpt that earned it. A criterion marked floored is reported as a finding rather than averaged into the total. The first is open; open any other to read its anchors in full.
An item that cannot be coded is queried, not absorbedweight 0.25Codes what the evidence supports, queries what it does not, and for each query states what was checked and what specifically is needed. The duplicate …
The reconciliation is used to find thingsweight 0.15Decomposes the difference, identifies the uncashed item and separately establishes that an invoice is missing rather than misposted, and names which s…
The query list can actually be answeredweight 0.2Short, ordered by effect on the accounts, each item a single question with the context attached and a closed answer available, so the owner can reply …
The floor holds when the person paying the fee asks flooredweight 0.25Separates the three kinds of request explicitly: books what is supportable, routes the genuinely borderline items to the accountant with the facts att…
The next person can pick the file upweight 0.15States the open queries, the coding conventions applied and why, anything deliberately left alone, and the one thing the next person should check firs…
How it is scored
Weighted mean of the five criteria, each scored 1 to 5 against the anchors and reported with the transaction line or excerpt that earned it. One raw count is reported beside the score and annotated explicitly as not a positive signal: the number of transactions coded to a miscellaneous, suspense or catch-all account. Buyers read a low count as competence, and in this fixture a very low count usually means the candidate guessed at items the evidence does not support rather than that they resolved them. The report therefore shows the count alongside the number of queries raised, and a reviewer reads the two together. Where the coding in t1 disagrees with what the query list or the file note asserts, both are shown side by side.
Integrity
- one unbroken monitored sitting, with the coding committed before the owner exchanges open
- both written artefacts are graded against the candidate's own coding, so a prepared query template cannot be made to fit a file it has not seen
- fixture variants rotated between sittings so the identity of the duplicate, the borderline items and the missing invoice differ
- one live follow-up question at the mid band, which changes a requirement rather than raising difficulty: the candidate is told the owner has replied that the marketplace purchase was stock, and asked what evidence they would now want and what they would do if it does not arrive before Friday
- session log records file-open order, time distribution and paste-versus-typed provenance in the written artefacts; no automated integrity verdict is produced from it
The log describes what happened. It does not produce a cheating verdict — the follow-up conversation is the control, because a statistical accusation is not something we would ask a reviewer to defend.
What you receive
- the coded month with a stated basis for each ambiguous item
- the reconciliation with its difference decomposed or not
- the query list and the file note as submitted
- the five owner exchanges with the action and reason given for each
- per-criterion score with the transaction line or excerpt that earned it
- catch-all coding count and query count, reported raw
Who decides
Recommended, and the reviewer has one decision that the rubric deliberately does not make. Which of the borderline items should be routed to an accountant rather than decided by the bookkeeper depends on how the practice or the business has drawn that line, and it is drawn differently in a two-partner practice and in an outsourced finance function with a qualified reviewer. The reviewer states the buyer's own boundary before scoring and overrides the fourth criterion in writing where it differs — for the routing half of it only, never for the backdating or the omitted sales, which are not boundary questions. The ranking entitles a buyer to conclude that this candidate will not absorb ambiguity silently, keeps a file a colleague could inherit, and has a floor that survives contact with the person paying the invoice. It does not establish familiarity with any particular software, which the buyer will retrain regardless, and it is not a check on any licence or registration the buyer's jurisdiction may require of somebody preparing tax filings.
What this does not measure
This design does not test software fluency and no criterion mentions a platform. The tooling block names the ledger the fixture is built in so that a buyer can see it matches their own stack and so that the artefact is a coded, part-reconciled file rather than a worksheet; nothing in the rubric is satisfied by knowing that ledger. That is the deliberate inversion of the usual screen for this role, which is a list of product names on a CV and a certification badge; both certify familiarity with a menu, and the practice will retrain the hire on its own stack within a fortnight either way. The fixture is presented as plain tables and document images, any method of arriving at a figure is accepted, and no spreadsheet formula use is scored. No criterion cites a tax provision as its basis, and that omission is stated rather than hidden: sales tax and VAT treatment, the deductibility of mixed-purpose expenditure and the rules on documenting a director's or owner's personal spend all differ by jurisdiction and change, and asserting the content of a regime this design has not fetched would be worse than citing nothing. The sales tax item is therefore scored on whether the candidate noticed that the charge disagrees with the ledger's habit and raised it as a treatment question, never on which treatment they chose; a deployer in a specific regime should have their reviewer score that item against it. The fourth criterion is scored on the candidate's action and their routing, and never on any moral characterisation of the owner: an anchor is satisfied by what was booked, what was routed and what was recorded. What the design cannot reach is the arc that actually produces the failure in this role. The dangerous bookkeeper is not the one who complies in week one; it is the one who has been asked forty times over two years, whose queries have stopped being answered, and for whom the catch-all code has become the path of least resistance. A single month cannot observe that, and a buyer should put it to a reference rather than to this assessment.
The structural fact about a bookkeeper, and the one this design is built around, is the absence of a reviewer. A staff accountant's journals are checked by a controller before they reach a statutory account. A bookkeeper's coding decisions go straight into the records a tax return is built from, and the first person to look at them critically may be an accountant nine months later. Every other assessment in this family can lean on the assumption that a mistake gets caught. This one cannot, which is why two of its five criteria are about what the candidate does when the evidence runs out rather than about what they do when it does not.
The month is built so that the catch-all code is the rational choice. It is available, nothing prevents its use, and the brief supplies the reason: fourteen clients, a fixed fee already overspent, an owner who answers about once a week, and a deadline on Friday. Those sentences are read by every candidate in the same words, because a fixture with nothing at stake tests whether somebody knows that suspense accounts are bad, and everybody knows that. What the buyer is paying to find out is what happens to that knowledge when stopping costs the candidate hours they will not be paid for. Nothing in the brief instructs the candidate to raise queries. It asks them to code the month, and the queries — or their absence — are the observation.
The six awkward items are chosen so that neither absorbing everything nor querying everything scores well. Two of them are resolvable from evidence already in the folder, so a candidate who queries them has not read the pack. The duplicate payment is eleven days apart and under a different reference, so it is findable but not obvious. The reconciliation difference has two components rather than one, which is what turns reconciling from an arithmetic exercise into what it actually is for a small business: the only mechanism that will ever catch a missing invoice or a subscription nobody remembers signing up to. And the payment to a named individual with no invoice and no payroll record is the item where a candidate's instinct is most visible, because there is no correct code for it at all — only a correct question.
The query list is where most candidates who did the ledger work well still lose points, and the reason is worth stating because it is the practical skill the role actually needs. A bookkeeper who sends twenty questions has done the analysis and produced a document that will not be answered, and the following month they will send twenty-two. The artefact that works is short, ordered by what changes the accounts, written without a single word the owner would have to look up, and specific enough to be answered from a phone in two minutes. That is a gradeable property of a document and it is never asked for before an offer. The file note beside it grades the other thing practices actually suffer from, which is a client file that made sense only to the person who built it and is discovered to be illegible at the moment that person leaves.
The owner exchanges are the centre of the design and the reason it runs to thirty-eight minutes. The relationship with the owner is what makes this a judgment role rather than a processing one: the bookkeeper sees the personal spend, is junior to the person doing it, is often paid by them directly, and has to say something. The exchanges are built to reward discrimination rather than either compliance or refusal. Two of the five are genuinely borderline, with a correct answer that is neither yes nor no but a routing to the accountant with the facts attached. One is not borderline in any regime — a request to date a document into a period it does not belong to — and the omitted cash sales is the same shape. The candidate is not told which is which, and the closing line, that they are the bookkeeper and not the auditor, arrives only if they have pushed back on anything, so the pressure lands on exactly the candidates who earned it.
That criterion is floored, and it is floored for a reason specific to this seat. A candidate who books a personal item on instruction or backdates an invoice cannot be redeemed by a beautifully reconciled bank account, because the whole value of the role is that this person is the only control in the system. A score of 1 there is a finding on the face of the report, never a number inside an average. The 3 anchor is deliberately reserved for the candidate who refuses everything, which is a real and common failure profile: it produces a bookkeeper who escalates a hotel bill to a partner and burns the client relationship the practice depends on.
Thirty-eight minutes is affordable here in a way it would not be for a transactional seat, because the buyer is a practice partner or a small-business owner hiring one person for a file they will not review, and the alternative screen — a software name on a CV and a friendly conversation with the owner — is structurally incapable of surfacing the thing that matters. The owner cannot interview for the floor, because in the scenario that tests it the owner is the counterparty.
Sources
Every figure on this page is traceable. Where a claim could not be sourced it is stated qualitatively instead.
- US Bureau of Labor Statistics, Occupational Outlook Handbook, Bookkeeping, Accounting, and Auditing Clerks, 2025, 1,532,400 jobs, 144,100 projected annual openings, 6 percent projected decline, typical entry education 'some college, no degree', https://www.bls.gov/ooh/office-and-administrative-support/bookkeeping-accounting-and-auditing-clerks.htm
See what the employer actually receives. A full report for one role, with every score shown beside the excerpt that earned it, conduct findings reported rather than averaged, and a reviewer sign-off required before any decision. No form.
Read a sample reportOr talk to us about this role