Taking an assessment rather than buying one? This page is written for employers. Here is the page for candidates.
Accounting and finance operations · Mid level
How to assess a Staff Accountant
A qualification proves the candidate was taught the standard and a competency interview proves they can narrate a close in retrospect, when the close has already resolved and the story has a tidy ending. Neither observes the decision the role actually turns on, which happens at six o'clock on day three with an unexplained difference sitting in an accrual account and a controller waiting for the pack: force it, park it in suspense and say nothing, or stop and write it up. Every candidate says in interview that they would flag it. The reconciliation shows which ones do. The Excel test is aimed still further away — syntax is trainable and increasingly automated, while the scarce judgment is materiality, which is the decision about which differences are worth anybody's time at all.
A staff accountant's year is a repeating calendar. Working days one to five or one to ten are the close: accruals and prepayments posted, intercompany balances agreed, the bank and the control accounts reconciled, journals supported, and a set of numbers handed upward with commentary attached. The rest of the month is the work that makes the next close survivable — clearing the items that were parked, chasing the cost centre that keeps miscoding, rebuilding a schedule that has drifted out of agreement with the ledger. Unlike a bookkeeper, this person is reviewed; unlike an FP&A analyst, they are accountable for whether the number is right rather than for what it implies.
The distinguishing competency is cut-off. Almost everything that goes wrong in a close is a timing question wearing a different costume: an invoice that arrived in the new period for goods received in the old one, a bonus that is legally uncertain but economically incurred, revenue recognised on despatch when the contract says on acceptance. The technical answer is usually available in a standard the candidate has been examined on. What is scarce is the willingness to apply it against pressure, because the correct treatment is very often the one that makes the month look worse and generates a conversation with an operations director. A weak staff accountant resolves that tension in the ledger, quietly. A strong one resolves it in an email.
The second differentiator is what the person does with a difference they cannot explain. Every reconciliation eventually produces one. The median response is to close the gap — post a rounding journal, net it against another difference, carry it forward with a note that says "to investigate" and never investigate. It is worth being precise about why that is dangerous rather than merely untidy: a forced balance destroys the only evidence that would have identified the cause, so the same difference recurs the following month having become harder to trace, and the fraud or systematic error that produced it survives another cycle. The top-quartile behaviour is unglamorous and easy to grade — the difference is quantified, its age is stated, what was checked is listed, and it is escalated before the pack goes rather than after.
The third is the commentary, which is the part of the role most likely to be treated as an afterthought and most likely to determine whether the accountant is trusted. Budget holders do not read management accounts; they read the sentence next to the variance. "Marketing £42k over budget" restates the number. "Marketing is £42k over, of which £38k is the Q3 campaign invoiced a month earlier than planned and will reverse in October; the residual £4k is recurring software we did not budget for" tells someone what to do. Writing the second version requires having actually found out, which is why it is the best single proxy in this family for whether the close was done or merely completed.
What a controller is trying to predict when they hire is therefore quite narrow and very consequential: does this person stop when the evidence runs out, and can they explain a number to someone who does not have the ledger open. Both are visible in an hour of work on a small, deliberately imperfect set of accounts. Neither is visible in a professional qualification, which is a floor worth verifying and a ranking worth ignoring.
What the job actually needs
- month-end close discipline
- cut-off and accrual judgment
- balance sheet reconciliation
- variance explanation in writing
- materiality
How people fail in this seat
- forces a reconciling difference to clear rather than reporting it
- books an accrual with no evidence behind the estimate
- explains a variance by restating the number instead of naming the cause
- finds a material error after the pack has gone out
- treats every difference as equally worth chasing
What most employers do instead
Qualification screen (ACCA, CIMA, CPA, part-qualified or finalist), a competency interview asking the candidate to describe a close they have run, and occasionally an Excel test on lookup and pivot syntax.
The assessment
About 40 minutes end to end.
The systems it runs in
A general ledger with a close checklist over it. The ledger is where the journal is posted, the entity's trial balance is drawn and the suspense account exists and is unblocked; NetSuite and Sage Intacct are the reference builds, and published job advertisements for this seat routinely pair one of them with a reconciliation tool. The checklist is the second surface and it is the one that changes what this task observes: in a close-management tool each reconciliation is an item that is prepared, evidenced with the supporting schedule attached, and signed off, so "what remains open" stops being a sentence in a note to the controller and becomes an item nobody has signed. That is the difference between a candidate who says they left a difference open and one who actually left it open where the controller will see it at 09:00. FloQast and BlackLine are the named reference builds for that layer, both of which publish integrations with the ledgers above. The schedules — accruals and prepayments, the bank reconciliation, the intercompany comparison — are worked in a spreadsheet, and the intercompany item is the one worth naming precisely: two lists of the same transactions from two entities, keyed differently, where the answer is not a total but the identification of which items do not appear on both sides.
- Oracle NetSuite
- Sage Intacct
- SAP
- Microsoft Dynamics 365 Business Central
- BlackLine
- FloQast
- Microsoft Excel
Any ledger that supports a manual journal, exposes a trial balance and has a suspense or clearing account available, paired with any checklist that records preparer, evidence and sign-off per reconciliation — including a shared workbook, which is what a large share of finance teams actually use. Rebuilt inside the buyer's instance where a sandbox exists; otherwise the pack ships as the ledger's own exports and the sign-off column is on the face of the checklist. No criterion mentions a platform and none is satisfied by knowing where a menu item lives.
What the candidate actually does
| Task | What happens |
|---|---|
| Day three of the close data_task · 18 min | A small set of deliberately imperfect accounts for one entity: a trial balance extract, a bank reconciliation, an accruals and prepayments schedule, and an intercompany balance the counterparty has confirmed at a different figure. The pressure is in the brief, in identical words for every candidate: it is 18:00 on working day three of a five-day close, the controller has asked for the pack by 09:00, and the operations director has already emailed asking why their cost centre looks so far over. Planted in the accounts are five things. A timing difference that looks like an error — an invoice in the new period for goods received in the old one. A duplicated supplier invoice sitting in accruals. Four small differences that are individually below any sensible materiality line and collectively still below it. An accrual carried forward from two months ago with no supporting evidence behind the estimate. And one difference of real size that cannot be resolved with anything in the pack. The candidate produces the reconciliations, posts what they have decided to post as journals in the ledger, and leaves what remains open as an unsigned item on the close checklist with the reason attached, so that an open difference is a state of the close rather than a remark in a note. The fork. The unresolved difference can be cleared in one journal. There is a suspense account available, the amount is small enough relative to the entity that nobody will query it, the reconciliation then ties, the pack goes at 09:00, and the controller is pleased. It will recur next month having become harder to trace, but that is somebody else's problem and, on a five-day close, it is next month's version of this same candidate. |
| The commentary and the note written_artifact · 12 min | Two pieces of writing produced back to back and graded against what the candidate actually did in t1, not against a model answer. The first is the variance commentary on the operations director's cost centre, written for somebody who does not have the ledger open and will act or not act on this paragraph alone. The second is the internal note to the controller about anything still open. The candidate is not told that a note to the controller is expected; the segment asks for the commentary and offers a free-text field for anything else they want to send. The fork. A commentary that restates the number — the cost centre is over by this much, driven by higher spend — is quick, is literally true, is what most management accounts contain, and never gets challenged. Naming the cause requires having actually found it, and naming the part that will reverse next month invites a follow-up question the candidate then has to answer. |
| Six cut-off decisions judgment_scenario · 10 min | Six short cases, three or four lines each, answered with a treatment and one line of reasoning. Goods received on the last day with the invoice still to come. A bonus that is legally uncertain and economically incurred. Revenue recognised on despatch where the contract specifies acceptance. A credit note promised verbally by a supplier and not yet issued. A legal fee for a matter that will conclude next quarter. A stock write-down the operations director says will be recovered by a sale already agreed. In four of the six, the correct treatment makes the month worse, and each card states who will be unhappy about it. The fork. Each card names a person who wants the other answer and gives a reason that is not stupid. The cheap path is available on every one of them and is defensible in the moment; the cost only appears at audit, or at the point where the reversal lands in somebody else's month. |
The mark scheme
Each criterion is scored 1 to 5 against written anchors, and every score is reported with the excerpt that earned it. A criterion marked floored is reported as a finding rather than averaged into the total. The first is open; open any other to read its anchors in full.
Stops when the evidence runs out flooredweight 0.25Leaves it open, states the amount and how long it has been there, lists what was checked and what was ruled out, and says what would resolve it. The r…
Materiality — chases what changes a decisionweight 0.2States a materiality position and applies it consistently in both directions: the four small differences are aggregated, described and left, and the r…
Cut-off treatment holds when the treatment is unwelcomeweight 0.2Applies the same principle across all six, separates what is documented from what has been asserted verbally, and where a case is genuinely finely bal…
The commentary names the cause, not the numberweight 0.25Splits the variance into its components, distinguishes the part that is a timing difference and will reverse from the part that is recurring and will …
What is still open reaches somebody before the pack doesweight 0.1Writes to the controller unprompted, states the amount, the effect on the reported result, what is being done and what is being asked for, before the …
How it is scored
Weighted mean of the five criteria, each scored 1 to 5 against the anchors and reported with the journal, schedule line or sentence that earned it. Two raw timings are reported beside the score, outside the weighted mean, because the gap between them is more informative than either: the clock time at which the candidate first identified the unresolvable difference, and the clock time at which they first wrote to anybody about it. A candidate who found it in six minutes and told nobody for thirty is a different hire from one who found it at twenty-eight and escalated at twenty-nine, and the weighted mean cannot see the difference. The value of differences left unexplained is also reported and annotated explicitly as not a negative signal, since in this fixture a residual is the correct outcome and a zero residual almost certainly means a plug.
Integrity
- one unbroken monitored sitting, with the ledger work committed before the writing segment opens
- both written pieces are graded against the candidate's own postings, so a prepared commentary cannot be made to fit
- fixture variants rotated between sittings so the size and location of the unresolvable difference and the identity of the duplicated invoice differ
- one live follow-up question at the mid band, which changes a requirement rather than raising the difficulty: the candidate is told the supplier has since confirmed the work slipped into the following month, and asked what that does to the accrual they booked and to the commentary they wrote
- session log records file-open order, time distribution and paste-versus- typed provenance; no automated integrity verdict is produced from it
The log describes what happened. It does not produce a cheating verdict — the follow-up conversation is the control, because a statistical accusation is not something we would ask a reviewer to defend.
What you receive
- the reconciliations and the journals as posted, with support stated
- the variance commentary and anything sent to the controller
- the six cut-off treatments with their one-line reasons
- per-criterion score with the schedule line or excerpt that earned it
- time-to-identification and time-to-escalation, reported raw
- the floored criterion reported as a finding where it scores 1
Who decides
Recommended. Two things need a person. The first is the materiality line: a difference that is immaterial in a group with a hundred million of revenue is a serious matter in a twelve-person subsidiary, and the reviewer sets the threshold the candidate is scored against rather than the rubric assuming one. The second is the finely balanced cut-off card, where a reviewer with the buyer's own accounting policy in hand may legitimately override the score. The ranking entitles a buyer to conclude that this candidate reconciles honestly under a deadline and can explain a number to somebody who does not have the ledger open. It does not establish that they know this buyer's chart of accounts, group policy or consolidation system, and it is not a substitute for verifying a professional qualification where one is required — the role page's position is that the qualification is a floor worth checking and a ranking worth ignoring.
What this does not measure
No criterion here rewards spreadsheet formula fluency, and none is satisfiable by knowing where a menu item lives. The fixture is a plain set of schedules, any method of arriving at a figure is accepted including paper and a calculator, and the reconciliations are scored on what was concluded and what was evidenced rather than on how the workbook was built. This is deliberate: the Excel syntax test is the conventional screen for this role and it measures the most trainable input in the job. Nor does any criterion cite a specific accounting standard as a basis, and that omission is stated rather than hidden. Cut-off treatment is governed by a different framework depending on where the buyer reports — IFRS, US GAAP, a local GAAP — and citing a provision the design has not fetched, or asserting that one applies to a buyer whose regime is unknown, would be worse than citing nothing. The cards are therefore written so that the reasoning is scoreable under any of them: what is documented against what is asserted, what is incurred against what is invoiced, the same principle applied consistently in both directions. A deployer reporting under a specific framework should have their reviewer score the cards against it. What the design does not reach is the rest of the month: it samples day three of a close and cannot observe whether this person clears the parked items in the quiet fortnight afterwards, which is the habit that makes the next close survivable. It also cannot observe behaviour across a full five-day close with a real controller applying real pressure over three consecutive evenings. Deployers should monitor the data task's distribution separately by group; extended time and assistive technology should be available on request.
A staff accountant is hired on a qualification and interviewed about a close they have already survived, which is to say about a story with a tidy ending. The decision the job actually turns on has no ending at the time it is made: it is six o'clock on day three, there is an unexplained difference in an account, the controller is waiting, and the candidate can either force it, park it silently, or stop and write it up. Every candidate says in an interview that they would write it up. This design exists to find out which ones do, and it is built around making the alternative genuinely attractive rather than obviously wrong.
The fixture gives the plug every advantage. The difference is small enough that nobody would query the journal. A suspense account is available. Clearing it makes the reconciliation tie, which makes the pack go on time, which is the thing the candidate has been told they are measured on in a brief that every candidate reads in the same words. Nothing in the brief asks them to flag anything or to escalate anything, because an instruction to raise ambiguities converts the exercise into a test of compliance with an instruction. The difference is simply there, unresolvable with the evidence supplied, and what the assessment observes is whether the candidate chooses to say so.
That criterion is floored. A candidate who plugs the balance and scores well on everything else must not surface as a strong hire, because the failure is not untidiness. Forcing a balance destroys the evidence that would have identified the cause, so the same difference returns the following month harder to trace, and whatever produced it — a systematic coding error, a duplicate payment, a fraud — survives another cycle. A score of 1 there is a finding printed on the face of the report, not a number folded into an average.
The second thing the fixture is built to see is materiality, which is the scarce judgment in this role and the one no Excel test has ever come near. Four of the planted differences are small, and they are small in both senses: individually trivial and, aggregated, still below any sensible line. A candidate who resolves all four while the real difference sits untouched has worked hard and misunderstood what they are for. A candidate who dismisses the material one alongside them has made the opposite error with the same reasoning. Scoring requires the reviewer's own threshold, which is why human review is recommended rather than advisory here: the correct answer genuinely differs between a large group and a small subsidiary, and a rubric that pretends otherwise would be ranking candidates against an invented policy.
The cut-off cards resample under pressure what the reconciliation observes once. Six short cases is nine or ten minutes and turns a single observation of technical judgment into six, which matters because the characteristic failure in this role is not ignorance of the treatment. Every card names somebody who wants the other answer and gives them a reason that is not stupid, and in four of the six the correct treatment makes the month look worse. That is the whole point: the technical answer is available in a standard the candidate has been examined on, and what is scarce is applying it when doing so generates a conversation with an operations director who outranks them. The scenario has to put something on the table or it is testing knowledge that everybody already has.
The commentary is graded against the journals the candidate actually posted, which does two jobs. It makes a prepared paragraph unfittable, because the variance in this fixture went somewhere the candidate could not have anticipated. And it exposes the specific dishonesty that management accounts quietly tolerate, where the words next to a number describe a cause the ledger does not support. Budget holders do not read management accounts; they read the sentence beside the variance, and the difference between restating the number and naming the cause is the difference between an accountant who is consulted and one who is copied in. Writing the second version requires having actually found out, which is why it is the best single proxy in this family for whether a close was done or merely completed.
The last rubric line carries the smallest weight and produces one of the two most useful artefacts in the report. Escalation before a deadline and escalation after it are different acts entirely — the first is a decision aid, the second is a report of a failure — and the weighted mean cannot see the distinction because both look like an email in the evidence pack. So the report shows two raw clock times outside the score: when the candidate found the difference, and when they first told anyone about it. The gap between them separates the accountant who stopped and raised it from the one who found it early, hoped, and mentioned it at the end.
Forty minutes is the budget and it is longer than the transactional designs in this family for a specific reason. This seat is hired in ones and twos by a financial controller rather than in tens by a shared-services manager, the cost of a bad hire is a set of statutory accounts built on a forced balance, and the reviewer time per candidate is affordable at that ratio. The ledger work needs eighteen minutes because a fixture small enough to complete in eight cannot contain both a materiality decision and an unresolvable difference without signalling which is which.
Sources
Every figure on this page is traceable. Where a claim could not be sourced it is stated qualitatively instead.
- US Bureau of Labor Statistics, Occupational Outlook Handbook, Accountants and Auditors, 2025, https://www.bls.gov/ooh/business-and-financial/accountants-and-auditors.htm
- US Bureau of Labor Statistics, Occupational Outlook Handbook, Bookkeeping, Accounting, and Auditing Clerks, 2025, https://www.bls.gov/ooh/office-and-administrative-support/bookkeeping-accounting-and-auditing-clerks.htm
See what the employer actually receives. A full report for one role, with every score shown beside the excerpt that earned it, conduct findings reported rather than averaged, and a reviewer sign-off required before any decision. No form.
Read a sample reportOr talk to us about this role