Taking an assessment rather than buying one? This page is written for employers. Here is the page for candidates.
Accounting and finance operations · Mid level
How to assess a FP&A Analyst
The modelling test is the closest thing in this family to a genuine work sample, and it still misses the half of the job that is adversarial. FP&A fails at the moment a commercial director pushes back on a forecast the analyst cannot fully defend, and the analyst either folds and rebuilds the model to the answer, or digs in on a number that is in fact shaky. Neither response is diagnosable from a spreadsheet, and the case presentation does not reproduce it either, because a hiring panel asks clarifying questions rather than applying commercial pressure with something at stake. The second miss is that modelling tests ship clean data. Real FP&A begins with a data set that is wrong — a restated cost centre, a duplicated month, a mapping that changed in April — and the ranking signal is whether the analyst notices before they build eight tabs on top of it.
FP&A is the point in this family where the work stops being about whether the number is right and starts being about what the number means and who has to be told. The analyst owns a budget cycle, a rolling forecast, a monthly variance pack, and — in the version of the role most employers now hire for — a relationship with one or two parts of the business who treat them as their finance person. The technical floor is modelling and the ceiling is influence, and hiring processes are calibrated almost entirely to the floor.
Hour to hour the job is unglamorous. Pulling actuals, mapping them to a forecast structure that has drifted, working out why a cost centre that has never spent more than £30k a quarter has spent £71k, and discovering that the answer is a recharge that used to sit somewhere else. Most of what looks like analysis is reconciliation and archaeology. The output is small: a variance commentary, a revised forecast, a one-page recommendation on whether to hire the two roles the sales director wants in Q3. The leverage is entirely in whether those small outputs are trusted.
Two habits separate the top quartile. The first is building models that other people can audit. A model with assumptions on a labelled input sheet, drivers that connect visibly to outputs, and no constant buried in the middle of a formula is not a matter of style — it is what allows a finance director to change one growth rate in a meeting and see the answer move, which is the only way a forecast ever gets stress-tested. Analysts who hard-code produce models that only they can operate, which feels like job security and is actually the reason their work gets ignored. The second habit is stating confidence. A strong analyst does not present a single number; they present a number, the two or three drivers it is most sensitive to, and a plain statement of what would have to be true for it to be wrong. That is what converts a forecast from a prediction into a decision aid.
The hardest part is the pressure. FP&A sits between a controller who wants the numbers conservative and a commercial leader who wants headroom, and both of them outrank the analyst. The characteristic failure is not dishonesty; it is gradual accommodation. A director says the pipeline conversion assumption is too pessimistic, the analyst has no strong evidence either way, and the assumption moves — not because it was wrong but because defending it was uncomfortable. Repeat that four times in a planning cycle and the plan has been negotiated rather than modelled, and nobody can point to the moment it happened. The behaviour that marks a strong hire is neither compliance nor stubbornness: it is insisting that if the number moves, the driver behind it moves too, and that the change is written down with whose view it reflects. That is a specific, observable, gradeable act, and it is the single most useful thing to test for.
There is also a discipline of proportion that experienced FP&A people have and juniors almost never do. Given a variance, the question is not "can I explain all of it" but "does explaining the rest change what anyone does". An analyst who spends two days resolving the final four percent of a variance while the budget holder waits has misunderstood what they are for. A hiring manager is trying to predict exactly this: whether the person will produce a defensible number on time, say honestly how sure they are of it, and hold the line on the one occasion in the quarter when holding the line matters.
What the job actually needs
- driver-based modelling
- variance diagnosis
- forecast honesty
- defending a number under commercial pressure
- one-page written recommendation
How people fail in this seat
- hard-codes assumptions inside formulas so nobody can find them
- presents a forecast with no sensitivity and no statement of confidence
- lets a senior stakeholder move the output without changing a driver
- spends three days on precision that does not change the decision
- builds on a data set they never sanity-checked
What most employers do instead
CV screen, a modelling or Excel case test, and a case interview or presentation to the hiring manager.
The assessment
About 55 minutes end to end.
The systems it runs in
A spreadsheet as the modelling surface — Excel or Google Sheets — over actuals exported from the ledger, with a planning tool named as the place the agreed forecast is ultimately published. That split is the real one: the planning platforms in this category are built around driver-based models with inputs separated from calculations, and Vena's own positioning is for teams who want that governance on top of native Excel modelling, which is an admission that the model is still built in a spreadsheet in most of these seats. So the fixture puts the model where the work happens and treats the planning tool as the system of record for what was published and when. Naming the spreadsheet changes what three things in this design observe. First, the fifteen months of actuals arrive as an export, not as a clean table: the April restatement means the cost-centre codes before and after do not join, so a lookup written against the current mapping silently returns nothing for the prior months, and a positional lookup written against column order returns the wrong category outright. Second, the duplicated month is only visible if the candidate checks that period labels are unique before trending on them — a two-minute check that no instruction asks for. Third, the driver sheet is a literal artefact: the second rubric criterion asks whether a finance director could change the conversion rate in the meeting and watch the answer move, and that is a property of where the number lives in the workbook, not of which function surrounds it. Fast-and-correct here means the candidate found the three defects and still had time to build the forecast; slow-and-correct means they found them and ran out of runway. Fast-and-wrong is eight tidy tabs built on a duplicated month, and it is the version that presents best at 14:00.
- Microsoft Excel
- Google Sheets
- Anaplan
- Workday Adaptive Planning
- Pigment
- Vena
- Planful
- Oracle NetSuite
- SAP
Any spreadsheet the candidate prefers, with the actuals supplied as a flat export so no buyer's ERP is assumed. Where the buyer runs a planning platform, the fixture states which one and asks the candidate to say what they would publish into it and what they would leave out — a written answer, not an operated one, because platform familiarity is the most trainable input in this seat and the corpus refuses to screen on it. No criterion rewards a keyboard shortcut, a particular function or a formatted output.
What the candidate actually does
| Task | What happens |
|---|---|
| The actuals nobody checked data_task · 20 min | Fifteen months of actuals for one business unit against the current forecast, plus a headcount schedule and a note of the assumptions the existing forecast was built on. The pressure is stated in the brief in identical words for every candidate: the forecast review is at 14:00, the commercial director has already circulated a version of these numbers that supports the case they intend to make, and the analyst's job today is to produce the variance analysis and a revised forecast for the next two quarters. The data has three defects, none of them flagged. One cost centre was restated in April and the prior months were never remapped, so a category shows a step change that is not a step change. One month appears twice under slightly different period labels, which flatters the trend. And a recharge that used to sit in another unit moved in, so a cost line has grown for a reason that has nothing to do with the business unit's own behaviour. The candidate produces the variance analysis and the revised forecast, and commits a pre-registered position before t3 opens: the number they will present, the two or three drivers it is most sensitive to, and what would have to be true for it to be wrong. That commitment is timestamped and cannot be edited. The fork. The data arrives looking finished, in the format the last analyst used, and building eight tabs on top of it is the fastest route to a forecast by 14:00. The duplicated month makes the trend look better, which is the direction everybody in the meeting wants it to look, so the defect that helps is the one least likely to be interrogated. |
| The one-page recommendation written_artifact · 10 min | One page on whether to approve the two roles the sales director wants in Q3, written for people who will not open the model. Graded against the candidate's own numbers rather than a model answer. The recommendation is written before the call, so what the candidate says to the director's face can be compared with what they wrote when nobody was pushing. The fork. A page that presents a single number with no sensitivity is quicker to write, reads as more confident, and is what most forecast packs contain. Stating what the number is most sensitive to invites exactly the challenge the analyst will then have to survive fifteen minutes later, and every candidate can feel that as they write it. |
| The commercial director live_call · 15 min | A live call with an AI counterpart playing a commercial director who outranks the analyst, is not hostile, and is good at this. They accept the variance work without argument. They then press on one assumption — the pipeline conversion rate the forecast rests on — and their case is not stupid: they have information from the field the analyst does not have, they point out correctly that the analyst's own basis for the assumption is thin, and they observe that the conservative version blocks two hires the business has already told candidates about. Midway they offer the accommodation that makes this a fork rather than a stand-off: leave the driver alone, just show the output at the higher figure for the meeting, and the two of them can revisit the driver properly next month. The counterpart never insists and never threatens; if the analyst holds, they accept it and say they will raise it with the CFO instead. The fork. The director is offering a way for the analyst to be helpful, avoid a confrontation with somebody senior, and keep an assumption they cannot strongly defend anyway. Digging in is also wrong here, and the scenario is built so that it is: the analyst's basis for the original number genuinely is thin, the director genuinely does hold information the analyst does not, and a candidate who simply refuses has mistaken stubbornness for rigour. |
| What goes to the CFO written_artifact · 10 min | The revised note after the call, graded against the transcript of a conversation the candidate cannot re-read and against the position they pre-registered before it. Whatever the analyst agreed, conceded, or held, this is the document that records it for somebody who was not in the room. The fork. The comfortable note reports the revised number as the forecast, which is true, easy, and erases the fact that it moved because a director asked. The note that records whose view the change reflects requires the analyst to write something the director may read. |
The mark scheme
Each criterion is scored 1 to 5 against written anchors, and every score is reported with the excerpt that earned it. A criterion marked floored is reported as a finding rather than averaged into the total. The first is open; open any other to read its anchors in full.
The data is checked before it is built onweight 0.25Finds all three, states what each does to the numbers, corrects or excludes them explicitly, and separates the movement caused by the recharge from th…
The model can be operated by somebody elseweight 0.15Drivers are on an input sheet, labelled, sourced where they have a source, and connected visibly to the outputs, so a finance director could change th…
A number is presented with its confidenceweight 0.2States the number, names the two or three drivers it is most sensitive to with the effect of each, and states plainly what would have to be true for i…
If the output moves, a driver moves and it is attributed flooredweight 0.25Declines to move the output without moving a driver, in terms; accepts the director's information as a reason to revise the driver rather than the res…
Proportion — analysis stops when it stops changing the decisionweight 0.15Resolves what changes the answer, explicitly says what was left unexplained and why it does not matter at this level of materiality, and the one page …
How it is scored
Weighted mean of the five criteria, each scored 1 to 5 against the anchors and reported with the model cell, transcript line or sentence that earned it. Three figures are printed side by side on the front of the report: the number the candidate pre-registered before the call, the number they defended on the call, and the number in the note that went to the CFO. Where those three differ, the report gives the transcript offset at which the position changed and what was said immediately before it — conceded ninety seconds after the two hires were mentioned, for instance — rather than allowing the movement to disappear into an average. Without the pre-registration a candidate who folds can always say afterwards that folding was the plan; with it, the concession is a timestamped event. The call is scored from the transcript, which is the reviewer's default and only view; audio is retained for dispute and is not a scoring surface, and no criterion is named clarity or professionalism.
Integrity
- one unbroken monitored sitting in the order t1, t2, t3, t4
- the pre-registered position is timestamped and locked before the call opens and cannot be edited afterwards
- the post-call note is graded against a transcript the candidate cannot re-read, which is the strongest anti-coaching control in the design: a memorised answer and an off-camera coach both fail it, because the note must match a conversation that went somewhere unplanned
- data pack variants rotated between sittings so the identity of the restated cost centre and the duplicated month differ
- one live follow-up question at the mid band, which invalidates an assumption rather than adding work: the candidate is told the recharge is being reversed from next quarter and asked what that does to their forecast and to the recommendation on the two hires
- session log records model build order, time distribution and paste-versus- typed provenance; no automated integrity verdict is produced from it
The log describes what happened. It does not produce a cheating verdict — the follow-up conversation is the control, because a statistical accusation is not something we would ask a reviewer to defend.
What you receive
- the model as built, with its input sheet
- the timestamped pre-registered position
- the one-page recommendation and the post-call note
- full call transcript with the accommodation offer and any concession time-marked
- the three headline figures side by side with the offset of any change
- per-criterion score with the cell, line or excerpt that earned it
Who decides
Required rather than recommended, and this is the one design in this family where that distinction is real. Two of its criteria can be satisfied by an output that a mechanical ranking will read as weak. A candidate who says the evidence behind the conversion assumption is thin, gives a range rather than a point, and declines to produce the precision the fixture invites is doing the job correctly, and an automated comparison against a confident single figure would rank them below a candidate who guessed well. The reviewer also has to decide whether the concession, where one occurred, was a legitimate revision or a capitulation, and the transcript offset is evidence for that decision rather than the decision itself. The ranking entitles a buyer to conclude that this candidate checks data before building on it, states confidence honestly, and can hold an assumption in front of somebody senior without either folding or refusing. It does not establish knowledge of the buyer's planning system, chart of accounts or industry drivers.
What this does not measure
The modelling task is scored on structure and on stated confidence, never on spreadsheet virtuosity: no criterion rewards a keyboard shortcut, a particular function, or a formatted output, and a model built plainly with visible inputs scores 5 on the second criterion while an elegant one with a constant buried in a formula scores 1. Any tool the candidate prefers is permitted. The call is scored from the transcript rather than the audio, which is the structural control on accent, dialect and speech variation — an instruction not to be biased is weaker than a design in which the biasing signal is not in front of the scorer — and no criterion is named clarity, presence or executive presence, which are the labels under which that filter conventionally re-enters a rubric for a role like this one. The counterpart applies pressure through argument and accommodation only; the fixture contains no interruption, no speed test and no requirement to think aloud at pace, because those disadvantage candidates for reasons that have nothing to do with FP&A. What the design does not reach is the arc that the role is actually judged on: a planning cycle is three months long, and the characteristic failure is not one concession but four small accommodations across a quarter, none of which was visible at the time. One call samples the behaviour and cannot observe its accumulation, and a buyer should put that to a reference. Fifty-five minutes is also the longest design in this family and that is a deliberate allocation rather than an oversight — this is the lowest-volume, highest-consequence seat of the six, hired individually rather than in cohorts, and the reviewer time it consumes is affordable at that ratio in a way it would not be for a payables queue.
The conventional screen for an FP&A analyst is a modelling test, and it is the closest thing in this family to a genuine work sample. It is also, by itself, the wrong half of the job. Modelling tests ship clean data and they ask the candidate to build. Real FP&A starts with a data set that is wrong in a way nobody has flagged and ends in a room with somebody senior who wants a different answer, and neither end of that is visible in a spreadsheet exercise or in a case presentation where a hiring panel asks clarifying questions rather than applying commercial pressure with something at stake.
So this design puts a defect in the data at the front and a director at the back, and it connects them with a lock. The data pack has three problems and none is announced: a cost centre restated in April with the prior months never remapped, a month that appears twice under slightly different labels, and a recharge that moved in from another unit. They are chosen to be different kinds of wrong. The restatement is visually obvious to anybody who plots the series. The recharge is findable only by somebody who asks what the cost line is composed of. The duplicated month is the interesting one, because it flatters the trend, and a defect that pushes the numbers in the direction everybody wants them to go is the defect least likely to be interrogated. A candidate who builds eight tabs on top of it produces a forecast that is confidently wrong and looks, from the outside, more finished than the correct one.
The pre-registration is the mechanism that makes the second half of the design scoreable. Before the call opens, the candidate commits their number, the drivers it is most sensitive to, and what would have to be true for it to be wrong. The commitment is timestamped and cannot be edited. Without it, folding under pressure is unfalsifiable — a candidate who moves their forecast can always say afterwards that they had been persuaded by better information, and there is no way to distinguish that from accommodation. With it, the movement becomes an event with a time attached, and the report can show what was said in the ninety seconds before it. Three numbers then sit on the front page: what was pre-registered, what was defended, and what went to the CFO. In this role the gaps between those three are more informative than any single score.
The director is built carefully, because the easy version of this scenario teaches the wrong lesson. If the counterpart is simply wrong and simply pushy, the assessment rewards refusal, and refusal is not the competency. So the director has a real case: they hold information from the field the analyst does not have, they are correct that the analyst's evidential basis for the conversion assumption is thin, and the conservative forecast blocks two hires the business has already begun to promise. They never threaten. Halfway through they offer the specific accommodation that produces the characteristic failure in this role: leave the driver alone, just show the output at the higher figure for this meeting, and revisit it properly next month. That is the moment the whole fifty-five minutes exists to observe. It is not dishonesty and nobody in the scenario experiences it as such. It is how a plan stops being modelled and starts being negotiated, one reasonable accommodation at a time, with no moment anybody can point to afterwards.
The behaviour that scores 5 is neither compliance nor stubbornness, and the anchors say so explicitly. It is the insistence that if the output moves, a driver moves with it, and that the change is recorded with whose view it reflects. That is a specific, observable act rather than a disposition, and it is gradeable from two artefacts: what the analyst says on the call, and what the note to the CFO records afterwards. The criterion is floored. An analyst who agrees to present a number their own model does not produce cannot be a strong hire because their variance analysis was excellent, since every number they produce afterwards inherits the same doubt.
The final note is graded against a transcript the candidate cannot re-read. This is the strongest anti-coaching control available in a written assessment: a prepared paragraph cannot describe a conversation that went somewhere unplanned, and neither can somebody feeding the candidate answers off-camera. It also measures precisely what the job requires, which is an accurate record rather than good prose on demand.
Human review is required here rather than recommended, and the reason is the proportion criterion. A candidate who reports that the evidence behind the key assumption is thin, gives a range instead of a point, and declines to manufacture precision is doing the job right, and a mechanical ranking will put a confident wrong figure above them every time. Somebody has to read the output and decide whether the uncertainty stated was honest or evasive, and no rubric can make that call from a distance.
Sources
Every figure on this page is traceable. Where a claim could not be sourced it is stated qualitatively instead.
- US Bureau of Labor Statistics, Occupational Outlook Handbook, Financial Analysts, 2025, 443,100 jobs, 29,500 projected annual openings, 7 percent projected growth to 2035, https://www.bls.gov/ooh/business-and-financial/financial-analysts.htm
See what the employer actually receives. A full report for one role, with every score shown beside the excerpt that earned it, conduct findings reported rather than averaged, and a reviewer sign-off required before any decision. No form.
Read a sample reportOr talk to us about this role