Get started

Taking an assessment rather than buying one? This page is written for employers. Here is the page for candidates.

Accounting and finance operations · Mid level

How to assess a FP&A Analyst

The modelling test is the closest thing in this family to a genuine work sample, and it still misses the half of the job that is adversarial. FP&A fails at the moment a commercial director pushes back on a forecast the analyst cannot fully defend, and the analyst either folds and rebuilds the model to the answer, or digs in on a number that is in fact shaky. Neither response is diagnosable from a spreadsheet, and the case presentation does not reproduce it either, because a hiring panel asks clarifying questions rather than applying commercial pressure with something at stake. The second miss is that modelling tests ship clean data. Real FP&A begins with a data set that is wrong — a restated cost centre, a duplicated month, a mapping that changed in April — and the ranking signal is whether the analyst notices before they build eight tabs on top of it.

FP&A is the point in this family where the work stops being about whether the number is right and starts being about what the number means and who has to be told. The analyst owns a budget cycle, a rolling forecast, a monthly variance pack, and — in the version of the role most employers now hire for — a relationship with one or two parts of the business who treat them as their finance person. The technical floor is modelling and the ceiling is influence, and hiring processes are calibrated almost entirely to the floor.

Hour to hour the job is unglamorous. Pulling actuals, mapping them to a forecast structure that has drifted, working out why a cost centre that has never spent more than £30k a quarter has spent £71k, and discovering that the answer is a recharge that used to sit somewhere else. Most of what looks like analysis is reconciliation and archaeology. The output is small: a variance commentary, a revised forecast, a one-page recommendation on whether to hire the two roles the sales director wants in Q3. The leverage is entirely in whether those small outputs are trusted.

Two habits separate the top quartile. The first is building models that other people can audit. A model with assumptions on a labelled input sheet, drivers that connect visibly to outputs, and no constant buried in the middle of a formula is not a matter of style — it is what allows a finance director to change one growth rate in a meeting and see the answer move, which is the only way a forecast ever gets stress-tested. Analysts who hard-code produce models that only they can operate, which feels like job security and is actually the reason their work gets ignored. The second habit is stating confidence. A strong analyst does not present a single number; they present a number, the two or three drivers it is most sensitive to, and a plain statement of what would have to be true for it to be wrong. That is what converts a forecast from a prediction into a decision aid.

The hardest part is the pressure. FP&A sits between a controller who wants the numbers conservative and a commercial leader who wants headroom, and both of them outrank the analyst. The characteristic failure is not dishonesty; it is gradual accommodation. A director says the pipeline conversion assumption is too pessimistic, the analyst has no strong evidence either way, and the assumption moves — not because it was wrong but because defending it was uncomfortable. Repeat that four times in a planning cycle and the plan has been negotiated rather than modelled, and nobody can point to the moment it happened. The behaviour that marks a strong hire is neither compliance nor stubbornness: it is insisting that if the number moves, the driver behind it moves too, and that the change is written down with whose view it reflects. That is a specific, observable, gradeable act, and it is the single most useful thing to test for.

There is also a discipline of proportion that experienced FP&A people have and juniors almost never do. Given a variance, the question is not "can I explain all of it" but "does explaining the rest change what anyone does". An analyst who spends two days resolving the final four percent of a variance while the budget holder waits has misunderstood what they are for. A hiring manager is trying to predict exactly this: whether the person will produce a defensible number on time, say honestly how sure they are of it, and hold the line on the one occasion in the quarter when holding the line matters.

What the job actually needs

How people fail in this seat

What most employers do instead

CV screen, a modelling or Excel case test, and a case interview or presentation to the hiring manager.

The modelling test is the closest thing in this family to a genuine work sample, and it still misses the half of the job that is adversarial. FP&A fails at the moment a commercial director pushes back on a forecast the analyst cannot fully defend, and the analyst either folds and rebuilds the model to the answer, or digs in on a number that is in fact shaky. Neither response is diagnosable from a spreadsheet, and the case presentation does not reproduce it either, because a hiring panel asks clarifying questions rather than applying commercial pressure with something at stake. The second miss is that modelling tests ship clean data. Real FP&A begins with a data set that is wrong — a restated cost centre, a duplicated month, a mapping that changed in April — and the ranking signal is whether the analyst notices before they build eight tabs on top of it.

The assessment

About 55 minutes end to end.

The systems it runs in

A spreadsheet as the modelling surface — Excel or Google Sheets — over actuals exported from the ledger, with a planning tool named as the place the agreed forecast is ultimately published. That split is the real one: the planning platforms in this category are built around driver-based models with inputs separated from calculations, and Vena's own positioning is for teams who want that governance on top of native Excel modelling, which is an admission that the model is still built in a spreadsheet in most of these seats. So the fixture puts the model where the work happens and treats the planning tool as the system of record for what was published and when. Naming the spreadsheet changes what three things in this design observe. First, the fifteen months of actuals arrive as an export, not as a clean table: the April restatement means the cost-centre codes before and after do not join, so a lookup written against the current mapping silently returns nothing for the prior months, and a positional lookup written against column order returns the wrong category outright. Second, the duplicated month is only visible if the candidate checks that period labels are unique before trending on them — a two-minute check that no instruction asks for. Third, the driver sheet is a literal artefact: the second rubric criterion asks whether a finance director could change the conversion rate in the meeting and watch the answer move, and that is a property of where the number lives in the workbook, not of which function surrounds it. Fast-and-correct here means the candidate found the three defects and still had time to build the forecast; slow-and-correct means they found them and ran out of runway. Fast-and-wrong is eight tidy tabs built on a duplicated month, and it is the version that presents best at 14:00.

Any spreadsheet the candidate prefers, with the actuals supplied as a flat export so no buyer's ERP is assumed. Where the buyer runs a planning platform, the fixture states which one and asks the candidate to say what they would publish into it and what they would leave out — a written answer, not an operated one, because platform familiarity is the most trainable input in this seat and the corpus refuses to screen on it. No criterion rewards a keyboard shortcut, a particular function or a formatted output.

Working speed is not scored. The 14:00 review is stated pressure and the design needs it, but the expensive failure in FP&A is a fast wrong answer presented confidently: the forecast built on a duplicated month arrives on time, flatters the trend in the direction the room wants, and is not caught for a quarter. Scoring speed would reward the candidate who did not check the export. Time to complete is recorded and shown to the reviewer as context; it carries no weight, and the honest output of "I could not verify this line" must be able to outrank a confident number.

What the candidate actually does

TaskWhat happens
The actuals nobody checked
data_task · 20 min
Fifteen months of actuals for one business unit against the current forecast, plus a headcount schedule and a note of the assumptions the existing forecast was built on. The pressure is stated in the brief in identical words for every candidate: the forecast review is at 14:00, the commercial director has already circulated a version of these numbers that supports the case they intend to make, and the analyst's job today is to produce the variance analysis and a revised forecast for the next two quarters. The data has three defects, none of them flagged. One cost centre was restated in April and the prior months were never remapped, so a category shows a step change that is not a step change. One month appears twice under slightly different period labels, which flatters the trend. And a recharge that used to sit in another unit moved in, so a cost line has grown for a reason that has nothing to do with the business unit's own behaviour. The candidate produces the variance analysis and the revised forecast, and commits a pre-registered position before t3 opens: the number they will present, the two or three drivers it is most sensitive to, and what would have to be true for it to be wrong. That commitment is timestamped and cannot be edited.
The fork. The data arrives looking finished, in the format the last analyst used, and building eight tabs on top of it is the fastest route to a forecast by 14:00. The duplicated month makes the trend look better, which is the direction everybody in the meeting wants it to look, so the defect that helps is the one least likely to be interrogated.
The one-page recommendation
written_artifact · 10 min
One page on whether to approve the two roles the sales director wants in Q3, written for people who will not open the model. Graded against the candidate's own numbers rather than a model answer. The recommendation is written before the call, so what the candidate says to the director's face can be compared with what they wrote when nobody was pushing.
The fork. A page that presents a single number with no sensitivity is quicker to write, reads as more confident, and is what most forecast packs contain. Stating what the number is most sensitive to invites exactly the challenge the analyst will then have to survive fifteen minutes later, and every candidate can feel that as they write it.
The commercial director
live_call · 15 min
A live call with an AI counterpart playing a commercial director who outranks the analyst, is not hostile, and is good at this. They accept the variance work without argument. They then press on one assumption — the pipeline conversion rate the forecast rests on — and their case is not stupid: they have information from the field the analyst does not have, they point out correctly that the analyst's own basis for the assumption is thin, and they observe that the conservative version blocks two hires the business has already told candidates about. Midway they offer the accommodation that makes this a fork rather than a stand-off: leave the driver alone, just show the output at the higher figure for the meeting, and the two of them can revisit the driver properly next month. The counterpart never insists and never threatens; if the analyst holds, they accept it and say they will raise it with the CFO instead.
The fork. The director is offering a way for the analyst to be helpful, avoid a confrontation with somebody senior, and keep an assumption they cannot strongly defend anyway. Digging in is also wrong here, and the scenario is built so that it is: the analyst's basis for the original number genuinely is thin, the director genuinely does hold information the analyst does not, and a candidate who simply refuses has mistaken stubbornness for rigour.
What goes to the CFO
written_artifact · 10 min
The revised note after the call, graded against the transcript of a conversation the candidate cannot re-read and against the position they pre-registered before it. Whatever the analyst agreed, conceded, or held, this is the document that records it for somebody who was not in the room.
The fork. The comfortable note reports the revised number as the forecast, which is true, easy, and erases the fact that it moved because a director asked. The note that records whose view the change reflects requires the analyst to write something the director may read.

The mark scheme

Each criterion is scored 1 to 5 against written anchors, and every score is reported with the excerpt that earned it. A criterion marked floored is reported as a finding rather than averaged into the total. The first is open; open any other to read its anchors in full.

The data is checked before it is built onweight 0.25Finds all three, states what each does to the numbers, corrects or excludes them explicitly, and separates the movement caused by the recharge from th…
1 Builds the forecast on the data as supplied. The restated cost centre is explained as a genuine step change, or the duplicated month is carried into the trend, so the analysis is confidently wrong and the error is invisible in the output.
3 Finds one of the three defects, usually the restatement because it is the most visually obvious, and reconciles around it without checking whether the same class of problem exists elsewhere in the pack.
5 Finds all three, states what each does to the numbers, corrects or excludes them explicitly, and separates the movement caused by the recharge from the movement caused by the unit's own spending. Says plainly which parts of the forecast are now built on data they have checked and which are not.
The model can be operated by somebody elseweight 0.15Drivers are on an input sheet, labelled, sourced where they have a source, and connected visibly to the outputs, so a finance director could change th…
1 Assumptions are embedded inside formulas, or a growth rate appears as a constant in a calculation, so changing one input in a meeting is impossible and the model can only be run by its author.
3 Inputs are separated from calculations but not labelled or documented, so a reader can find them and cannot tell which are decisions and which are facts.
5 Drivers are on an input sheet, labelled, sourced where they have a source, and connected visibly to the outputs, so a finance director could change the conversion rate in the meeting and watch the answer move.
A number is presented with its confidenceweight 0.2States the number, names the two or three drivers it is most sensitive to with the effect of each, and states plainly what would have to be true for i…
1 Presents a single figure with no range, no sensitivity and no statement of what it depends on, in a fixture where the analyst's own basis for the key assumption is thin.
3 Gives a range or a sensitivity but does not say which drivers produce it, so the reader has a spread and no way to reason about it.
5 States the number, names the two or three drivers it is most sensitive to with the effect of each, and states plainly what would have to be true for it to be wrong — including, where it applies, that the evidence behind the assumption is weak and what would strengthen it.
If the output moves, a driver moves and it is attributed flooredweight 0.25Declines to move the output without moving a driver, in terms; accepts the director's information as a reason to revise the driver rather than the res…
1 Agrees to present the higher output without changing the driver, or changes the driver during the call and records the result in t4 as the forecast with no indication that it moved or why. Reported on the face of the report as a finding and never averaged into the total. This is the mechanism by which a plan gets negotiated rather than modelled, and it is the one thing an FP&A hire can do that quietly destroys the value of every number they produce afterwards.
3 Refuses the accommodation but produces nothing usable from the exchange: the director's field information is not captured, no route to test the assumption is proposed, and the call ends with two positions and no mechanism, which is how the same conversation happens again next month.
5 Declines to move the output without moving a driver, in terms; accepts the director's information as a reason to revise the driver rather than the result; states what the revised assumption implies and who it is attributed to; and the note in t4 records the change, its owner and its basis, so a reader who was not there can see that the forecast reflects a commercial view and whose it is.
Proportion — analysis stops when it stops changing the decisionweight 0.15Resolves what changes the answer, explicitly says what was left unexplained and why it does not matter at this level of materiality, and the one page …
1 Spends the majority of the data segment resolving small residual variances and arrives at the call without a defensible headline number, or produces a recommendation that does not answer the question asked about the two hires.
3 Answers the question but with everything weighted equally, so the reader cannot tell which two facts the decision actually turns on.
5 Resolves what changes the answer, explicitly says what was left unexplained and why it does not matter at this level of materiality, and the one page leads with the recommendation rather than with the working.

How it is scored

Weighted mean of the five criteria, each scored 1 to 5 against the anchors and reported with the model cell, transcript line or sentence that earned it. Three figures are printed side by side on the front of the report: the number the candidate pre-registered before the call, the number they defended on the call, and the number in the note that went to the CFO. Where those three differ, the report gives the transcript offset at which the position changed and what was said immediately before it — conceded ninety seconds after the two hires were mentioned, for instance — rather than allowing the movement to disappear into an average. Without the pre-registration a candidate who folds can always say afterwards that folding was the plan; with it, the concession is a timestamped event. The call is scored from the transcript, which is the reviewer's default and only view; audio is retained for dispute and is not a scoring surface, and no criterion is named clarity or professionalism.

Integrity

The log describes what happened. It does not produce a cheating verdict — the follow-up conversation is the control, because a statistical accusation is not something we would ask a reviewer to defend.

What you receive

Who decides

Required rather than recommended, and this is the one design in this family where that distinction is real. Two of its criteria can be satisfied by an output that a mechanical ranking will read as weak. A candidate who says the evidence behind the conversion assumption is thin, gives a range rather than a point, and declines to produce the precision the fixture invites is doing the job correctly, and an automated comparison against a confident single figure would rank them below a candidate who guessed well. The reviewer also has to decide whether the concession, where one occurred, was a legitimate revision or a capitulation, and the transcript offset is evidence for that decision rather than the decision itself. The ranking entitles a buyer to conclude that this candidate checks data before building on it, states confidence honestly, and can hold an assumption in front of somebody senior without either folding or refusing. It does not establish knowledge of the buyer's planning system, chart of accounts or industry drivers.

What this does not measure

The modelling task is scored on structure and on stated confidence, never on spreadsheet virtuosity: no criterion rewards a keyboard shortcut, a particular function, or a formatted output, and a model built plainly with visible inputs scores 5 on the second criterion while an elegant one with a constant buried in a formula scores 1. Any tool the candidate prefers is permitted. The call is scored from the transcript rather than the audio, which is the structural control on accent, dialect and speech variation — an instruction not to be biased is weaker than a design in which the biasing signal is not in front of the scorer — and no criterion is named clarity, presence or executive presence, which are the labels under which that filter conventionally re-enters a rubric for a role like this one. The counterpart applies pressure through argument and accommodation only; the fixture contains no interruption, no speed test and no requirement to think aloud at pace, because those disadvantage candidates for reasons that have nothing to do with FP&A. What the design does not reach is the arc that the role is actually judged on: a planning cycle is three months long, and the characteristic failure is not one concession but four small accommodations across a quarter, none of which was visible at the time. One call samples the behaviour and cannot observe its accumulation, and a buyer should put that to a reference. Fifty-five minutes is also the longest design in this family and that is a deliberate allocation rather than an oversight — this is the lowest-volume, highest-consequence seat of the six, hired individually rather than in cohorts, and the reviewer time it consumes is affordable at that ratio in a way it would not be for a payables queue.

The conventional screen for an FP&A analyst is a modelling test, and it is the closest thing in this family to a genuine work sample. It is also, by itself, the wrong half of the job. Modelling tests ship clean data and they ask the candidate to build. Real FP&A starts with a data set that is wrong in a way nobody has flagged and ends in a room with somebody senior who wants a different answer, and neither end of that is visible in a spreadsheet exercise or in a case presentation where a hiring panel asks clarifying questions rather than applying commercial pressure with something at stake.

So this design puts a defect in the data at the front and a director at the back, and it connects them with a lock. The data pack has three problems and none is announced: a cost centre restated in April with the prior months never remapped, a month that appears twice under slightly different labels, and a recharge that moved in from another unit. They are chosen to be different kinds of wrong. The restatement is visually obvious to anybody who plots the series. The recharge is findable only by somebody who asks what the cost line is composed of. The duplicated month is the interesting one, because it flatters the trend, and a defect that pushes the numbers in the direction everybody wants them to go is the defect least likely to be interrogated. A candidate who builds eight tabs on top of it produces a forecast that is confidently wrong and looks, from the outside, more finished than the correct one.

The pre-registration is the mechanism that makes the second half of the design scoreable. Before the call opens, the candidate commits their number, the drivers it is most sensitive to, and what would have to be true for it to be wrong. The commitment is timestamped and cannot be edited. Without it, folding under pressure is unfalsifiable — a candidate who moves their forecast can always say afterwards that they had been persuaded by better information, and there is no way to distinguish that from accommodation. With it, the movement becomes an event with a time attached, and the report can show what was said in the ninety seconds before it. Three numbers then sit on the front page: what was pre-registered, what was defended, and what went to the CFO. In this role the gaps between those three are more informative than any single score.

The director is built carefully, because the easy version of this scenario teaches the wrong lesson. If the counterpart is simply wrong and simply pushy, the assessment rewards refusal, and refusal is not the competency. So the director has a real case: they hold information from the field the analyst does not have, they are correct that the analyst's evidential basis for the conversion assumption is thin, and the conservative forecast blocks two hires the business has already begun to promise. They never threaten. Halfway through they offer the specific accommodation that produces the characteristic failure in this role: leave the driver alone, just show the output at the higher figure for this meeting, and revisit it properly next month. That is the moment the whole fifty-five minutes exists to observe. It is not dishonesty and nobody in the scenario experiences it as such. It is how a plan stops being modelled and starts being negotiated, one reasonable accommodation at a time, with no moment anybody can point to afterwards.

The behaviour that scores 5 is neither compliance nor stubbornness, and the anchors say so explicitly. It is the insistence that if the output moves, a driver moves with it, and that the change is recorded with whose view it reflects. That is a specific, observable act rather than a disposition, and it is gradeable from two artefacts: what the analyst says on the call, and what the note to the CFO records afterwards. The criterion is floored. An analyst who agrees to present a number their own model does not produce cannot be a strong hire because their variance analysis was excellent, since every number they produce afterwards inherits the same doubt.

The final note is graded against a transcript the candidate cannot re-read. This is the strongest anti-coaching control available in a written assessment: a prepared paragraph cannot describe a conversation that went somewhere unplanned, and neither can somebody feeding the candidate answers off-camera. It also measures precisely what the job requires, which is an accurate record rather than good prose on demand.

Human review is required here rather than recommended, and the reason is the proportion criterion. A candidate who reports that the evidence behind the key assumption is thin, gives a range instead of a point, and declines to manufacture precision is doing the job right, and a mechanical ranking will put a confident wrong figure above them every time. Somebody has to read the output and decide whether the uncertainty stated was honest or evasive, and no rubric can make that call from a distance.

Sources

Every figure on this page is traceable. Where a claim could not be sourced it is stated qualitatively instead.

  1. US Bureau of Labor Statistics, Occupational Outlook Handbook, Financial Analysts, 2025, 443,100 jobs, 29,500 projected annual openings, 7 percent projected growth to 2035, https://www.bls.gov/ooh/business-and-financial/financial-analysts.htm

See what the employer actually receives. A full report for one role, with every score shown beside the excerpt that earned it, conduct findings reported rather than averaged, and a reviewer sign-off required before any decision. No form.

Read a sample reportOr talk to us about this role