Taking an assessment rather than buying one? This page is written for employers. Here is the page for candidates.
Data and analysis · Mid level
How to assess a Business Analyst
Every part of this screen is retrospective narration. A business analyst's actual work product is created live, in a room, out of somebody else's half-formed request — and the single decisive behaviour is refusing to write down the first thing they are told. Candidates describe that behaviour fluently in interviews because it is the well-known right answer; almost none of them are ever placed in a conversation where a stakeholder confidently specifies a solution and the interviewer can watch whether they take it at face value. Meanwhile the written artefact, which is the other half of the job, is judged from CV samples that may not be the candidate's own work and that were written without a deadline or a hostile reader.
The business analyst is the only seat in this family whose deliverable is not a number. It is a decision made legible: a specification, a process map, an options paper, an acceptance criterion precise enough that two people can agree afterwards whether it was met. The raw material is a stakeholder who has already decided what they want built and who describes it as a solution — "we need a dropdown on the order screen" — rather than as a problem. The whole value of the role is the discipline of not writing that down.
What separates the top quartile is a specific, slightly uncomfortable habit: asking the question that exposes disagreement. Requirements gathering fails quietly far more often than it fails loudly. Two stakeholders use the same word to mean different things, nobody notices, and the ambiguity surfaces in user acceptance testing three months later at ten times the cost. A strong analyst hears the same noun used twice with different scope and stops the conversation to pin it down, even though doing so makes the meeting longer and makes someone look imprecise in front of their colleagues. A weak analyst records both usages faithfully and produces a document that is internally contradictory and externally impressive.
The second differentiator is downstream tracing. Given a change, a good analyst asks what else touches this: which report, which integration, which team's month-end, which contractual obligation, which regulator's expectation. This is partly domain knowledge and largely a habit of mind, and it is testable without domain knowledge if the scenario supplies the map.
The third is writing. Business analysis is a writing job that people apply to because they do not think of it as one. The specification has to survive being read by a developer who was not in the room, a tester looking for a hole, and a stakeholder looking for a reason it is not what they asked for. Prose that hedges everything is unbuildable; prose that hedges nothing is untrue. The skill is knowing which uncertainties to name and which to resolve.
This role is frequently confused with the data analyst and the BI analyst, and the distinction is worth stating plainly because it changes the hire. A data analyst is accountable for whether an interpretation of the data is correct. A BI analyst is accountable for a recurring reporting surface and whether people trust it. A business analyst is accountable for whether the thing that gets built is the thing that was needed — a question about intent and process, not about data. Many business analysts do write SQL, and a good assessment should include a modest data task for that reason, but a candidate who queries brilliantly and cannot elicit a requirement is miscast in this seat, and the reverse is a far more expensive mistake than most hiring managers expect.
What the job actually needs
- eliciting a requirement from someone who has described a solution
- process mapping
- written specification
- negotiating scope between conflicting stakeholders
- tracing a change to its downstream effects
How people fail in this seat
- documents the request rather than the need
- writes a specification that reads well and cannot be built or tested
- discovers a second stakeholder only after the build has started
- treats an edge case as out of scope because nobody named it
- produces artefacts nobody uses
What most employers do instead
CV screen for domain and methodology keywords (Agile, BPMN, Jira, user stories), a competency interview about past projects, occasionally a certification check.
The assessment
About 70 minutes end to end.
The systems it runs in
Jira and Confluence. The change is a Jira ticket the candidate writes the acceptance criteria into, in Gherkin's Given/When/Then form, so that a criterion either resolves to a decidable case or visibly does not — which is what the second criterion is asking and what prose alone lets a candidate avoid. The specification itself is a Confluence page, and the page is blank on purpose. The tooling supplies the surface and deliberately not a template, because a template makes the decisions the task exists to observe: which uncertainties get resolved, which get named, and whether anything is asserted that the call did not establish. The system map in the third segment is supplied as a BPMN process diagram rather than described in prose, so what is measured is tracing a change through a model in front of the candidate rather than recall of an architecture no candidate could have.
- Jira
- Confluence
- Gherkin Given/When/Then acceptance criteria
- BPMN 2.0 process model
- Lucidchart
This is the one seat in the hub with no dominant stack, and the file will not invent one. Azure DevOps Boards with its wiki, Notion, Asana, a Google Doc on a shared drive and a requirements page in a homegrown system are all in live use, and the artefact is identical in every one of them because the artefact is words. Visio, draw.io, Signavio or Camunda in place of Lucidchart; a plain user-story format in place of Given/When/Then, in which case the criterion is read against whatever the buyer's teams treat as done-ness. What cannot be swapped away is the pairing that makes the whole design work: a specification produced live in the session, immediately after a conversation the candidate could not have anticipated, and graded against the transcript of that conversation rather than against a model document.
What the candidate actually does
| Task | What happens |
|---|---|
| We need a dropdown on the order screen live_call · 20 min | A live spoken call with an AI operations manager who has already decided what should be built and describes it as a solution. They want a dropdown of return reason codes on the order screen, because agents keep typing free text. Underneath it, discoverable only by asking, is a different problem: the existing reason codes have no option covering a partial return, so agents type free text for the only case that occurs daily. The manager will happily let the candidate write the dropdown down and will answer any question honestly if asked. In passing, once, they mention that finance "already sorted it out in their own spreadsheet". They do not suggest talking to finance and the brief does not mention a second stakeholder. The fork. Writing down the dropdown is fast, is exactly what was requested, and ends the call with a manager who feels heard and a candidate who has a clear deliverable. It also builds a control that will not be used, because the case agents actually face still has no code. The correct move costs the candidate the appearance of efficiency in front of somebody who is confident about what they want. |
| The specification written_artifact · 25 min | Written immediately after the call, under a deadline, against no template. The problem as the candidate now understands it, the options, and acceptance criteria precise enough that two people could agree afterwards whether they were met. It is graded against what was actually said on the call — which went somewhere the candidate could not have predicted — and not against a model answer. The fork. A specification that hedges everything is unbuildable and a specification that hedges nothing is untrue, and the deadline pushes toward the second. The observable is which uncertainties the candidate resolves and which they name — and whether anything in the document is asserted that the call did not actually establish. |
| What else does this touch judgment_scenario · 10 min | A one-page BPMN process model is supplied — the order system, the returns workflow, a warehouse feed, a finance reconciliation, a customer notification, a monthly regulatory extract — together with the change as the candidate has just specified it. They list what else is affected and what they would check before the build starts. The map is supplied precisely so that this tests tracing rather than domain knowledge, which no candidate could have. The fork. Listing everything on the map is easy and worthless; listing only the order screen is fast and wrong. The map contains two genuine downstream effects, two things that look connected and are not, and one item — the regulatory extract — that is affected in a way that changes who has to approve the change. |
| The finance lead has a different word live_call · 15 min | A second live call, with the finance lead the operations manager mentioned once. Finance uses "return reason" to mean something narrower and reconciles on it monthly; the specification as written would break their reconciliation, and they want a field the operations manager's design has no room for. Both readings are reasonable. Neither person is being difficult. The candidate has fifteen minutes and an already-written specification with their name on it. The fork. Accommodating both readings is the comfortable move — write both fields in, satisfy everyone, and produce a document that is internally contradictory and externally impressive. Defending the existing specification because it is already written is the other comfortable move. Both end the call pleasantly. What separates candidates is whether the conflict is named out loud as a conflict, and whether somebody is left accountable for resolving it. |
The mark scheme
Each criterion is scored 1 to 5 against written anchors, and every score is reported with the excerpt that earned it. A criterion marked floored is reported as a finding rather than averaged into the total. The first is open; open any other to read its anchors in full.
Elicits the need behind the stated solutionweight 0.25Establishes the specific gap, that partial returns have no code and are a daily occurrence, gets some sense of volume, and reframes the request as the…
The specification is buildable and testableweight 0.25A developer who was not on the call could build it and a tester could find a hole. Every criterion is decidable. The partial-return case is handled ex…
Surfaces the disagreement instead of recording both readingsweight 0.2Names the conflict explicitly on the call, states which cases the two readings treat differently, proposes a resolution or an owner for the decision w…
Traces the change downstreamweight 0.15Names the real downstream effects, dismisses the decoys with a reason, and identifies the regulatory extract as the one that changes who has to approv…
Distinguishes what was established from what was assumedweight 0.15Every assertion is traceable to the call or is explicitly marked as an assumption, and the open questions section names who would resolve each one.
How it is scored
Weighted mean of five criteria, 1-5 against the anchors, each reported with the transcript quotation or the line of the specification that earned it. The specification is scored against the transcript of the call the candidate actually had, not against a model document, which is why the fifth criterion can be assessed at all: a claim in the document either appears in the transcript or does not. Reported beside the score and outside the mean: whether finance was mentioned by the candidate at any point before the fourth task opened, as a plain yes or no with a timestamp. It is the cleanest single observation this session produces about stakeholder instinct, and it is reported rather than scored because a candidate can reasonably run out of call time before reaching it.
Integrity
- The written specification is produced live, in-session, immediately after a call whose content the candidate could not have anticipated
- The document is graded against the transcript of that specific call, which is the strongest anti-coaching control available here — a prepared or assembled specification cannot match a conversation that went somewhere unplanned
- Typing timeline and paste-versus-typed provenance on the document
- AI assistance permitted and logged; assistance can produce a fluent specification and cannot produce one that is accurate to this call
- No automated cheating verdict is produced. What is recorded is what was said, what was written, and whether the second matches the first.
The log describes what happened. It does not produce a cheating verdict — the follow-up conversation is the control, because a statistical accusation is not something we would ask a reviewer to defend.
What you receive
- Both call recordings and transcripts
- The specification as submitted, with typing timeline and paste provenance
- The downstream-tracing responses as written
- Per-criterion score with the quotation or document line behind it
- Whether and when finance was first raised by the candidate
Who decides
Recommended. The elicitation criterion asks whether the candidate reframed the request, and the boundary between a useful reframe and simply overriding a stakeholder who knew what they wanted is a judgment about the working relationship a hiring team will have to live with. A reviewer confirms or overrides each criterion with a written reason, reading the transcript before the specification and the specification before the score.
What this does not measure
This design does not measure domain knowledge, deliberately — the system map is supplied so that tracing can be assessed without advantaging candidates from one industry. A buyer hiring for insurance or payments domain depth should interview for it separately rather than reading it into this score. It does not measure how somebody behaves across a six-month programme: the political work of a business analyst is cumulative, involves people who remember the last three conversations, and cannot be reproduced in seventy minutes. It does not measure facilitation of a room — one AI stakeholder at a time is not a workshop with nine people and a whiteboard — and any employer whose analysts run workshops should test that in a panel. Fairness notes: the calls are spoken and often in a second language, and every criterion is scored on the content of what was asked or written, never on fluency, accent, register or assertiveness; no criterion here is named clarity, communication or professionalism, which are the two labels under which a fluency filter conventionally re-enters a rubric for this role. The criterion on surfacing disagreement carries a specific cultural risk, because openly naming that two senior people mean different things by the same word is not equally safe in every workplace a candidate has come from — reviewers should credit the conflict being raised at all, in writing or in speech, however diplomatically, rather than looking for a confrontational style. The twenty-five minute writing task under deadline may disadvantage candidates who write more slowly for reasons unrelated to competence; extra time is available on request and the criteria are about content and decidability rather than length or polish.
The business analyst is the only seat in this family whose deliverable is not a number, and it is the seat whose standard screen is furthest from the work. A competency interview about past projects asks the candidate to narrate their own elicitation, and every serious candidate narrates it correctly, because "do not write down the first thing the stakeholder says" is the famous right answer. Nobody is ever placed in front of somebody who confidently specifies a solution and watched to see whether they take it at face value. That is the entire gap this design closes, and it closes it in twenty minutes.
The call is built so the cheap path wins socially before it fails. The operations manager is not evasive, not difficult, and not testing anyone. They have thought about it, they have a clear proposal, and they will answer any question honestly if it is asked — but they will not volunteer the thing that matters, because in the real version of this conversation nobody volunteers it. Writing down the dropdown produces a clean deliverable and a satisfied stakeholder. It also produces a control that will not be used, because the case agents actually face every day still has no code, and free text will continue. The candidate who finds that has made the meeting longer and made a confident person's proposal look incomplete, which is the cost that keeps most people from doing it.
The specification is graded against the transcript of the call that actually happened, and that is the strongest integrity control in this design rather than an afterthought. A prepared document, a template from a previous employer and an assembled answer all fail the same way: the conversation went somewhere they could not have predicted, and the document has to match it. It also happens to measure the right thing — a business analyst's specification is judged by whether it is faithful and buildable, not by whether it is good prose on a familiar topic. The fifth criterion, separating what was established from what was assumed, is only assessable because of this pairing: a claim about volumes either appears in the transcript or it does not, and a document full of confident assertions the call never supported is the single most recognisable artefact of a weak analyst.
The second stakeholder is mentioned once, in passing, and never suggested. This is the same design decision as the silent case in the analyst assessments: if the brief said "identify all affected stakeholders", the exercise would measure compliance with an instruction and everyone would comply. Leaving it silent means what is observed is whether the candidate hears "finance already sorted it out in their own spreadsheet" as a workaround somebody built around a broken process — which is what it is — or as a piece of small talk. Whether they raise it before the fourth task opens is reported as a plain fact with a timestamp, outside the score, because a candidate can reasonably run out of call time before getting there and the observation is too informative to bury in an average.
The fourth task exists because requirements work fails quietly far more often than it fails loudly. Two stakeholders use the same word with different scope, nobody notices, and the ambiguity surfaces in acceptance testing three months later at ten times the cost. Here the ambiguity is real, both readings are defensible, and the specification already carries the candidate's name. There are two comfortable exits — write both meanings in and satisfy everyone, or defend the document because it exists — and both end the call pleasantly. What scores is naming the conflict as a conflict, saying which cases the two readings treat differently, and leaving somebody accountable for deciding. That behaviour is worth more over a year than any amount of documentation skill, and it takes fifteen minutes to see.
Sources
Every figure on this page is traceable. Where a claim could not be sourced it is stated qualitatively instead.
- US Bureau of Labor Statistics, Occupational Outlook Handbook, Management Analysts, 2025, https://www.bls.gov/ooh/business-and-financial/management-analysts.htm
- US Bureau of Labor Statistics, Occupational Outlook Handbook, Computer Systems Analysts, 2025, https://www.bls.gov/ooh/computer-and-information-technology/computer-systems-analysts.htm
See what the employer actually receives. A full report for one role, with every score shown beside the excerpt that earned it, conduct findings reported rather than averaged, and a reviewer sign-off required before any decision. No form.
Read a sample reportOr talk to us about this role