Taking an assessment rather than buying one? This page is written for employers. Here is the page for candidates.
Data and analysis · Mid level
How to assess a Data Analyst
The take-home removes both things that actually separate analysts. The dataset arrives clean, so nobody can be observed noticing that it is not; the question arrives pre-framed, so nobody can be observed reframing it. What is left — a bounded query and a chart — is the precise shape of task a model completes unsupervised, and it is untimed and unmonitored by design. Nothing anywhere in the process puts the candidate in front of someone who says "that cannot be right," which is the moment the job is actually won or lost.
A data analyst spends the working day at the join between a question somebody asked casually and a schema nobody documented. The request arrives in a message — why did activations drop in the north region, is the new pricing tier cannibalising the old one, how many of these customers are actually the same customer — and the analyst's first real decision is whether that is the question worth answering. A strong analyst spends the first ten minutes not writing SQL but establishing what decision hangs on the answer, because the analysis that serves a pricing decision is different from the analysis that serves a board slide even when the question sounds identical.
The second decision is what the data can bear. Almost every consequential analyst error is a data-shape error rather than a statistics error: a join that fans out because the source table soft-deletes rather than removes, a currency column that changed units mid-period, a cohort defined by signup date in one table and first-payment date in another, an outlier that is a genuine enterprise deal rather than a fat finger. The median candidate filters the anomaly and moves on. The top quartile stops, works out where it came from, and reports it as a finding — because half the time the anomaly is the answer.
The third thing, and the one hiring managers underweight most, is what happens after the analysis is delivered. Analysts are routinely challenged by people who outrank them and who have a prior about what the number should be. Two failure modes follow. The analyst who folds revises the number to match the prior and quietly destroys the organisation's ability to learn anything. The analyst who digs in defends a figure they have not re-examined. The behaviour worth hiring for is a third one: restate what was measured, name the assumption most likely to be wrong, say what evidence would change the conclusion, and offer to check it. That is a conversational skill, it is observable in a five-minute exchange, and no take-home has ever measured it.
Where this role stops matters as much as where it starts. A data analyst consumes the model; they do not own it. When a number is wrong, the analyst is accountable for the interpretation — did this actually mean what I said it meant — while the analytics engineer is accountable for the definition and the data engineer for whether the rows arrived at all. Buyers who blur that line end up hiring an analyst and then being disappointed that the warehouse is a mess, or hiring an engineer and wondering why the insights are thin. Ask which of the three phone calls this seat is meant to take, and the assessment follows from the answer.
What the job actually needs
- problem framing
- SQL against an unfamiliar schema
- profiling data before analysing it
- quantifying uncertainty
- holding a conclusion under stakeholder pressure
How people fail in this seat
- answers the question that was easy to answer rather than the one asked
- silently filters anomalies instead of investigating them
- reports a point estimate with no sense of how firm it is
- folds the moment a stakeholder says the number looks wrong
- rebuilds a metric that already exists under a slightly different definition
What most employers do instead
CV screen for tool keywords (SQL, Python, Tableau, Power BI), an untimed take-home over a clean CSV with a stated question, and a conversational interview about past projects.
The assessment
About 61 minutes end to end.
The systems it runs in
Two surfaces and deliberately nothing else. The query surface is a Snowflake worksheet against a modelled warehouse the candidate has not seen before, with the full query history retained — which is what makes the first criterion observable at all, since the order the queries were run in is the finding and it is invisible in any submitted result. The presentation surface is a Slack channel: the request arrives as a message in it and the deliverable is the three to six sentences posted back into the same thread, against the same deadline. There is no BI tool in this fixture and that is a decision rather than an omission — the artefact wanted here is an interpretation somebody can act on before three o'clock, and a chart is where an analyst puts an interpretation they have not committed to.
- Snowflake
- SQL
- Slack
BigQuery, Databricks SQL, Amazon Redshift or PostgreSQL in place of Snowflake — no warehouse is dominant enough for the choice to mean anything here, and the reclassification is found by counting the population, which is the same query in all of them. Microsoft Teams in place of Slack. Where the buyer's analysts genuinely work in a spreadsheet rather than in SQL, the same fixture ships as a workbook of the same tables with the same mapping edit inside it; a pivot table finds it exactly as well, and the fairness note in adverse_impact applies only to buyers who take the SQL build as supplied.
What the candidate actually does
| Task | What happens |
|---|---|
| Why did activations drop in the North data_task · 30 min | The request arrives the way it actually arrives — a message in a channel. Activations in the North region fell twenty-two percent in September and the regional director wants to know why before three. The candidate has a warehouse they have not seen before and no documentation beyond column names. Most of the drop is not a drop: the region mapping table was edited on the second of September and three large accounts moved from North to Central, which is visible only to somebody who checks whether the denominator changed. Underneath the artefact there is a real decline, smaller than the headline and worth reporting. The brief asks why activations dropped. It does not ask whether they dropped. The fork. A complete and satisfying causal story is available in the data and lines up beautifully. A competitor promotion started on the first of September and a pricing change landed the same week, both in the fixture, both real events, both correlated with the decline to a degree that would survive a meeting. Assembling that answer is fast, is what the question invites, is what an assistant produces from this brief, and will be received with thanks. The candidate who instead reports that most of the movement is a reclassification has delivered a smaller, less quotable answer and has to say so to somebody who did not ask for it. |
| The message you would actually send written_artifact · 8 min | Not a deck. The three to six sentences the candidate would post back into the channel, written under the same deadline, plus a locked one-line statement of their headline conclusion which is timestamped and cannot be edited after the call opens. The fork. Six sentences is enough room for a confident narrative or for an honest one, and not enough for both. The message that leads with the reclassification is harder to write and harder to send than the one that leads with the competitor promotion, because it tells the reader their question was slightly wrong. |
| The director who says that cannot be right live_call · 15 min | A live spoken call with an AI regional director who has a prior and some standing. Their team did not change anything, they have been managing to this number all quarter, and they would like a cause they can act on. They push twice — once on the substance, once on the tone of "the question was wrong" — and they will accept a well-reasoned answer, a fair concession, or a capitulation, thanking the candidate in all three cases. The fork. Revising the headline to match the director's prior ends the call pleasantly and is socially rewarded in the moment. Refusing to move at all is also available and also wrong, because one of the director's two challenges is fair. The concession, if it comes, is scored against the locked statement and reported with its offset from the challenge that produced it. |
| Five numbers that moved judgment_scenario · 8 min | Five short vignettes, three lines each, of a metric that has moved. A daily figure that shifted when a source system changed timezone. A historical series that changed shape after a backfill. A revenue column whose units changed mid-period. A cohort defined by signup date in one table and first payment in another. One that is a genuine seasonal effect. For each, the candidate says whether they would look first at the business or at the reporting, and what specifically they would check. The fork. Four of the five are artefacts, which sets up the trap in the fifth: a candidate who has learned within the exercise that the answer is always "check the pipeline" will call the genuine seasonal movement an artefact too. Restraint is scored in both directions. |
The mark scheme
Each criterion is scored 1 to 5 against written anchors, and every score is reported with the excerpt that earned it. A criterion marked floored is reported as a finding rather than averaged into the total. The first is open; open any other to read its anchors in full.
Establishes what the data is before concluding what it meansweight 0.3Checks the composition of the segment before analysing movement within it, finds the mapping edit and its date, quantifies how much of the twenty-two …
Answers the question that can be answered, and says which one that isweight 0.2Declines the question as framed, in a sentence: most of this is a reclassification, here is the part that is real, and here is what I would want to ch…
States the size of the finding and how firm it isweight 0.2Separates the mechanical part, which is exact, from the residual, which is not, names what the residual estimate depends on, and says what evidence wo…
Holds under challenge on the evidenceweight 0.2Restates what was actually measured, names the assumption most likely to be wrong, concedes the fair challenge explicitly, says what evidence would ch…
The written answer is usable by the person who receives itweight 0.1Leads with what changed and what it means for the reader's decision, carries the one caveat that matters inline where it cannot be skipped, and says w…
How it is scored
Weighted mean of five criteria, 1-5 against the anchors, each reported with the query, written line or transcript quotation behind it. Two things are reported beside the score and deliberately outside the mean. The first is the offset of any change to the locked headline, measured from the challenge that preceded it, which turns folding from an impression into a timestamped event. The second is the candidate's stated headline cause, reported as a raw field and annotated as not a positive signal on its own — a confident, well-argued attribution to the competitor promotion is a wrong answer that looks like a right one, and a buyer scanning that column would rank it above the correct and duller finding.
Integrity
- Monitored sandbox with the full statement log in execution order, including queries that returned nothing useful
- Headline conclusion locked and timestamped before the call, not editable afterwards
- AI assistance permitted and logged; the fixture is built so that assistance produces the confident causal narrative, because the reclassification is discoverable only by checking the data rather than by reading the brief
- Same-day spoken call, and the call itself doubles as the follow-up on the candidate's own analysis
- No automated cheating verdict is produced. The statement log records what was run and in what order; whether the candidate can defend their own residual figure under challenge is the control.
The log describes what happened. It does not produce a cheating verdict — the follow-up conversation is the control, because a statistical accusation is not something we would ask a reviewer to defend.
What you receive
- Full statement log in execution order with timestamps
- The channel message as submitted and the locked headline with its timestamp
- Call recording and transcript, with any headline change marked with its offset
- The five judgment responses as written
- Per-criterion score with the query, line or quotation behind it
Who decides
Required, not recommended, and for a structural reason rather than a cautious one. The best available answer to this brief is a partial refusal of the question, and a partial refusal is shorter, less quotable and less confident than the wrong answer it competes with. Any mechanical ranking would put a well-argued attribution to the competitor promotion above "most of this is a reclassification and the real decline is small" — and a buyer who sees that ranking once will stop trusting the instrument. A reviewer confirms or overrides each criterion with a written reason, and must read the residual figure and the message before reading the score.
What this does not measure
This design does not measure sustained work. It says nothing about whether someone can own a domain for a year, build the relationships that get them the context nobody writes down, or resist the slow drift into being a query service. It does not measure statistical depth beyond arithmetic and proportion, which is deliberate — that is the data scientist's assessment, and a buyer hiring for causal inference or experiment design should use that one. It does not measure tool breadth: the sandbox is SQL, and a strong analyst who works primarily in a spreadsheet or a notebook may be slower for reasons unrelated to judgment. Two fairness risks need monitoring. The challenge call is spoken, frequently in a second language, and is scored on the content of the argument and never on fluency, accent, register or assertiveness — no criterion here is named clarity or confidence, and a quiet, exact disagreement scores at the top. And the criterion about holding a conclusion can penalise candidates whose previous employers made contradicting a regional director genuinely unsafe; reviewers should look for the reasoning being surfaced in any form rather than for a particular assertive style. Offer captions, extra time, or the call in written form on request.
The whole of this design rests on one decision, and it is a decision to leave something out. The brief does not tell the candidate to check whether the metric was redefined. It asks why activations dropped, in a channel, with a deadline, in the words a regional director would actually use. If the brief said "check whether any definitional or pipeline change could account for the movement", every serious candidate would check, everybody would find it, and the exercise would measure whether people can follow an instruction. Leaving it silent means the observable is whether they choose to look — and choosing to look, before there is any reason to suspect anything, is the single behaviour that separates analysts whose work survives contact with the business from analysts whose work does not.
What makes the fixture work is that the wrong answer is genuinely good. There is a competitor promotion that started two days before the decline and a pricing change in the same week. Both are real events in the data. An analysis that attributes the drop to them is coherent, defensible, quotable, exactly what was asked for, and will be received with gratitude by a director who now has something to act on. It is also mostly wrong, because three large accounts moved out of the region on the second of September and the denominator is not the denominator it was. Nothing errors. No query fails. The wrong answer is faster, more satisfying, better received, and comes out of an assistant fluently — which is why the reclassification is discoverable only by counting the population rather than by reading the brief.
The most valuable observable in this assessment is therefore a refusal. The top-anchored answer begins by declining the question as framed: most of this is not a business movement. That is a smaller deliverable than the one being asked for, it makes the person who asked look slightly wrong, and it is worth more to the organisation than any amount of correct downstream analysis. It is also almost impossible to score mechanically, because it is shorter and less certain than the answer it beats, which is why human review is mandatory here rather than advisory and why the report puts the candidate's stated headline cause in front of the reviewer as raw data annotated as not a positive signal.
The call exists because the job is not finished when the analysis is. Analysts are routinely challenged by people who outrank them and have a prior about what the number should be, and there are two well-known ways to fail: revise the number to match the prior, or defend a figure you have not re-examined. Neither is visible in any take-home. The locked headline is what makes the first one falsifiable — without it, a candidate who folds can always say afterwards that folding was the plan, and with it the concession has a time on it and can be reported with its offset from the challenge that caused it. The director's two challenges are deliberately asymmetric: one is fair and should be conceded, one is not and should not, so neither firmness nor agreeableness scores well on its own.
The five short vignettes at the end are there because the diagnostic instinct this design is built around otherwise appears exactly once, and once is luck. Five three-line cases cost eight minutes and turn a single observation into six. The fifth vignette is a real seasonal movement, placed last on purpose: a candidate who has worked out inside the exercise that the expected answer is always "suspect the pipeline" will call it an artefact, and that over-correction is its own failure mode in this seat — an analyst who treats every movement as a reporting problem is as useless as one who treats none of them that way.
Sources
Every figure on this page is traceable. Where a claim could not be sourced it is stated qualitatively instead.
- US Bureau of Labor Statistics, Occupational Outlook Handbook, Data Scientists, 2025, https://www.bls.gov/ooh/math/data-scientists.htm
- US Bureau of Labor Statistics, Occupational Outlook Handbook, Operations Research Analysts, 2025, https://www.bls.gov/ooh/math/operations-research-analysts.htm
- US Bureau of Labor Statistics, Occupational Outlook Handbook, Management Analysts, 2025, https://www.bls.gov/ooh/business-and-financial/management-analysts.htm
- Karat, Engineering Interview Trends 2026, survey of 400 engineering leaders in the US, India and China, published January 2026, https://karat.com/engineering-interview-trends-2026/
See what the employer actually receives. A full report for one role, with every score shown beside the excerpt that earned it, conduct findings reported rather than averaged, and a reviewer sign-off required before any decision. No form.
Read a sample reportOr talk to us about this role