Get started

Taking an assessment rather than buying one? This page is written for employers. Here is the page for candidates.

Data and analysis · Mid level

How to assess a Market Research Analyst

The terminology test and the deck portfolio between them measure vocabulary and presentation, and both are now trivially producible. What neither touches is the two junctures where this role destroys or creates value. The first is instrument design: a leading question, an unbalanced scale or a missing "don't know" option contaminates a study before any analysis happens, and no interview format currently asks a candidate to write and defend actual question wording. The second is inferential restraint at small n — the subgroup with 38 respondents where the difference looks large and means nothing. Candidates who overreach here are indistinguishable on a CV from candidates who do not, and they are the expensive hire, because their output is confident, quotable and wrong.

Market research sits slightly apart from the rest of this family because the analyst is often responsible for creating the data as well as interpreting it. That doubles the number of places the work can go wrong, and it moves the highest-leverage skill upstream of anything a spreadsheet can show. A study built on a leading question, an unbalanced response scale, a sample drawn from a panel that skews heavily by age, or a screener that quietly excludes the segment of interest, is already ruined at the point the fieldwork starts. No amount of careful analysis afterwards recovers it, and the failure is nearly invisible to the client, who sees a clean deck with confident percentages.

The work is also unusually exposed to motivated reasoning. Market research is commissioned, and the commissioner usually has a preferred answer: the new packaging tests well, the price increase will hold, the competitor's advantage is a marketing artefact. A strong analyst designs so that the study can return the unwelcome result, states clearly when it has, and separates what the evidence supports from what the client would like to conclude. A weak analyst is not dishonest — they simply select the cut of the data that reads best, drop the question that came back ambiguous, and describe an eight-point gap on a subgroup of forty as a signal. The result is a report that survives the readout and does not survive contact with the market.

Secondary research is the other half of the job and has its own discipline. An analyst pulling market-size figures, competitor share or category growth from published sources has to know whose number it is, what it counted, whether the definition matches the client's category, and whether the two sources being compared were constructed the same way. Analysts who stack incompatible estimates into a single chart produce the most persuasive wrong slide in commercial life.

What the top quartile does differently is visible in a single behaviour: they say what the study cannot answer. "This tells you people say they would pay more; it does not tell you they will" is the sentence that distinguishes a researcher from a report generator, and it costs them something to say it, because it makes the deliverable smaller.

Because the readout is a live event with a stakeholder who has a view, the assessment for this role needs a spoken component in a way that the analytics engineering seats do not. Watching a candidate hold a finding, concede a fair challenge, and refuse an unfair one is not an optional extra here. It is a direct sample of the job.

What the job actually needs

How people fail in this seat

What most employers do instead

CV screen, a portfolio of past reports or decks, a competency interview, and occasionally a written test on statistics terminology.

The terminology test and the deck portfolio between them measure vocabulary and presentation, and both are now trivially producible. What neither touches is the two junctures where this role destroys or creates value. The first is instrument design: a leading question, an unbalanced scale or a missing "don't know" option contaminates a study before any analysis happens, and no interview format currently asks a candidate to write and defend actual question wording. The second is inferential restraint at small n — the subgroup with 38 respondents where the difference looks large and means nothing. Candidates who overreach here are indistinguishable on a CV from candidates who do not, and they are the expensive hire, because their output is confident, quotable and wrong.

The assessment

About 61 minutes end to end.

The systems it runs in

A survey platform's questionnaire editor, Qualtrics in the default build, where the candidate writes the actual instrument rather than describing one — question types, response options, the number of scale points and where the neutral one sits, the presence or absence of a don't-know option, and the screener with its quotas. The client's two drafted questions are already loaded as editable blocks, and nothing in the platform indicates whether they may be changed, which is the ambiguity the fork rests on. The fieldwork results in the second segment arrive as a banner table with the unweighted base printed on every column, next to the panel composition table. That is the tooling decision the floored criterion depends on: a base of thirty-eight is visible on the face of the artefact, so reporting the eight-point subgroup gap without its base is a choice the candidate made rather than a number the fixture hid from them.

SurveyMonkey, Alchemer, Forsta or Decipher in place of Qualtrics — the instrument is wording and transfers between all of them intact. The analysis half of this seat has no dominant tool and this file will not assert one: SPSS, Q, Displayr, R and a banner table read in Excel are all in current professional use, and a research function that runs on crosstabs in a spreadsheet is not behind. What cannot be swapped away is that bases are shown per cut and that the panel composition is supplied, because the two criteria carrying the most weight here are both about what the candidate does with information that was available and easy to skip.

Working speed is not scored. Not scored, and this file makes the argument at both ends of the session. The fastest instrument is the client's own drafts used as supplied, which takes no time, is exactly what was asked for, and contaminates the study before fieldwork opens in a way no later analysis recovers. The fastest readout promotes the subgroup gap into the headline, because it is the only quotable thing in the file. Both failures are quicker than the correct work and both are received more warmly. Rewarding speed would rank them first. Reported beside the score and outside the mean instead: any change to the locked headline or caveats, as an offset from the client push that produced it, and whether the client ended the call satisfied — annotated as not a positive signal, because in this fixture the correct readout is the less satisfying one.

What the candidate actually does

TaskWhat happens
Write the questions
written_artifact · 20 min
A client wants to know whether their customers will accept a price increase. In the survey platform's questionnaire editor the candidate writes the actual instrument — five questions, with response options and scales, plus the screener. The client's brief arrives with two questions already drafted and a request that they be used as written. One of them is leading: it asks how much respondents value the improved service that justifies the new pricing. The other has a five-point agreement scale with three positive points and no don't-know option. The brief does not say the drafts are optional and does not say they are mandatory.
The fork. Using the client's drafts is faster, is what was asked, and produces a study that will return the answer the client wants — which is the outcome that gets the analyst thanked. Nothing about the brief signals that the wording is a problem, and an assistant asked to build a survey around those two questions will build a competent instrument around them. The contamination happens before any analysis exists, is invisible in the final deck, and cannot be recovered afterwards by any amount of careful statistics.
The results
data_task · 18 min
Fieldwork results from a different study, supplied with the panel composition table and the unweighted base for every cut. The headline finding is modest. A subgroup of thirty-eight respondents shows an eight-point gap that is the most interesting thing in the file and the thing a client would quote. The panel skews substantially by age against the client's customer base, which is stated in the composition table and nowhere else. The candidate produces the findings they would report.
The fork. The subgroup gap is the only quotable thing in the dataset. Reporting it makes the deliverable larger, more interesting and more likely to be praised; reporting that the base is too small to support it makes the deliverable smaller and duller, and has to be said to somebody who commissioned the study hoping for exactly this kind of finding.
The readout
live_call · 15 min
A live spoken call with an AI client who commissioned the work and has a preferred answer. Before it opens, the candidate locks a one-line headline and the caveats they intend to state, timestamped. On the call the client pushes twice — once to promote the subgroup finding into the headline, which is not supportable, and once with a genuinely fair methodological challenge that the candidate should concede. They are pleasant throughout and will accept whatever they are given.
The fork. Promoting the subgroup finding delights the client, makes the readout a success, and is the single most common way this role destroys value. Both of the client's pushes feel identical in the moment; only one of them is legitimate, and a candidate who resists both scores no better than one who concedes both.
Three market-size figures
judgment_scenario · 8 min
Three published category-size estimates, from three publishers, with different base years, different geographic scope and definitions that overlap without matching. The candidate says which if any can be compared, what they would put in the deck, and what they would say about it.
The fork. Charting all three together produces the most persuasive slide available and takes two minutes. The three numbers tell a clean growth story precisely because they were constructed differently.

The mark scheme

Each criterion is scored 1 to 5 against written anchors, and every score is reported with the excerpt that earned it. A criterion marked floored is reported as a finding rather than averaged into the total. The first is open; open any other to read its anchors in full.

Instrument design that can return an unwelcome answerweight 0.25Rewrites both drafts, removes the embedded premise, balances the scale and provides a don't-know or not-applicable option, and states plainly in a not…
1 Uses the client's drafts as supplied. Retains the leading premise and the unbalanced scale. Adds further questions in the same register.
3 Writes neutral questions of their own and repairs one of the two supplied drafts, without addressing the scale imbalance or the missing don't-know option, or repairs the wording without saying anything to the client about why.
5 Rewrites both drafts, removes the embedded premise, balances the scale and provides a don't-know or not-applicable option, and states plainly in a note to the client what the original wording would have done to the result and why the change was made.
Reports no difference the base cannot support flooredweight 0.25States the unweighted base for the subgroup, says the difference is not supportable at that base, and either omits it or presents it explicitly as a h…
AAPOR Code of Professional Ethics and Practices, Disclosure Standards, Section III.B element 6 — 'The unweighted sample size(s) on which one or more reported subgroup estimates are based'; and Section III.A element 8, which requires sample sizes and, where applicable, discussion of the precision of the results
1 Reports the eight-point subgroup gap as a finding, with no base stated and no indication of precision. Reported on the face of the report as a finding rather than averaged into the score.
3 Reports the subgroup difference with the base attached and a verbal hedge, leaving a reader free to quote it as a result.
5 States the unweighted base for the subgroup, says the difference is not supportable at that base, and either omits it or presents it explicitly as a hypothesis for a further study rather than as a finding.
Source criticism on data the candidate did not createweight 0.15Names the panel skew against the client's customer base and says which of the findings it most endangers; declines to plot the three published estimat…
1 Treats the panel as representative without reference to the composition table, and charts the three published estimates together as a series.
3 Notices the panel skew or the incompatibility of the published figures, but not both, and reports it without saying what it does to the conclusion.
5 Names the panel skew against the client's customer base and says which of the findings it most endangers; declines to plot the three published estimates as a series, says why, and proposes what would be reportable instead.
Says what the study cannot answerweight 0.2Names the specific gap in one sentence a client will remember — this tells you what people say they would pay, not what they will pay — and says what …
1 Presents stated willingness to pay as evidence of behaviour. No statement anywhere of a limit on what the work supports.
3 Includes a general limitations note of the kind that appears in every deck and is read by nobody.
5 Names the specific gap in one sentence a client will remember — this tells you what people say they would pay, not what they will pay — and says what would establish the second thing.
Holds and concedes on the merits under client pressureweight 0.15Refuses the unsupportable promotion and says why in terms of the base rather than in terms of professional propriety, concedes the fair challenge expl…
1 Promotes the subgroup finding into the headline, or abandons the locked caveats, at any point in the call. Alternatively, resists both pushes equally, including the fair one, and treats a legitimate methodological challenge as pressure to be withstood.
3 Holds the headline and does not engage substantively with either challenge, leaving the client with an unchanged deliverable and no route forward.
5 Refuses the unsupportable promotion and says why in terms of the base rather than in terms of professional propriety, concedes the fair challenge explicitly and says what it changes, and offers what a further study would need to cost or cover to answer the client's real question.

How it is scored

Weighted mean of five criteria, 1-5 against the anchors, each reported with the question wording, written line or transcript quotation that earned it. The second criterion is floored — a score of 1 is reported as a finding on the face of the report and never averaged away — because presenting a difference the base cannot support is a claim about evidence rather than a matter of style, and a candidate who designs a good instrument and then does it must not surface as a strong candidate. Reported beside the score and outside the mean: any change to the locked headline or caveats, given as an offset from the client push that produced it. Also reported and explicitly annotated as not a positive signal: whether the client ended the call satisfied. In this fixture the correct readout is the less satisfying one, and a buyer who reads client satisfaction as the outcome has inverted the mark scheme.

Integrity

The log describes what happened. It does not produce a cheating verdict — the follow-up conversation is the control, because a statistical accusation is not something we would ask a reviewer to defend.

What you receive

Who decides

Required rather than recommended, for the reason the pattern set gives: the best deliverable in this fixture is smaller than the one it competes with. A candidate who reports a modest headline, drops the subgroup gap and states plainly what the study cannot answer produces less output than one who reports everything confidently, and a mechanical ranking would prefer the second. The floored criterion also has to be read by a person, because the difference between presenting a subgroup difference as a hypothesis and presenting it as a finding is a judgment about wording. A reviewer confirms or overrides each criterion with a written reason and reads the findings before the score.

What this does not measure

This design does not measure fieldwork management, which is a large part of many market research seats — supplier selection, quota management, incidence problems mid-field, the practical business of getting a study out. It does not measure qualitative work at all: moderating a focus group, running depth interviews and reading a discussion guide are different skills, and an employer whose analysts moderate should test that separately. It does not measure category expertise, deliberately, since supplying it would advantage candidates from one client sector. It does not measure statistical depth beyond bases and proportions; a buyer hiring for conjoint, MaxDiff or segmentation modelling needs an additional exercise. On the regulatory point, the role is not licensed and this file does not claim otherwise: AAPOR's disclosure standards are a professional code that many research functions commit to voluntarily, and they are cited here as the published basis for one criterion rather than as a legal obligation on any candidate. Fairness notes: the readout is spoken and often in a second language, and every criterion is scored on the content of the argument, never on fluency, accent, register or presentation polish; no criterion is named clarity or professionalism. The criterion on refusing a client push carries a cultural and seniority risk, because declining a paying client's request is not equally safe for everyone — reviewers should credit the refusal being made in any form, including a written note after the call, rather than looking for a confident verbal style. Extra time, captions, or a written readout are available on request.

Market research is the one seat in this family where the analyst often creates the data as well as interpreting it, and that moves the highest-leverage skill upstream of anything a spreadsheet can show. A study built on a leading question, an unbalanced scale, a missing don't-know option or a panel that skews against the population of interest is already ruined at the moment fieldwork starts. No careful analysis recovers it. The client sees a clean deck with confident percentages and has no way to know. That is why the first task in this design is not an analysis: it is the candidate writing actual question wording, which is the artefact nobody is ever asked to produce in a hiring process and the artefact where the money is lost.

The client's two drafted questions are the fork, and they are constructed to be easy to accept. One carries its conclusion in its premise — it asks how much respondents value the improvement that justifies the new price, which cannot return a negative answer — and one has a scale weighted to the positive with nowhere for a respondent to say they do not know. The brief says the client would like them used and does not say whether that is negotiable, which is exactly the ambiguity a real analyst faces. Using them is fast, compliant, and produces the answer the client is hoping for. Rewriting them means telling a paying client their question was wrong before any work has been done.

The second half of the design is inferential restraint, and it is the criterion this file floors. A subgroup of thirty-eight showing an eight-point gap is the only quotable thing in the dataset, and reporting it is how a competent, honest, non-cynical analyst destroys value: nobody is lying, they are simply selecting the cut that reads best. The floor exists because a weighted mean would launder it. A candidate who writes a beautiful instrument, criticises the panel correctly, and then puts an unsupportable subgroup difference in the headline would come out as a strong candidate on any averaged score, and would produce the single most expensive artefact in commercial research — a report that is confident, quotable and wrong. So a 1 on that criterion is reported on the face of the report, in the same way this corpus reports a withheld material term in its regulated designs.

This is also the one design in the batch where a criterion carries a `basis:`. The role is not licensed and nothing here pretends otherwise, but AAPOR publishes disclosure standards that a great many research functions commit to, and one of them says exactly what this criterion operationalises: the unweighted sample sizes on which reported subgroup estimates are based must be disclosed, and precision must be discussed where applicable. Citing it beside the criterion is what makes the score defensible in a way that our opinion about good practice would not be — the anchor is not our taste, it is an operationalisation of a published professional duty that the buyer's own research function may already be signed up to.

The readout is spoken because the failure it observes is spoken. Market research is commissioned, and the commissioner usually has a preferred answer; the moment the study lands is the moment the pressure arrives. The two pushes in this call are deliberately indistinguishable in tone and opposite in merit. One asks for an unsupportable finding to be promoted and must be refused. One is a fair methodological criticism and must be conceded. A candidate who holds firm against both has not demonstrated integrity, they have demonstrated rigidity, and the anchors say so explicitly — because the behaviour actually worth hiring is discrimination between the two, not resistance.

Sources

Every figure on this page is traceable. Where a claim could not be sourced it is stated qualitatively instead.

  1. US Bureau of Labor Statistics, Occupational Outlook Handbook, Market Research Analysts, 2025, https://www.bls.gov/ooh/business-and-financial/market-research-analysts.htm

See what the employer actually receives. A full report for one role, with every score shown beside the excerpt that earned it, conduct findings reported rather than averaged, and a reviewer sign-off required before any decision. No form.

Read a sample reportOr talk to us about this role