Taking an assessment rather than buying one? This page is written for employers. Here is the page for candidates.
Outbound sales development · Entry level
How to assess a Sales Development Representative
The hiring-manager role-play fails for a reason peculiar to this seat. The job is defined by talking to someone who does not want to talk to you, and the one person who cannot reproduce hostile indifference for four minutes is the manager who has already decided they like the candidate and is telegraphing the answer in their tone. Second, the role-play almost always ends when the candidate "books the meeting", which is scored as the win — while in production a meeting the AE disqualifies in three minutes is worse than no meeting, and is the single most expensive habit in the function because the cost moves downstream and is blamed on marketing. Third, with 16 percent of a median team leaving each year by promotion, the manager is settling two different bets in one interview: can this person survive cold calls now, and will they be the AE we promote in eighteen months.
An SDR's day is a funnel of attempts that mostly do not connect. The Bridge Group's 2025 cohort reports a median of 112 activities a day — 44 phone, 41 email, 19 LinkedIn, 8 text or other — converting into 4.1 quality conversations. Everything up to those four is throughput that tooling already measures and management already manages. The entire scarce skill of the role lives inside them, and specifically inside the twenty seconds after a stranger says "who is this?"
What separates the top quartile is not energy, and it is certainly not resilience described in the abstract. It is three narrow behaviours. The first is the response to the reflexive brush-off — "we're all set", "send me an email", "now's not a good time" — where the median candidate either apologises and hangs up or argues, and the strong one earns another sentence by saying something specific enough to be worth hearing. The second is qualification discipline. A quality conversation only counts if the prospect actually matches the criteria, and the whole weekly incentive structure pushes towards booking anyone who agrees. A rep who books eight qualified meetings is worth more than one who books fourteen of which five are junk, but only the second number is visible in week one, and the first is the one that determines whether the SDR gets promoted. The third is the follow-up. A note that references the two things the prospect actually said, in their words, is a different artefact from a templated recap, and it is the difference between a meeting that holds and one that quietly disappears.
What the hiring manager is really predicting is survival plus trajectory. This is a feeder seat by design: nearly a sixth of a median team leaves each year by promotion into an AE role, so the person being hired is simultaneously being evaluated as next year's closer. Those are genuinely different bets — plenty of people are excellent at opening and never learn to close, and some of the best closers are mediocre prospectors — and most processes conflate them because they have only one conversation in which to decide.
Two constraints belong on any advert for this seat. Score the objection handling, the qualification decision and the written follow-up; never the accent, dialect or manner of speech, which are proxies for protected characteristics and have no place in an outbound mark scheme. And resist scoring volume: activity counts are managed after hire and are a poor proxy for the four minutes that matter.
What the job actually needs
- surviving the first brush-off without collapsing or bulldozing
- earning the next thirty seconds by being specific
- qualifying honestly against criteria under pressure to book
- writing a follow-up that quotes the prospect
- consistency across a high-volume day
How people fail in this seat
- books an unqualified meeting to make the weekly number
- reads the next script line over the prospect's answer
- folds at the first no or argues past it
- sends a template with a merge field
- treats research as optional because activity is what is counted
What most employers do instead
A CV that carries almost no signal because the pool is early-career by design, a phone screen, a drive or personality questionnaire, and a final-round role-play run by the hiring manager.
The assessment
About 22 minutes end to end.
The systems it runs in
A CRM lead record worked from a sequencer step, which is the pairing this seat actually sits in: the step tells the candidate who to call and the lead record is where the call has to end up. The fixture opens on the lead — the company, the contact, the lead status, the activity history showing two previous unanswered attempts — and the brief's three ICP criteria are fields on that record which are blank until the candidate fills them. After the call the candidate logs the call outcome, updates the lead status, sets the next step with a date, and composes t2's follow-up as the reply step in the sequence rather than in a blank document.
- Salesforce Sales Cloud
- HubSpot Sales Hub
- Outreach
- Salesloft
- Apollo.io
- Pipedrive
Any CRM holding a lead or contact object with a status, an activity log and a next-step field, paired with any sequencer or cadence tool holding the step the call came from. The fixture needs those four things and nothing vendor-specific beyond them, so it is rebuilt against the buyer's own instance and their own ICP field names where they can provide a sandbox. A buyer whose SDRs work leads and a buyer whose SDRs work contacts get the same exercise with the object renamed.
What the candidate actually does
| Task | What happens |
|---|---|
| The cold open live_call · 7 min | The candidate is given a one-page brief thirty seconds before dialling — who they work for, what it does in two sentences, and three ICP criteria that make a meeting worth booking. The AI prospect picks up mid-task, unenthusiastic, and issues a reflexive brush-off inside the first fifteen seconds. It is configured not to warm up on tone; it warms up only on specificity, and it will admit one real, buried problem if the candidate earns a second exchange by saying something that could only be said to this prospect. The fork arrives around minute four. The prospect offers the meeting unprompted — "fine, put something in the diary, send me an invite" — while two of the three ICP criteria are still unasked, and the true answers to both are disqualifying: they do not own the budget line, and there is no timeline. The prospect will not volunteer either. Booking immediately ends the call successfully in under five minutes and looks like a win on every count the candidate can see. |
| The note that has to survive the week written_artifact · 6 min | Immediately after the call, with the transcript not available to them, the candidate writes the follow-up the prospect will actually receive. Scored on whether it quotes what the prospect said rather than what the script says, and on whether every commitment in it appears in the call. A note claiming an agreement the transcript does not contain is the written form of the same failure t1 tests. It is composed as the reply step in the sequence, and the lead record has to be closed out with it — call outcome, lead status, the ICP fields the candidate did or did not establish, and a dated next step. What the record says and what the note claims then become two artefacts read against one transcript. |
| Three conversations, one diary judgment_scenario · 6 min | Three short written summaries of conversations from the same notional day, against the same ICP criteria. One is a clear fit. One is a clear miss dressed up as interest. One can buy but not this quarter. For each, the candidate books, nurtures or disqualifies, names the single fact that decides it, and states what would have to change. Run in writing rather than as a fourth call because the judgment is separable from the pressure and costs a fifth as much to observe. |
The mark scheme
Each criterion is scored 1 to 5 against written anchors, and every score is reported with the excerpt that earned it. A criterion marked floored is reported as a finding rather than averaged into the total. The first is open; open any other to read its anchors in full.
Survives the brush-off without collapsing or bulldozingweight 0.2Within one sentence of the brush-off, says something specific enough that it could not have been said to a different company, earns a further exchange…
Asks the qualifying question whose answer could cost the meetingweight 0.3Asks the disqualifying question after the meeting has already been offered, states plainly what is missing, and either redirects to the person who own…
Listens rather than waits to talkweight 0.15At least two questions are visibly built on the prospect's own words, and one follows a hesitation rather than a stated fact.
Follow-up quotes the prospect and claims only what was agreedweight 0.2Quotes or closely paraphrases two things the prospect actually said, states the next step and who owns it, and omits everything not agreed.
Qualification judgment on the written casesweight 0.15Separates cannot-buy from not-now, names the one deciding fact in each case, and states what would have to be true to revisit it.
How it is scored
Weighted mean of the five criteria, each scored 1-5 against the anchors above, reported with the transcript excerpt that earned it. The fork in t1 is timestamped by the system, so every reviewer looks at the same forty seconds. A score of 1 on the qualification criterion is reported as a named risk rather than folded into the average, because it is the failure mode that costs the most downstream and it is invisible in every other step of a normal process.
Integrity
- live unscripted branching so a memorised answer does not fit
- written artefact cross-checked against the candidate's own call transcript
- seeded ICP criteria rotated per candidate
- timing anomalies between call end and first keystroke
- one follow-up question about a specific moment in their own call
The log describes what happened. It does not produce a cheating verdict — the follow-up conversation is the control, because a statistical accusation is not something we would ask a reviewer to defend.
What you receive
- call transcript with timestamps
- audio recording
- follow-up email
- judgment answers with reasoning
- per-criterion score with excerpt
- marked fork timestamp and which branch the prospect took
Who decides
Required before any rejection. The system marks the fork timestamp; the reviewer listens to the two minutes either side of it and confirms or overrides the qualification score with a one-line written reason. Also sample ten percent of the passes, not only the failures — the characteristic error of an automated sales rubric is rewarding a fluent candidate who booked a junk meeting, and that error only surfaces if someone reads the passes. Reviewers should be told the anchors before they hear the audio, in that order, because hearing the call first anchors on delivery.
What this does not measure
This design does not measure accent, dialect, first language, speech rate, disfluency, voice pitch, vocabulary range or perceived confidence, and no criterion admits them by another name — there is deliberately no "energy", "presence" or "professionalism" criterion, which is where they normally re-enter an outbound mark scheme. Every anchor is written to be decidable from the text of the transcript alone. Deployers should use that: rescore a sample from transcript only and compare to the audio-based score, because systematic disagreement means delivery is leaking into the judgment. Candidates who are fluent but not native in the assessment language may take longer to build a specific sentence under time pressure, so do not score words per minute, filler words or dead air, and run the assessment in the language the job is actually done in. The live call is a reasonable-adjustment surface: publish how to request extra written time or a text-based variant of t1 before the assessment starts, not in an appeals process afterwards. Monitor pass rates by criterion rather than only overall, since adverse impact in sales rubrics concentrates in one or two criteria and averages out of a headline number.
The whole design is built around a single claim from the role file: the scarce skill of this seat lives inside roughly four real conversations a day, and every existing screening step observes either the throughput around them or a sanitised imitation of them. So the assessment spends its budget almost entirely inside one of those four, and spends the rest checking whether the candidate can tell a good one from a bad one afterwards.
The fork in t1 is the reason this design exists. A hiring-manager role-play ends when the candidate books the meeting, and books it because the manager wants them to. Here the prospect offers the meeting early, warmly, and for the wrong reasons — the fastest, most rewarding-looking path through the call is to take it and hang up four minutes ahead of schedule. Taking it is not a failure of nerve or knowledge; it is the exact production behaviour that fills an account executive's calendar with meetings that die in three minutes and get blamed on marketing. The candidate who asks the budget-ownership question *after* the booking has been offered, knowing the answer might remove it, is demonstrating the one thing the role file says separates the top quartile and the one thing nobody currently observes before hire. Because the prospect never volunteers the disqualifying facts, the score is not a matter of interpretation: either the question was asked or it was not, and the transcript settles it.
The second reason for the design is the pairing of t1 and t2. A follow-up note is cheap to fake in isolation and impossible to fake against a transcript the candidate cannot see. Grading the note for whether its commitments appear in the call catches the same over-claiming instinct in a different medium, and a candidate who is coached through the call by someone off-camera will usually produce a note that does not match it.
Cost was a first-class constraint, not an afterthought. This seat runs roughly eight requisitions a year per twenty-person team at the median attrition the role file cites, and an assessment that only survives on a shortlist will quietly stop being run. Twenty-two minutes is seven of live call, six of writing, six of written judgment and three of brief and setup. The judgment task exists precisely because it buys the qualification signal a second time at a fifth of the cost of a second call — and it separates a candidate who qualified well once under pressure from one who understands why.
What this does not attempt is the promotion bet. The role file is right that hiring managers settle two questions in one interview — survives the cold call now, becomes the AE in eighteen months — and this design answers only the first. Nothing in twenty-two minutes predicts whether someone learns to close. Deployers should say so out loud rather than let a strong score be read as both.
Sources
Every figure on this page is traceable. Where a claim could not be sourced it is stated qualitatively instead.
- The Bridge Group, SDR Models, Motions & Metrics 2025 Research Report (351 B2B companies; 78 percent North America, 83 percent B2B SaaS; median revenue $47M, median ASP $50K): 40 percent median attrition split 13/11/16 involuntary/voluntary/promotion, 1.9 year median tenure, 60 percent of reps at quota, median 112 daily activities producing 4.1 quality conversations, 3.0 month ramp, https://www.bridgegroupinc.com/research/2025-sdr-models-metrics-report-the-bridge-group
- US Bureau of Labor Statistics, 2018 Standard Occupational Classification Definitions (contains no 'sales development' or 'business development' occupation), https://www.bls.gov/soc/2018/soc_2018_definitions.pdf
See what the employer actually receives. A full report for one role, with every score shown beside the excerpt that earned it, conduct findings reported rather than averaged, and a reviewer sign-off required before any decision. No form.
Read a sample reportOr talk to us about this role