Taking an assessment rather than buying one? This page is written for employers. Here is the page for candidates.
Sales and revenue · Lead level
How to assess a Sales Manager
This is the one role in the hub where the status quo has been measured in published research and found wanting: studying sales workers across 214 firms, Benson, Li and Shue find firms prioritise current sales performance in promotion decisions at the expense of observable characteristics that better predict managerial performance. The selection signal (personal attainment) and the job (making eight other people better) are different, and an interview cannot separate them because a strong seller interviews well about selling. Concretely, nobody watches the candidate inspect a deal — the daily core of the role, in which a rep insists a deal closes this month and the manager has six minutes to find out whether that is true without the rep going defensive. Nor does anyone watch them coach: hand a candidate a mediocre call recording and most produce a list of ten faults, which is exactly the response that makes coaching useless.
The first-line sales manager's week is not selling. It is a pipeline review on Monday, six one-to-ones, a forecast call where they have to defend a number upward, two deals they are pulled into, an escalation, and — if they are any good — three hours of listening to their reps' calls. The output is other people's numbers, produced through conversations in which the manager has authority but very little direct control. Everything they can actually change happens in the gap between what a rep believes about a deal and what is true.
What separates the top quartile is inspection and coaching, and both are specific, learnable behaviours that look nothing like selling. Inspection is the ability to ask about a deal in a way that surfaces the missing evidence rather than triggering a defence: not "is it going to close?" but "who else has to sign and when did you last speak to them?" A median manager accepts the confident answer, and their forecast is materially wrong every quarter, which destroys their standing with their own leadership more thoroughly than a missed number does. Coaching is the ability to select one thing. Nearly every new manager reviews a call and returns a list, and a list changes nothing; the manager who says "in the next five calls, do only this" produces measurable movement in a fortnight.
The third thing, and the one that most often ends a first-line manager's tenure, is the willingness to be direct with a person they like. Underperformance in a sales team is visible in a dashboard within a month and is routinely allowed to run for two quarters because the conversation is unpleasant. A manager who can hold that conversation early, specifically, and without theatre keeps the team's respect; a manager who cannot ends up either carrying the underperformer or firing them abruptly, and both destroy the team's trust in the number.
What the hiring manager is really trying to predict is whether this person can make a mid-tier rep better, and whether their forecast can be believed. Neither is contained in the candidate's own attainment history, which is the primary thing every current process measures. It is also worth naming what should not be measured here: seniority in this seat attracts proxies — years of tenure, industry pedigree, the size of the logo they last carried — that stand in for managerial ability without predicting it, and a mark scheme that reproduces them is simply the CV screen wearing a different hat.
What the job actually needs
- deal inspection that separates a real forecast from a hopeful one
- coaching from evidence rather than anecdote
- delivering unwelcome performance feedback without losing the person
- honest forecasting upward
- territory and quota arithmetic
How people fail in this seat
- takes the deal over instead of coaching the rep through it
- forecasts optimistically to protect the team and burns their own credibility
- coaches by retelling their own selling stories
- avoids performance-managing the likeable underperformer
- spends their time with the top rep who needs them least
What most employers do instead
Internally, promotion of the highest-attaining rep. Externally, a CV of team numbers, a competency interview, and a "first 90 days" presentation to the leadership team.
The assessment
About 50 minutes end to end.
The systems it runs in
Three surfaces, and the seat is the movement between them. A conversation intelligence library holds the recorded discovery call in t1, so the candidate coaches from a timestamped transcript with the moments they cite linked back to it rather than from an impression of a call — the single item they choose to coach has to point at evidence someone else can open. A forecast and pipeline inspection view holds the rep's deals: stage, amount, close date, forecast category and last activity date, which is where t3 happens, because the three-week-old last activity and the absent approval process are visible on the record before the rep says a word about them. Salesforce documents the mapping between opportunity stage and forecast category that makes the commit view mean something, and the fork in t3 is whether the candidate changes it against the rep's protest. The performance note in t4 is written into the coaching or one-to-one record attached to the rep — the place a second manager would look, and the reason the note has to say plainly whether this was coaching or the first step of a performance process.
- Gong
- Chorus by ZoomInfo
- Clari
- Salesforce Sales Cloud
- HubSpot Sales Hub
- Microsoft Dynamics 365 Sales
Any call recording library with transcripts, any pipeline view with a commit or forecast flag the manager can change, and any durable record of a one-to-one. Teams without conversation intelligence run deal reviews off notes and a spreadsheet, and the fixture supports that with the transcript supplied directly; what it will not do is score coaching from a summary of a call, because the whole distinction the design draws is between coaching from evidence and coaching from anecdote, and removing the evidence removes the distinction. Rebuilt against the buyer's own forecast categories and their own performance-management language where they can provide a sandbox — the second of those matters, because whether a note reads as coaching or as a first written warning is a question about the buyer's own process.
What the candidate actually does
| Task | What happens |
|---|---|
| One thing to coach judgment_scenario · 12 min | A transcript of a rep's discovery call, mediocre in the ordinary way. The rep pitched over the buyer twice, never asked who else has to approve, accepted a stated requirement without testing it, and offered a discount unprompted near the end. There are at least eight identifiable faults. The candidate may submit exactly one thing they will coach, the evidence from the transcript that supports it, what they will ask the rep to do differently in the next five calls, and how they will know in a fortnight whether it worked. The single-item constraint is the fork. The reflex is a list, and a list changes nothing. |
| The coaching conversation live_call · 16 min | The AI rep is the difficult kind — likeable, agreeable, and hard to reach. They accept every point warmly and convert each one into a reason it was not really about them, mostly that the lead was poor. Two things are hidden. They are anxious about their number and read any feedback as the opening of a performance process, which surfaces as deflection rather than as a stated fear. And they will accept an offer to have the manager join the next call and handle it, gratefully, if one is made. Three forks. Coaching by retelling the candidate's own selling stories against getting the rep to hear the evidence in their own call. Taking the deal over against leaving the rep owning it. And softening until nothing specific is agreed against leaving the conversation with one behaviour, a date and a way to check. |
| The deal inspection live_call · 10 min | The same rep, on a different footing. They insist a deal closes this month. Hidden facts, all of which they will confirm if asked the right way and none of which they volunteer — the only contact is a manager who has said they are keen but has never mentioned an approval process, there has been no conversation about procurement or legal, and the last substantive contact was three weeks ago. The candidate has ten minutes. The fork is between questions that invite a confident answer and questions that surface the missing evidence, and between accepting the rep's confidence and changing the forecast against the rep's protest. The forecast category on that deal is the candidate's to change, so whether they changed it is a logged field event rather than a claim made in a debrief. |
| The performance note written_artifact · 12 min | The written record of both conversations, entered into the coaching record attached to the rep where a second manager would look for it, addressed to the rep and copied to the candidate's own manager. It must state what was observed with the evidence, what was agreed, by when, how it will be measured, and — the fork — whether this is coaching or the first step of a performance process, said plainly enough that the rep is not surprised later and a second manager could pick it up. The cheap version is encouraging and unfalsifiable. |
The mark scheme
Each criterion is scored 1 to 5 against written anchors, and every score is reported with the excerpt that earned it. A criterion marked floored is reported as a finding rather than averaged into the total. The first is open; open any other to read its anchors in full.
The coaching conversation leaves the rep owning one changeweight 0.3The rep articulates the problem in their own words before the candidate has finished making it, one specific behaviour is agreed for a named set of up…
Selects one thing to coach and evidences itweight 0.2Picks one high-leverage behaviour, quotes the turn in the transcript where it happened, states the specific alternative action for the next five calls…
Inspection surfaces the missing evidenceweight 0.2Establishes that nobody who can approve has been spoken to and that there has been no contact for three weeks, does it without the rep becoming defens…
Does not take the deal overweight 0.1Keeps the deal with the rep, and any involvement is defined as observing a named behaviour with the rep leading.
The written note is specific, dated and unsurprisingweight 0.15States the evidence, the single agreed change, the date, the measure, and says explicitly which kind of conversation this was — in terms the rep could…
Coaches from the rep's evidence rather than their own historyweight 0.05Every point is anchored to something the rep actually said or did not say in their own call.
How it is scored
Weighted mean of the six criteria, 1-5 against the anchors, each attached to the excerpt that earned it. The weighting is deliberately lopsided and the reason is worth stating on the result. Sixty-five percent of the mark sits on coaching behaviour — selecting one thing, running the conversation, not taking over, and using the rep's evidence rather than the candidate's own — because the output of this job is other people's output, and the published research on promotion in sales organisations is that firms over-weight personal selling performance at the expense of characteristics that better predict managerial performance. A design that split its weight evenly between selling-adjacent and coaching behaviours would reproduce that error with a rubric attached. The candidate's own selling ability is not scored anywhere in this design.
Integrity
- the rep's anxiety, the missing approval process and the three-week gap are absent from the brief and surface only under specific prompts
- the call transcript in the first task varies per candidate, with the eight faults distributed differently
- the coaching conversation runs against whichever single fault the candidate selected, so a shared answer to the first task does not transfer
- keystroke and paste timing on the performance note
- a follow-up question asking which turn of the rep's call the candidate would quote back to them and why
The log describes what happened. It does not produce a cheating verdict — the follow-up conversation is the control, because a statistical accusation is not something we would ask a reviewer to defend.
What you receive
- the selected coaching point with the candidate's cited transcript evidence
- full transcripts of both live conversations, with any offer to take the deal over marked
- the performance note as written
- per-criterion score with the excerpt that earned it
Who decides
Required rather than recommended, for a specific reason. This assessment scores a candidate on how they handle a subordinate, and the norms for that vary more between companies than anything else in this corpus — directness that is correct at one employer is a cultural failure at another. A reviewer who manages sellers at the deploying company reads both transcripts and the note in full for every shortlisted candidate. Three checks in particular. Whether the rep's deflection about lead quality was in fact partly true in the scenario as run, which changes what a good response looks like. Whether a candidate who scored low on inspection did so because they had already decided to spend the time on coaching, which is a defensible allocation and not a failure. And whether the performance note's framing — coaching or performance process — matches how the deploying company actually escalates, since the anchors describe a default. Overrides are written with a reason and stored with the score. No candidate is rejected on the composite alone, and the composite must never be compared against the candidate's own historic attainment, which this design deliberately does not measure.
What this does not measure
Nothing here scores accent, dialect, fluency, pace or manner of speech, and nothing scores authority, command, or how naturally the candidate sounds like a manager. That last exclusion needs emphasis in this seat above all others, because "leadership presence" and "executive presence" are the standard descriptors used in management hiring and they are among the most reliable known routes for accent, gender, age and class to enter a mark scheme with a professional justification attached. They are absent here and must not be added. Every anchor above points at an event — which fault was selected, what the rep said back, whether a date was set, what the note contains. The design also excludes, deliberately, the thing the status quo measures most: the candidate's own selling ability and attainment history. No task asks the candidate to sell, and no criterion rewards it. Nor is tenure, industry pedigree or the size of a previously carried logo visible to the assessment at all. The written note is judged on whether the rep and a second manager could act on it, not on idiom or register, and non-native phrasing is not a deduction. What fifty minutes cannot see is management. A first-line sales manager's performance is a team's numbers over two to four quarters, produced through hundreds of small conversations, and it depends heavily on things this design cannot stage — whether the manager keeps coaching in a bad quarter, how they allocate time across eight reps rather than one, whether they can recruit, and whether their forecast holds up when their own manager is applying pressure. This session observes one coaching conversation, one inspection, and the record that follows. Deployers should back-test scores against the change in attainment of the manager's mid-tier reps and against forecast accuracy at two and four quarters, and should retire any criterion that does not predict either. It should not be used as the sole basis for an internal promotion decision, since the whole argument for the design is that a single measure should not be. The live-call format may disadvantage candidates with speech, hearing or anxiety-related disabilities. Text-based runs of both conversations, with the same hidden facts and release conditions, must be available on request, scored on the same anchors and unmarked on the result. Extra time on the transcript review and the note must be available without disclosure.
This is the one design in the family where the candidate never sells anything, and the omission is the argument. The role page cites research on promotion in sales organisations finding that firms prioritise current sales performance at the expense of observable characteristics that better predict managerial performance. If that is true — and it is the best-evidenced claim on any page in this corpus — then an assessment for this seat has exactly one job, which is to observe the managerial behaviours directly instead of inferring them from a selling proxy. Every design decision below follows from that, including the lopsided weighting, which puts nearly two-thirds of the mark on coaching and puts none at all on the candidate's ability to run a buyer conversation.
The output of a sales manager is other people's output, so the monitored simulation has to put another person in front of them. The AI rep is deliberately the hardest kind to coach, which is not the hostile one. It is the likeable one. They agree with everything, they are warm, they convert each point into a reason it was not really about them, and they have an excuse — the lead was poor — that is plausible enough to be difficult to dismiss and comfortable enough for both parties to settle on. The role page identifies willingness to be direct with a person you like as the thing that most often ends a first-line manager's tenure, and this rep is built to make agreeableness the path of least resistance throughout the conversation. A candidate can have a pleasant sixteen minutes, leave with mutual warmth and nothing agreed, and score a 1.
The single-item constraint on the first task is the design's cleanest instrument. The role page observes that handing a candidate a mediocre call recording produces, from most of them, a list of ten faults — which is exactly the response that makes coaching useless, because a rep cannot change ten things and will change none. The transcript here contains at least eight real faults, and the candidate may name one. Watching which one they choose is more informative than watching them find all eight: the discount offered unprompted and the missing approval question are high-leverage, and the interruptions are real but minor. The coaching conversation then runs against whichever fault they selected, so a poor selection is not merely marked down in the abstract — the candidate spends sixteen minutes coaching the wrong thing, which is what the error costs on the job.
The inspection is short on purpose. Ten minutes is roughly the real budget for this conversation in a pipeline review, and the whole skill is compressed into question design: the role page's formulation is that inspection means asking about a deal in a way that surfaces missing evidence rather than triggering a defence, which is the difference between "is it going to close?" and "who else has to sign, and when did you last speak to them?" The scenario is loaded so the answers exist and are findable, and the deal is not closing. What separates the top anchor is not discovering that, which most competent candidates manage. It is changing the forecast position while the rep is still protesting, and telling them precisely what would move it back.
The performance note carries the last of the weight because a coaching conversation with no record is a conversation that did not happen. The specific fork built into it — say plainly whether this is coaching or the first step of a performance process — targets the failure the role page names as routine: underperformance visible in a dashboard within a month, allowed to run for two quarters, and then ended abruptly by a manager who never wrote down what they had already said. A note that a rep could read without surprise, and that a second manager could pick up cold, is a small artefact and an unusually honest one.
Sources
Every figure on this page is traceable. Where a claim could not be sourced it is stated qualitatively instead.
- US Bureau of Labor Statistics, Occupational Outlook Handbook, Sales Managers, 2025 (650,100 jobs; 47,300 annual openings; 4 percent growth; median pay $148,270; work experience in sales typically required), https://www.bls.gov/ooh/management/sales-managers.htm
- Alan Benson, Danielle Li and Kelly Shue, Promotions and the Peter Principle, NBER Working Paper 24343 (published in Quarterly Journal of Economics 134(4), 2019); microdata on sales workers at 214 firms, https://www.nber.org/papers/w24343
See what the employer actually receives. A full report for one role, with every score shown beside the excerpt that earned it, conduct findings reported rather than averaged, and a reviewer sign-off required before any decision. No form.
Read a sample reportOr talk to us about this role