Get started

Taking an assessment rather than buying one? This page is written for employers. Here is the page for candidates.

Outbound sales development · Entry level

How to assess a Appointment Setter

The trial day is scored on the exact number the role's failure mode inflates. Appointments set is what the setter is paid on and what the trial measures, and the entire risk of the seat is that this number can be raised simply by lowering the bar — with the cost landing on a closer's calendar days later, where nobody traces it back to hiring. A process that ranks candidates on sets is therefore selecting for the failure mode directly. The script read fails for a different reason: fluency is not the scarce skill, because the script is supplied and the setter is usually forbidden to improvise on product. The job turns on the moment the prospect asks a question the setter is not permitted to answer — "how much is it?" — where the choice is between an invented number, a stonewall, and a clean bridge back to the appointment. Nobody currently observes that moment, and show rate, the only metric that matters commercially, materialises long after the hiring decision.

An appointment setter is not a junior SDR, and the distinction is worth being precise about because it changes what should be assessed. An SDR owns a qualification judgment and hands over a briefed opportunity. A setter owns a calendar slot. They are usually working a consumer or very small business list — home services, insurance and financial appointments, high-ticket coaching and training, medical and dental, solar, trades — and they are almost always explicitly prohibited from discussing price, terms or product depth, because those are the closer's to control. The work is dialling, a short opener, three or four checklist questions, and a booking.

That prohibition is the whole shape of the role, and it produces a skill that appears nowhere else in the hub. A prospect who is warming up will ask a substantive question within the first minute, and the setter has to decline to answer it in a way that increases rather than reduces trust. Three responses are possible and only one is acceptable: guess (and burn the closer, who now opens by contradicting their own colleague), refuse flatly (and lose the prospect), or bridge — acknowledge the question, say honestly why the specialist rather than they will answer it, and convert the interest into a time. The median candidate guesses. The top quartile bridges without a pause, and they do it four or five times a day.

The second separator is the quality of the commitment obtained. "Yeah, Thursday sounds fine" and "Thursday at 2, and I've just sent the invite to your email — can you confirm you've got it, and will Sam be there too?" produce very different show rates, and show rate is the number the business actually runs on. A setter with a high set rate and a poor show rate is costing money; one with fewer sets that nearly all happen is the reason the closing team hits its number. The confirmation message is a real artefact here and worth reading: it is short, it is written under time pressure between calls, and it either contains the specifics that make an appointment stick or it does not.

What the hiring manager is really predicting is show rate and closer-side disqualification rate — two numbers that appear a week or more after the call and that no current screening step approximates. The setter seat also carries the highest per-hire volume and the shortest tenure in this family, which puts a hard practical ceiling on assessment length: anything that cannot be run on every applicant will not be run at all.

What the job actually needs

How people fail in this seat

What most employers do instead

A recorded script read or voice audition, a typing and CRM-entry test, and frequently a paid trial day or first week scored on appointments set.

The trial day is scored on the exact number the role's failure mode inflates. Appointments set is what the setter is paid on and what the trial measures, and the entire risk of the seat is that this number can be raised simply by lowering the bar — with the cost landing on a closer's calendar days later, where nobody traces it back to hiring. A process that ranks candidates on sets is therefore selecting for the failure mode directly. The script read fails for a different reason: fluency is not the scarce skill, because the script is supplied and the setter is usually forbidden to improvise on product. The job turns on the moment the prospect asks a question the setter is not permitted to answer — "how much is it?" — where the choice is between an invented number, a stonewall, and a clean bridge back to the appointment. Nobody currently observes that moment, and show rate, the only metric that matters commercially, materialises long after the hiring decision.

The assessment

About 16 minutes end to end.

The systems it runs in

Three surfaces the setter moves between all day: a dialer working a call list, a shared booking calendar showing the closers' available slots, and the CRM contact record where the outcome of the call is recorded. The checklist questions in the brief are fields on that record. At the end of t1 the candidate sets a call disposition, books the slot against a named closer, and records who will be present — which is where the passing mention of a partner either lands or does not. t2 is composed and sent as the appointment confirmation attached to that booking, so a confirmation that contradicts the slot the candidate actually booked is visible as a contradiction rather than only as prose. t3's handover line is written into the appointment record itself, in the field the closer reads before dialling, under whatever character limit the buyer's own system imposes.

Any dialer plus any calendar plus any contact record carrying a disposition or outcome code and an appointment object. Consumer and small-business setting desks run on lighter stacks than enterprise sales, and the fixture assumes the lighter one; a buyer on an all-in-one platform where the dialer, calendar and record are one screen gets the same exercise with fewer windows. It is rebuilt against the buyer's own disposition codes and appointment fields where they can provide a sandbox, because the disposition list is the one part of this stack that is genuinely local.

Working speed is scored. Setters are paid per booking and write their confirmations in the gap before the next dial, so pace is part of the seat rather than an artefact of the test — t2 is already capped at four timed minutes for that reason. Time to a completed record is reported and scored only against the booking-quality criteria. A slot booked in three minutes with the second decision-maker unasked scores below a slower booking that names them, and a confirmation written fast that contains a price the setter was not permitted to give fails outright regardless of how quickly it was produced. Speed never stands alone, and the candidate is told this before they dial.

What the candidate actually does

TaskWhat happens
The question you are not allowed to answer
live_call · 6 min
The candidate gets a five-line brief: who they are calling for, four checklist questions, the calendar slots available, and one explicit prohibition — they may not discuss price, terms or product detail, because that is the specialist's job. The AI prospect is a consumer on a list, mildly irritated at first, who thaws within a minute and then does the thing every real prospect does: asks what it costs. It presses twice. On the second press it offers the trade openly — "just give me a rough idea and I'll book something in" — which is a booking available immediately in exchange for one invented number. A third of the way through the call the prospect mentions a partner who handles this sort of decision, once, in passing, and never mentions them again. At the end the prospect gives a soft yes: "yeah, sometime next week works." Accepting that sentence as an appointment ends the call on target.
The confirmation written between calls
written_artifact · 4 min
Four minutes, timed, to write the confirmation the prospect receives — the real constraint of the seat, where this message is composed in the gap before the next dial. Scored on the specifics that make an appointment stick, and failed outright if it contains a price or product claim the setter was not permitted to make. This is the artefact that determines show rate, and it is the one artefact of the role that nobody currently reads before hiring. It is composed as the confirmation attached to the booking the candidate actually made, so a time, a closer or an attendee that disagrees with the appointment record shows up as a contradiction rather than only as prose.
Marking your own sets
judgment_scenario · 4 min
Three appointments notionally set by the candidate earlier that day, each described in four or five facts. One will show. One will probably not — the slot clashes with something the prospect mentioned. One should never have been set, because the person cannot make the decision. The candidate marks each, names the fact behind the mark, and writes the one- line handover note into the appointment record itself, in the field the closer reads before dialling. Pay is per booking; the honest mark reduces today's count.

The mark scheme

Each criterion is scored 1 to 5 against written anchors, and every score is reported with the excerpt that earned it. A criterion marked floored is reported as a finding rather than averaged into the total. The first is open; open any other to read its anchors in full.

Bridges the out-of-scope question rather than guessing or stonewallingweight 0.3Treats the question as legitimate, says plainly why the specialist rather than they will answer it, gives the one fact they are permitted to give, con…
1 Invents a price, a range, or a product claim the brief does not contain — including hedged forms such as "most people pay around".
3 Declines to answer and moves on, without acknowledging that the question was reasonable or explaining why they cannot answer it.
5 Treats the question as legitimate, says plainly why the specialist rather than they will answer it, gives the one fact they are permitted to give, converts the interest into a time, and repeats this without irritation or evasion when pressed a second time.
Quality of the commitment obtainedweight 0.3Specific day and time, the second decision-maker mentioned earlier is named and invited, the channel and expected duration are stated, and the prospec…
1 Records "sometime next week" or an equivalent soft yes as a booked appointment.
3 Secures a specific day and time and nothing further.
5 Specific day and time, the second decision-maker mentioned earlier is named and invited, the channel and expected duration are stated, and the prospect is asked to confirm they have received the invitation.
The confirmation messageweight 0.2Time with timezone, who attends on both sides, what the call will and will not cover, one sentence in the prospect's own terms about why they agreed, …
1 Absent, or contains a figure or product claim outside the setter's remit, which the closer will have to contradict.
3 States the time and how to join.
5 Time with timezone, who attends on both sides, what the call will and will not cover, one sentence in the prospect's own terms about why they agreed, and how to move it if they need to.
Honest self-audit and handoverweight 0.2Distinguishes the set that will not show from the set that should not have been made, names the deciding fact in each, and the handover note puts the …
1 Marks all three as likely to show, or writes handover notes that omit the known risk.
3 Identifies the weakest of the three.
5 Distinguishes the set that will not show from the set that should not have been made, names the deciding fact in each, and the handover note puts the risk in front of the closer rather than burying it.

How it is scored

Weighted mean of the four criteria, 1-5 against the anchors, each score reported with its excerpt. Two absolute flags sit outside the average and are reported by name because averaging hides them: any invented price or product claim in t1 or t2, and any soft yes recorded as a booking. Both are cheap to detect in a transcript and both are the specific way this seat loses money.

Integrity

The log describes what happened. It does not produce a cheating verdict — the follow-up conversation is the control, because a statistical accusation is not something we would ask a reviewer to defend.

What you receive

Who decides

Required before rejection and recommended before hire, but scoped tightly so it survives volume: the reviewer reads the confirmation message and the ninety seconds of transcript around the second price press, which the system marks. That is under three minutes per candidate. Reviewers confirm or override the bridging score with a written reason. Because this seat is usually hired in cohorts, review the whole cohort's flags together before any individual decision — an invented price appearing in most transcripts usually means the brief did not make the prohibition clear, which is a design fault, not thirty candidate faults.

What this does not measure

Nothing here scores accent, dialect, first language, speech rate, disfluency or perceived warmth, and there is no rapport, energy or telephone-manner criterion, which in setter hiring is the usual disguise for exactly those things. Every anchor is decidable from the transcript text. This matters more in this seat than almost anywhere else, because voice auditions are the incumbent screening method and they select on speech characteristics almost exclusively. Two specific risks to monitor. First, written English proficiency carries twenty percent of the score through t2, and the four-minute timer is deliberately tight; deployers whose setters write in a second language should extend the timer and score content only, never grammar or spelling, and should never score t2 with a general writing-quality model. Second, the honest self-audit in t3 rewards a behaviour that some sales cultures actively train out; tell candidates in advance that accuracy is scored above volume, because the fork should be a real temptation, not a trick. Offer a text or chat variant of t1 on request for candidates with speech, hearing or anxiety-related conditions, and publish how to ask for it before the assessment opens. Track pass rates per criterion; the bridging criterion is where impact will show first if it shows at all.

The appointment setter's assessment is deliberately the shortest in the sales hub, because the seat has the highest per-hire volume and the shortest tenure of any role in it, and an assessment that cannot be run on every applicant is one that will be run on none. Sixteen minutes is six of call, four of writing, four of written judgment and two of brief. Every design decision after that was made under that ceiling.

The task set is narrower than the SDR's next door, and it should be. An SDR owns a qualification judgment against buyer criteria and hands over a briefed opportunity; a setter owns a calendar slot and is usually forbidden from discussing the very thing the prospect most wants to talk about. That prohibition produces the skill this design is mostly built to observe, and it appears nowhere else in the hub. Three responses to "how much is it?" are available. Guessing produces an immediate booking and a closer who opens their call by contradicting a colleague. Refusing flatly loses the prospect. Bridging — acknowledging the question, being honest about why the specialist will answer it, and turning the interest into a time — is the only one that works, and the role file's claim is that the median candidate guesses. The AI prospect is built to make guessing pay: it presses twice and then offers the trade explicitly, so the cheap path is not merely available, it is handed over. Nothing in a voice audition or a script read reaches that moment, because a script read has no prospect in it.

The second half of the design goes after show rate, which is the number the business actually runs on and the number that materialises a week after every hiring decision has been made. Show rate is manufactured in two places: the last forty seconds of the call, and the confirmation message. So the AI gives a soft yes at the end — the sentence that inflates set counts across the entire industry — and mentions a second decision-maker exactly once, early, where a candidate who is listening will catch it and a candidate running a script will not. Then t2 makes the candidate write the message that either makes the appointment stick or does not, under the four-minute pressure it is genuinely written under.

The third task exists because of a specific claim in the role file: the standard paid trial day scores candidates on appointments set, which is precisely the number the role's failure mode inflates. So a process that ranks on sets selects for the failure. Marking your own sets inverts that. The candidate is paid per booking, has three, and is asked to say which of them is worthless — with the answer written into the note a closer will read. A candidate who marks all three green is not lying so much as demonstrating the exact instinct that fills a closing team's calendar with people who do not turn up.

What this design does not measure is dialling volume, resilience across a full day, or CRM speed. All three are real parts of the seat, all three are managed and measured after hire, and none of them is what separates the setter who costs money from the one who makes it.

Sources

Every figure on this page is traceable. Where a claim could not be sourced it is stated qualitatively instead.

  1. US Bureau of Labor Statistics, Occupational Employment and Wage Statistics, National employment and wage data by occupation, May 2025: Telemarketers (41-9041), 58,430 jobs, median annual wage $37,370, https://www.bls.gov/news.release/ocwage.t01.htm
  2. US Bureau of Labor Statistics, 2018 Standard Occupational Classification Definitions (contains no 'appointment setter' occupation), https://www.bls.gov/soc/2018/soc_2018_definitions.pdf

See what the employer actually receives. A full report for one role, with every score shown beside the excerpt that earned it, conduct findings reported rather than averaged, and a reviewer sign-off required before any decision. No form.

Read a sample reportOr talk to us about this role