Taking an assessment rather than buying one? This page is written for employers. Here is the page for candidates.
Administrative and business support · Entry level
How to assess a Administrative Assistant
This is the role where the typing test is actually still administered, and it is the clearest instrument-to-job mismatch in the corpus. Words per minute stopped being a constraint on this work when the work stopped being transcription; what constrains it now is accuracy across a batch while being interrupted, and format discipline in documents that somebody senior will forward onward without checking. A multiple-choice question about where the mail-merge command lives measures recall of a menu. The interview is worse than useless rather than merely weak, because it actively selects the wrong way: it rewards candidates who narrate their organisation fluently, which is a presentation skill, over candidates who are organised, which is visible only in output. The two are correlated closely enough to be confused and loosely enough to produce a bad hire.
An administrative assistant supports a team rather than a person. That single difference from the executive assistant role changes almost everything about what good looks like. The EA has one principal and deep context; the administrative assistant has eleven people who all believe their request is the only one outstanding, no visibility into which of them is under real pressure, and no delegated authority to tell any of them to wait. The work arrives from every direction at once — a room booking, a purchase requisition, a document to format before a two o'clock, a visitor to be added to a system, a spreadsheet of attendance to reconcile, an expense claim missing receipts — and none of it individually is hard.
The difficulty is entirely in the aggregate. A hundred small tasks a week, each with a low individual failure cost and a compounding collective one, executed while being interrupted roughly every fifteen minutes. This is why "attention to detail" is the most-used and least-useful phrase in the job adverts for this role: nobody is bad at detail on one task in a quiet room. The question is whether accuracy survives the eleventh interruption, and whether the person has built the small mechanical habits — reading back the date, confirming the room as well as the time, checking that the attachment is the version that was asked for — that make accuracy independent of concentration.
The second thing that separates the top quartile is document discipline, which hiring managers consistently underrate because it looks like formatting. A large share of an administrative assistant's output is a document somebody more senior will send to a client, a board or a regulator without reading it closely. If the headers are inconsistent, the numbering restarts, the table breaks across a page badly or the template is last year's, the document comes back and is redone — usually by the senior person, at eight times the cost, and usually without feedback, so it happens again next month. The strong performer produces work that goes out unchanged. That is a measurable property of an artefact and it is never measured before hiring.
The third is the one that turns a good administrative assistant into someone a team cannot function without: noticing that a process is wrong and saying so. These roles inherit instructions from people who have not done the task themselves — a form that asks for information the requester already holds, a routing rule that sends approvals to someone who left, a recurring report nobody opens. A weak assistant executes it, correctly, forever. A strong one executes it and then writes two sentences to whoever owns it explaining what breaks and what they would change. Doing that requires a specific kind of confidence in a role that is structurally junior to everyone it serves, and it is far more predictive of value than any of the traits the interview is trying to elicit.
What a hiring manager is really trying to predict is failure rate at volume, and whether the small failures get caught by the assistant or by somebody else. That is why the useful assessment is not one careful task but a realistic batch: six or seven requests arriving together, two of which conflict, one of which is underspecified, and one of which contains an instruction that will produce a visibly wrong result if followed exactly. What comes back — what was completed, what was flagged, what was quietly corrected, and how it was written up — is a direct sample of the job, and it takes less time to grade than the interview it replaces.
What the job actually needs
- accuracy at volume under interruption
- document and format discipline
- scheduling across many diaries
- following a process and noticing when it is wrong
- plain written English
How people fail in this seat
- books the meeting without the room or the joining link
- produces a document that is right in content and wrong in format so it is redone by someone senior
- answers a question with data they did not verify
- works the queue in arrival order regardless of consequence
- follows a broken instruction silently
What most employers do instead
CV screen, a conversational interview about being organised and proactive, and a typing test or a Microsoft Office multiple-choice test.
The assessment
About 28 minutes end to end.
The systems it runs in
Microsoft 365, because that is where this seat's seven requests actually live: Outlook mail and shared calendars with free/busy visible, room and resource booking with capacity on the room record, the staff directory that shows the departed approver as departed, a SharePoint document library with version history and modified dates, Word working against the supplied board template, and Excel for the sixty-row list. Naming the stack changes what four of the seven observe, because in every case the contradicting evidence is on a screen the candidate has to choose to open. The two unavailable attendees are unavailable in free/busy, not in the request. The room seats four on the room record, not in the email. The approver has left according to the directory, not according to the process document. And the policy library holds three files where the one not named final was modified most recently, which is visible only in the library's own modified column. Executing the request as written never requires opening any of them. The spreadsheet task is the one worth being precise about, because it is the part of this seat most often hired on a claim in a CV. The sixty-row list is supplied as an export carrying the defects real exports carry: one person present twice under two spellings, dates in two formats so a sort produces nonsense, one row outside the period, and a total row that disagrees with the rows above it. What is required is a summary rebuilt from the rows, for the period the requester named — which means recognising that the pre-computed total is unfit before using it, normalising the dates before filtering on them, and deduplicating on something other than the name string. Column order differs between sitting variants, so a lookup written positionally returns the wrong person and one written against the header survives; Microsoft's own documentation draws exactly that distinction between VLOOKUP's column index number and XLOOKUP's lookup and return arrays. This is not a syntax test. The candidate may reach the figure by filter, by pivot, by formula or by counting carefully, and no criterion reads which. It is a fixture built so the shortcut produces a different answer from the correct method, which is the only thing that makes a spreadsheet task worth scoring at all. Fast and correct is a candidate who spends thirty seconds establishing that the source is defective and then answers the question that was actually asked; slow and correct is one who rebuilds everything before noticing the period; fast and wrong is the total row, in ten seconds, quoted in a meeting.
- Microsoft 365
- Microsoft Outlook
- Microsoft Excel
- Microsoft Word
- SharePoint
- Google Workspace
- Google Sheets
Google Workspace is the standing swap and the fixture is maintained in both: Gmail, Google Calendar with free/busy and calendar resources, the directory, Drive with revision history, Docs against the same template, and Sheets for the list. Where a buyer runs a specific document library, ticketing or procurement system, the equipment request is rebuilt against their own approval route, which is the request most worth localising because approval matrices are where these organisations actually differ. No criterion is satisfied by knowing where a menu item lives in any of these, and the design contains no typing test.
What the candidate actually does
| Task | What happens |
|---|---|
| The nine o'clock batch written_artifact · 14 min | Seven requests arriving together from six different people, with the shared calendars, the room list, the staff directory and the document library available. The pressure is stated in the brief in identical words for every candidate: all seven are wanted today, three of the requesters are more senior than the role, and two of the seven cannot both be done in the time available. Book a review for six people next Tuesday in a named room, where two attendees are visibly unavailable and the room seats four. Format an appendix into the current board template, from a draft built on last year's template with inconsistent headings and numbering that restarts. Send the quarterly attendance figure to a regional manager, who states the figure in their request, where the underlying sheet gives a different one. Raise a replacement equipment order following the process as written, where the approval step routes to somebody the directory shows has left. Retrieve the final version of a policy for a director, where the library holds three candidate files and the one not named final is the most recently modified. Set up a visitor for a date the requester has written as a weekday that does not fall on that date. And one request that is complete, unambiguous and simply has to be done. The candidate returns the work and a short list of what was done, what was changed and what was raised. The fork. Every one of the six problem requests can be executed exactly as written. Doing so is faster, is defensible in the ordinary sense that the assistant did what they were asked, and produces six results that are wrong in ways that will surface later and be absorbed by somebody else. Nothing in the brief asks the candidate to check anything, question anything or flag anything; it asks them to complete the requests. Raising the mismatch is therefore a choice the candidate makes rather than an instruction they follow. |
| The list that does not add up data_task · 8 min | A sixty-row list of the kind that actually lands on this desk — a training attendance export, or a claims log — from which a one-line summary has been requested by a specific person for a specific purpose stated in the request. The list contains the same person twice under two spellings of their name, a total row that does not agree with the rows above it, dates in two different formats so that a sort produces nonsense, and one entry dated outside the period the summary covers. The candidate produces the summary that was asked for. The fork. The total row is right there and the requester asked for one number. Using it takes ten seconds and produces a figure that will be quoted in a meeting. Rebuilding the total from the rows takes four minutes and produces a different figure, which the assistant then has to explain to somebody who did not ask for an explanation. |
| The end-of-day note written_artifact · 6 min | A short note to the team covering what is finished, what was changed and why, and what is waiting on whom. It sits in its own quiet segment after the clock on the batch has stopped, and it is graded against what the candidate actually did in t1 and t2 rather than against a model answer. The fork. A note that says everything is done is shorter, sounds better, and is the note most people write. It is also the mechanism by which six silent corrections become invisible, so that the broken approval route and the disagreeing attendance figure survive into next month unchanged. |
The mark scheme
Each criterion is scored 1 to 5 against written anchors, and every score is reported with the excerpt that earned it. A criterion marked floored is reported as a finding rather than averaged into the total. The first is open; open any other to read its anchors in full.
Accuracy survives the batchweight 0.25Every returned item is internally consistent with the sources available, and the pattern of checking is visible rather than lucky: dates read back aga…
The conflict and the ambiguity are surfaced rather than guessedweight 0.2States the conflict concretely, proposes the option they would take and why, and asks a closed question that can be answered in one word. The two requ…
The instruction that does not work is not silently executedweight 0.2Completes the request by the workable route and tells whoever owns the process what is broken, in two sentences, with what they would change. On the a…
The document could be sent on without being redoneweight 0.2Consistent throughout, matching the current template, and would go to a board or a client unchanged. Any change made to the supplied content, as oppos…
The handover note lets somebody else actweight 0.15Separates finished, changed and waiting; names the person each waiting item sits with and what is needed; and is short enough that a busy reader finis…
How it is scored
Weighted mean of the five criteria, each scored 1 to 5 against the anchors and reported with the returned item or the sentence that earned it. One raw count is reported beside the score and annotated explicitly as not a positive signal: how many of the seven requests were returned complete. A buyer scanning the report will read seven-of-seven as throughput, and in this batch a candidate who completed all seven has almost certainly executed at least two instructions that could not correctly be executed. The report also lists which of the six planted problems were raised, which were silently corrected and which were missed, because those three responses are different findings and a single score conceals them.
Integrity
- one unbroken monitored sitting, with the batch committed before the data task opens
- the end-of-day note is graded against the candidate's own returned work, so a prepared handover template cannot be made to fit
- batch variants rotated between sittings so the identity of the broken instruction, the conflicting booking and the ambiguous document differ
- one live follow-up question at the entry band, testing ownership of the reasoning rather than difficulty: the candidate is asked to say in their own words why they changed one thing they changed, and what they would have done if the requester had insisted
- session log records file-open order, time distribution across the seven requests and paste-versus-typed provenance; no automated integrity verdict is produced from it, and the follow-up question is the control
The log describes what happened. It does not produce a cheating verdict — the follow-up conversation is the control, because a statistical accusation is not something we would ask a reviewer to defend.
What you receive
- the seven returned items as submitted, including the formatted document
- the one-line summary and any statement about the source data
- the end-of-day note
- a table of the six planted problems against what the candidate did with each
- per-criterion score with the returned item or excerpt that earned it
- completed count, annotated as not a positive signal
Who decides
Recommended, and light. The one thing a reviewer should decide rather than the rubric is how much upward flagging this buyer actually wants: a lean team with a senior manager who reads everything is well served by an assistant who raises the broken approval route, and a large administrative function with a documented change process may prefer it logged rather than emailed. The reviewer confirms or overrides the third criterion with a written reason. The ranking entitles a buyer to conclude that this candidate's accuracy holds across a realistic batch, that they verify rather than assume, and that their documents go out unchanged. It does not establish speed at genuine volume, and it says nothing about how they behave in week thirty.
What this does not measure
There is no floored criterion in this design, and the omission is deliberate rather than an oversight. Flooring exists for conduct failures that must never be averaged away — a disclosure, a redirected payment, a forced balance — and every failure available in this fixture is a work-quality failure that a good manager fixes with feedback in a fortnight. Marking one of them as a disqualifying finding would misrepresent what the assessment saw. The design also does not contain a typing test and does not score words per minute, which is the conventional screen for this role and the clearest instrument-to-job mismatch in this corpus: transcription speed stopped constraining this work when the work stopped being transcription, and a timed keying test is the single most likely component to disadvantage candidates for reasons unrelated to the job. Throughput is scored, and narrowly: the count of requests returned complete inside the box, read only in combination with whether they were returned right, since the fastest seven are the seven executed exactly as written. Nothing anywhere rewards keystrokes. No criterion reads register, idiom or fluency: the writing criteria are satisfied by the presence of specific content — what changed, who each item waits on, what is needed — and a blunt note containing all of it outscores a polished one that omits the owners. Format discipline is scored against a supplied template rather than against taste, so the anchor can be pointed at. What the design does not reach is the aggregate that actually defines the role: a hundred small tasks a week, interrupted every fifteen minutes, sustained for months. Seven requests in fourteen minutes samples the judgment and the checking habits, not the endurance, and a buyer who needs to know about the endurance should ask a reference rather than lengthening the assessment. Candidates who need extended time or assistive technology should have it, and the batch is deliberately sized so that nobody who works methodically runs out of time on the first four items.
The conventional screen for this role is a typing test and a conversation about being organised. Both measure something, and neither measures this job. Words per minute constrained the work when the work was transcription; what constrains it now is whether accuracy survives the eleventh interruption, and whether the documents this person produces go onward without being rebuilt by somebody paid eight times as much. The interview is the worse of the two instruments, because it does not merely miss the target — it selects against it. Talking fluently about organisation is a presentation skill, being organised is visible only in output, and the two correlate closely enough to be confused and loosely enough to produce a bad hire.
So the design is a batch, not a task. Nobody is bad at detail on one request in a quiet room, which is why a single careful exercise ranks nothing. Seven requests arriving together from six people, three of them senior to the role, with two of them unable to both be finished today, is the smallest fixture that reproduces the actual constraint. Six of the seven contain a planted problem and every one of those problems is executable: the meeting can be booked in a room that does not seat the attendees, the equipment approval can be routed to somebody who has left, and the regional manager's own figure can be sent back to them. Following the instruction is faster, and it is defensible in the shallow sense that the assistant did what they were asked. It is also how these roles fail — not visibly in week one, but by degrees, in corrections absorbed by other people.
The brief does not tell the candidate to check anything. That is the single most important design decision in the file. An instruction to raise ambiguities would convert the batch into a test of compliance with an instruction, and every candidate would comply. Leaving it out and planting the cases silently means what gets observed is whether they choose to raise it, which is the competency actually being bought. The report then distinguishes three responses that a single score would flatten: raised, silently corrected, and missed. Silently corrected is the interesting one, because it looks like competence and produces an organisation where the same broken approval route quietly costs somebody fifteen minutes every week forever.
The commercial pressure is stated in the same words for every candidate — all seven wanted today, three requesters more senior, two of them impossible to finish together. Without it there is no cost to flagging everything, and a fixture in which flagging is free measures nothing about a role whose difficulty is entirely in the aggregate. With it, the candidate has to decide which of the six problems is worth the friction, and the difference between a 3 and a 5 on the conflict criterion is whether they hand the whole problem back or return it with a recommendation and a closed question.
The data task is eight minutes and does one specific job: it puts a plausible answer directly in front of the candidate. The total row is present, the requester asked for one number, and using it takes ten seconds. Rebuilding the total from the rows produces a different figure and obliges the assistant to explain something nobody asked them to explain. That is the same judgment the batch tests, sampled a second time in a different medium, which matters because a judgment observed once is close to luck.
The end-of-day note sits in its own quiet minutes after the batch clock has stopped, and the separation is deliberate. If a candidate produces a poor handover while still holding six half-finished requests in their head, a buyer cannot tell whether they cannot write it or simply ran out of time. Those are different findings that lead to different decisions, and the design refuses to conflate them. Grading the note against the candidate's own returned work also does the anti-coaching job: a memorised handover template cannot describe six decisions it has never seen.
Twenty-eight minutes is the budget and the arithmetic is unforgiving here. This is the highest-volume role in the largest occupational group in the published statistics, hired by a line manager rather than a specialist recruiter, usually several at a time, and the realistic competitor is not a better assessment but no assessment at all. Fourteen minutes buys the batch, eight buys the second sample of the same judgment, six buys the artefact that shows whether any of it reached another human being. There is no live call, which is a real omission rather than a scoping convenience: how this person sounds to a visitor or on a telephone is not observed, and a buyer who cares about it should watch for it in the interview this assessment is meant to shorten rather than replace.
Sources
Every figure on this page is traceable. Where a claim could not be sourced it is stated qualitatively instead.
- US Bureau of Labor Statistics, Occupational Outlook Handbook, Secretaries and Administrative Assistants, 2025, 3,515,600 total jobs of which 1,889,800 excluding legal, medical and executive, 314,400 projected annual openings, https://www.bls.gov/ooh/office-and-administrative-support/secretaries-and-administrative-assistants.htm
- US Bureau of Labor Statistics, Occupational Outlook Handbook, General Office Clerks, 2025, about 2.6 million jobs and 249,000 projected annual openings, https://www.bls.gov/ooh/office-and-administrative-support/general-office-clerks.htm
See what the employer actually receives. A full report for one role, with every score shown beside the excerpt that earned it, conduct findings reported rather than averaged, and a reviewer sign-off required before any decision. No form.
Read a sample reportOr talk to us about this role