Work samples beat interviews. It's not close.
Every founder believes they're a good judge of people. That belief survives contact with almost any amount of contrary evidence, because bad hires get explained away — "they interviewed so well" — without anyone noticing that the sentence is the diagnosis.
They interviewed well. That's the problem. Interviewing well is a skill. It correlates with confidence, polish, likability, and preparation. For most roles it correlates weakly with doing the actual job — and for client-facing roles it's actively treacherous, because the same performance skills that charm an interviewer can mask the judgment gaps that later torch a client relationship.
What the research keeps finding
Industrial-organizational psychology has spent the better part of a century measuring which selection methods actually predict job performance. The landmark work here is the meta-analytic tradition associated with Frank Schmidt and John Hunter, who synthesized decades of studies across thousands of employees, and whose findings have been refined and re-analyzed many times since. The details shift between analyses; the ordering is stubborn:
- Work-sample tests and job-knowledge tests sit at or near the top. Watching someone do a slice of the job predicts how they'll do the job. Revolutionary, I know.
- Structured interviews — same questions, defined scoring criteria, every candidate — perform respectably. Structure is doing the heavy lifting: it turns a conversation into a measurement.
- Unstructured interviews — the default "let's chat for 45 minutes" — sit far down the list. They add noise, anchor on first impressions, and reward performance over substance.
The precise coefficients vary by study and era, and honest researchers argue about them. What nobody has managed to overturn is the ranking. Evidence of work beats testimony about work, every time anyone measures.
The uncomfortable version: the hiring method most small companies rely on almost exclusively — vibes at 45 minutes' distance — is one of the weakest measurements available, applied at the highest-stakes moment in company building.
Why this bites agencies hardest
An agency's product is judgment delivered through communication: the QBR narrative, the scope-creep boundary, the save-the-account call, the email that de-escalates a furious client at 8:47am. When you hire an account manager or CSM through unstructured interviews, you are testing exactly the surface — articulate, warm, confident in conversation — that a weak candidate can rehearse, while leaving the substance — what they'd actually write to Dana, CEO, currently furious — completely unmeasured.
Meanwhile the resume pile got worse. In 2026, every application is AI-polished; the artifacts that used to leak signal (clumsy writing, generic cover letters) have been sanded away. Testimony has never been cheaper to fake. Evidence has never mattered more.
"But won't strong candidates refuse assessments?"
Some will — disproportionately the ones whose main asset was interviewing. What strong candidates actually resent is the black hole: applying into silence, four rounds of repetitive conversations, ghosting. A tight async process — clear expectations, 40 minutes of real work, a decision within 72 hours — reads as respect. The best people extend more effort to processes that look like the company has its act together, not less.
Two design rules keep it fair and humane: keep total candidate time under an hour, and give everyone an answer, fast. If you waste four hours of every applicant's evening, you deserve the Glassdoor reviews you'll get.
Build one work sample tonight
- Pick the moment of truth — the recurring situation where judgment separates great from adequate in this role. For an AM: the angry client email. For sales: the objection reply. For an EA: the calendar collision with two "urgent" executives.
- Simulate it in 15 minutes or less. Realistic inputs, one deliverable, written form.
- Write the rubric before you see responses. Four or five criteria, weighted. (Ours for the angry-client email: ownership without groveling, diagnosis with numbers, concrete next step, senior tone, commercial protection.) Scoring before reading is what keeps the halo effect out.
- Score everyone against the rubric, then rank. Interview only the top of the ranking — and now your interview is a structured conversation about their demonstrated work, the one format where interviews genuinely help.
Or let the machine do all of it. MeetTheFive builds the whole funnel for your role — branded page, judgment gate, AI-graded work samples against a rubric we write with you, and a ranked top five who book your calendar directly. You read zero resumes. Live in 48 hours, $3,000 flat.
Try the demo — you play the candidate → or book an intro call