Est.

Reference Check Questions That Surface Real Performance Data

Structured reference checks uncover hiring red flags that generic calls consistently miss.

Features Editor · · 9 min read
Cover illustration for “Reference Check Questions That Surface Real Performance Data”
Interview Design · September 23, 2026 · 9 min read · 1,987 words

What a failed reference check costs

The Department of Labor puts the cost of a bad hire at roughly 30% of that employee's first-year wages. SHRM's estimate runs wider and higher: between half and double the annual salary, depending on role and how long the mistake takes to unwind. Nearly three out of four employers admit to having made one anyway, even though roughly 87% of employers run a reference check before every offer they extend. Sit those two facts next to each other and the real problem shows itself: checking a box and gathering evidence are not the same activity, and most recruiters on these calls are doing the former while convinced they're doing the latter.

For an entry- to mid-level role, the average financial hit from a bad hire runs around $17,000. For a senior position, that figure climbs past $240,000. SHRM's survey puts average recruiting costs alone, before counting any losses tied to the hire's actual performance, at several thousand dollars for nonexecutive roles and near the tens of thousands for executive ones.

Stacking recruitment spend on top of lost productivity, team disruption, client damage, and the cost of running the search again brings a failed executive hire to $500,000 in documented costs under conservative accounting. A five-minute reference call is supposed to catch that damage before it happens. Most calls don't, because most of them confirm start dates, collect a warm and empty endorsement, and hang up. It's a formality wearing a check's clothes, not a check. It's a formality wearing a check's clothes.

Reference checks as a predictive tool

The Office of Personnel Management's assessment guidance positions reference checks as stronger predictors of future performance than resume credentials alone. Anyone still screening primarily on resume credentials should sit with that ranking for a moment: a well-run reference check beats two of the signals most hiring pipelines lean on hardest.

It does not beat cognitive ability testing, and OPM never claims it does. Schmidt and Hunter's meta-analytic work on personnel selection backs that up: structured methods, cognitive ability tests and structured interviews chief among them, predict job performance far more reliably than informal, unstructured evaluation. An unstructured reference check is informal evaluation wearing a formal name, whatever the recruiter running it wants to believe.

Sequencing matters too, and this is where most companies get it backwards. Reference checks work best late in a multiple-hurdle process, used to separate a small pool of finalists, not to screen a wide field of applicants. Running the check at the wrong stage, with the wrong questions, collapses its value toward zero. The tool works. Most people just deploy it badly, treating the final formality of hiring as an afterthought instead of the last real chance to catch what interviews missed.

Diagram: The Real Cost of a Bad Hire. Visualizes: Show the financial escalation of a failed hire across two role levels, contrasting the entry-to-mid-level cost ($17,000) against the senior-level cost ($240,000), which climbs to $500,000 in total…

Why generic questions produce useless answers

Asking a reference to "describe" someone's performance produces a vague answer almost by design, because the prompt itself invites subjective, unfalsifiable language. "Great to work with." "Really dependable." None of that is signal. It's filler that lets both parties hang up feeling like something happened.

Candidate-supplied references are a curated list, and treating them as anything else is the first mistake most recruiters make. Nobody hands a hiring manager the name of the supervisor who fired them. These calls sample friendly witnesses, not opinion at large, and every question that assumes neutrality is wasted before it's even asked.

Unstructured checks compound the error. Without pointed questions, references quietly skip over attendance issues, interpersonal conflict, and reliability concerns, especially when volunteering anything negative feels socially awkward or legally risky. A reference praises attitude, dodges any real evaluation of technical capability, and months later the new hire's actual weaknesses map precisely onto the ground that reference never touched. The pattern repeats often enough that it stops looking like coincidence.

The single most predictive question

OPM's guidance names one question as the most predictive in the entire check: "Would you rehire this person?"

Its power sits in how cleanly the answer splits into two camps. An immediate, unhedged "yes" is one camp. Everything else, silence, a subject change, "I'd have to think about that," any hedge whatsoever, belongs in the other, and every answer landing there earns a follow-up before the call ends.

A formal "no" isn't automatically a verdict on performance. Plenty of companies bar rehiring anyone who left without the required notice, regardless of how well that person actually did the job. Don't take "no" at face value: ask what the policy is first, and separate the policy from the performance before drawing any conclusion.

When the answer comes back hedged, the useful follow-up is direct: ask what would need to be true for the answer to flip, or what the reference would want to know before deciding either way. That question does more work than the original one. It reopens a conversation the reference was trying to close.

Questions that force specific, quantified performance evidence

Swap every "tell me about" for a prompt that demands a specific example, a number, or a direct comparison. "Tell me about" invites narrative, and narrative can be shaped, softened, or rehearsed in the car on the way to the call.

A generic prompt like "What were their biggest accomplishments?" gets a generic answer back. Ask instead: "Could you share a specific example of an achievement where you saw their direct impact, maybe with numbers or metrics attached?" That version doesn't leave room for a vague, feel-good non-answer. It demands a data point.

Once an example appears in the conversation, the real diagnostic work happens in the follow-up: "Was that a project they led independently, or was it a team effort?" That single distinction, independent ownership against team contribution, is something almost no resume states honestly, and a reference is often the only source who can answer it with any accuracy.

For roles where impact needs to translate into dollars, ask directly: "Can you provide a specific metric that improved as a direct result of the candidate's leadership?" That question moves the conversation off opinion entirely and onto something checkable against later performance.

Competency scoring rounds this out. Ask the reference to rate the candidate on each role-critical competency, one through ten, and require a behavioral example behind every rating given. Anchor first with the broad version, "How would you rate their overall job performance on a scale of one to ten?", and let that baseline number, paired with what follows it, calibrate whatever impression a resume and an interview already built.

Behavioral questions that reveal how the candidate operates

Behavioral interviewing rests on one premise: past behavior predicts future behavior more reliably than self-reported strengths under interview conditions. The same premise applies to a reference check, except now the person answering actually watched the behavior happen instead of performing it under interview conditions.

The framing has to shift accordingly. Ask the reference to narrate a situation, not render a verdict on character. "Tell me about a specific time this candidate handled a difficult challenge. What was the situation, and what did they actually do?" forces recall of an event instead of a summary judgment. "What would you say was their single most significant contribution to the team or company?" tends to surface what the reference genuinely valued, which is often more honest than whatever they think the hiring manager wants to hear.

Management style deserves its own question, separate from competence: "What kind of management or support helped them perform at their best?" This surfaces working-style data, how someone responds to autonomy against oversight, and predicts fit with a specific team better than any competence rating on its own.

Feedback receptivity closes this out. Ask bluntly: "What has the candidate's response to feedback and critique looked like?" A reference who hesitates here, or pivots straight back to unrelated praise, is very often flagging something they'd rather not say out loud.

The closing question that opens the real conversation

One question, asked right before thanking the reference and ending the call, tends to reveal more than everything before it: "What haven't I asked that you think I should know?"

Its power comes from what it removes. Throughout the call, the reference has been answering questions, not volunteering an assessment, and every unasked topic has stayed technically off the table. This question hands over explicit permission to raise anything at all, without requiring the caller to guess the right topic first.

It surfaces attendance patterns never mentioned earlier, friction with a specific colleague, a performance improvement plan the candidate never disclosed, a departure considerably less voluntary than the resume implies. None of it was hidden out of malice. It simply never fit inside the shape of the questions that came before.

After asking, go quiet and stay there. Don't fill the pause with reassurance or rush toward a closing pleasantry. The silence does the work: the reference needs room to decide, unprompted, whether whatever they've been holding is worth saying out loud.

Reading evasion: what hedged, vague, and scripted answers signal

Certain phrases carry more weight than their flat tone suggests. "They worked well enough." "No major problems." These sound like clearance, but the absence of a specific positive is not an endorsement, and treating it as one is a common and costly error.

Watch for a reference who praises attitude and personality at length while consistently sidestepping any evaluation of actual output or technical skill. Repeated across multiple questions, that pattern is rarely accidental. Scripted-sounding answers deserve equal scrutiny, especially when the reference can't describe the candidate's day-to-day responsibilities or name specific colleagues they worked alongside.

Beyond word choice, watch the shape of the conversation itself. A reference who can't produce one specific example, despite being asked directly more than once, may simply have a thin relationship with the candidate, or may be leaving something out on purpose. There's no reliable way to tell which from the outside. Small inconsistencies can arise from differing recollections, but significant gaps in title, scope, or tenure are a different matter, and they warrant a direct follow-up before moving forward.

A candidate who cannot produce a single reference from a former direct supervisor is its own signal, independent of anything said on any individual call. It limits how much real performance insight is available at all, and it often points toward relational friction or performance problems the candidate would rather keep buried.

A suspicious reference often reveals itself through vague or oddly rehearsed answers, an inability to describe the candidate's actual work environment, and a general unfamiliarity with team context that a genuine coworker would have picked up without trying.

When evasion persists past the first attempt, try one more specific behavioral question. If the pattern holds after that, document it and move on. Persistent vagueness, tracked and recorded, is a data point in its own right, not an inconclusive call that gets thrown away.

Former employers can generally share more than most hiring managers assume going into the call, often including basic employment facts and documented performance information, though what is permissible varies by jurisdiction and employer policy.

Anything touching a protected class, age, family status, health status, religion, national origin, stays off the table entirely, regardless of how carefully the question gets worded around it. No phrasing makes an off-limits topic safe to ask about.

In practice, employers in many regions keep formal reference responses narrow, sticking to basic verification, largely out of concern about defamation exposure, even where the law would let them say more. That caution is why legal protections around good-faith employment references matter: in many jurisdictions, employers who share factual, documented information with a legitimate hiring party have meaningful defamation protections. Behavioral questions built around documented, observable performance sit squarely inside that protection, which is a large part of why they hold up better in practice than questions chasing character judgments or personal opinion.

Sources

  1. Reference Checks: Questions, Process & Best Practices (2026) - Pin
  2. distantjob.com
  3. frontlinesourcegroup.com
  4. opm.gov
  5. testpartnership.com
  6. dartmouth.edu
Filed underInterview Design

More in Interview Design