Acquiring Talent in the Age of AI


The technical interview quietly turned into an open-book exam, and just about everyone brought the book.

The numbers aren’t subtle. In Resume Genius’s 2026 Job Seeker Insights Report, 78% of job seekers use AI somewhere in their search and 22% use it live, during the interview itself — feeding answers to the person on the other end of the call in real time. Recruiters have noticed. Per Gartner, 72.4% of recruiting leaders have gone back to in-person interviews specifically to fight fraud, and their analyst’s line about remote IT roles is bleak: expect “at least half of the applications” to be fake. One hiring lead described candidates “acing Socket.io quizzes” and then, asked to explain their own answers, having nothing — “parroting answers without real understanding.”

So the industry did what industries do: it treated the symptom. Google banned AI tools in interviews and hauled candidates back on-site; McKinsey and Cisco followed. Sundar Pichai told Lex Fridman that Google now wants “at least one round of in-person interviews” to make sure people have “the fundamentals.” I understand the reflex. I also think it’s a tell — a very expensive, very human tell — that we’re still trying to win a game whose rules already changed.

Because here’s the thing nobody flying candidates across the country wants to say out loud: the problem isn’t that they’re using AI. It’s that our test still measures the wrong thing. We built the whole hiring apparatus around detecting fast, correct code production — and code production is exactly the part that just got cheap. Banning the tool in the room doesn’t fix that. It just hides it for forty-five minutes.

Knowledge stopped being the scarce thing

Start where a product owner starts: what are you actually paying for when you hire an engineer?

For twenty years the honest answer was knowledge and the speed to apply it — someone who’d internalized the language, the framework, the patterns, and could turn them into working software faster than the next person. LeetCode is a proxy for exactly that. It made sense when the job was writing code line by line.

That’s the world that ended. The models hold the knowledge now — more of it, more fluently, than any candidate you’ll ever interview. Which means knowledge stopped being the bottleneck. What’s scarce now is everything around the knowledge: whether someone can think sequentially, decompose a fuzzy problem, and communicate what they actually want clearly enough that a machine — or a teammate — can execute it. Knowing things is table stakes. Operationalizing what’s known is the job.

intent [in·tent] · noun

Everyday — what you mean to do; the aim behind the action.

In agentic development — the precise, structured expression of what to build and why, complete enough that an agent can execute it and you can verify the result. It's the real work product now; the code is just the exhaust.

If intent is the work product, then the interview question isn’t “can you produce the code?” The machine can produce the code. The question is “can you produce the intent — and would you catch it when the machine gets it confidently wrong?”

We’re grading the wrong exam

This reframes the on-site panic entirely. Dragging people back into a room to watch them write a binary search by hand isn’t rigor — it’s nostalgia with a travel budget. You’ve confirmed they can do the one thing you no longer need them to do unassisted.

The uncomfortable part is that the open-book format was accidentally telling you the truth. One recruiter reframed the modern interview as “an open-book environment” where the real question is whether you’re measuring “independent thinking, or AI fluency.” That’s the whole ballgame in one sentence. The candidate parroting ChatGPT and the candidate who is genuinely excellent both have the book open. On a coding-puzzle exam, they look identical. The difference only shows up when you stop grading the answer and start grading how they got there — the questions they asked, the intent they formed, the moment they disagreed with the model and turned out to be right.

Voltaire supposedly had the hiring rubric four centuries early: judge a candidate by their questions, not their answers. In the age of AI the answers are commodity. The questions are the signal.

Nate B Jones made a version of this point in a recent video on the splitting AI job market. I’m paraphrasing him from memory, so take the wording as mine and the idea as his:

Paraphrasing Nate B Jones: in the months ahead a candidate won’t be judged on their LeetCode score, but on how well they can use their imagination to push models to their limits — their ability to learn, to ask questions, to conduct.

That reframing — from execution to intent, from answers to questions — is the entire basis for how I’d hire and grow engineers now. It splits into two jobs: finding new people who already think this way, and growing the ones you’ve already got. Let’s take them in order.

Finding it: interview for intent, not output

If you can’t test execution anymore, test the thing execution was always a proxy for: judgment.

judgment [judg·ment] · noun

Everyday — the ability to weigh a messy situation and make a sensible call.

In agentic development — the scarce human input the machine can't supply: deciding what's worth building, sensing when confident-looking output is subtly wrong, and knowing what to trust versus what to verify. Exercised under uncertainty; the last thing you can't delegate.

In this world, judgment is the job description. The model brings limitless knowledge and tireless execution; you bring the call on what’s worth doing and whether what came back is actually right. It shows up as taste — what to build, what to cut — as verification, catching the plausible-but-wrong answer before it ships, and as knowing the edge of your own competence: where to lean on the model and where to keep your hands on the wheel. None of that survives a coding puzzle. All of it surfaces the moment you watch someone work.

So hand the candidate a real, deliberately-vague problem, give them the tools they’d actually use on the job — including the AI — and watch how they operate. You’re not scoring the artifact. You’re scoring the process. A battery I’d actually run:

  • Given a vague ask, what do you ask before writing anything? Hand them fog — “our onboarding is a mess, fix it” — and see whether they start typing or start interrogating. Which users? Bounce where? Is “fixed” a number? The good ones turn ambition into a spec before they touch the keyboard. This one question tells you more than a whole LeetCode round.
  • Walk me through a time the model was wrong and you caught it. This is the verification muscle, and it’s the one that atrophies first. Someone who can’t remember catching the model has never really been reading its output.
  • What have you disagreed with an LLM about — and who won? You’re listening for independent judgment. “I don’t really disagree with it” is a red flag wearing a smile.
  • How do you decide what not to delegate to AI? Taste is knowing where the boundary is. No boundary means no taste.
  • Where do you deliberately keep your hands on the keyboard? The people who stay sharp know exactly which skills they refuse to let rot, and why.
  • Teach me something you learned this month. Learning velocity is the single most durable trait you can hire for right now — and whether they can explain it cleanly is a bonus signal on communicating intent.
  • When you’ve got five things half-done at once, how do you keep them straight? Conducting is parallel by default now — multiple agents, multiple threads, all mid-flight. The strong ones don’t white-knuckle it in working memory; they’ve built a system — externalized state, checkpoints, a way to reload context and know exactly where each thread stands. Ask them to walk you through how they actually stay oriented. “I just keep it all in my head” is the answer of someone who hasn’t yet run more than one thing at a time.

Notice what none of these measure: raw output. Every one of them measures judgment under uncertainty — which is the scarce thing, and the thing your current pipeline is worst at seeing.

Growing it: the talent is already on your payroll

Here’s the part most leaders miss while they’re busy rewriting the interview loop. The same shift that breaks hiring also means the person you need might already be three desks over — because this skill does not map onto seniority. Some of your most decorated engineers are brilliant at the instrument and mediocre in front of the orchestra. And some quiet mid-level dev nobody fast-tracked turns out to be an exceptional conductor: clear-headed, precise, allergic to ambiguity. You promoted people for execution. The axis of value just rotated ninety degrees, and your org chart hasn’t been told.

And be honest about what you’re actually teaching, because it’s harder than a new framework. Conducting isn’t a bolt-on to engineering skill — it’s a fusion of two disciplines that used to live in different people. You need someone’s technical aptitude to stay sharp enough to know when the model is confidently wrong, and at the same time you’re teaching them to be an exceptional manager: delegating with clear intent, trusting work they didn’t type themselves, reviewing output instead of producing it. Engineer and manager, blended into one person. Most training tracks pick a lane — you’re deep technical or you’re on the management ladder. The job now needs both at once, and almost no one is being coached that way.

So grow it on purpose:

  • Embrace the tools — officially. If you ban AI in the interview and then expect people to use it all day, you’re testing for the wrong thing and lying about the job. Make tool fluency a first-class, coached skill, not a dirty secret.
  • Teach intent-writing as a craft. Turning a fuzzy ambition into a legible spec is trainable, and almost nobody practices it deliberately. Make people do it, review it, and rewrite it — the way we used to run code review.
  • Make review the taught skill. The habit that keeps judgment alive is active verification: read the output, reconstruct the intent, disagree with the model out loud. Pair on it. Reward the engineer who caught the subtle bug over the one who shipped ten unread features.
  • Measure leverage, not tickets. The person who ships ten features at ten dollars of inference each is worth more than the one who ships fifteen at three hundred. If your dashboard only counts volume, you’re rewarding the exact behavior that rots a codebase while the burndown chart smiles.

None of that shows up on a velocity chart. That’s precisely why you have to go out of your way to see it, name it, and pay for it.

And take the flip side seriously, because it’s the whole reason this matters. Point these tools at someone who ships without verifying and they pass unverified, half-understood output downstream, where it becomes everyone else’s cleanup cost — except now at machine speed. You’re not accumulating tech debt anymore; you’re shipping it at a rate no human team could have produced on its own. The burndown chart has never looked better while the foundation quietly rots.

And the worst of it never touches a code file.

Even more important is when the tech debt you ship is context. A bad line of code is a local bug — someone eventually finds it and fixes it. A bad assumption baked into your company’s harness — the shared layer of context, rules, and memory every agent reads from — gets silently loaded into every conversation every engineer has with the system from that point forward. The error doesn’t sit in one file waiting to be caught. It fans out into every future prompt, compounding, magnifying itself exponentially across everyone who touches the harness. One person’s unverified answer quietly becomes everyone’s wrong answer.

That’s the real reason review is the job now, not overhead. In a world of shared context, an unverified output isn’t a local mistake — it’s a distributed one, and it’s contagious. Hire and grow the people who feel that in their bones.

The real test

The good news buried in all of this: the skill is learnable, top to bottom. The tier someone’s in isn’t about IQ — it’s about whether they’re deliberately practicing the new craft or defending the old one. That’s true of the person across the interview table, and it’s true of the person you see in the mirror.

So maybe retire the whiteboard binary search. Try this instead: give a candidate a real problem and the same AI you’d give an employee, and spend the hour watching how they think — the questions, the intent, the moment they catch the model being confidently wrong. You’ll learn more in that hour than in a month of LeetCode scores.

Which leaves me with the question I actually want answered: if you dropped your best-scoring candidate and your most “productive” current engineer into that room — same vague problem, same tools, one hour — are you confident you’d still rank them the same way? Because if the answer is no, your interview loop and your promotion ladder are both measuring a job that no longer exists.


This is Part 1 of Leadership in the Agentic Era. Coming next:

  • 2 · Leveraged Time — from soloist to conductor: multiplying your intent across a team of agents.
  • 3 · Lean on the Tools — why review is the new bottleneck, and how specs, evals, and tests let you review lighter.
  • 4 · Measure the Right Thing — the scrum metrics that lie in an agentic world, and what frontier teams track instead.
  • 5 · The Incongruent Org — what breaks when half your team moves up the stack and half doesn’t.
  • 6 · Leading the Change — empathy, foundations, and a shared language for the transition.

Comments