
AI skills are now the hardest thing for employers to find anywhere in the world. ManpowerGroup’s 2026 Talent Shortage Survey, based on fieldwork across 39,063 employers in 41 countries, found that AI model and application development and AI literacy have overtaken every traditional engineering and IT discipline as the hardest skills to source (ManpowerGroup). That scarcity has changed who gets hired, but it hasn’t changed what actually predicts whether someone can do the job well once they’re in it.Â
Most interview processes for AI roles are still built around the wrong signal. They test whether a candidate can build something that works once, in a controlled setting, with no real users and no real consequences. That’s a demo skill. It’s not the same skill as building something that keeps working after a thousand real users hit it in ways nobody planned for.Â
The Portfolio Test Everyone Gets WrongÂ
A polished GitHub repo with a working notebook is the easiest thing in AI hiring to fake confidence around, and the hardest thing to actually learn from. Plenty of genuinely capable people spent the last two years building impressive demos that never carried real production traffic, never triggered an on-call page at 2am, and never forced an uncomfortable conversation with finance about a monthly bill that grew faster than anyone expected.Â
What actually predicts on-the-job success looks different. Research into what separates strong AI hires from weak ones found that a documented portfolio of real work ranks highest as a predictor, production experience in a relevant context ranks second, and formal certifications rank last (HeroHunt). That ordering is worth sitting with. The market has quietly stopped rewarding credential accumulation and started rewarding evidence of things actually shipped and maintained.Â
The Questions That Actually Separate CandidatesÂ
The interview questions that reveal the most rarely ask a candidate to build something from scratch on a whiteboard. They ask a candidate to reconstruct a real failure. “Walk me through a production AI system that broke — what actually happened?” tends to separate people fast, because someone who has only ever worked in demos has no failure to walk through.Â
Other questions that do real work: how would you design memory so an agent can recall context from three weeks ago without it becoming unreliable or expensive? What happens in your system when a tool call fails halfway through a multi-step task — does it retry, roll back, or fail silently? How do you catch a model that’s producing plausible-sounding wrong answers before a customer does? None of these have a single correct answer, but the way a candidate reasons through them says more than any take-home assignment.Â
Red Flags Worth Taking SeriouslyÂ
A few patterns show up often enough in AI hiring that they’re worth naming directly. A candidate claiming a decade of experience specifically with AI agents is worth a second look, since the category has only existed at production scale for two to three years — the claim itself is a signal about how carefully someone represents their background. Heavy reliance on no-code tools without any underlying grasp of infrastructure, cost, or failure modes is another, since it usually means the person has automated the easy 80% of a workflow and has no plan for the hard 20%.Â
The clearest red flag is simpler than either of those: an inability to talk comfortably about something that went wrong. Anyone who has actually run an AI system in production has a failure story, because production is where failure modes that never show up in testing finally surface. A candidate with a flawless account of every project they’ve touched is either unusually lucky or hasn’t been doing this long enough to have hit the wall yet.Â
Why This Matters More Now Than a Year AgoÂ
The stakes on getting this evaluation right are rising fast. Job postings mentioning agentic AI skills grew 986% between 2023 and 2024, and industry projections put task-specific AI agents inside 40% of enterprise applications by the end of 2026, up from under 5% just a year earlier (HeroHunt). That pace of adoption means the cost of a bad AI hire compounds faster than it used to — a system built by someone who never learned to think about failure modes doesn’t just underperform, it becomes a production liability that someone else has to discover and fix later.Â
None of this means credentials or academic background don’t matter at all. It means they’re not what separates a candidate who can demo AI from one who can be trusted to ship it. The teams getting AI hiring right have mostly converged on the same habit: they spend interview time on real failure, not hypothetical success, because that’s the only place the difference actually shows up.Â
That habit is easy to state and harder to practice, because it requires interviewers who have their own production scars to know which answers ring true. A hiring panel that has never shipped anything past the demo stage will struggle to tell a well-rehearsed answer from a genuinely lived one. The most reliable fix isn’t a better question bank — it’s making sure someone in the room has actually been paged at 2am for the kind of failure they’re asking the candidate to describe.Â



