TL;DR: Two of hiring's standard filters — the portfolio and the take-home — now measure something that takes hours to fake. Hiring managers in PainHunt's HR data have already worked this out and named what replaces them: live conversation that probes judgment. What they don't have is a way to run that consistently, which is the actual product. This is a tooling gap in the one part of hiring everyone agrees still works.
The evidence
PainHunt's HR and recruiting categories together hold 1,022 high-scoring signals (10+/15), average intensity 7.4/10, sourced from Mastodon (231), Medium (196), Reddit (180), BlueSky (129), AppStore (57) and HackerNews (56). 415 of those signals touch candidate assessment, portfolios or interviewing — it's one of the largest recurring themes in the category.
The sharpest cluster sits in Recruitment Assessment, average score 13.0/15 with intensity 8.2/10 — among the highest-scoring domains in the entire dataset this round. The reporters are hiring managers, design leaders and HR professionals. What they describe:
- AI-generated portfolios make it impossible to verify if candidates actually did the work themselves.
- AI can generate polished portfolios and case studies in hours, so visual work samples no longer indicate capability.
- Take-home exercises can be completed with AI, rendering them ineffective as assessment tools.
- Traditional portfolio reviews cannot assess real-time judgment or thinking under pressure.
- Candidates with polished presentations cannot explain alternative solutions or decision rationale when asked.
- No systematic way to conduct "room test" style interviews that probe deeper than the presented artifact.
- Live interview conversations are subjective and difficult to scale consistently across multiple interviewers.
- The framing that recurs in the data: an economic inversion — artifacts are now cheap to produce, and harder-to-fake conversations are the only reliable signal.
The requested fixes describe a product with unusual precision: a live interview question bank with follow-up probes designed to test unprepared thinking, candidate response recording and analysis for debriefing across multiple interviewers, real-time coaching that suggests follow-up questions to the interviewer, and structured interview frameworks with consistent follow-ups.
Note what is not being asked for by the people closest to the problem. AI-detection for portfolio authenticity appears in the data, but it sits alongside a diagnosis — the artifact is no longer the evidence — that makes detection the wrong layer to fix.
Why now
The cost of producing a convincing work sample collapsed inside about two years, and hiring processes were built on the assumption that it was high. A portfolio was a proxy for months of real work; a take-home was a proxy for a few days of it. Both proxies broke at once, and neither had a fallback.
The response so far has mostly been detection — tools that claim to spot AI-generated work. The data suggests practitioners have already moved past that, and they're right to: detection is an arms race with a well-funded opponent and a brutal false-positive cost, since accusing a real candidate of fabrication is a hiring disaster.
What's left is the interview, and the interview was never engineered. It's the least instrumented part of a hiring stack that has otherwise been automated end to end — ATS, sourcing, scheduling, assessment platforms. The one step everyone now agrees carries the signal is the one with no tooling behind it, and the complaint in the data is exactly that: it's subjective and inconsistent between interviewers.
The wedge
Instrument the conversation. Don't try to detect the artifact.
- A probe bank, not a question bank. The value isn't the opening question — it's the second and third: "what did you rule out and why," "what would you do with half the time," "what broke." Prepared answers and generated ones both thin out under follow-up. Organized by role and by the specific claim being tested.
- Real-time follow-up suggestions for the interviewer. The data asks for this directly. Most interviewers are practitioners doing this a few times a quarter, and they run out of probes before the candidate runs out of prepared material. Coaching the interviewer is more tractable than scoring the candidate.
- Structured capture for honest debriefs. Recording and structured notes so multiple interviewers compare the same dimensions instead of trading impressions. This is the named fix for the consistency complaint, and it's also the part that survives a legal review — provided consent and retention are handled properly.
- Calibration across interviewers. Show where two interviewers rated the same signal differently. Consistency is the complaint; making inconsistency visible is the smallest change that addresses it.
- Sell to the team that stopped trusting take-homes. They're already looking for a replacement and already believe the diagnosis — that's a much shorter sale than convincing someone the problem exists. Design and product hiring show up first in this data.
Risks and honest caveats
- Recording interviews is legally loaded. Consent requirements vary by jurisdiction, retention creates discovery risk, and anything resembling automated candidate scoring runs into hiring-discrimination law — EU AI Act obligations and NYC Local Law 144 both bite here. This constrains the product's shape before the first line of code, and it's a compliance burden, not a footnote.
- "Structured interviewing" is an established practice, not a new idea. The research is decades old and the incumbents (HireVue, Karat, BrightHire and others) are funded. What's genuinely new is the collapse of artifact-based filters, which changes the urgency rather than the concept — and urgency is a weaker moat than novelty.
- Better interviews may be a training problem. If interviewers improve after a workshop and stop needing prompts, the software is a wedge into a services business. Worth knowing before pricing it as SaaS.
- The buyer is diffuse. Hiring managers feel the pain; HR or talent acquisition holds the budget; neither owns interview quality outright. That mismatch is a common reason good hiring tools stall.
- Assisting the interviewer can look like scoring the candidate. The compliant version helps a human ask better questions and stops short of recommending a decision. That line is easy to state and easy to cross under customer pressure for "a score," and crossing it converts a tooling product into a regulated one.
How to validate this further
Read the HR and recruiting signals in the Pain Point Browser — with 415 assessment-related signals there's enough to separate the design-hiring version from the engineering one before choosing. For the adjacent wedge aimed at contract and freelance hiring rather than full-time roles, see verifying contractor skill and proof of work. Then take the legal constraints seriously early, and pressure-test the buyer question with the Idea Validator and how to validate a startup idea.