
John Goleby is the CEO and co-founder of Askable, a human intelligence company. A serial entrepreneur, he cofounded the agency Orange, where Askable grew out of the incubator in 2017. Askable gives enterprises like Atlassian, Canva and Toyota the human insight to build better products. Now Askable Labs, its research arm, gives frontier AI labs the rare human data no model can generate for itself. Backed by a $14M Series A from Airtree, Askable made a quiet bet years ago that the rest of the world is only now catching up to: real people are the one input you can’t fake.
I spend most of my week in two kinds of rooms. In one, enterprises are rolling out AI and sprinting to keep up with it. In the other, the frontier labs are building it, and wrestling with a harder question: how do you know the answer is actually true? Two very different problems. But the more time I spend in both rooms, the more they look like one problem wearing different clothes.
The enterprise version is the one most people have noticed by now. We’ve wired every part of the product stack for speed. PMs draft specs with copilots, designers run options in Figma, developers ship with Cursor and whatever tool landed last week. But the one wire that carries what customers actually think is still unplugged. We’ve built a race car and painted over the windscreen. The team here has laid that case out in detail: the old research cycle was built for quarterly shipping, and no amount of speeding it up fixes a model that assumes the question always comes before the evidence. I won’t relitigate it. I want to talk about the room next door, because I think it tells us where all of this is heading.
The harder problem the labs are chasing
In the other room, nobody is worried about speed. The frontier labs can generate an answer to almost anything in a second. What keeps them up at night is whether the answer is true. They’re pouring enormous effort into models that don’t just sound right but are right, grounded in something real instead of confidently inventing it. That’s the actual frontier now. Raw capability is close to handled. Trust isn’t.
And here’s what took me a while to see. Both rooms are circling the same question: is the AI grounded in something real, and can a human check it? Speed without that just means you ship the wrong thing with more confidence. Intelligence without that is just a very fluent guess.
The labs are the clearest signal we have of where this is heading. When the people building the most capable models in the world decide the hard part is grounding and trust rather than raw capability, the rest of us should take the hint. Most of us aren’t training frontier models. We’re making product decisions with AI every day, with far less grounding than we’d like to admit. Grounding the labs’ models is their problem to solve. Grounding yours is very much within reach, and that’s the room I want to talk to.
I’ve come to think this is the defining question of the next few years. And for the rest of us, it’s a question about the evidence already sitting inside our own companies.
The evidence already exists
Here’s the part that surprises people. The evidence to ground all of this mostly already exists. Companies are soaked in human signal. Customers tell you what they need on sales calls, what confused them in support tickets, where they hesitate in onboarding, why they’re walking away in churn notes. The problem was never that evidence is scarce. It’s that it can’t move. A support ticket helps support and dies there. A sales objection helps one account executive and dies there. A research transcript helps the researcher and ends up in a readout nobody opens twice.
So when an AI reaches for something to ground a decision, it usually finds nothing it can use, and does what it always does in a vacuum: it produces something plausible. That’s the failure mode both rooms are fighting. Not too little intelligence. Too little grounding.
What it takes to ground AI in evidence you can trust
Across hundreds of these conversations I’ve landed on what evidence actually has to do before an AI, or a person, should lean on it.
It has to be captured as the work happens, not only when someone kicks off a project. The richest signal is in the sales calls, support chats and churn interviews already happening every day, and most of it evaporates.
It has to be structured, not just transcribed. A transcript tells you what someone said. Real evidence holds onto what happened: what they tried, where they hesitated, what they meant, which segment they’rein, whether the same pattern shows up elsewhere.
And it has to trace back to a real person. This is the one I care about most. It’s what the labs are wrestling with at model scale, and it’s the piece you can get right in your own product today. AI will happily summarise anything you hand it, and the more fluent it gets, the more it matters that you can see exactly what it’s working from: which customer, which session, which clip, how strong the signal really is. Without that, you haven’t grounded anything. You’ve just added another layer of confident-sounding output, which is the last thing any of us needs more of.
Evidence at your fingertips
Evidence has to be there the moment you need it, not a day later. It only changes a decision while the decision is still open. If reaching it means stopping, switching tools and hunting through a repository, the call will be made long before you surface the answer. It has to meet you in the flow of the work, not wait for you to come looking.
That is the whole reason we’ve built what we’ve built at Askable. Not a faster research tool, there are enough of those. An evidence layer: verified human evidence, structured and traceable, sitting one question away from wherever you’re already working.
Ask it directly inside Askable with Ask AI, or bring the same evidence into the AI tools your team already lives in through MCP, so what your customers actually said stays connected to the specs, designs and builds happening right now, and to the model answering questions along the way. Same evidence, wherever the decision is being made.
Here’s what it looks like in practice. Take a question every team building with AI is wrestling with right now: how much should you let the AI act on its own before a person signs off? You can argue it out fromopinion, or you can ask what people actually say. When we put that question to people across Australia, the UK and the US, the answer was strikingly consistent. They’ll happily let AI do the work, but they want to check it before it counts. One participant put it plainly:
“I don’t feel that I could be fully confident that it was mistake-free. And I think that is because of past experiences… with things like AI hallucinations, it basically purporting fake information to be correct and confidently doing so. So for me, it’s just the knowledge that I have checked and double-checked.”
— Kim, 42, AU
Every piece of that is checkable. I can open the session, confirm Kim was a verified participant, and read the full conversation around what she said. The model isn’t inventing a plausible answer from whatever happened to be in the prompt. It’s pointing at a real person who really said it, and I can go and look. You don’t take its word for it, you check it. That, to me, is the whole game.
