In May 2025, a London dev studio backed by Microsoft and valued at $1.5 billion filed for bankruptcy.
What’s undisputed: the company, Builder.ai, had inflated its 2024 revenue by roughly 300%, claiming $220 million against an actual figure closer to $50 million, and substantial human engineering work sat behind an AI assistant marketed to clients as largely autonomous. What’s disputed: the viral version of the story, that hundreds of engineers were literally posing as a chatbot. Investigative reporting from The Pragmatic Engineer pushed back hard on that specific detail after talking to former employees directly.
Sit with that for a second. Journalists with direct access to former staff couldn’t fully agree on what happened inside one AI dev studio, after the fact, with months to investigate. If that’s how hard it is to verify after a collapse, it’s not something a founder resolves on a sales call.

Why this evaluation is different from a normal agency search
Software projects fail for boring, well-documented reasons even without AI in the picture. The Standish Group’s 2024 CHAOS research puts unclear requirements at the top, cited in 39% of failures, with scope creep close behind at 33%. PMI’s own tracking finds scope creep showing up in 52% of projects generally. None of that requires an AI claim to go wrong. It’s the baseline risk of hiring anyone to build anything.
AI execution adds a second layer on top of that baseline, and it’s the layer standard vendor diligence, portfolio, references, pricing, was never built to catch: how much of the delivery is genuinely AI-driven, reviewed by whom, and accountable to whom when the AI gets it wrong. A portfolio doesn’t answer that. A case study doesn’t either. You have to ask.
Seven questions that separate real from claimed
On what the AI actually does
What does your AI touch, and what does a human always review?
A strong answer names specific categories, on this and off that. A weak answer says “our AI handles everything” and treats the follow-up as an inconvenience.
Can I talk to a current client about your delivery process, not just your output?
A strong studio offers this without hesitation, because the process is the thing they’re confident about. A weak one redirects you to a polished case study PDF instead of a person.
On what happens when it breaks
What happens when the AI gets something wrong in a compliance-critical flow?
Listen for a specific escalation path, a named role, a documented step. “It rarely happens” isn’t an answer; it’s a hope, and hope isn’t a control.
Who is accountable if something ships broken, and how is that documented?
A name and a process, or a shrug toward “the team.” If nobody can point to a document that says who owns what, nobody actually does.
On regulated-industry specifics
What’s your track record with regulated fintech builds specifically?
General software portfolio experience doesn’t transfer cleanly. KYC, payments, and lending logic fail in ways a generic SaaS build never surfaces, quietly, and usually after launch.
How do you handle KYC, AML, or payment edge cases in your spec process?
This needs a concrete answer before a line of code gets written, not a promise to figure it out during QA, because by QA the architecture is already set.
On the contract itself
What does your pricing model assume, and what happens if scope changes mid-build?
This one matters more than it sounds like it should. A fixed-price contract with no room for scope discussion creates a structural incentive to under-specify the estimate and quietly cut corners once reality diverges from it, because the studio eats every hour past the quote. A paid discovery phase before the fixed price gets set is usually the difference between a studio that priced the real job and one that priced a guess.

The failure mode a checklist alone won’t catch
Most evaluation checklists are built to catch a vendor that’s lying. That’s the less common problem.
The harder one to catch is a dev studio that genuinely believes its own pitch. Nobody sets out to overstate what their AI does. A team ships something that mostly works, calls it AI-native because that’s the market’s confident word right now, and nobody in the room has actually tested where the AI’s judgment ends and a person quietly steps in to save it. We’ve written before about this specific gap under the name positioning debt: the distance between what a pitch claims about the architecture and what the architecture can demonstrate under a real question. It’s rarely intentional fraud. More often it’s ambition that outran verification, on both sides of the table.
That changes what evaluating a studio actually means. It’s not only interrogating them for deception. It’s asking them to demonstrate, not assert, in a way that would surface the gap even if they don’t know it’s there themselves. Ask to see the actual escalation record from a past compliance-sensitive incident, not a description of the process. Ask which specific decision in a past build the AI made without a person touching it, and how that gets proven after the fact, not just claimed. A studio telling the truth can usually produce this. A studio that’s confidently wrong often can’t, and doesn’t realize it until someone asks.
The answer patterns worth distrusting
A few phrases are worth treating as a prompt to dig deeper, not a reason to walk away outright on their own: “our AI handles everything,” reluctance to name a specific human reviewer, no regulated-industry client willing to talk, and pricing that never moves regardless of what the build turns out to require.
None of these alone are disqualifying. Stacked together, they’re the same pattern investors missed the week before the last collapse, the one where the story only came apart under investigation nobody ran until it was too late to matter.
Where we land on our own checklist
Applying this to ourselves without dressing it up: senior engineers own the specification and compliance boundary before AI executes inside it, and that split is checkable, not just claimed. A cross-border payment platform we engineered for a client in Lithuania had compliance built into the same build as payments and onboarding, not audited in afterward. For a fintech SME in Australia, restructuring off a monolith onto microservices got delivery 50% faster and cut run-rate by 35%, evidence the speed came from how the team and system were rebuilt, not from swapping in a faster autocomplete.
If you’re vetting a partner for a fintech MVP and want a second, technical opinion before you sign, NeoBank Labs works with teams operating in the US, EU, and Australian markets.
Take a look at how these engagements play out in practice in our case studies, or get in touch to talk through your evaluation checklist.
This is part of an ongoing series on production AI in regulated industries.