We use analytics cookies to see how the site is used. Privacy policy

All posts8 October 2026

AI-native or vibe coded? How to tell before you sign

Ask for evidence, not the label. Six questions that tell an AI-native supplier from a vibe-coding shop before you sign, and what a real answer sounds like.

7 min readbusiness

Ask to see the work behind the word. A supplier that builds well with AI can show you a specification from a past job, a named reviewer on every merged change, tests that fail when the feature they were written for breaks, and written limits on what the AI may touch. A vibe-coding shop can show you a demo and a speed multiplier. Either will call itself AI-native, so the label settles nothing.

AI-native delivery means AI does much of the coding while a named engineer reads and answers for every change; vibe coding is the same generation with nobody in that chair. Both use the same tools, often the same models in the same week. What differs is whether a person is accountable for what ships, and whether the supplier can prove it before you sign.

Ten of eleven agencies now say AI-native

We read the homepages of eleven UK software agencies and studios in September 2026, while rewriting our own, and ten of them called themselves AI-native or AI-first on the homepage or in the page title. So the word no longer separates anyone. All it tells you is that a supplier has noticed the market, which you knew already.

Speed claims have gone the same way. Three of those sites sold delivery as a multiplier, five or ten times faster, and only one said faster than what. The best measurement we know of is less exciting. A working paper published through the US National Bureau of Economic Research in 2026 followed more than half a million GitHub developers. Once they took up autonomous coding agents, their commits rose by about 240% and the software they released by about 30% (Demirer, Musolff and Yang, 2026). The study is observational and sees only public open-source work, and two of the authors do consulting work for Microsoft. It still matches our experience on current builds: the widest changes are the machine's, and the pace is set by the person reading them.

Every supplier uses AI now, so asking whether they do wastes a question. Ask what they can show you about the parts AI does not do. Those are the parts investors work through in technical due diligence, a year or two later and with more at stake.

Six questions, and what a real answer sounds like

All six fit in a first call. A supplier that works this way answers most of them from memory and the rest with a link.

AskA real answerWorry if you hear
Can I see a specification from a past job?A redacted document with acceptance criteria, what is out of scope, and what the change must leave alone"We work from your brief and iterate"
Who reviewed the last change you merged?A person's name, and a review record you could open"The AI reviews its own output"
Can you show me a test failing without its change?A recent change reverted on a branch, and the test written for it failingA coverage percentage
How do you mark the code AI wrote?A line in every commit message, so the split is one search away"We don't track that"
Where does the AI run, and what can it see?Under business accounts whose terms rule out training on your code, with no production credentials and your customers' personal data kept out, all in the contractA list of tool brands
Who merges to the main branch?A named engineer, every time"The pipeline takes care of it"

If you only have time for one, ask for the specification. On one platform we build, a specification carried a fourth list after the files to create, change and delete: the files to leave alone. That list exists because a tool told to add a feature will happily tidy three neighbouring things nobody asked about, and every one of those tidy-ups is a change somebody then has to read. A supplier who has never needed that list has probably not been reading the diffs closely enough to want one.

Where the AI runs, and what it can touch

During a founder's public twelve-day experiment in July 2025, an AI agent working in development deleted data from a production database in the middle of a declared code freeze. Replit, the platform it ran on, explained that its apps had kept development and live customer data in a single database, and announced separate ones three days later (Replit, 2025). The rollback worked and the data came back (The Register, 2025). Many of the AI failures that make the news look like this one. The agent misbehaved, but the damage came from its reach: nothing stood between the tool's workspace and the customers' records.

So ask where the AI works and what it can reach from there. The answer you want is dull: the tools run on business accounts whose terms rule out training on your code, nowhere near a production credential, and your customers' personal data never reaches them.

The strongest version is written down as a decision. A client marketplace we build runs several AI services of its own, and the tool that releases code to production contains no AI at all, by design and in writing. A release is a typed command checked against a strict version format. The matching tag and the built image both have to exist before anything moves, and production needs an approved person asking from the right place, with that approval checked again when they confirm. No model is anywhere near it. We publish how we build with AI and what always stays with a person, and that release path is the shape we mean.

The test that fails without its change

A green test suite is the easiest thing in software to fake without meaning to. Generated code tends to arrive with generated tests, written from the same understanding of the task, and the two pass together because they agree with each other.

METR, an AI research group, put AI-written patches that had already passed a coding benchmark's automated tests in front of maintainers from three of the projects involved. Even after allowing for how inconsistent maintainers' own decisions are, about half would have been turned away (METR, March 2026). The sample was small, and the agents had one attempt with no chance to answer review comments, so read it as a direction rather than a rate. The direction is bad enough, given that every one of those patches had already passed the tests.

The check a buyer can run is blunt. Pick a change from the supplier's recent work, ask them to take it out on a branch, and run the suite. The test written for that change should fail. If nothing goes red, the tests were decoration, and the supplier learns it in the same minute you do. Investors run the same check a year later, and a green suite that proves nothing is one of the flags that stall a round.

It is also the check that separates a team that reads its agents' work from one that forwards it. First on our list of the jobs coding agents still get wrong is work an agent says it has done, and a test that cannot fail is a version of the same problem.

Who merges, and how you would find out

The quiet risk in AI-assisted delivery is review thinning out as volume rises. Faros AI, which sells engineering analytics, reported in September 2026 that across telemetry from 22,000 developers, the number of pull requests merged with no review at all went up 76% as its customers leaned harder on AI (Faros AI, 2026). That figure is a rise, not the share of pull requests left unreviewed, and it comes from a vendor's own customers. It also describes a failure no buyer could spot from outside.

Two artefacts let you see it: a review record with a person's name against every merge, and a line in each commit marking where AI was involved, so anyone can count the split. Merging sits on a short list of actions we never hand to a tool, alongside infrastructure changes and new dependencies, and a small company's AI use policy should carry the same list. Investors ask for this evidence too, which makes the provenance checks diligence now runs on AI-written code the buyer's questions with a later deadline.

Save one question for the end of the call, because it catches out suppliers whose review is only a word on a slide: which change did you last send back, and why? Review with teeth leaves a trail of rejected work, usually with the reason written beside it. A supplier who answers with a speech about velocity has told you nobody sends anything back.

Start a conversation

Tell us what you're building or deciding.

Tell us about a product you want built, or a decision about AI you need to make, and we'll set up a 30-minute call. You'll get a straight answer on what it would take, including when the answer is that you don't need us.

What do you need?
The first call is free, and there is no sales pitch.

What happens after you get in touch

  1. We reply within one working day

    By email, to arrange a time for the call.

  2. A free 30-minute call

    With a senior engineer, not a salesperson, about what you are building or deciding.

  3. The fee in writing first

    If a first step is worth taking, you get its exact fee in writing before any work starts.