Vibe coding is asking an AI for software, running whatever comes back, and shipping it if it seems to work. AI-assisted engineering also has AI writing a lot of the code, but inside an engineering process: a written spec before anything is built, tests that have to pass, a review on every change, and a person deciding what ships. The first is fine for a prototype. The second is how you build software that handles other people's money and data.
We run four products and a set of client platforms this way. Here's how we draw the line.
What vibe coding actually is
The term gets used loosely, so it's worth being precise.
Vibe coding is a way of working where the output is judged by feel. You describe what you want in plain language, an AI produces code, you run it, and if the screen looks right you keep going. Nobody reads the code closely. There's no agreed definition of "done" beyond "it seems to work". When something breaks, you describe the symptom and ask the AI to fix it.
That's not a criticism. It's a genuinely useful way to find out whether an idea has legs. You can get from nothing to a clickable thing in an afternoon, and for a lot of questions, a clickable thing is all you need.
The problem is what it doesn't give you:
- No record of intent. The "spec" lives in a chat window. A week later, nobody can say what the code was supposed to do, only what it does.
- No proof it works. "It seems to work" covers the paths you happened to click. It says nothing about the paths you didn't.
- No second opinion. The same system that wrote the code is the one deciding it's fine.
- No owner. If nobody chose to ship a change, nobody is accountable when it matters.
What AI-assisted engineering is
AI-assisted engineering keeps all the speed of AI writing code and wraps it in the parts of software engineering that make speed safe.
In our studio, AI coding agents (we use Claude Code) do a large share of the typing across our codebases. But every change goes through the same steps:
- A written spec. Every piece of work starts as a ticket in Linear that says what's wrong or missing, what "done" looks like, the files involved, and how we'll check it.
- Tests that must pass. The agent runs the relevant tests while it builds. The full suite runs in CI, and CI is the authority, not the agent's own opinion.
- Review on every change. Nothing ships unreviewed, in any product.
- Stronger review where the risk is. Anything touching money, logins, signing, or keeping one customer's data away from another's gets a separate, adversarial review whose only job is to find how it could go wrong.
- People deciding what ships. People decide what to build, write and approve the specs, own every decision about customers and money, and choose when something launches.
The AI is fast. The process is what makes it safe to be fast. We wrote about the full workflow in how a small studio ships 100+ changes a week.
Side by side
| Vibe coding | AI-assisted engineering | |
|---|---|---|
| Starting point | A prompt in a chat | A written spec in a ticket |
| Definition of done | "It seems to work" | Agreed acceptance criteria, checked |
| Testing | Clicking around | Scoped tests while building, full suite in CI |
| Review | None, or the same AI | Every change reviewed; adversarial review on high-risk code |
| Who decides what ships | Whoever pressed the button | A person, on purpose |
| Where decisions live | Chat history | The ticket |
| Good for | Prototypes, throwaway tools, exploring an idea | Production software, especially anything with money or customer data |
When vibe coding is fine
We don't think vibe coding is bad. It's a tool with a range.
It's the right choice when the cost of being wrong is close to zero:
- Prototypes you'll show to five people to see if they lean in.
- Internal one-off scripts where you can see the output and check it by eye.
- Design explorations where the point is the screen, not the code behind it.
- Learning what's possible before you commit to building it properly.
The honest test: if this thing quietly did the wrong thing for a week, would anyone lose money, data or trust? If the answer is no, vibe away.
We do a version of this ourselves before building. One of our rules is to put the page up before you build the thing: an early-access page tells us who wants it before we write production code. Fast, rough, disposable work has a place. It's just not the place where customers' money lives.
Why vibe coding fails for software that moves money
Most of our products touch money. Handl handles agency invoicing, payments and multi-currency reporting. Mortar handles budgets, purchase orders and client payments for interior designers. That's where "it seems to work" stops being good enough, for a few reasons.
The failure cases are invisible on the happy path. Money software goes wrong in the corners: a rounding difference that only shows on certain totals, a currency conversion using the wrong day's rate, two actions landing at the same moment, a tax amount applied to a line that shouldn't carry it. You won't see any of that by clicking through a demo. You find it by writing down what should happen and then checking it, deliberately.
The decisions matter more than the code. When we added tax handling to project costs in Handl, the hard part wasn't writing the calculation. It was deciding that agencies who reclaim GST or VAT should see margin on the net figure, that it should be a setting, and that it should ship switched off so nobody's numbers changed overnight. When we extended multi-currency reporting across Handl, the decision was to convert every entry at the rate on the date it happened, not today's rate. An AI can implement either choice in minutes. It can't tell you which one your customers need. That's a product decision, and it has to be written down before anyone builds.
Trust doesn't come back easily. If an invoicing tool emails your client something you didn't expect, you switch it off. That's why our rule is that anything that emails your client is opt-in, and new money settings ship switched off. Those are engineering decisions as much as product ones, and a process that skips specs and review has nowhere to enforce them.
Multi-tenant data raises the stakes. Every SaaS product holds many customers' data in one system. Keeping each customer's data visible only to that customer is the kind of requirement that has to be checked on every change that touches it. That's exactly what our adversarial review is for.
What an engineering process around AI looks like in practice
If you're a founder or CTO deciding how to get software built with AI, here's what to look for, in us or anyone else.
Specs come first, and they're written down
The single biggest difference is whether intent is captured before code is generated. A clear ticket is what lets an AI agent build the right thing first time. It also means the decision outlives the conversation. If you can't find where a decision was made, it wasn't really made.
If you're writing requirements for an AI product yourself, our technical specifications checklist for founders is a good place to start.
The AI has limits
More agents isn't more output. We run at most three AI builders at once, merge in a fixed order, and have each one stop at "pushed" rather than wandering off to improve things nobody asked for. Limits are what keep parallel work from colliding.
Review is tiered by risk
Spending heavy review on every button colour would slow everything down. Spending it nowhere is how you end up writing an apology email. So every change is reviewed, and the changes that touch money, auth, signing or customer data get a second, adversarial pass. That second pass is a separate, stronger model whose only job is to find how the change could go wrong.
Shipping and launching are separate
Merged code isn't announced code. A feature can sit in the product switched off until the help docs, the product page and the feature itself all agree. That gap is where people check the work before customers see it.
What we've seen
Running this way for a while, a few things stand out.
The speed is real. In a busy week our studio log reports well over a hundred merged changes across our products and client work. But the number comes from the system, not the AI on its own. Deploys that take about nine minutes mean small, checked changes go out one at a time instead of waiting in a big batch.
The quiet weeks are a feature. After a week where we shipped project costs and retainer auto-pay in Handl, both money features, we deliberately spent the next week on a second pass rather than piling new features on top. A vibe-coding workflow has no reason to slow down. An engineering process does.
And the AI doesn't replace judgement. The team is small, and everyone on it is doing work AI can't: design, client relationships, deciding what matters, and deciding what ships.
FAQ
Is vibe coding bad?
No. It's a fast way to explore an idea or build a throwaway prototype. It becomes a problem when the result is shipped to customers as production software without specs, tests or review, especially if it handles money or personal data.
Does AI-assisted engineering mean there's no human review?
The opposite. Every change is reviewed before it ships, and anything touching money, logins or customer data gets a stronger adversarial review. People write the specs, make the product decisions and choose what ships.
Can a vibe-coded prototype become a real product?
It can be the starting point. We'd treat it as a working sketch: write proper specs for what it should do, add tests, and rebuild the risky parts through a reviewed process before real customers rely on it.
How do I tell whether a studio is vibe coding?
Ask where specs live, what has to pass before a change merges, who reviews code that touches money or customer data, and who decides what goes live. If the answers are vague, that tells you something.
If you want a product built with AI inside a real engineering process, that's what we do.


