Skip to content
Dazlab
Reviewing AI-Written Code That Handles Money — illustration
AI Workflows

Reviewing AI-Written Code That Handles Money

AI-written code that handles money should get a different review from the rest of your codebase: a separate, adversarial review whose only job is to find how the change could go wrong, on top of the normal review every change gets. We tier review by risk. Money, logins, signing and anything that keeps one customer's data away from another's get the heavy review; everything else gets a lighter one once it works. Nothing ships unreviewed.

This is how we do it across our own products, several of which move real money every day.

Why AI-written code needs risk-tiered review

AI coding agents are good at writing code that looks right. That's both the benefit and the risk.

Code that looks right passes a casual read. It compiles, the tests the agent wrote pass, the demo works. But money software rarely fails on the obvious path. It fails on the third decimal place, on the day the exchange rate moved, on the moment two people press the same button. Those cases don't show up unless someone goes looking for them on purpose.

At the same time, giving every change the deepest possible review doesn't work either. Most changes are copy, layout, a new filter on a list. Spending the heavy review on all of them would slow the whole studio down and make reviewers numb to the changes that matter.

So we split it.

The tiers

TierWhat's in itReview
High riskMoney (amounts, tax, currency, payments, invoices, payouts), authentication and logins, signing, multi-tenant data boundaries, anything that acts on a customer's behalfNormal review plus a separate adversarial review. Money changes also get a separate review before release.
StandardFeatures, reports that read data without changing it, integrations that don't move money, workflow changesReview once it works, by a lighter, cheaper reviewer
Low riskCopy, styling, layout, test-only changes, internal toolingReview once it works

Two rules sit across all of it:

  • Every change is reviewed before it ships. The tiers decide how hard, not whether.
  • When in doubt, go up a tier. If a "UI change" happens to touch the component that shows an invoice total, it's a money change.

What adversarial review means

The adversarial reviewer is a separate, stronger AI model, set up with one job: assume the change is wrong and find out how. It's not the agent that wrote the code marking its own homework, and it's not a general "looks good to me" pass.

A person then reads its findings and decides what's real and what ships. The reviewer's output is input to a decision, not the decision.

We think of it as the difference between proofreading and an audit. A normal review asks "does this do what the ticket says?" An adversarial review asks "what would have to be true for this to lose someone money, show them someone else's data, or let the wrong person in?"

What reviewers look for

None of the following is exotic. These are the well-known ways money software goes wrong, and they're the questions we expect a high-risk review to ask every time.

Rounding

  • Where does rounding happen, and how many times? Rounding each line and then summing gives a different total from summing and then rounding once.
  • Does the total on the screen match the total on the PDF, the email, and the accounting system it syncs to, to the cent?
  • Are amounts stored in a form that doesn't drift (whole minor units, or a proper decimal type) rather than floating-point numbers?
  • When tax is calculated, is it applied only to the lines that should carry it?

When Handl's QuickBooks sync went live, the standard we set was that invoices push across matching to the cent, with tax only on taxable lines. That's the kind of sentence that should be in a spec, and the kind of thing a reviewer should check against.

Currency

  • Which currency is each number in, and is that obvious to the person reading it?
  • When converting, which rate is used, and from which date?
  • Does a report mix converted and native figures without saying so?

These are product decisions as much as code questions. In Handl, reports convert each time entry, cost and payment at the rate on the date it happened, and every report labels whether its figures are native or converted. A reviewer's job is to check that new code follows those decisions, not to make them up on the spot.

Races and double actions

  • What happens if the same action fires twice: a double click, a retried request, a webhook delivered again?
  • What happens if two things change the same record at the same moment, like a payment arriving while someone edits the invoice?
  • Can an automated action (a reminder, an auto-payment) fire while a person is in the middle of something that should pause it?

These show up in product design too. In Mortar, purchase orders have a guard so you can't order the same item twice by accident. In Handl, a client asking a question about an invoice line doesn't pause reminders, but auto-pay waits, so nobody gets charged mid-question. Each of those is a deliberate rule, and each needs checking every time the surrounding code changes.

Data leaking between customers

Every multi-tenant SaaS product keeps many customers' data in one system. The single most important property is that each customer only ever sees their own.

Reviewers ask:

  • Is every query scoped to the right customer, including the new one added in this change?
  • Do new endpoints, exports, search results, AI tools and background jobs respect that scope, not just the main screens?
  • Can a client portal user see only what that client is allowed to see?

AI features raise the stakes here, because an AI assistant that can look things up needs the same boundaries as the rest of the app. When we added lookups to Handl's client portal chat, the rule was read-only, and only what that client can already see.

Defaults and side effects

  • Does this change alter anyone's numbers without them choosing it?
  • Does it send anything to a customer's client that they didn't expect?

Our rules here are simple: new money settings ship switched off, and anything that emails your client is opt-in. A review on a money change checks both.

What lighter review looks for

Standard and low-risk changes still get a real review, just a cheaper one. It checks that the change does what the ticket says, that nothing unrelated was rewritten, that the relevant tests exist and pass, and that it reads and behaves consistently with the rest of the product.

Then CI runs the full test suite, and CI is the authority. A reviewer saying "looks fine" doesn't override a failing test.

How we do it

A few practices make tiered review work in a small studio.

The ticket says what tier it is. Every change starts as a written spec in Linear. If it touches money or tenant data, that's known before a line of code is written, and the expected behaviour (which rate, where rounding happens, what's switched off by default) is in the spec. Reviewers check against something written down, not against their memory of a conversation.

Money gets a quieter week when it needs one. After a week where we shipped project costs and retainer auto-pay in Handl, we deliberately spent the following week on a second pass, adding tax and currency handling, rather than piling new features on top. The weekly change count dropped a lot. That was the point.

Shipping isn't launching. Money features can be merged and sit switched off until the product, the docs and the pricing page all agree. That gap gives review room to work without holding up everything else.

Small batches. Deploys take about nine minutes, so changes go out one at a time. A small change is far easier to review properly than a big batch, and that matters most for money.

We describe the wider workflow in how a small studio ships 100+ changes a week, and the principles behind it in how we build. The broader argument for working this way is in AI-assisted engineering vs vibe coding.

If you're hiring someone to build with AI

Ask them:

  • Which changes get a stronger review, and who decides?
  • Is the reviewer independent of whatever wrote the code?
  • Where are money rules (rounding, currency, tax, defaults) written down?
  • How do they check that one customer can never see another's data?
  • What has to pass before a change can merge?

A studio that has thought about this will answer quickly and specifically.

FAQ

Do you review every AI-written change?

Yes. Every change is reviewed before it ships, in every product. Risk decides how deep the review goes, not whether it happens.

What is adversarial code review?

A review whose only goal is to find how a change could go wrong: the rounding case, the race, the data leak. We use a separate, stronger model for it on high-risk changes, and a person decides what to act on.

Which code counts as high risk?

Anything involving money, authentication, signing, keeping one customer's data separate from another's, or software acting on a customer's behalf. If a change is borderline, it goes up a tier.

Isn't heavy review on everything safer?

In theory. In practice it slows everything down and dulls attention, so the changes that matter get the same tired look as a copy change. Tiering puts the deepest attention where mistakes cost the most.

If you need a product that handles money built with this kind of review, talk to us.

Let’s Work Together

Dazlab is a Product Studio

Our products come first. Consulting comes second. Whichever path you take, you’ll see how a small team can deliver outsized results.