SEtt457jheedr

How Much Should AI Be Allowed to Decide in Your Product? A PM’s Decision Tree

Every product manager I know is under the same pressure right now: add AI. A chatbot, auto-replies, smart recommendations, auto-approvals, something. The roadmap review has a new standing question, and it is some version of “where’s the AI?”

Meanwhile, the writing about AI in products has split into two camps that do not help you on a Tuesday afternoon. One camp publishes governance frameworks and ethics principles, which matter but rarely tell you whether your specific feature should ship. The other camp publishes demos, which make everything look shippable. Nobody hands you a way to decide which features can be AI-driven and which cannot.

I have spent a major chunk of my career across engineering and product in fintech: insurance, low-code automation, international lending, and global payments. In fintech, you learn quickly that the interesting question about any automated decision is not whether the system can make it. The interesting question is what happens when it gets it wrong, and how that failure differs from the human failure it replaced. Because the human process was never clean either. A loan officer has a bad afternoon. A fraud analyst develops a blind spot. The bar was never zero errors, and pretending otherwise is how you end up defending the status quo by accident.

What automation changes is not the error rate. It is the error shape. A tired analyst is wrong one case at a time, in scattered directions, and can be asked what they were thinking. A miscalibrated model is wrong ten thousand times in the same direction before lunch, and there is nobody to ask.

That reframe is the whole decision tree. Let me walk through it.

Stop asking “can AI do this?”

The honest answer to “can AI do this?” is yes, almost every time you ask it in a roadmap review. Models can draft the email, classify the ticket, score the applicant, and approve the refund. If capability is your bar, everything clears it, which means capability is useless as a bar.

The question that actually sorts a feature is: what does it cost when the model is wrong, and who absorbs that cost?

A model that is right 95 percent of the time sounds impressive until you multiply the other 5 percent by your transaction volume. At scale, “occasionally wrong” is not an edge case. It is a standing population of users who got the wrong answer today, and another one tomorrow. Your job as a PM is not to hope that population is small. Your job is to design what happens to them.

Three questions tell you almost everything you need to know about that population. They sort a feature into one of three buckets. A fourth question does not sort at all. It vetoes, and it can kill a feature in any of the three: what does the feature cost when the model is right? Take the three in order first.

Question 1: Can the user see and correct the mistake?

This is the most underrated property in AI product design. When the model’s output lands in front of a user who can inspect it and fix it before anything happens, the error cost collapses. The wrong answer becomes a bad first draft, and users are remarkably tolerant of bad first drafts.

AI drafting an email is the canonical example. The model writes, the user reads, the user edits or deletes, the user sends. Every mistake gets caught by the person with the most context and the most incentive to catch it. That review step looks like friction on a roadmap slide, but what you actually built is an error-correction layer, and you got it for free.

Now compare that to AI silently reordering a merchant’s payout schedule, or auto-archiving support tickets it scores as resolved. The user cannot see the decision, so the user cannot correct it. Errors accumulate quietly until they surface as churn, or as a very uncomfortable audit.

If the answer to Question 1 is yes, you have already handled most of the risk in this feature.

Question 2: Is the mistake reversible?

Some wrong answers can be undone with a click. Some cannot be undone at all.

A bad content recommendation is fully reversible in the softest sense: the user scrolls past it, and the cost was two seconds of mild annoyance. A wrongly sent auto-reply is mostly reversible: you follow up and correct the record, slightly embarrassed. But money that moved is a different animal. In payments, I learned to treat irreversibility as its own risk class. An auto-issued refund is not a suggestion. Funds moved, ledgers updated, and in some corridors you are not getting that money back with an “undo” button. The same goes for deleted data, canceled accounts, and anything that touches a third party who is under no obligation to cooperate with your rollback.

Reversibility is also a design choice, not just a property you inherit. A holding period or a staged execution step can convert an irreversible action into a reversible one. That is often the most valuable work you can do on an AI feature. Building the undo beats improving the model.

If the mistake is irreversible and the user cannot catch it first, you have left the green zone no matter how good your eval numbers look.

Question 3: Will someone demand an explanation?

Some decisions come with an audience. A customer who was denied, a regulator who is auditing, a legal team that has to defend the decision in writing. If the answer to “why did the system decide this?” is ever going to be required, “the model scored it below threshold” is not going to satisfy anyone, and in regulated domains it may not even be legal.

Loan decisions are the classic case. In lending, an adverse decision generally obligates you to tell the applicant why. Account rejections, benefit denials, fraud flags that freeze someone’s funds: these all carry an explanation debt, and the debt comes due at the worst possible time. Working on lending data across dozens of countries taught me to treat the explanation requirement as a functional requirement, as hard as any latency budget, and one that varies by jurisdiction in ways your model does not know about.

If a feature carries explanation debt, the bar is not “the AI is usually right.” The bar is “we can show our work, every time, to someone hostile.”

The three buckets

Run a feature through the first three questions, and it lands in one of three places. Here is what each one actually looks like, not on a roadmap slide, but on a Tuesday.

AI decides. A recommendation engine surfaces the wrong next article. The user scrolls past it and forgets it existed within the second. Nobody escalates a bad summary. Nobody appeals a suggested tag. The output is visible, the mistake is forgettable, and no one is coming to ask why. Drafted replies, search ranking, and summarization all live here for the same reason. Ship it, measure it, move on.

AI suggests, human decides. A model flags a transaction as likely fraud and drops it into a queue. An analyst opens it, reads the pattern, and decides in fifteen seconds whether to freeze the account or clear it. The model did the sorting. The analyst owns the call, and owns defending it later if it was wrong. That division of labor is the whole bucket: AI-drafted emails a person edits before sending, suggested refund amounts an agent approves, generated code sitting behind a pull request nobody merges blind. The model does the volume. The human owns the commit.

AI never touches it. An auto-rejected loan application. An account closed overnight with no one to call. A refund issued at scale with no hold, so by the time anyone notices the pattern, the money is already gone in a hundred directions. These are not edge cases a better model fixes next quarter. Somewhere behind each one is a person owed an explanation that nobody wrote down. Only a change to reversibility or the explanation requirement moves a feature out of this bucket. A better model does not.

Notice something about all three: the model itself can be identical. AI that writes a refund recommendation is a suggestion feature. The same model wired straight into the payments API is a gated one. The model did not change. The blast radius did.

That is the sorting done. Blast radius decides the bucket. Now the veto.

The veto: what does it cost when it is right?

The three questions price the wrong answers. This one prices the right ones, and it kills more AI features than any incident review does.

Traditional software has the property that the marginal user costs nothing. AI does not. Every query has a bill attached, which means an AI feature can be entirely safe, fully reversible, perfectly explainable, and still not survive its second year, because the unit economics never closed. McKinsey put it well in a 2026 piece on agentic systems: tokens are not value; tokens are the bill. Per-token prices have fallen for two years straight, and cost per completed task has still gone up, because an agentic workflow can burn a thousand times the tokens of a single call chasing an answer it is not sure of yet.

Right is also not one price. It depends on who is checking the work. I can build an entire application with AI, start to finish. A software engineer can do the same thing. The two builds do not cost the same, on either invoice. An engineer scopes a precise request, catches a bad output early, and moves past it. Someone without that judgment accepts the bad output, hits a wall later, and re-prompts from scratch, sometimes more than once, and each retry adds to the bill. The token cost and the cleanup cost trace back to the same gap: whether the operator can tell a correct output from a confident one. That gap shows up on the way in and on the way out.

This is where the buckets argue with each other in a useful way. “AI suggests, human decides” is not only the safer design. It is usually the cheaper one, twice over. The human review layer means the feature does not need the largest model to clear the bar, and it means someone qualified is catching the expensive mistakes before they compound into a rebuild. A fully autonomous feature has to be right on its own, every time, and buying that last few percent of accuracy without a backstop is where the margin goes.

So ask what a feature costs per completed job, not per call, and ask it before you ship rather than when finance asks. If it cannot generate a comfortable multiple of its own compute cost, the bucket it landed in does not matter.

“AI suggests” is not the consolation prize

Here is the part I want to push back on hardest, because I see the framing everywhere: the idea that human-in-the-loop is a temporary compromise, a training-wheels phase before the real, fully autonomous feature ships.

Usually it is the opposite. The suggestion design is the better product, full stop. It captures most of the value, because the model still does the heavy lifting on volume. It also protects the one thing autonomy cannot easily give back: taking a decision away from a model and handing it to a human reads to users as a downgrade, even when it is the safer call, so the direction only really moves one way. And it builds the audit trail and user trust that a fully automated version would need anyway. Many features should graduate from suggests to decides as evidence accumulates. Plenty never should, and that is fine.

The PMs I have watched ship AI features that survive their second year are not the boldest ones. They are the ones who priced the wrong answers from day one: what the mistake costs, who absorbs it, whether it can be undone, and who will demand the explanation. Then they priced the right answers, because a feature can clear all three and still be too expensive to keep.

So the next time the roadmap review asks “where’s the AI?”, bring the tree. Can the user see and correct the mistake? Is it reversible? Will someone demand an explanation? Sort your features into the three buckets, defend the sorting, then let the fourth question veto whatever cannot pay for itself per completed job. Because “can AI do this?” was never the question. The question is what you are willing to let it be wrong about, and what you are willing to pay for it when it is right.

Leave a Comment