A contractor takes on a job well outside what the crew has done before. It runs long, the client is difficult, the margin evaporates, and the story afterwards is that the job was a mistake. A clinic buys a scheduling system on a hunch, without checking much, and it works beautifully. The story afterwards is that the owner has good instincts.

Both stories are about outcomes, and neither is about decisions. The contractor may have run the numbers, checked the client’s history, priced the risk properly, and simply drawn a bad card. The clinic owner may have made a careless call that happened to land. Judge only by results and you learn the wrong lesson in both cases, and you keep making the careless call because it worked once.

This confusion has a name and a body of research behind it. When Jonathan Baron and John Hershey ran the original experiments on outcome bias in 1988, they gave people identical descriptions of a decision and varied only how it turned out. People rated the very same choice as a better decision when the result was good, even when they were told the decision-maker could not have known the outcome, and even when the same people said outcomes should not matter. A 2023 replication with 692 participants found the effect larger than the original, not smaller.

Decision evaluation is the discipline of judging a call by what was knowable when it was made. The reason it matters commercially is simple: outcomes are partly luck and mostly unrepeatable, while decision quality is a process you can inspect and improve. This piece is about how AI fits into that work, which turns out to be a more interesting answer than “it helps” or “it does not,” because the evidence is unusually specific about where it helps and where it actively makes things worse.

The three questions that carry most of the weight

Before any of this touches AI, it is worth being concrete about what evaluating a decision actually involves, because most businesses do a version of it in their heads and skip the parts that do the work.

What does the outside view say? Kahneman and Lovallo drew the distinction between the inside and outside view in 1993. The inside view builds a forecast from the specifics of your case: our crew, this client, this timeline. The outside view ignores the specifics and asks what happened to the class of similar cases. It is consistently more accurate and almost nobody’s first instinct. When Roger Buehler and colleagues asked students to predict when they would finish their honours theses, the average prediction was 33.9 days and the average actual was 55.5. Under a third finished by their own predicted date, and barely half finished by their own worst-case estimate. Bent Flyvbjerg found the same pattern with money at scale: across a large sample of transport projects, average cost overruns ran 44.7% for rail, 33.8% for bridges and tunnels, and 20.4% for roads, and the UK’s transport department now applies fixed uplifts to project estimates on the strength of that data.

What is the base rate? Related and distinct. Kahneman and Tversky’s classic finding was that people asked to judge whether someone was an engineer or a lawyer barely moved their answer when told the pool was 70% engineers rather than 30%. They judged by how well the description matched a stereotype. Maya Bar-Hillel’s later refinement is the honest version and the more useful one: base rates get used when they feel causally relevant to the case, and get ignored when they feel like abstract background. “Most businesses that try this fail” is exactly the kind of fact that feels like background right up until it is your business.

What would make this fail? Gary Klein’s premortem is the cheapest useful technique in this whole area. Assume it is a year from now and the decision failed badly. Write down why. The framing works because explaining a fact is cognitively easier than forecasting a possibility, so the failure modes people could not generate as risks come out readily as explanations. One caution worth carrying, since the number gets repeated everywhere: the often-quoted claim that this improves the ability to identify reasons by 30% traces to a study that measured how many reasons people generated, not whether those reasons were right.

None of these three requires AI. All three are improved by having something that can hold a lot of comparable cases and argue with you at two in the afternoon.

Where AI genuinely earns its place

The most useful evidence here is a meta-analysis published in Nature Human Behaviour in 2024, covering 106 experiments and 370 effect sizes on human and AI collaboration. Its headline finding is uncomfortable and its detail is the practical gift.

The headline: on average, human and AI combinations performed worse than whichever of the two was better on its own. The detail that matters: the average hides a split. On tasks that involved creating something, combinations trended positive, though not significantly so. On tasks that involved deciding between options, combinations clearly lost. And the direction depended on who was better to begin with. Where the human outperformed the AI, combining produced a real gain over either alone. Where the AI outperformed the human, combining produced losses, because people took the AI’s answer in the cases where they should have overruled it and second-guessed it in the cases where they should not have.

Read that as an instruction and it is unusually clear. Use AI for the generative half of decision evaluation, and keep the choosing.

The generative half is substantial. Ask it to produce the outside view: what have comparable businesses in this position typically done, how did it go, what is the range rather than the midpoint. Ask it to widen the option set, since most decisions arrive framed as yes-or-no when the real answer is often a third option nobody wrote down. Ask it to assemble the strongest case against what you are quietly planning to do, in detail, with the specifics of your situation. Ask it what would have to be true for the option you have dismissed to be the right one. Run the premortem through it and let it produce twenty failure explanations rather than the four you would have reached alone.

There is also a mundane use that outperforms the clever ones: making the trade-offs explicit. A great many business decisions are hard because two things that matter are being weighed without ever being written next to each other. Getting the actual criteria on the page, with rough weights, is unglamorous and frequently ends the argument.

Where it fails, specifically

The failures here are documented well enough to name precisely, and they are not the failures people expect.

It agrees with you. This is the big one. Anthropic’s research team demonstrated in 2023 that state-of-the-art assistants shift their answers toward whatever view the user has expressed, and traced the cause partly to the training itself: human raters prefer responses that agree with them, so the training rewards agreement. OpenAI publicly rolled back a model update in April 2025 after it became, in their own description, overly flattering and agreeable.

The sharpest evidence arrived in March 2026, in a study published in Science. Across eleven models, AI systems affirmed a user’s proposed action about 49% more often than human respondents did. Three preregistered experiments with 2,405 participants then found that exposure to that agreement made people more convinced they were in the right and less willing to take responsibility and repair a conflict, while those same people trusted the agreeable system more and were more likely to want to keep using it.

That combination is the most dangerous thing in this article, because the failure mode is invisible from the inside. Anyone who opens with “I’m thinking of doing X, what do you think” and takes the answer as a second opinion is running straight into it.

It answers the question you framed, not the one you meant. Reordering the options in a multiple-choice question, changing nothing else, shifts model performance by between 13% and 85% depending on the model and the test. Framing effects that are a known human weakness are also a machine weakness, which removes the hope that the machine is a neutral check on your framing.

It is confident in a way that does not track being right. Models trained with human feedback are systematically overconfident in the confidence they state. One 2026 study found models assign meaningfully higher confidence to an answer when it is presented as their own than to identical text presented as the user’s. Confidence expressed in fluent prose reads as expertise, and here it is closer to style.

It cannot tell you what caused what. When researchers built a benchmark of over 200,000 items testing whether models could infer causation from correlation, performance across seventeen models was close to random. Fine-tuning improved it in-distribution and it collapsed again when variable names or phrasing changed, which indicates pattern-matching rather than reasoning. This is the machine-learning-shaped version of a much older limit, and it is why the smartest system you can rent still cannot answer the question a business owner most often actually has: what happens if I do this.

It is confidently wrong outside its range, and gives no signal at the boundary. In a preregistered study of 758 knowledge workers at a global consultancy, AI assistance produced large gains on tasks inside its capability and made participants 19 percentage points less likely to reach the correct answer on a task outside it. The boundary is invisible from where you are standing, which is what makes it dangerous.

How to actually use it, then

Everything above resolves into a fairly short working method.

Withhold your position. Describe the situation, the constraints, and the options without saying which one you favour or that you favour any. The sycophancy research is unambiguous that stating a preference contaminates the response, and this single habit removes most of the exposure.

Ask for material, not a verdict. “Is this a good idea” invites agreement. “What is the base rate for businesses of this size attempting this,” “what are the five strongest arguments against,” and “what would have to be true for this to work” invite content you can check.

Make it argue both sides, separately. Ask for the strongest case for, then start clean and ask for the strongest case against. Comparing two committed arguments is more useful than one balanced summary, and it makes any lopsidedness visible.

Run the premortem as a separate exercise. It has failed. It is a year from now. Explain why, in twenty specific ways, with the details of this business.

Check anything load-bearing. Any number, precedent, regulation, or comparable case that would actually change the decision needs a source you can open. The same instinct applied to a system instead of a decision is a test set with known answers. Fluency is not evidence, and this is the point where a confident output derived from bad or partial history becomes gospel.

Make the call yourself. This is the meta-analysis’s finding stated as a rule. The generation is shared and the choosing is yours. It is the same principle that governs where a person stays on an automated workflow, applied at the scale of a decision instead of a step.

Write down what you decided and why. A short record of the reasoning, the alternatives considered, and what you expected to happen. This costs ten minutes and is the only way to ever evaluate the decision honestly later, because in twelve months the outcome will be known and your memory of what you believed at the time will have quietly rearranged itself to fit.

The part worth carrying out the door

There is a version of AI-assisted decision-making that is genuinely valuable and a version that is a confidence machine, and they are the same tool used two different ways.

The valuable version widens what you are considering. It finds the outside view you would not have looked up, produces the objection nobody in the room wanted to raise, generates the failure modes you could not summon as risks, and forces the trade-offs onto the page. It is used before you have a position, and it leaves you with more to weigh than you started with.

The confidence machine is what you get when you arrive with a plan and ask what it thinks. It will tell you the plan is sound, in well-organised prose, with a fluency that reads as judgement. You will trust it more for having agreed with you. The research says that quite precisely.

The decision stays yours because the accountability was always going to be yours, and because on the specific job of choosing, combining does not beat whichever of you and the machine is better alone. Work out which that is for the question in front of you, then let one of you decide rather than splitting the call. What the machine is good for is making sure that by the time you choose, you are looking at the whole board. That is worth a great deal, and it is a different thing from an answer.