Cheku is currently in public beta. Please report any issues to support@cheku.ai.

Apr 15, 2026

"AI hallucinates tax law" is a product problem, not a user problem

A partner at a mid-sized firm told us recently: “I tried ChatGPT for a tax question last year. It quoted a section of the Finance Act that doesn’t exist. I haven’t touched an AI tool since.”

That is an entirely reasonable response. And the standard industry reply — “you need to prompt it better”, “you need to verify the output”, “that’s just how LLMs work” — is an entirely unreasonable one.

Here is the uncomfortable truth the AI industry needs to sit with: if your product hallucinates tax law, that is a product defect. Not a user defect. Not a prompting defect. A product defect.

What a hallucination actually is

Large language models are trained to produce plausible-sounding text. That is their entire job. When you ask one an open question — “what is the capital disposals adjustment under CIR?” — it will generate something that sounds like the answer, whether or not it has seen the answer.

Sometimes the plausible answer happens to be correct. Often the plausible answer happens to be confidently, beautifully, authoritatively wrong. The model does not know the difference. It has no internal signal that says “I’m guessing here.” That is not a bug in a specific model — it is the nature of the technology.

So when a generic AI chatbot invents CFM96461 (a section that doesn’t exist) and attributes it to HMRC, it is not malfunctioning. It is doing exactly what it was built to do. The bug is deciding that a pure language model is an appropriate tool for tax research in the first place.

The generic-chatbot approach is the problem

Most AI tools accountants have tried so far follow the same pattern:

  1. User types a tax question into a box.
  2. Model is asked to answer the question from its training data alone.
  3. Model produces a confident, fluent, possibly-fictional answer.
  4. Accountant is expected to verify every citation.

Step four is where the whole thing falls apart. The partner above has to fact-check the AI against HMRC manuals — which is the exact thing the AI was supposed to save them doing. The accountant is doing the work twice: once to ask, once to check. And the check is harder than the original question, because the AI’s output sounds right.

That is not a productivity tool. That is a productivity tax.

What a well-built tax research AI actually looks like

The problem is solvable, but not by making the language model “smarter.” It is solvable by changing the product architecture around the model.

A tax research AI should:

  1. Never answer from memory. The model’s training data is not a source. Full stop.
  2. Retrieve before it generates. Every question should trigger a live search across a curated, trusted corpus — HMRC manuals, ATO rulings, the actual statutes, your own engagement files — before the model writes a word.
  3. Quote, don’t paraphrase (in the citations). The reference shown to the user should be the literal text from the source, with a link back. The model’s job is synthesis and explanation; the source’s job is authority.
  4. Refuse gracefully. If the corpus does not contain an answer, the correct response is “I could not find authoritative guidance on this — here is what I searched and what I found closest to it”, not a confident fabrication.
  5. Surface the chain of reasoning. The user should see every lookup the AI made, in order. If something looks off, they can spot it in seconds instead of re-doing the entire search.

This is the difference between generating an answer and finding an answer. The former hallucinates. The latter cannot, because the output is bounded by what was retrieved.

How Cheku is built

When you ask Cheku about the capital disposals adjustment in the Corporate Interest Restriction, it does not reach into the model’s memory. It:

  1. Searches the HMRC Corporate Finance Manual and the relevant legislation in TIOPA 2010.
  2. Pulls up the engagement files for the specific client, if you have asked about one.
  3. Quotes the exact text from CFM96460 — the real section — and links you to it.
  4. Explains how that section applies to your client’s specific disposals.
  5. Shows every tool call it made along the way.

If the guidance does not exist, Cheku will tell you so. It will not invent CFM96461 to make the answer feel complete. We have chosen “I don’t know” over “here is something that sounds right” every single time, at the product level.

That is a design constraint, not a feature we might add later. It is the thing the product is for.

What this means when you evaluate AI tools

When a vendor demos an AI research tool, ask these three questions:

  • “Show me where every number and citation in that answer came from.” If the answer is “the model”, walk away.
  • “What happens when you ask it something the corpus doesn’t cover?” The right answer is a graceful refusal. Watch carefully.
  • “Can I click the citation and land on the exact paragraph in the exact source?” If the citation is a free-text string the model wrote, it is not a citation. It is decoration.

Accountants are trained to distrust assertions without sources. That professional instinct is exactly right for evaluating AI tools, and the industry should meet you on your terms — not ask you to lower them.

Hallucinations are not an unavoidable cost of using AI in professional services. They are a symptom of using the wrong architecture for the job. The right architecture is out there. It is worth holding vendors to it.