# The Certainty Ladder — a starter kit for your own AI product

Paste this whole file into Claude (or any capable model) and it will help you run the
same exercise on your product. It came from Crak, a mahjong coach, but nothing here is
about mahjong.

---

## The problem this solves

AI features answer. That is what they do. Ask one a question and it produces something
fluent and confident, whether or not it has any basis for it.

We caught our own product doing it. Our mahjong coach displayed a single "best tile" on
2,004 of 2,441 decisions where several tiles were tied at the identical top score — the
tie was being broken on an internal id. The UI looked certain. The evidence did not
support certainty. Nobody had lied; the layout had a slot for one answer, so one answer
appeared.

That failure mode is not specific to games. Any product with a single-answer slot in the
UI will manufacture an answer to fill it.

## The wrong fix

We tried twelve methods to find a real best answer. Eleven were stopped by a pass mark
we had written down before the run. The lesson was not "try harder." It was that no
single rule could rank the options every time, because often the right answer genuinely
depends on the user's goal — and picking one silently means picking their goal for them.

## The fix: answer at the rung you can defend

Stop asking "what is the answer." Ask "how sure can I be, right now?" Then answer at
that level and say which level it is.

Our six rungs, from most to least certain:

1. **One answer** — "Do this." Only when the evidence points to exactly one.
2. **A short list** — "Any of these three is fine." When they are genuinely equivalent.
3. **It depends on your goal** — "Going for X, do this. Going for Y, do that."
4. **A safe default** — "This one is hard to regret." No winner, but one rarely hurts.
5. **How they differ** — "This costs you A. That one gains you B." No recommendation.
6. **Not enough evidence** — "I cannot tell you on this turn." Said out loud, never hidden.

**Every rung has real value.** A clear answer is useful. So is knowing that your answer
depends on your goal. So is being told there is not enough evidence, before you act on a
guess. Rung 6 is not a failure state — it is the honest product working correctly. If you
design the ladder so rung 1 is "the real product" and the rest are consolation prizes,
you have rebuilt the original problem with more steps.

## The five steps

**A. Name the claim.** Write down what your product wants to say, and what evidence
would actually justify it. Be specific enough that you could check it.

**B. Build the ladder.** List the weaker statements underneath that claim that are still
true and still useful. This is the hard, valuable part. Most teams skip it and end up
with "answer" or "error."

**C. Measure the rung.** For each case, work out which rung the evidence really reaches.
Not which one you hope for — which one it reaches.

**D. Say that, and label it.** Give the answer at that rung, and tell the user how sure
you are. The label is not a disclaimer; it is part of the answer.

**E. Never let the screen force it.** A layout with one answer-shaped slot will
manufacture an answer. Design the surface for the honest rung, including the lowest one.

## What made it checkable

Every number our coach shows is recalculated from the rules, not generated by a language
model. Run it again and you get the same answer, exactly. The models proposed approaches,
wrote code, attacked their own assumptions, and ran thousands of checks — they built the
thing. They never decide what it says to a user.

The practices that kept us honest, which you can copy directly:

- **The pass mark goes in first.** Commit the test set, its fingerprint, and the number
  needed to pass before you look at results.
- **Publish the zeros.** One of our methods returned 0 usable answers across all 96
  decisions against a bar of 4. We reported the zero.
- **Never move a bar.** One factor had to change the answer in 3 of 16 sealed decisions.
  It changed 2, so nothing shipped from it.
- **Strike your own best number if it does not mean what it looks like.** Our largest
  figure measured how coarse our own model was, not anything about the game. We ruled in
  writing that we may not cite it.
- **Print the unflattering half.** Of 2,376 comparisons, only 41 could name a clear
  winner. The rest showed real structure without one — and saying so is the product.

---

## Prompt: run this on my product

> I want to apply the certainty-ladder method above to my own AI feature.
>
> My product is: [describe it in two sentences]
> The single answer it currently gives users is: [describe the answer slot]
> The evidence it actually has when it answers is: [describe your inputs]
>
> Help me do this, one step at a time, and push back where I am fooling myself:
>
> 1. Name the strongest claim my feature makes, and what evidence would genuinely
>    justify it.
> 2. Build my ladder: the weaker statements below that claim that are still true and
>    still useful to my users. Aim for four to six rungs. Make sure each one is
>    independently worth shipping — no consolation prizes.
> 3. For each rung, tell me what I would have to measure to know I had reached it.
> 4. Draft the actual user-facing wording at each rung, including the lowest one.
> 5. Tell me where my current UI would force an answer that the evidence does not
>    support, and what to change.
>
> Then help me write down a pass mark for the whole thing before I run any numbers.

---

Crak / PulsePoint internal science fair / Josh Stewart
