The math was honest. The words were not.
Hedge is a daily calibration trivia game I build solo; the architecture lives in the case study. This one is about a bug that no test could catch, because nothing was broken: the app's scoring kept its promise the whole time, and the app's sentences quietly did not. Companion to The second half of the promise.
Hedge is built on one claim: it is not about being right, it is about knowing how much you know. The scoring has always kept that promise. An honest 50% earns the same fifteen points whether the answer turns out to be true or false. There is no way to farm the math by being lucky. That part has been correct since the first commit and I have never had to touch it.
A scoring table that never lied
But a player does not read the scoring table. A player reads sentences. And last week I finally read the sentences in the order a new player meets them.
Onboarding, slide one: "Nail a bold call for big points; get cocky and miss, and it stings." The first coachmark of the first round, at the exact instant the player picks their first bet: "Higher confidence pays more when right and hurts more when wrong." And when someone commits an answer without ever touching the confidence dial, a sheet appears asking "Forgot to set your confidence?"
Read them in a row
Every one of them described the bet in terms of being right. Two of them spent both halves of a sentence on it: bold and correct pays, bold and wrong hurts. Not one of them said the third thing, the thing the whole app exists to say, which is that betting fifty percent is a legitimate move and scores real points. The nudge went further and implied the opposite: you did not choose fifty, you forgot.
None of that was a decision. Nobody sat down and decided to celebrate accuracy over self-knowledge. It accumulated, one reasonable sentence at a time, because celebrating correctness is the default gravity of games and resisting it takes deliberate effort in every single sentence you write. The math was honest. The words had quietly drifted.
The gravity of games
The same gravity shows up in the parts of the app that feel best. The top celebration tier fires for five right answers at ninety-nine percent and calls you a prophet. The rarest badge is the one for a perfect round at maximum confidence. The moments with confetti and sound and a crown are, almost without exception, moments of being right. The badge for matching your confidence to your accuracy, which is the actual skill the app teaches, is one tile in a wall of thirty-six and it does not make a sound.
I am not sure that is wrong, exactly. Being right at ninety-nine percent IS the hardest thing to do here, and it deserves the crown. But it means the loudest signals in my app are all pointed at the half of the promise that was never in doubt. It's a problem, and I'm reckoning with how to resolve it.
Where the doubt actually happens
Here is why it matters commercially and not just philosophically. A new player meets a genuinely hard question in their first ninety seconds. Nothing they have read gives them permission to say "no idea." So they guess, they miss, they conclude the app is too hard for them, and they leave. The calibration chart that would have explained everything does not unlock until forty answers, which is eight days away. The moment of doubt and the moment of explanation were separated by more than a week.
So the fix had to land in the moment, not in a tutorial. Onboarding got shorter, not longer: the slide that repeated its own title lost the repetition and spent the space on permission instead. The first-bet coachmark now ends "when you don't know, fifty percent isn't giving up, it's the right answer." The nudge asks "did you mean to bet fifty percent?" and leads by saying that fifty is a real bet, not a skipped one.
And two new lines now appear once each, inside a round, at the only moment they could possibly be believed: the first time you hedge and miss, the app tells you that is exactly why it barely cost you. The first time you bet big and miss, it tells you what the same miss would have cost if you hedged. Both quote your own bet and score, generated from the same scoring function the round itself uses, so they can never drift from the math the way the rest of the copy did.
What the fix cost
Six surfaces, a couple hundred words changed, and one new column pair in the database. No new screens. Nothing got longer. The whole thing ships over the air.
The reason it took three months to notice is that nothing was broken. Every sentence was true. The tests passed the entire time, because there is no test for "this app subtly disagrees with itself." You find that by reading your own words in the order a stranger reads them, and being willing to hear that the thing you built to teach honest uncertainty had been quietly teaching people to be sure.