← Hedge

Case study

Hedge: How It's Built

A no-backend calibration game for iPhone, iPad, and Mac, synced through the player's own iCloud and backed by a self-hosted LLM content pipeline.

Stack: Expo SDK 55 · React Native 0.83 · React 19 · TypeScript · expo-sqlite · CloudKit · Node + Hono · SQLite · Gemini + Anthropic APIs

Platform: iPhone · iPad · Mac (Apple silicon) · Backend for the game: none of ours; on-device SQLite plus the player's own private iCloud · Status: Live on the App Store (v2.4.0)

Home screen: a 14-day streak, today's score orb, and the Gut Check pack orb peeking alongside it Round score: 81 out of 100, with a per-question breakdown of confidence against outcome The Pulse calibration chart: accuracy versus confidence, with a per-category breakdown below

The idea

Most trivia rewards luck and process-of-elimination. Hedge rewards self-knowledge. Every day you get five true/false questions, and for each one you set a confidence level (50, 60, 70, 80, 90, or 99 percent) before locking in.

Scoring is Brier-based: a wrong answer at 99 percent is brutal, a hedge at 50 percent is neutral, and being right is only worth the most if you were also confident. Over many rounds, a per-confidence calibration chart reveals whether you are systematically over- or under-confident, and in which categories. The trivia is the surface. The real product is a personal calibration instrument.

Two modes feed it. The daily round is five questions, about 90 seconds, played once and locked. Gut Check is the deep dive: a curated pack of 40 to 50 questions on one topic, played in a single sitting, that ends with a focused read on how your gut performs there and that you can replay to chase mastery.

What makes it interesting

This is a small app with a deliberately large amount of system behind it. The parts I am most happy with:

A shared TypeScript core that prevents drift

Three surfaces (the game, an authoring app, and a server) all import one package, @hedge/shared, which owns the canonical types, the Brier scoring function, the SQLite schema and migration ledger, the Zod validators, and the round scheduler. The scoring math and data model are physically impossible to drift between client and server because there is only one copy.

A nice constraint: the database interface is defined structurally to match expo-sqlite's async shape, so the shared package has zero React Native runtime dependency. That lets the test suite wrap better-sqlite3 in the same interface and run pure logic under vitest with no simulator in the loop.

Scoring that protects its own signal

Gut Check: calibration as a deep dive

The daily round measures you five questions at a time. Gut Check is the focused version, and its report splits one number into two. Calibration asks whether your confidence was honest: did your 80 percent bets actually land around 80 percent of the time? Conviction asks whether your gut committed: did you bet hard on the ones you knew and hedge the ones you didn't, or play everything safe in the middle?

The first version computed both straight from a Murphy decomposition of the Brier score, calibration as reliability and conviction as resolution. It read well in a textbook and badly in the app. Reliability is a squared error, so a visibly overconfident run still scored in the nineties, and the resolution term collapsed to zero the moment a player bet a single confidence level, which a short focused run does constantly. The shipped version keeps the decomposition under the hood but reports two numbers built to be robust and legible instead. Calibration is the gap the chart already shows you, the weighted distance of your bets from the perfect-calibration line, so the figure and the picture finally agree. Conviction is a Brier skill score: how far your bets beat a blind coin flip. Both still read from the same confidence-and-outcome pairs the daily chart uses, so the two modes share one scoring core.

iCloud sync without a backend

v2.4.0 made Hedge multi-device (iPhone, iPad, Mac) without giving it a server. Sync runs through the player's own CloudKit private database: storage counts towards their iCloud quota, the data lives under their Apple ID, and it is unreadable by me by construction. The App Store privacy label stays "Data Not Collected" because nothing is.

The design leans almost entirely on properties the data model already had:

The native layer is a deliberately thin, stateless Swift bridge (raw CKDatabase operations; even the change token round-trips through JS and lives in SQLite), so every merge decision is pure TypeScript in the shared package, covered by the same vitest suite as the scoring. Settings travel separately as last-writer-wins key-values, and because the app keeps its no-push-notification posture, a key-value "doorbell" write after each push nudges other live devices to pull, with foreground pulls as the convergence guarantee.

One hardening lesson shipped days after launch: dev builds and store builds sync to different CloudKit zones, and the local database survives swapping binaries, so sync state now records which zone its bookkeeping was earned against and invalidates itself on a mismatch. The recovery is a full re-push and re-pull, which the idempotent merge makes safe, and the fix healed every affected device in minutes with no user interaction. The bridge itself is extracted and published as an open-source package, expo-cloudkit-bridge, with the merge patterns documented for reuse.

A self-hosted LLM content pipeline

Content is its own system here, not a folder of JSON, for two reasons. The first is originality: most trivia apps draw from the same shared question banks (free ones like OpenTDB, plus commercial licensed sets), so they end up asking the same questions as each other. Generating and verifying an in-house bank means Hedge's questions are written for it, not recycled from the well everyone else uses. The second is volume: a year of daily rounds is roughly 1,800 questions, and generated candidates only yield 30 to 50 percent through review.

A pluggable generator, judged across vendors

Hand-curating the first Gut Check pack turned up a list of generation failures, and the worst was true/false balance. Opus 4.8 skewed hard toward true statements: only about a quarter of a batch came back false against a target of half, and a prompt rewrite aimed squarely at producing more false questions plateaued around a third. The other failures (drifting off the requested topic, telegraphing the answer) had the same shape. Every one of those rules was already in the prompt. The model just was not following them tightly.

The first fix tried to work around the skew instead of solving it. Each generated question comes with an "alternate", a truth-flipped rephrasing of the same fact, so a surplus true question can be swapped to its false twin and rebalance a batch for free. But Opus produced a usable alternate for only about a third of its questions, and not always a good one, so the lever ran short exactly when a batch was most lopsided. That reframed the problem from "build more balancing machinery" into "find a model that follows the balance rule on its own."

So I ran a bake-off instead: three models, one identical prompt, Opus 4.8 against Gemini 3.5 Flash and GPT-5.5. The metrics favored one model and reading the actual questions favored another. GPT-5.5 scored 95 percent on the alternate-phrasing metric but wrote flat, voiceless trivia and made the same "easy answer over the sharp false" mistake already on my list. Gemini Flash wrote questions with voice, the oblique near-miss shapes I had been hand-selecting, hit a clean 50 percent true/false split in every category, and is free. On the metrics alone, GPT-5.5 wins; reading twenty of its questions is what ruled it out.

The migration stayed small: a generation-only seam that routes by model id and keeps Anthropic prompt-caching intact, with the critic, deduplication, and source verification left on Anthropic. That makes the pipeline cross-vendor by construction, one model writing and a different vendor judging, which removes a shared-blind-spot failure mode for free.

The honest part came at sign-off. The pre-build check tested source reliability on one category, history, and concluded Gemini cited real URLs. The live run generated a sports batch, and the verifier dropped 16 of 20 questions. The facts were true, but Gemini had invented plausible deep links on official sites (nba.com, masters.com) that 404. History leans on Wikipedia, which it cites accurately; sports leans on official-site links, which it guesses. Two things held: the verifier did exactly its job and the bank was never polluted, and a one-paragraph prompt nudge to prefer Wikipedia took the drop from 16 of 20 back to 5 of 20, on a free model. A single-category check was one data point dressed up as a conclusion.

Ship content without App Store review

The game delivers new questions through Expo over-the-air updates: a push to main reaches players in about 30 seconds with no re-review. Native builds (new modules, icons) are a separate, deliberate manual step. Getting this right meant learning some sharp edges the hard way, including pinning runtimeVersion after a fingerprint policy silently orphaned older builds from the update stream, and adding path filters so documentation and tooling changes never burn a build minute.

Data-driven theming and "special days"

A composable, dark-first theme system (independent accent, background, and motif layers) sits on top of the base theme. A specific calendar date can also become a special day: a custom theme plus greeting that travels to the game as data in the content file, so a holiday look ships through the same update as questions, with no hardcoded packs and no rebuild. The end-of-round celebration escalates by performance tier, up to a perfect-game finale with confetti cannons firing in sequence from the bottom corners.

Architecture

hedge/
  apps/
    game/             # Player app, ships to the App Store. Fully on-device.
    admin/            # Question-authoring app. Thin HTTP client, never published.
  packages/
    shared/           # Canonical types, scoring, DB schema, validators, scheduler
  services/
    admin-server/     # Node + Hono + SQLite authoring backend (self-hosted)

Tech stack

AreaChoice
App frameworkExpo SDK 55, React Native 0.83, React 19, TypeScript
Navigationexpo-router (typed routes)
AnimationReanimated + Gesture Handler; calibration chart hand-drawn with Animated primitives (no Skia dependency by choice)
Persistenceexpo-sqlite on device; append-only migration ledger in shared
SyncCloudKit private database through a hand-written stateless Swift bridge (open-sourced as expo-cloudkit-bridge); NSUbiquitousKeyValueStore for settings
Authoring serverNode + Hono + better-sqlite3, self-hosted, reached over Tailscale
GenerationMultiple providers (Gemini + Anthropic) behind one seam, Zod validation, prompt-caching; an Anthropic critic for dedup, plus fetch-and-judge source verification
Toolingpnpm workspaces, vitest, EAS Build + Expo OTA

Engineering practices

Status

Live on the App Store as v2.4.0: v2.0.0 added Gut Check on top of the daily round, v2.3.0 brought the game to iPad and Mac, and v2.4.0 added iCloud sync across all of them; between binaries, features, fixes, and content ship over the air. Roughly 980 verified daily questions are scheduled out several months at exactly five per day, alongside six curated Gut Check packs of about 300 more, all generated and vetted through the pipeline above. For scale: OpenTDB, the long-running community trivia database, lists 5,298 verified questions; Hedge's solo, pipeline-generated and source-checked bank is already around 1,560 and climbing. Single-developer project, end to end: product, data model, scoring design, sync architecture, content tooling, and release engineering.

The calibration and scoring primitives are factored so one core powers both the daily round and Gut Check, and could power a planned sibling app (an active-recall study tool) without touching any of the trivia-specific code.

This page is the evergreen picture. Engineering events (campaigns, incidents, bake-offs) are documented as they happen in the build log.