Sure Enough
Twelve two-way calls. Answer each one, then say how sure you are — 50 is a coin toss, 100 is certain. There is no score. At the end every call you rated 90 gets checked against how often you were actually right, and the whole run is plotted against the line of perfect calibration. The output is the relationship between what you said and what happened, which cannot be seen on any single question and is invisible without a dozen of them. Options are deliberately binary so the scale has a real floor and “no idea” is an honest 50. Runs alone or with a room agreeing each call out loud. Five looks ship with it — the activity’s own, plus a racetrack, an observatory, a summit and a winter market, drawn in the file itself with no images to load.
- Problem
- Being wrong is survivable. Being wrong while certain is what breaks plans, and it is invisible from the inside — certainty feels identical whether or not it is warranted. On any single question a confident miss is bad luck and everyone has a story explaining it. Only across a run does the pattern stop being deniable, and almost nobody keeps that record.
- Who it helps
- Anyone who estimates, forecasts, or signs off on a plan — operations leads, project and change teams, quality reviewers — and anyone teaching judgement rather than knowledge.
- In training
- The score is not the point and saying so up front changes how people play. Push them to use the bottom of the scale: a room that never says 50 is not being careful, it is avoiding admitting it does not know. Run it with the answers hidden until the end if the group is competitive, because seeing a result mid-run drags the next few ratings toward it. Finish on the 90-and-100 line — that single number does more than the whole curve.
- Skills shown
- Calibration measurement, bucketed confidence analysis, SVG chart generated from live data, binary framing chosen so the probability floor is meaningful, in-browser editing with persistent export.
How to edit
- Click Edit this activity in the corner.
- Each call needs exactly two options and the index of the right one. Two, not four — with two, chance is 50% and the bottom of the confidence scale means something.
- Twelve calls is about the minimum for a curve worth reading. Below eight, one unlucky answer swings a whole bucket.
- Mix genuinely hard calls with ones that only look hard. A set where the obvious answer is always wrong teaches distrust, not calibration.
- The why matters more than the answer. Write why the confident guess was tempting.
- Click Save my version for a copy with your sets built in.