RuleScore Methodology
What RuleScore measures
- Whether the text reads as executable contract prose at all.
- Whether the structural minimums exist: an operative clause, an outcome binding, a deadline, a checkable source.
- Whether the controlling source is named and bound to resolution.
- Whether operative terms and thresholds are defined.
- Whether timing, cutoffs, boundaries, rounding, and revisions are clear.
- Whether contingency and adjudication language creates avoidable discretion.
What it does not measure
- The probability that the event occurs.
- The expected value or price of a contract.
- Whether a venue or trader is legally correct.
- Whether a structurally complete contract makes real-world sense — the rubric judges structure and known hazard patterns, not meaning.
- Every operational, oracle, or real-world failure that text cannot reveal.
The three stages
1 · Validity gate
Text statistics (function-word density, letter ratio, repetition, word-length profile) must read as English contract prose. Failing input is unratable: no score, no letter, an explicit reason. Zero of the 2,265 real market texts in the calibration corpus trip this gate.
2 · Structural-evidence floors
Missing structure sets a minimum score — no operative clause, no outcome binding, no deadline, nothing checkable. Floors are removed by adding the missing structure, never by rewording; text can never buy credit by adding words.
3 · Hazards, gated safeguards & the contest prior
Language Risk (LRS): calibrated hazard detectors plus safeguards that earn credit only when they answer a hazard actually present. Contest Prior (CPR): the measured base rate of escalated disputes for the contract's subject family, published beside the language score — subject risk, never a drafting criticism.
An unfalsifiable qualifier is published beside the rating when the text names no checkable record or relies only on press consensus. It is not hidden inside the number.
The rating scale
A worse notch has a higher in-sample escalated-dispute share in the historical calibration corpus (strictly monotone, 1.8% to 48.5%). These band shares were used to choose the cutoffs; they are not prospectively validated failure probabilities. The labels describe settlement clarity, not credit quality or investment merit.
| Rating | Score | Risk label | Records | Language-failure share | Escalated-dispute share |
|---|---|---|---|---|---|
| BR-Aaa | 0–9 | Minimal | 397 | 0.3% | 1.8% |
| BR-Aa | 10–15 | Low | 880 | 1.1% | 5.9% |
| BR-A | 16–23 | Moderate | 249 | 1.6% | 7.2% |
| BR-Baa | 24–49 | Elevated | 547 | 5.3% | 24.1% |
| BR-Ba | 50–100 | High | 66 | 19.7% | 48.5% |
Shares are computed on the graded 2026 backtest universe (disputed records + never-disputed controls + live top-volume markets). Contract text that fails the validity gate is unratable and appears in no band.
Structural-evidence floors
A floor is a minimum score set by missing structure. It renders in every rating's risk ledger as an actionable line-item, exactly like a hazard.
| Missing evidence | Floor | Why |
|---|---|---|
| No operative resolution clause | 55 | The text never states how or when it resolves; nothing binds settlement. |
| No outcome binding | 30 | Resolution is mentioned but no condition is bound to a specific outcome (Yes/No/named result). |
| No deadline evidence | 20 | No date, expiry reference, or time anchor — an open-ended contract cannot settle No. |
| Nothing checkable (anchor tier 0) | 18 / 14 | No bound source, official record, or URL; 18 for bespoke text, 14 for machine templates, plus a +12 tier penalty. Historically the proposer wins ~75–86% of disputes on such contracts by default. |
| Press-anchored geopolitics or mention market | 26 | Press consensus (or nothing) as the only anchor in the two highest-dispute families (36.6% / 32.7% escalation rates). |
Evidence base
The methodology was calibrated against a generated 2026 corpus containing 1,595 disputed-market records plus 672 never-disputed high-volume control records and 95 live top-volume markets (2,265 graded texts in all). Brierly classifies 80 census records as hard-failure outcomes: 44 50/50 resolutions and 36 escalated overturns. For calibration, 50/50 records whose rules contain an explicit 50-50 void clause are excluded from the positive class — those are correctly drafted contracts whose stated contingency executed — leaving 57 language-failure records, alongside 184 contested-but-held records. The detailed case studies in the Dispute Database are sourced explanatory examples, not the entire census.
RuleScore rated 1,498 records; 97 lacked retrievable rules text. The committed export is text-free, derived from a private source corpus, and includes one apparent duplicate, so it does not independently establish 1,595 unique markets. On this corpus, in sample: AUC 0.876 for language-failures vs never-disputed controls and 0.862 for all escalated disputes vs controls (paired-bootstrap improvement over the prior engine +0.057, 95% CI +0.035 to +0.079); 19.3% of language-failure records rate in the lower-risk tier; 15.5% of never-disputed controls rate elevated — concentrated in language families textually identical to documented failures.
Reproducibility
The same normalized title, rules, designated source, and venue produce the same score under the same methodology version. Rating pages publish the engine version and SHA-256 receipt of that canonical rating input. Rules are retrieved from the complete public fields available from each venue, not a marketing snippet, and Kalshi's series-level designated settlement sources are read into the rating input. No model call is used to produce the deterministic score.
Known limitations
- The calibration is Polymarket/UMA-heavy. Band shares are consistent across the census's two halves, but the cutoffs were selected on the same period; true out-of-time validation remains necessary.
- The language-failure positive class is small (57 records), so sparse detector weights are lift-informed priors, not tight estimates; the calibration should be re-run as the failure census grows.
- A text-only rubric cannot detect every operational or real-world failure. Some failed contracts are textually identical to contracts that settled cleanly — that ceiling is priced as an honest elevated-control share, not hidden.
- A structurally complete but nonsensical contract can earn a mid grade: the floors guarantee garbage never reaches the lower-risk tier, and no more. The rubric judges structure and known hazard patterns, not sense.
- Non-English input is unratable by design — an English-calibrated rubric grading other languages would be noise presented as precision.
- The public export cannot independently reproduce source collection, classification, deduplication, or rules-text regeneration.
- Unretrievable, empty, or non-prose rules text is unratable and should receive human review.
Technical detector details
Hazard weights are 9·log₂ of the measured lift against the control sample, with sign constraints. A safeguard may lower risk only when it addresses a hazard actually present — boilerplate answering no live hazard earns nothing, and adding a genuine safeguard can never raise the score.
| Detector | Points | What it detects |
|---|---|---|
| Unnamed backup source | +42 | An undefined fallback lets settlement move to whichever source appears convenient after the fact. |
| Perceptual predicate | +35 | Resolution turns on a perceptual human act no named official record captures. |
| Speech-act threshold | +35 | A qualitative sufficiency test is used without a binding definition. |
| Open-ended list | +32 | An operative definition uses an unbounded list or carve-out. |
| Self-contradiction | +30 | The same condition is bound to opposite outcomes in different sentences — the contract cannot be executed as written. |
| Observability predicate | +22 | Resolution requires interpreting what is visible in media or footage rather than a recorded official fact. |
| Undefined subject | +20 | The operative subject is an undefined “something” — no checkable proposition exists. |
| Covert discriminator | +19 | Settlement depends on a distinction that no public record resolves. |
| Non-prose padding | +18 | A substantial share of the text does not parse as contract prose — filler surrounds the operative language. |
| Sole-consensus source | +13 | Press consensus is the only primary source, leaving no controlling document to consult. |
| Joint confirmation | +13 | Requires confirmation from both parties — settlement is hostage to the least cooperative party's communications. |
| Source hierarchy without precedence | +11 | Multiple resolution sources with no stated order of precedence — when they disagree, either side can claim its source controls. |
| Equivalent-language clause | +10 | “Equivalent/similar language” invites litigation over paraphrase. |
| Mutual agreement | +9 | Proving both sides' assent is an evidentiary threshold press reports rarely settle cleanly. |
| Completion undefined | +8 | Hinges on whether a match is “completed” without defining retirement, walkover, or default handling. |
| Initial-print rule on an officiated result | +6 | “Corrections after expiration ignored” on an officiated result pays the scorer's error, not the official outcome. |
| Judgment predicate | +5 | The operative clause relies on discretion without an accompanying definition. |
| Decimal threshold without rounding rule | +5 | The threshold sits at the data's own precision with no rounding rule — a one-tick revision flips the outcome. |