Market NewsRuleScoreRuleReviewRuleDraftAboutContact
Evidence & company
Dispute DatabaseMethodologyPolicy & RegulationAPISystem Status

RuleScore Methodology

RuleScore is a deterministic review of event-contract rules. It measures the risk that settlement may become ambiguous or disputed. It does not predict whether the event will happen, decide who should win a dispute, or replace legal review.
Plain-English summary: a good grade must be earned. RuleScore first checks that the text reads as contract prose at all (otherwise it is unratable — no score is published), then requires structural evidence — an operative clause, an outcome binding, a deadline, a checkable source — before a low score is possible, then scores drafting patterns associated with documented settlement failures. It separates language risk from subject-contest risk and publishes a reproducible letter rating from BR-Aaa to BR-Ba.

What RuleScore measures

  • Whether the text reads as executable contract prose at all.
  • Whether the structural minimums exist: an operative clause, an outcome binding, a deadline, a checkable source.
  • Whether the controlling source is named and bound to resolution.
  • Whether operative terms and thresholds are defined.
  • Whether timing, cutoffs, boundaries, rounding, and revisions are clear.
  • Whether contingency and adjudication language creates avoidable discretion.

What it does not measure

  • The probability that the event occurs.
  • The expected value or price of a contract.
  • Whether a venue or trader is legally correct.
  • Whether a structurally complete contract makes real-world sense — the rubric judges structure and known hazard patterns, not meaning.
  • Every operational, oracle, or real-world failure that text cannot reveal.

The three stages

1 · Validity gate

Text statistics (function-word density, letter ratio, repetition, word-length profile) must read as English contract prose. Failing input is unratable: no score, no letter, an explicit reason. Zero of the 2,265 real market texts in the calibration corpus trip this gate.

2 · Structural-evidence floors

Missing structure sets a minimum score — no operative clause, no outcome binding, no deadline, nothing checkable. Floors are removed by adding the missing structure, never by rewording; text can never buy credit by adding words.

3 · Hazards, gated safeguards & the contest prior

Language Risk (LRS): calibrated hazard detectors plus safeguards that earn credit only when they answer a hazard actually present. Contest Prior (CPR): the measured base rate of escalated disputes for the contract's subject family, published beside the language score — subject risk, never a drafting criticism.

An unfalsifiable qualifier is published beside the rating when the text names no checkable record or relies only on press consensus. It is not hidden inside the number.

The rating scale

A worse notch has a higher in-sample escalated-dispute share in the historical calibration corpus (strictly monotone, 1.8% to 48.5%). These band shares were used to choose the cutoffs; they are not prospectively validated failure probabilities. The labels describe settlement clarity, not credit quality or investment merit.

RatingScoreRisk labelRecordsLanguage-failure shareEscalated-dispute share
BR-Aaa0–9Minimal3970.3%1.8%
BR-Aa10–15Low8801.1%5.9%
BR-A16–23Moderate2491.6%7.2%
BR-Baa24–49Elevated5475.3%24.1%
BR-Ba50–100High6619.7%48.5%

Shares are computed on the graded 2026 backtest universe (disputed records + never-disputed controls + live top-volume markets). Contract text that fails the validity gate is unratable and appears in no band.

Structural-evidence floors

A floor is a minimum score set by missing structure. It renders in every rating's risk ledger as an actionable line-item, exactly like a hazard.

Missing evidenceFloorWhy
No operative resolution clause55The text never states how or when it resolves; nothing binds settlement.
No outcome binding30Resolution is mentioned but no condition is bound to a specific outcome (Yes/No/named result).
No deadline evidence20No date, expiry reference, or time anchor — an open-ended contract cannot settle No.
Nothing checkable (anchor tier 0)18 / 14No bound source, official record, or URL; 18 for bespoke text, 14 for machine templates, plus a +12 tier penalty. Historically the proposer wins ~75–86% of disputes on such contracts by default.
Press-anchored geopolitics or mention market26Press consensus (or nothing) as the only anchor in the two highest-dispute families (36.6% / 32.7% escalation rates).

Evidence base

The methodology was calibrated against a generated 2026 corpus containing 1,595 disputed-market records plus 672 never-disputed high-volume control records and 95 live top-volume markets (2,265 graded texts in all). Brierly classifies 80 census records as hard-failure outcomes: 44 50/50 resolutions and 36 escalated overturns. For calibration, 50/50 records whose rules contain an explicit 50-50 void clause are excluded from the positive class — those are correctly drafted contracts whose stated contingency executed — leaving 57 language-failure records, alongside 184 contested-but-held records. The detailed case studies in the Dispute Database are sourced explanatory examples, not the entire census.

RuleScore rated 1,498 records; 97 lacked retrievable rules text. The committed export is text-free, derived from a private source corpus, and includes one apparent duplicate, so it does not independently establish 1,595 unique markets. On this corpus, in sample: AUC 0.876 for language-failures vs never-disputed controls and 0.862 for all escalated disputes vs controls (paired-bootstrap improvement over the prior engine +0.057, 95% CI +0.035 to +0.079); 19.3% of language-failure records rate in the lower-risk tier; 15.5% of never-disputed controls rate elevated — concentrated in language families textually identical to documented failures.

Reproducibility

The same normalized title, rules, designated source, and venue produce the same score under the same methodology version. Rating pages publish the engine version and SHA-256 receipt of that canonical rating input. Rules are retrieved from the complete public fields available from each venue, not a marketing snippet, and Kalshi's series-level designated settlement sources are read into the rating input. No model call is used to produce the deterministic score.

Known limitations

  • The calibration is Polymarket/UMA-heavy. Band shares are consistent across the census's two halves, but the cutoffs were selected on the same period; true out-of-time validation remains necessary.
  • The language-failure positive class is small (57 records), so sparse detector weights are lift-informed priors, not tight estimates; the calibration should be re-run as the failure census grows.
  • A text-only rubric cannot detect every operational or real-world failure. Some failed contracts are textually identical to contracts that settled cleanly — that ceiling is priced as an honest elevated-control share, not hidden.
  • A structurally complete but nonsensical contract can earn a mid grade: the floors guarantee garbage never reaches the lower-risk tier, and no more. The rubric judges structure and known hazard patterns, not sense.
  • Non-English input is unratable by design — an English-calibrated rubric grading other languages would be noise presented as precision.
  • The public export cannot independently reproduce source collection, classification, deduplication, or rules-text regeneration.
  • Unretrievable, empty, or non-prose rules text is unratable and should receive human review.
Technical detector details

Hazard weights are 9·log₂ of the measured lift against the control sample, with sign constraints. A safeguard may lower risk only when it addresses a hazard actually present — boilerplate answering no live hazard earns nothing, and adding a genuine safeguard can never raise the score.

DetectorPointsWhat it detects
Unnamed backup source+42An undefined fallback lets settlement move to whichever source appears convenient after the fact.
Perceptual predicate+35Resolution turns on a perceptual human act no named official record captures.
Speech-act threshold+35A qualitative sufficiency test is used without a binding definition.
Open-ended list+32An operative definition uses an unbounded list or carve-out.
Self-contradiction+30The same condition is bound to opposite outcomes in different sentences — the contract cannot be executed as written.
Observability predicate+22Resolution requires interpreting what is visible in media or footage rather than a recorded official fact.
Undefined subject+20The operative subject is an undefined “something” — no checkable proposition exists.
Covert discriminator+19Settlement depends on a distinction that no public record resolves.
Non-prose padding+18A substantial share of the text does not parse as contract prose — filler surrounds the operative language.
Sole-consensus source+13Press consensus is the only primary source, leaving no controlling document to consult.
Joint confirmation+13Requires confirmation from both parties — settlement is hostage to the least cooperative party's communications.
Source hierarchy without precedence+11Multiple resolution sources with no stated order of precedence — when they disagree, either side can claim its source controls.
Equivalent-language clause+10“Equivalent/similar language” invites litigation over paraphrase.
Mutual agreement+9Proving both sides' assent is an evidentiary threshold press reports rarely settle cleanly.
Completion undefined+8Hinges on whether a match is “completed” without defining retirement, walkover, or default handling.
Initial-print rule on an officiated result+6“Corrections after expiration ignored” on an officiated result pays the scorer's error, not the official outcome.
Judgment predicate+5The operative clause relies on discretion without an accompanying definition.
Decimal threshold without rounding rule+5The threshold sits at the data's own precision with no rounding rule — a one-tick revision flips the outcome.
RuleScore ratings are independent editorial opinions about settlement risk. They are not event predictions, trading recommendations, legal advice, or guarantees of a dispute-free settlement.