Testing by betting when the bets are real
An e-value audit of 84 pre-registered market-efficiency hypotheses
E-values are usually motivated by a metaphorical gambler betting against a null hypothesis. Here the gambler is not a metaphor. The null is a bookmaker's posted price, the stake is a stake, and the wealth process is the e-process. The audit covers 4,141 walk-forward out-of-fold bouts scored by models refit at 32 quarterly origins, 1,787 of which carry an archived closing price, first analysed with paired log-loss and Benjamini-Hochberg and then re-asked as bets.
-
A multiplicity correction that was not licensed
The 84 slices are overlapping subsets of one pool. Membership correlations across the 3,486 segment pairs run from −0.49 to +0.84, half of them negative, and 83 pairs share no bouts at all. Benjamini-Hochberg needs independence or PRDS; neither is established. On the 1,191-bout discovery window the lab actually searched, BH rejects four: one in the model's favour and three against it. Benjamini-Yekutieli and e-BH, whose guarantees do hold here, reject none.
-
"Failed to replicate" and "not yet" are different findings
A segment the fixed-n analysis wrote off at p = 0.30 is still growing on the wealth scale: +0.0143 nats per bout against +0.0174 in the window that selected it, an e-value of 4.81 against a threshold of 20. It is not refuted. It is 100 bouts short at fair odds and 253 at the book's, which is seventeen months, or three and a half years. Optional continuation turns a refutation into a schedule.
-
Pricing the freedom a post-hoc rule used
A threshold rule chosen after the shape was visible, paid for by a mixture martingale over the 21 cuts it searched, clears the threshold at 31.28. Widen the mixture to the 21 cuts it could have searched in the opposite direction and it falls to 15.64. Add the other disagreement statistic it could have used and it falls to 10.43. The ladder is the result, not any rung of it.
At the prices the book actually offered, the same three mixtures pay 8.47, 4.23 and 2.82: a bettor paying the margin never clears the threshold at all.
Before any of it was used, the instrument was validated. Under a true null with prices and stakes held and only outcomes resampled, 40,000 replications give an expected e-value of 0.985 plus or minus 0.018 per segment, with none of the 84 above its own alpha. And the headline result is negative, stated as negative in the abstract: betting the model against the book across the whole priced pool gives e = 3.6 x 10-4 at fair odds and 3.1 x 10-8 at real ones. The paper is not about beating a market. It is about what can be honestly claimed inside a search that mostly failed.
The workshop's organisers include Grunwald and Ramdas, who built the machinery the paper uses.