Prigodskii

Roman Prigodskii Research

Testing by betting

Two sole-authored papers, three submissions, all under review at NeurIPS 2026 workshops. Both grew out of Vertex MMA, and both are about the same discomfort: an evaluation that reports what it found, and not what it could have found.

Author
Sole, on both
Status
Under review
Subject
A deployed forecaster and its own instrument
Data
4,141 walk-forward bouts, 1,787 of them priced
Paper one Under review Decisions due 29 September

Testing by betting when the bets are real

An e-value audit of 84 pre-registered market-efficiency hypotheses

E-values are usually motivated by a metaphorical gambler betting against a null hypothesis. Here the gambler is not a metaphor. The null is a bookmaker's posted price, the stake is a stake, and the wealth process is the e-process. The audit covers 4,141 walk-forward out-of-fold bouts scored by models refit at 32 quarterly origins, 1,787 of which carry an archived closing price, first analysed with paired log-loss and Benjamini-Hochberg and then re-asked as bets.

  1. A multiplicity correction that was not licensed

    The 84 slices are overlapping subsets of one pool. Membership correlations across the 3,486 segment pairs run from −0.49 to +0.84, half of them negative, and 83 pairs share no bouts at all. Benjamini-Hochberg needs independence or PRDS; neither is established. On the 1,191-bout discovery window the lab actually searched, BH rejects four: one in the model's favour and three against it. Benjamini-Yekutieli and e-BH, whose guarantees do hold here, reject none.

  2. "Failed to replicate" and "not yet" are different findings

    A segment the fixed-n analysis wrote off at p = 0.30 is still growing on the wealth scale: +0.0143 nats per bout against +0.0174 in the window that selected it, an e-value of 4.81 against a threshold of 20. It is not refuted. It is 100 bouts short at fair odds and 253 at the book's, which is seventeen months, or three and a half years. Optional continuation turns a refutation into a schedule.

  3. Pricing the freedom a post-hoc rule used

    A threshold rule chosen after the shape was visible, paid for by a mixture martingale over the 21 cuts it searched, clears the threshold at 31.28. Widen the mixture to the 21 cuts it could have searched in the opposite direction and it falls to 15.64. Add the other disagreement statistic it could have used and it falls to 10.43. The ladder is the result, not any rung of it.

The ladder, and the same three mixtures charged the book's margin

At the prices the book actually offered, the same three mixtures pay 8.47, 4.23 and 2.82: a bettor paying the margin never clears the threshold at all.

Before any of it was used, the instrument was validated. Under a true null with prices and stakes held and only outcomes resampled, 40,000 replications give an expected e-value of 0.985 plus or minus 0.018 per segment, with none of the 84 above its own alpha. And the headline result is negative, stated as negative in the abstract: betting the model against the book across the whole priced pool gives e = 3.6 x 10-4 at fair odds and 3.1 x 10-8 at real ones. The paper is not about beating a market. It is about what can be honestly claimed inside a search that mostly failed.

The workshop's organisers include Grunwald and Ramdas, who built the machinery the paper uses.

Paper two Under review Decision due 22 September

Measure the instrument first

Detection floors for model selection and for LLM judges

An evaluation reports that a candidate improved a metric. It almost never reports the smallest improvement it could have detected, and those are different claims. This paper measures the second one with a construction that needs no distributional assumption: the null lever, the same recipe refit under a different random seed. It changes nothing real, so everything it produces is instrument noise.

  1. The noise floor is measurable, and it is not small

    On a 3,087-row walk-forward pool under five seeds, ten null comparisons manufacture effects up to 0.00207 nats. Re-seeding alone reproduces 80% of the only lever this leg ever shipped, and the pipeline had 56% power against its own shipped result.

  2. Buying seeds walks to a floor, not past it

    The floor is an asymptote at 0.00288 nats. Fifteen seeds gets within 2% of it and no budget gets past it, because 63% of the paired noise lives beyond any seed budget. On four degrees of freedom that share is bounded only to [17%, 83%], so it is reported as an estimate and not as a finding.

  3. A floor belongs to the instrument that measured the effect

    Recomputed on the 793-row pool where this project's one large win was actually measured, the floor is 0.0119 rather than 0.00362. The win is 4x its own floor, not 13x another pool's. The paper corrects that error in its own earlier draft, in print.

  4. The protocol transplants

    On a deployed LLM classifier, two identical reruns disagree on 5.6% of items, and twenty of them move measured agreement by 2.6 points. The floor is 2.9 points. Same recipe, different instrument.

Eleven candidates, one candidate scored three ways, and the lever that shipped

The shaded band is the one-sided 80% minimum detectable effect at 0.00362 nats. Every lever inside it is a result the instrument cannot tell apart from re-seeding, including the one that shipped.

All 84 hypotheses of the audit, at the wealth each one earned

Each mark is one pre-registered hypothesis. The horizontal axis is the e-value at fair odds, the vertical axis the same bet charged the bookmaker's margin, both on a log scale. Mark size grows with the number of bouts in the segment. The dotted rules mark e = 20, which rejects at the 5% level for a single pre-specified hypothesis. For a family of 84 the cutoff that is actually licensed is e-BH's, e = 1,680 at one rejection, and nothing here comes near it. The single mark above the horizontal rule is also the most favourable of five walk-forward seeds: its five-seed median is 12.3, below the rule.

Questions, or a reason to build the next one.