Skip to content

Part XV — Multiple Testing

Part XIV showed that a search is part of the estimator, so every statistic computed after one was computed as though it had not happened, and it closed by naming what it had not done: turn the count of candidates into an explicit price, so that a reported statistic could be corrected rather than merely quarantined. This part does that, and the arithmetic turns out to be the easy half. Each correction below is exact, provable, and takes a single input — the number of hypotheses examined — which is the one quantity known with certainty at the moment a research process begins and unrecoverable from any dataset afterwards. The part therefore divides at its midpoint. The first three pages price a search whose width is assumed known and measure what the price is. The fourth establishes that it is not known. The last two stop asking for it and resample the family instead.

The dependencies run in file order and the first page is the vocabulary for the five that follow. Multiple Comparisons establishes the two quantities a family makes available — an expected count that no dependence structure can move and a probability that dependence moves by a factor of twenty — and the \(2\times2\) table on which five defensible error rates can be defined; choosing among them is the decision that Bonferroni Correction and False Discovery Rate then execute, the first pinning the count of false rejections and the second their share, so those two pages are one choice read twice rather than a strict and a lenient version of one procedure. Data Snooping Bias is the hinge: it removes the input all three assumed, and everything after it is a response to that removal. White's Reality Check and Hansen's SPA Test are a pair, the second repairing a defect the first measures, and reading either alone misses that the repair engages only on families the first page's failure does not describe. Where this part stops is worth stating too: it derives no sampling distribution for a maximum and no expected best of \(N\) nulls, which is Sampling Distributions; it inverts no power function into a span of calendar time, which is Statistical Power; it builds no resampling scheme from first principles and does not treat the bootstrap's inconsistency for a sample extreme, which is Bootstrap Methods; it does not establish the modification a resampling scheme needs before it can test rather than estimate, which is Bootstrap Tests; it corrects nothing for repeated looks at one hypothesis, which is the optional stopping of Martingales; and it assigns no probability to a hypothesis and computes no posterior over how many candidates could plausibly be real, which is Part XVI.

One failure runs through the part and it is never an error of arithmetic: every procedure below does exactly what its derivation promises, and is then handed an input describing the write-up rather than the search, or reports an average over runs that do not occur. An expected count of false rejections holds at \(2.4990\), \(2.4979\) and \(2.4702\) against a prediction of \(2.50\) made without looking at the dependence, while the probability of at least one runs \(0.9234\), \(0.6209\), \(0.1917\) over the identical draws and the textbook formula reads \(0.9231\) throughout. Two published shortcuts for a correlated family's effective size return \(37.99\) and \(26.00\) where the honest answer is \(19.10\), and return those same two numbers at a level where it has become \(22.79\). Bonferroni delivers a family-wise rate of \(0.0044\) when asked for \(0.05\), and the refinement everybody debates buys \(0.0021\) of power where calibrating to the family's actual dependence buys \(18.4\) points. Correcting fifty variants raises the Sharpe a strategy must truly possess to \(0.8031\) and drops the power against a genuine \(0.30\) edge to \(0.0525\). The false discovery rate is controlled at \(0.1029\) on a family where \(84.18\%\) of realizations have a false discovery proportion of exactly zero and \(5\%\) exceed \(0.7500\), and rising dependence drives its average down to \(0.0387\) while the worst realization rises to \(1.0000\). An analyst who stops at the first success and honestly reports the trial count achieves \(0.1579\) against a nominal \(0.05\); one who admits to fifty variants out of a thousand achieves \(0.6326\). The diagnostic built to detect all of this returns \(0.5046\) on average and anything from \(0.0305\) to \(0.9831\) on a single family with nothing whatever to overfit. A champion with a true Sharpe of \(0.80\), conclusive alone, is rejected on \(0.7333\) of occasions once ninety-nine candidates that could not possibly win are placed beside it. And the test built to repair that recovers all of it when the ninety-nine lose money and \(-0.0133\) of it when they merely break even, which is what a parameter grid contains.

Topics

Topic Focus
Multiple Comparisons The expected number of false rejections as \(m_0\alpha\) under arbitrary dependence against a family-wise probability confined only to \([\alpha,\min(1,m_0\alpha)]\), measured at \(2.4990\), \(2.4979\), \(2.4702\) against \(2.50\) while the probability ran \(0.9234\), \(0.6209\), \(0.1917\) and the independence formula held at \(0.9231\), reading \(1.0000\) where two hundred correlated tests deliver \(0.2547\); dependence leaving the count's mean untouched and inflating its spread \(10.6\)-fold; a correlated family's effective size as a function of the level as well as the correlation, \(19.10\) at \(\alpha=0.05\) and \(22.79\) at \(\alpha=0.01\) for one family at \(\rho=0.5\), against Cheverud–Nyholt's \(37.99\) and Li–Ji's \(26.00\) at both, since a function of the correlation matrix alone cannot depend on a threshold it never sees; and five defensible error rates on identical draws reading \(0.044999\), \(1.0000\), \(1.0000\), \(0.3445\) and \(1.0000\), with the exceedance rate refusing to fall over three successive tightenings because a shrinking threshold shrinks the denominator as fast as the numerator
Bonferroni Correction The union bound as an axiom rather than an assumption, giving a family-wise guarantee under every joint distribution, and Šidák's exact independent threshold exceeding it by a ratio tending to \(1+\alpha/2\)\(1.0206\), \(1.0253\), \(1.0259\) as \(m\) runs \(5\) to \(5000\), so the gap closes absolutely as the family grows; the guarantee kept with increasing waste as dependence rises, the realized rate falling \(0.0439\), \(0.0423\), \(0.0289\), \(0.0128\), \(0.0044\) while a calibrated threshold held at \(0.05\) throughout and the procedure's own power sat unmoved at \(0.760\); Šidák's refinement buying \(0.0021\) of power against calibration's \(18.4\) points at \(\rho=0.95\); Holm rejecting a superset of Bonferroni's hypotheses on every realization with \(0\) counterexamples in a million draws for a gain of half a point, and Hochberg's extra dependence assumption yielding a strictly larger set on \(0.00003\) to \(0.00012\) of draws; and the corrected bar over twenty-four years demanding a true Sharpe of \(0.5077\), \(0.8031\) and \(1.0747\) at \(m=1\), \(50\) and \(10{,}000\) while power against a genuine \(0.30\) edge collapses \(0.4304\), \(0.0525\), \(0.0016\)
False Discovery Rate The expectation of a ratio with a random denominator rather than a weaker bound on a probability, controlled at \(m_0q/m\) with the bound attained — \(0.1029\), \(0.0978\), \(0.0987\), \(0.0901\), \(0.0700\) against \(0.0998\), \(0.0995\), \(0.0980\), \(0.0900\), \(0.0700\) — and coinciding with the family-wise rate exactly when every null is true; \(0.7049\) of real effects found against Bonferroni's \(0.1963\), bought with a family-wise rate of \(0.9978\) against \(0.0464\); the controlled average describing no realization as discoveries thin, a family with two real effects returning a mean of \(0.1029\) while \(84.18\%\) of runs have a proportion of exactly zero, half discover nothing, and the ninety-fifth percentile is \(0.7500\); dependence moving mean and tail in opposite directions, the average falling \(0.0897\) to \(0.0387\) while the worst realization rises \(0.2330\) to \(1.0000\) and the median collapses to zero; the distribution-free repair costing a harmonic \(7.4855\) that makes its bar for the best single test \(7.49\) times stricter than Bonferroni's while keeping \(56.5\%\) of the power; and Storey's estimator reading \(0.9873\) against a truth of \(1.00\) and reclaiming \(0.00\) extra discoveries
Data Snooping Bias Selection bias as the expected maximum under the null, the closed form predicting \(0.3443\) to \(1.0124\) against measured maxima of \(0.3359\) to \(1.0099\) while every one of those champions was worth zero to three decimals out of sample, and overstating by \(0.16\) to \(0.46\) of a Sharpe on a family correlated at \(0.7\) so that the deflated Sharpe ratio sets double the honest hurdle on the grids it is used for; a wide search changing the winner's identity rather than only its statistic, the probability that a genuine \(0.50\) Sharpe wins its own selection falling \(1.0000\), \(0.7160\), \(0.4283\), \(0.2392\), \(0.1106\), \(0.0448\) while the reported figure climbed to \(1.0152\) and the selection's true value fell to \(0.0222\); a correction applied honestly to a stopping rule delivering \(0.2555\), \(0.1579\), \(0.0462\) against nominal \(0.10\), \(0.05\), \(0.01\), and a write-up admitting to fifty variants of a thousand delivering \(0.6326\); and the overfitting probability unbiased at \(0.5046\) across families and carrying a standard deviation of \(0.1702\) within them, ranging \(0.0305\) to \(0.9831\) on data with nothing to overfit, because \(12{,}870\) splits of sixteen blocks are a split count where a reader sees a sample size
White's Reality Check A composite null over the whole family whose least favourable configuration is the origin, imposed by recentring each candidate on its own mean while its variances and covariances survive, so the correction is read from the family's joint behaviour instead of assumed about it; the resampling scheme deciding whether the test works at all, an iid bootstrap rejecting a true null \(0.3550\) and \(0.6225\) of the time on AR(1) data where stationary block schemes held at \(0.0725\), \(0.0700\), \(0.0875\), \(0.0625\) and all three agreed at \(0.0525\) on independent data — twelvefold damage against the threefold a single mean suffers, because a maximum compounds twenty understated variances; exact size at \(0.0500\) where Bonferroni and Holm deliver \(0.0300\), converted into detection rates of \(0.1100\), \(0.2667\), \(0.5233\), \(0.7633\) against \(0.0567\), \(0.1733\), \(0.3967\), \(0.6367\); and a candidate's contribution depending on its second moments alone, so a champion at a true Sharpe of \(0.80\) conclusive at \(0.0023\) decays to \(0.0529\) with ninety-nine worthless siblings and to \(0.0882\) with ninety-nine losing a Sharpe point a year — the same damage from candidates that could not conceivably have won
Hansen's SPA Test Studentization stopping the maximum from being won on volatility rather than evidence, and a conditional recentring whose threshold needs \(A_n\to\infty\) with \(A_n/\sqrt n\to0\) — two limits admitting infinitely many sequences and no data, so the honest output is a bracket; the bracket's upper end reproducing the Reality Check at \(0.6817\), \(0.3740\), \(0.0426\), \(0.0026\) against \(0.6967\), \(0.3759\), \(0.0466\), \(0.0027\), so the test can never be the weaker one; the consistent p-value migrating across that bracket as the horizon grows, \(0.5511\) in \([0.2835,0.6817]\) at five years and \(0.0008\) in \([0.0008,0.0426]\) at twenty-five while the discard threshold shrinks only from a Sharpe of \(-0.89\) to \(-0.42\); the repair total on a family of loss-makers, holding at \(0.0015\), \(0.0012\), \(0.0016\), \(0.0017\) and rejecting \(1.0000\) of the time where White fell to \(0.7333\), discarding \(0.9879\) of the siblings; and the repair vanishing as they approach break-even, the gain running \(+0.1333\), \(+0.0933\), \(+0.0333\), \(+0.0333\), \(-0.0133\) while the discarded fraction falls \(0.9885\) to \(0.0199\), because a candidate with a true mean of zero belongs to the null and must be retained