Skip to content

Part XII — Hypothesis Testing

Part XI built estimators and error bars for them and closed by observing that every guarantee any of them carries is conditional on a description of the world asserted before the calculation and never checked afterwards. This part rearranges the quantifiers rather than the machinery. The question stops being how precisely a number can be pinned down and becomes whether it differs from something named in advance, the output collapses from an interval to one bit, and the same sampling distributions do the same work — which is why the fourth page of this part proves that a confidence interval and a family of tests are one object seen from two sides. What the rearrangement adds is a second way to be wrong, and a vocabulary for pricing it: a level, a size, a power, a p-value. What it discovers under that vocabulary is that the arithmetic is almost never the problem, that the reference distribution is almost always correct for some question, and that nothing in any output records whether that question was the one being asked.

The dependencies run in file order and the first five pages are one argument split five ways, because they were one page and separating them is what makes each nameable. The Hypothesis Testing Framework owns the map — a hypothesis as a subset, a test as a function to \(\{0,1\}\), a size as a supremum — and everything after it spends that vocabulary; Test Statistics owns what goes into the map and p-values what comes out of it, so those three read as a chain. Type I and Type II Errors and Statistical Power are a pair — the first owns the two rates and the frontier between them, the second owns one of those rates read as a function of the alternative — and reading either alone misses that \(\alpha\) is calibratable and \(\beta\) is not, which is the fact that motivates the last two pages. Likelihood Ratio Tests is the one optimality theorem in the part and is prerequisite to nothing after it. The final four are two pairs of test families, ordered by how much they assume: Parametric Tests and Nonparametric Tests assume a family and a group respectively, Permutation Tests and Bootstrap Tests generate the null instead of looking it up, and the last is the first without a group to stand on. Where this part stops is worth stating too: no maximum is corrected for the size of the search that found it and no family of strategies is accounted for, which is Part XV; nothing is regressed on anything, which is Part XIII; no non-nested models are ranked and no complexity penalized, which is Part XIV; no Bayes factor is computed and no posterior probability assigned to a hypothesis, which is Part XVI; no resampling scheme is built from scratch and no plug-in theorem proved, which is Bootstrap Methods; no interval is constructed, which is Confidence Intervals; and no limit theorem is proved anywhere, because Part VII proved them and this part spends them.

One failure runs through the part and it has a single shape: the arithmetic is correct, the reference distribution is the correct distribution for some problem, and the problem is not the one on the desk. A pooled proportion test verified at \(p=0.500\) and reading \(0.0517\) has a true size of \(0.0834\) attained at \(p=0.009\). A sample mean of Cauchy data has power \(0.0624\) at ten observations and \(0.0557\) at ten thousand, because its null \(95\)th percentile only moves from \(6.430\) to \(6.176\). A genuine Sharpe-\(0.30\) edge tested on twenty-four years produces a median p-value of \(0.1375\) and a \(5\)th-to-\(95\)th range of \(0.00176\) to \(0.8555\). Repairing a size from \(0.1985\) to \(0.0500\) discards \(63.8\%\) of the apparent power and no report shows the second row. The course's own weekday tests carried \(8.20\%\) power against the effect they were hunting, and closing that gap would take \(661\) years of Mondays. Three standard statistics for one hypothesis read \(2.09\), \(4.95\) and \(3.56\) against a critical value of \(3.8415\). A nominal \(5\%\) \(F\)-test fires on \(44.06\%\) of samples whose variances are identical by construction. A sign test rejects at \(7.09\times10^{-9}\) and reports that a profitable strategy loses money. An iid shuffle builds a null \(18\%\) too narrow and rejects \(9.67\%\) of strategies that have no edge at all. And an iid bootstrap is perfectly sound on a GARCH series at \(0.0600\) and wrong by a factor of three on an AR(1) at \(0.1520\). In every case the code is right, the p-value is a valid answer, and the gap between the question answered and the question asked appears in no field of the output.

Topics

Topic Focus
The Hypothesis Testing Framework A hypothesis as a subset of a model and a test as a function to \(\{0,1\}\) whose rejection region is the primitive, the size as a supremum over the whole null set so a pooled proportion test verified at \(p=0.500\) has a true size of \(0.0834\) at \(p=0.009\), the asymmetry that lets a test reject and never accept because only one hypothesis is computable, the duality returning a confidence set to an endpoint disagreement of \(0.0\) while non-uniqueness leaves the equal-tailed variance interval \(8.06\%\) wider than necessary at \(n=21\), and a valid \(\chi^{2}\) volatility test whose power sits below its own size throughout \([0.843,\,1.000]\) and detects a halving \(25.28\%\) of the time against an unbiased twin's \(34.24\%\)
Test Statistics A computable null law and a stochastic ordering as the only two requirements, pivotality as what makes a critical value a constant so a raw difference calibrated at one volatility has size \(0.0001\) at half of it and \(0.6184\) at four times while its studentized twin holds \(0.0484\) to \(0.0494\), monotone likelihood ratio as a property of the family rather than the statistic, the Cauchy sample mean carrying power \(0.0624\) at \(n=10\) and \(0.0557\) at \(n=10{,}000\) against a median reaching \(1.0000\) by \(n=100\), and a four-by-five grid of standard statistics against standard departures that is \(0.05\) almost everywhere off the diagonal
p-values The smallest level at which the data rejects, hence a statistic computed entirely under the null, super-uniformity at every level simultaneously as the actual validity requirement with uniformity the case where it binds, a lattice null making a nominal \(5\%\) binomial test size \(0.0107\) at \(n=10\) while the mid-\(p\) repair reads \(0.0547\) and is no longer level \(0.05\), a genuine Sharpe-\(0.30\) edge giving a median p-value of \(0.1375\) against the lesson's realized \(0.135\) with a \(5\)th-to-\(95\)th spread of \(0.00176\) to \(0.8555\) and a \(31.50\%\) chance of clearing the bar, and four findings whose ordering by p-value exactly reverses their ordering by effect size because only \(n\delta^{2}\) is identified
Type I and Type II Errors Two expectations under two different measures that combine into no accuracy without a prior nobody supplies, one convex frontier generated by the likelihood ratio on which a true Sharpe of \(0.30\) over twenty-four years buys \(43.04\%\) power at the conventional level and would need \(\alpha=0.2650\) for \(80\%\), the cost-minimizing level as a likelihood-ratio threshold set by the cost ratio and prior odds running from \(0.001\) to \(0.600\) with \(0\) of \(15\) cells near \(0.05\), the different currencies a desk pays for the two errors, and a size repaired from \(0.1985\) to \(0.0500\) that silently discards \(63.8\%\) of the apparent power
Statistical Power Power as a function on the alternative whose infimum is the size, so quoting it as a number fixes an undisclosed effect, the inversion cancelling the volatility for a Sharpe ratio and returning \(68.7\) years of daily data for \(80\%\) power on the course's own \(0.30\) against the twenty-four it has, an audit of the course's five weekday tests finding \(8.20\%\) to \(9.22\%\) power against a tradeable \(2\) bp effect and \(34{,}356\) Mondays needed to close it, observed power as a strictly decreasing function of the p-value at a measured rank correlation of \(-1.0000\), and a filtered genuine \(0.30\) Sharpe reported as \(1.1708\) at three years with \(8.07\%\) of significant findings carrying the wrong sign
Likelihood Ratio Tests The one optimality theorem in the part, worth \(0.8037\) power where a trimmed mean gets \(0.7588\) and a valid one-observation test \(0.0677\), generalization by two maxima keeping asymptotic optimality and losing the finite-sample guarantee, Wilks calibrating by projecting a Gaussian onto a subspace so the family cancels, the Wald, score and likelihood-ratio statistics reproducing the course's Kupiec p 9.51e-01 and its underflow to \(0.00\) from a \(2.287\times10^{-18}\) that 1 - cdf cannot represent while reading \(2.09\), \(4.95\) and \(3.56\) at \(250\) observations, and a boundary null putting \(60.52\%\) of its mass at exactly zero so the tabled cutoff delivers size \(0.0156\)
Parametric Tests Every test a statistic divided by an estimate of its own standard error, so two failures are possible and only one has a limit theorem, the pooled two-sample \(t\)-test reading size \(0.2819\) or \(0.0006\) according to which group was labelled first and \(0.3415\) against \(0.0001\) at a wider ratio, Welch holding \(0.0490\) to \(0.0506\) throughout at degrees of freedom falling from \(298\) to \(49.8\), the \(F\)-test climbing monotonically in kurtosis to \(0.4406\) at the course's measured \(11.4\) against a delta-method prediction of \(45\%\) while Brown–Forsythe holds \(0.0497\), and a self-standardized $
Nonparametric Tests The rank vector uniform on the symmetric group under exchangeability so the null law depends on \(n\) alone, the transform changing the functional as well as the reference so on the course's own trade log the \(t\)-test rejects at \(7.62\times10^{-3}\) in the profitable direction while the sign test rejects at \(7.09\times10^{-9}\) and always reports the strategy losing money, a sample-size ratio of \(1.058\) at the normal against the theoretical \(\pi/3\) repaid at \(0.638\) under \(t(3)\), Pearson's correlation test carrying size \(0.6657\) on independent series with \(2\%\) joint outliers where Spearman holds \(0.0533\), and Mann–Whitney tracking the \(t\)-test to \(0.2023\) against \(0.2105\) under dependence because ranks drop the assumption that was not binding
Permutation Tests A group of transformations under which the null leaves the law invariant, the observed statistic's rank among its own transformed values uniform by the group property so the size holds at \(0.0498\), \(0.0503\), \(0.0503\) and \(0.0483\) across a normal, a Cauchy mixture, a lognormal and an exponential where the pooled \(t\) wanders to \(0.0260\), the group encoding the hypothesis so the shift null isolates the \(0.07\) of the course's \(0.30\) Sharpe that timing contributes against a null mean of \(+0.23\), relabelling invariant only under equality of distributions so the raw-difference test reads \(0.0015\) and \(0.2535\) under unequal variances, and an iid shuffle building a null \(18\%\) too narrow that rejects \(9.67\%\) of edgeless strategies where a circular shift holds \(0.0433\)
Bootstrap Tests Resampling estimating the law of an error rather than of a statistic under a null, so the raw tail proportion is the percentile interval inverted and all three constructions sit at twice nominal on a skewed statistic at \(0.1025\), \(0.1005\) and \(0.1100\), studentizing buying an order of accuracy visible as \(0.0795\) against \(0.0675\) at \(n=21\) and invisible by \(n=126\), the iid bootstrap legitimate for the mean of a GARCH series at \(0.0600\) because its dependence lives in the squares and wrong by three times on an AR(1) at \(0.1520\) where a stationary bootstrap restores \(0.0520\), and the one failure no resampling repairs, since a statistic chosen after looking at the data has its selection priced nowhere in the sample