Part XVI — Bayesian Statistics¶
Part XV corrected statistics for the width of the search that produced them, and closed by naming the assumption none of its six pages had examined: that a hypothesis is a thing one rejects or fails to reject. It pointed at a treatment that instead assigns a probability to the hypothesis itself, with the number of candidates entering as a prior over how many could plausibly be real rather than as a divisor. This part is that treatment, and it delivers on the promise more completely than expected in one respect and not at all in another. The protection really is structural rather than procedural: a marginal likelihood ratio is a martingale under the null, so looking after every one of two thousand trials drives a p-value below \(0.05\) on \(0.4670\) of honest runs while the Bayes factor exceeds \(19\) on \(0.0410\) against Ville's bound of \(0.0526\), and no correction was applied to obtain that. What arrives with it is a new input, unrecorded and consequential, and the part is largely an account of where that input does its damage.
The dependencies run in file order and the first page is the vocabulary for the six that follow. The Bayesian Framework establishes what the disagreement actually is — which quantities may carry a distribution, not how conditioning works — and proves that exchangeability forces a prior to exist whether or not anybody writes one down, which makes Prior Distributions a question about which rather than whether. Posterior Distributions is the hinge: it establishes what a posterior is worth when the model is right and what it is worth when the model is wrong, and the second half is the assumption every later page inherits. Conjugate Priors and Bayesian Updating are a pair, the first supplying the closed forms and the second applying them one observation at a time, and reading either alone misses that the closure making the update cheap is also what fixes the tail behaviour deciding every conflict. Bayesian Model Comparison and Bayesian Prediction are the two things the machinery is built to produce, and both turn out to be exact calculations resting on choices made for convenience. Where this part stops is worth stating too: it takes no posterior mode and minimizes no expected loss, which are Maximum A Posteriori Estimation and Bayesian Estimation; it quotes no credible interval beside a confidence interval and proves nothing about coverage, which is Confidence Intervals; it does not construct the exponential-family conjugate prior from scratch, which is Exponential Families; it derives neither the Laplace approximation to the evidence nor the M-open distinction, which are Information Criteria and Model Averaging; it builds no chain whose stationary distribution is a posterior and runs no expectation-maximization, which is Part XVII; and it applies none of this to a live signal, a portfolio or an option, which is Part XVIII.
One failure runs through the part and it is never an error of arithmetic: every calculation below is exact, and every one of them is exact conditional on a prior that was chosen and a model class that was never checked, neither of which appears anywhere in the output. Three analysts who each decline to assume anything about a volatility, working in three coordinate systems, report posterior medians of \(1.1575\), \(1.3787\) and \(1.0166\) times the truth and risk-limit probabilities of \(0.2921\), \(0.4257\) and \(0.1952\). A prior a course lesson calls agnostic turns out to assert a median absolute Sharpe of \(10.7282\), an \(0.8498\) chance of exceeding three, and central annual outcomes from \(-589.13\%\) to \(+583.96\%\), where a prior twenty times tighter says \(0.7475\). A posterior narrows at the same rate whether it is approaching the truth or a projection of it: serially correlated days read as independent give a posterior standard deviation \(2.032\) times too small and coverage of \(0.6626\), while a Gaussian model asked for a one-per-cent daily loss on Student-\(t\) returns answers \(-2.2823\%\) against a truth of \(-2.6216\%\) with a posterior standard deviation of \(0.0847\%\) — a bias four times its own stated uncertainty — and covers on \(0.1542\) of runs while looking exactly as confident as the correct case that covers on \(0.9480\). A conjugate normal prior meeting an observation twenty standard errors away holds the answer \(2.0000\) standard errors off the data and reports the same \(0.9487\) it reports when nothing is wrong, where a Student-\(t\) prior of identical scale concedes, its pull peaking at \(0.3778\) and falling to \(0.1893\). Updating daily on overlapping twenty-day windows, on data with no serial correlation at all, performs \(981\) updates and reports a posterior standard deviation of \(0.7139\) basis points against a true \(3.1810\), an honest effective sample size of \(49.4\) and coverage of \(0.3388\). At a p-value pinned to exactly \(0.05\) a Bayes factor reads \(2.0790\) for the alternative or \(0.0341\) against it — \(29.30\) to one for the null — on identical data, and the most favourable alternative in existence caps the evidence at \(6.83\) while the honest bound of \(2.4560\) makes \(p=0.05\) a posterior probability of \(0.7107\) rather than anything like \(0.95\). And the predictive that repairs all the parameter uncertainty there is supplies \(1.0085\) of width where a heavy-tailed truth needed \(1.1389\), breaching its one-per-cent limit on \(0.0165\) of days while every calculation inside it is correct.
Topics¶
| Topic | Focus |
|---|---|
| The Bayesian Framework | The four textbook cases of Bayes' rule as one statement about densities against a dominating measure, computed by one normalising line that is never told which case it is in, with two-point and grid formulations of one question rising together from a prior of \(0.30\) to \(0.9860\) and \(0.9998\) as the sample runs \(250\) days to \(16{,}000\); conditioning on a win count rather than a return series costing exactly \(1-2/\pi\) of the Fisher information, the posterior variance ratio measuring \(1.5762\), \(1.5728\), \(1.5720\) and \(1.5718\) against a predicted \(\pi/2=1.5708\); the likelihood principle exactly, nine wins in twelve trades giving a posterior mean of \(0.7142857143\) and \(P(\theta>1/2)=0.9538556763\) under both a fixed and an inverse design while the p-value reads \(0.0730\) under one and \(0.0327\) under the other; Ville's inequality holding under optional stopping at \(0.0410\) and \(0.0075\) against bounds of \(0.0526\) and \(0.0101\) where the p-value fires on \(0.4670\); and de Finetti's mixing measure recovered at \(0.2009\), \(0.1071\), \(0.0488\), \(0.0156\) and \(0.0015\) from data that never named it, its neglect inflating the variance of a trade count \(34.55\)-fold |
| Prior Distributions | Flatness surviving only affine reparameterization, so uniformity commits to a coordinate system rather than expressing ignorance, against Jeffreys' \(\sqrt{\det I(\theta)}\) which transforms exactly as a density and selects one measure in every chart; three analysts declining to assume anything reporting posterior medians of \(1.1575\), \(1.3787\) and \(1.0166\) times the truth and risk-limit probabilities of \(0.2921\), \(0.4257\) and \(0.1952\), the gap against Jeffreys falling \(+35.62\%\), \(+12.78\%\), \(+3.59\%\), \(+1.02\%\) as \(n\) runs \(5\) to \(100\) so the coordinate problem is a small-sample problem and small samples are where a new strategy lives; a prior read in the units it constrains rather than the units it was written in, the lesson's agnostic \(N(0,100\text{bp})\) implying a median absolute Sharpe of \(10.7282\), probabilities of \(0.9502\), \(0.8498\) and \(0.5302\) of exceeding one, three and ten, and annual outcomes from \(-589.13\%\) to \(+583.96\%\) where three basis points gives \(0.7475\); and a hierarchical prior estimating its own width, costing \(1.8540\) times the best fixed rule at the boundary because a truncated moment estimator returns \(0.9339\) instead of \(1.0000\), and repaying it at \(0.5936\) once the family spreads |
| Posterior Distributions | The posterior available immediately up to a constant and the constant the only expensive step, so the mode is free and every expectation, probability and quantile is not; quadrature over \(361{,}201\) evaluations and self-normalised importance sampling agreeing on a non-conjugate two-parameter posterior at \(5.5687\) against \(5.5682\) basis points and log normalisers of \(817.3645\) against \(817.3662\), at an effective sample size ratio of \(0.114\) that decays geometrically with dimension; Bernstein–von Mises measured, total variation falling \(0.0293\), \(0.0082\), \(0.0015\), \(0.0004\) from \(n=25\) to \(10{,}000\) while a prior worth \(120\) pseudo-observations still moves the posterior mean \(42.63\) basis points at \(n=1000\); and the same theorem replaced under misspecification by one of identical shape and different variance, dependence giving a posterior standard deviation \(2.032\) times too small with coverage of \(0.6626\) restored to \(0.9458\) only by importing the long-run variance, and a Gaussian fitted to Student-\(t\) returns putting the one-per-cent daily loss at \(-2.2823\%\) against \(-2.6216\%\) with a posterior standard deviation flat at \(0.0847\%\) across models whose coverage runs \(0.1542\) to \(0.9480\) |
| Conjugate Priors | Conjugacy as the closure property of the exponential family rather than an algebraic accident, the update adding sufficient statistics and incrementing a count so hyperparameters read as pseudo-observations and prior strength becomes commensurable with sample size; Diaconis–Ylvisaker linearity characterising the family and fixing the shrinkage weight before any data arrives; the catalogue's two least glamorous entries carrying most of the value, normal–inverse-gamma raising the critical value \(1.9600\) to \(2.7764\) and coverage \(0.8765\) to \(0.9495\) at five observations, and one Dirichlet pseudo-count converting a \(-\infty\) log-likelihood into \(-2.3979\) on a transition an eight-visit MLE calls impossible; a mixture of conjugate priors remaining conjugate with weights updated by marginal likelihood, retaining \(0.1121\) of an observed two-basis-point edge and \(0.9569\) of a thirty-two-basis-point one where a single normal of identical total variance retains \(0.7006\) of both; and the tail deciding every conflict, a normal prior's pull growing \(0.1000\) to \(2.0000\) standard errors at a fixed posterior standard deviation of \(0.9487\) while a Student-\(t\) prior's peaks at \(0.3778\), falls to \(0.1893\) and widens to \(1.0091\) as it concedes |
| Bayesian Updating | Sequential updating reproducing the batch posterior exactly at every step and invariant to order, both following from the likelihood's factorization rather than from Bayes' rule, with batch, chronological and shuffled updates over four hundred days agreeing to \(2.060\times10^{-18}\); the posterior mean a martingale, its expected revision measuring \(0.0005\), \(-0.0032\), \(0.0001\) and \(-0.0027\) basis points with correlations to its own level of \(-0.0042\), \(0.0005\), \(-0.0016\) and \(-0.0019\), so a belief cannot forecast the direction of its own revision, and the variance decomposition holding at \(15.9990\) to \(15.9842\) against a prior variance of \(16.0000\); a daily update on overlapping twenty-day windows destroying all of it on data with no serial correlation whatever, \(981\) updates reporting a posterior standard deviation of \(0.7139\) basis points against a true \(3.1810\), an honest effective sample size of \(49.4\) and coverage of \(0.3388\), worsening to \(7.663\), \(16.0\) and \(0.2013\) at sixty days; and exponential discounting buying an effective window \(1/(1-\lambda)\) at a permanent floor of \(\sigma\sqrt{1-\lambda}\), measured at \(10.0002\), \(17.3205\) and \(31.6228\) basis points, the best discount walking \(1.00\), \(0.99\), \(0.97\), \(0.97\) as a regime shift grows from eight to a hundred basis points while an undiscounted belief concedes only after \(295\), \(452\), \(494\) and \(498\) days |
| Bayesian Model Comparison | The marginal likelihood a density over datasets, verified at \(1.0000000000\), so Occam's charge is the normalization constraint rather than an added term and is exact instead of asymptotic; the penalty growing \(0.0004\), \(0.0049\), \(0.0255\), \(0.1680\), \(0.9438\) nats with the sample while BIC's \(O(1)\) gap runs \(-0.3332\) to \(3.0709\) nats — a factor of twenty in the odds — so a BIC difference should rank and never be exponentiated, and the balance point $ |
| Bayesian Prediction | The predictive as the likelihood integrated against the posterior and the plug-in as the same integral with the parameter-uncertainty term deleted rather than approximated, that term being exactly \(1/(n+1)\) of predictive variance — \(0.1667\), \(0.0909\), \(0.0323\), \(0.0099\), \(0.0010\) — so restoring it gives coverage of \(0.9502\), \(0.9492\), \(0.9507\), \(0.9505\), \(0.9495\) where the plug-in delivers \(0.8520\), \(0.9049\), \(0.9370\), \(0.9465\), \(0.9491\) at a width ratio falling \(1.5518\) to \(1.0017\); the same integral in counts giving a beta-binomial over-dispersed by \(1.3587\) at twenty trades of history, so a fifth-percentile shutdown trigger breaches on \(0.0935\) of months against \(0.0393\) and a rule meant to fire once in twenty months fires once in eleven; and the whole construction pricing uncertainty about the parameter and having no term for uncertainty about the model, a Gaussian predictive on Student-\(t\) returns needing \(1.1389\) of extra width where parameter uncertainty supplies a constant \(1.0085\), breaching its one-per-cent limit on \(0.0165\) of days against a promised \(0.0100\) and putting the loss at \(-2.3310\%\) against a truth of \(-2.6495\%\) |