Part X — Foundations of Statistics¶
Part IX generated data from laws that were written down, and every failure it found was a simulation grading its own homework. This part reverses the arrow. The law is not written down, the data arrived without a receipt, and the question is what can be said about the mechanism from its output — which is a strictly harder problem than the one probability solves, because the forward map from a law to a sample is a function and the backward map is not. Many laws produce the sample in hand, one of them fits it perfectly and is certainly wrong, and every technique in the eight parts that follow is a disciplined way of choosing among the rest and stating what the choice cost. This part builds the objects that choice is made out of: a population, a sample, the law of a statistic computed from it, the set of candidate laws a model narrows to, the summaries that can stand in for the data without loss, and the split of an estimator's error into a piece that averages away and a piece that does not.
The dependencies run in file order and the first page is load-bearing for everything after it, since it fixes what a population is and every later estimand is a functional of one. Population vs Sample and Descriptive Statistics are a pair — one says what the sample is an output of and the other what a summary of it summarizes — and Sampling Distributions is the page both of them are owed, because every number the second page prints is a draw whose law the third derives. Statistical Models, Statistics and Sufficiency and Exponential Families are a chain and read in order: a model is a set of candidate laws, sufficiency asks what may be discarded without loss relative to that set, and the third names the one class where the discarding is cheap and answers the question the second raises and declines. Bias and Variance needs the second and third and closes the part. Where this part stops is worth stating too: no estimator is constructed and no property of one is catalogued, so consistency and efficiency are Part XI and every interval is Confidence Intervals; no hypothesis is tested and no critical region drawn, which is Part XII; nothing is regressed on anything, which is Part XIII; no prediction error is decomposed and nothing overfits, which is Bias–Variance Tradeoff; no maximum is corrected for the search that found it, which is Part XV; no parameter is given a prior, which is Part XVI; no model is fitted by any algorithm, which is Part XVII; and no limit theorem is proved anywhere, because Part VII proved them and this part spends them.
One failure runs through the part and it has a single shape: the number is a correct statement about the sample and it is read as a statement about the population, and the gap between the two is not noise, does not shrink with \(n\), and is never reported by the procedure that produced the number. A path whose sample mean is unbiased at every horizon covers the ensemble mean with a nominal \(95\%\) interval \(30.4\%\) of the time after two hundred and fifty years, and the coverage worsens as the sample grows while the estimate visibly settles down. A survivorship filter installs a bias of \(0.01526\) that is identical at five sample sizes while its \(t\)-statistic climbs to \(79.80\). A sample excess kurtosis climbs from \(3.93\) to \(68.38\) instead of converging on the \(11.41\) the course published, and an explained-variance share of \(63\%\) sits against a floor of \(11.4\%\) that nine independent series manufacture with nothing in common. A twenty-bar rolling mean of pure noise has a lag-one autocorrelation of \(0.9503\) and a \(t\)-test that rejects \(67.9\%\) of the time. A nominal \(95\%\) interval for a variance covers \(0.4566\) under a realistic tail while the \(t\) interval printed beside it holds \(0.9507\). Two ARMA coefficients significant at \(p=2.4\times10^{-4}\) scatter over four times their own standard errors while the one combination the data determines stays pinned at \(-0.0828\). A one-percent loss threshold computed from a running sum and sum of squares is exact under normality and breached on \(1.454\%\) of days once the tail thickens. A two-component mixture's log-likelihood climbs without bound, so the estimate every implementation returns is a property of its variance floor. And the most biased of three volatility estimators wins on correlation and on mean absolute error and loses to the one it beat as soon as three of them are averaged. In each case the arithmetic is correct, the estimator does exactly what its derivation promised, and the promise was about a population the sample was never a fair draw from.
Topics¶
| Topic | Focus |
|---|---|
| Population vs Sample | The inverse map from data to law failing to be a function, a population as a data-generating process where observing every existing datum can still leave one draw, stationarity buying convergence while only ergodicity buys the limit being a number, a path-fixed drift whose iid standard error is too narrow by \(1.84\) at twenty-five years and \(5.01\) at two hundred and fifty as coverage falls to \(0.304\), a survivorship bias flat at \(0.01526\) across a four-hundredfold change in universe size while its \(t\)-statistic climbs to \(79.80\), and twenty-one-day overlap collapsing \(6{,}118\) rows to \(324\) with a naive test that false-positives \(67.0\%\) of the time |
| Descriptive Statistics | Every summary as a functional evaluated at \(\hat F_n\) rather than \(F\), the mean and median of an overnight gap differing by fifty-four percent and of a levered book differing in sign, a sample excess kurtosis ceilinged at \(n-3\) and climbing \(3.93\) to \(68.38\) instead of converging on \(11.41\), the analysis-of-variance identity holding to \(4.55\times10^{-13}\) as pure arithmetic while its explained share is floored at \(1/N\) so that nine sectors' \(63\%\) is fifty-two points of real co-movement, and a rolling mean of independent noise carrying autocorrelation \(1-k/w\), \(314\) independent values in \(6{,}281\) rows, and a two-sigma threshold that fires \(23.0\%\) of the time at min_periods of three |
| Sampling Distributions | A statistic as a random variable and the standard error as the only summary of its law, a drift's precision pinned at \(0.039\) across a hundred-and-thirty-fold change in observation count while a volatility's improves elevenfold on the same data, one Helmert rotation delivering \(\chi^{2}_{n-1}\) and an independence of \(\bar X\) and \(s^{2}\) that characterizes the normal and holds nowhere else, a nominal \(95\%\) variance interval covering \(0.4566\) and deteriorating with \(n\) beside a \(t\) interval holding \(0.9507\) under every tail, and a Sharpe of \(0.30\) carrying an interval of \([-0.1050, 0.6988]\) while the best of fifty nulls averages \(0.4506\) and the winner's own interval covers the truth \(28.0\%\) of the time |
| Statistical Models | A model as a set of candidate laws on the space of paths rather than of points, the parametric–nonparametric split priced as a rate against a correctness, an AR(1) at \(\rho=0.995\) with a \(138\)-day half-life putting both stationarity tests into rejection on \(1.000\) of samples while a \(34\)-day half-life on one year is called integrated \(80.5\%\) of the time and is wrong, an ARMA sum of squares flat from \(1.0018357\) to \(1.0037447\) along the cancellation curve while refits scatter \(\hat\phi\) over \([+0.0702,+0.2967]\) and pin the lag-one autocorrelation at \(-0.0828\), and a misspecified iid fit consistent for the unconditional \(18.5\%\) and breaching a one-percent loss threshold \(1.504\%\) of the time |
| Statistics and Sufficiency | A statistic as a function of the data alone, sufficiency as a conditional law free of the parameter and therefore a property of a pair rather than of a statistic, the factorization theorem converting an unverifiable claim into a syntactic check on the likelihood, two twelve-point samples differing elementwise by \(1.3158\) producing log-likelihoods identical to \(3.55\times10^{-15}\), the order statistics sufficient for every iid model and compressing nothing, and a reduction that costs forty-six percent of achievable precision under \(t(3)\) and turns an exact one-percent loss threshold into one breached \(1.454\%\) of days |
| Exponential Families | One functional form absorbing the normal, Poisson, gamma and Bernoulli with the log-odds arriving as a natural parameter rather than a choice, the log-partition function generating \(\mathbb{E}[T]\) and \(\mathrm{var}(T)\) by differentiation and delivering concavity with them, the exponent's additivity fixing the sufficient statistic's dimension for every \(n\) so the same two numbers that are exact for a normal cost \(0.325\) of log-likelihood under \(t(3)\), Pitman–Koopman–Darmois with the support condition the uniform on \([0,\theta]\) violates, and a mixture whose log-likelihood climbs to \(-514.22\) without bound so that forty EM starts agree only because a variance floor stops them |
| Bias and Variance | Mean squared error splitting into a constant and a wobble with no cross term and a third term appearing the instant the target stops being a parameter, the Bessel correction derived from an \((n-1)\)-dimensional residual subspace and undone by a square root at \(c_4(21)=0.98758\) for a volatility running a quarter-point low forever, an \(n+1\) divisor beating the unbiased one on error at every sample size, a Parkinson estimator biased \(1.90\) points low winning on correlation and mean absolute error and losing past \(K=2.93\) assets, and shrinkage falling monotonically to complete shrinkage while Ledoit–Wolf moves a mean-variance book from \(0.078\) to \(0.080\) against \(1/N\)'s \(0.351\) |