Skip to content

Part VII — Asymptotic Theory

Part VI stacked variables into a vector and computed with them at a fixed sample size; this part lets the sample size run. Everything here trades an exact statement nobody can make for an approximate one everybody uses: the distribution of a sample mean is unknowable without knowing the distribution it came from, and the central limit theorem replaces it with a normal that needs only two numbers. The trade is extraordinarily good and it is a trade. What is surrendered is any statement about a finite \(n\) — every theorem in this part is a claim about a limit, none of them names a sample size, and the sample sizes finance actually has are small enough that the difference is measurable in every direction the reader will care about.

The dependencies run in file order with one exception worth knowing. The Weak Law of Large Numbers introduces convergence in probability, The Strong Law of Large Numbers adds almost-sure convergence and the contrast that names it, and The Central Limit Theorem adds convergence in distribution — so the three modes accumulate across the first three pages and everything after assumes all of them. The exception is that Continuous Mapping Theorem is logically prior to both The Delta Method and Slutsky's Theorem, since the delta method is Slutsky applied to a Taylor expansion and Slutsky is the mapping theorem applied to a two-argument map; the numbering follows how often a practitioner reaches for each rather than how the proofs depend on one another, and a reader who wants the logical development should take 06 before 04 and 05. Where this part stops is worth stating too: the inequalities the first proof consumes are established in Variance and are not re-derived, and the limsup/liminf vocabulary the second page needs is Sequences and Infinite Series; nothing is indexed by time and no dependent sequence gets its own limit theorem, so ergodic averages and the effective sample size that dependence costs are Random Processes; no estimator's properties are named, which is Properties of Estimators, and no interval is constructed, which is Confidence Intervals; no test or critical region is built, which is Part XII; the resampling machinery the first page licenses is Bootstrap Methods; and the laws whose tails put them outside these theorems are parameterised in Heavy-Tailed Returns and Extreme Value Theory.

One failure runs through the part and it has a single shape. Every theorem here is an unconditional promise about \(n\to\infty\) being used to license a decision at the \(n\) on hand, and in every case the theorem supplies no rate, so the check that would justify the substitution has to come from somewhere else and almost never does. The equity premium converges and after twenty-five years still has a one-in-thirty-six chance of the wrong sign; a nominal \(5\%\) test is exact at six thousand observations while a nominal \(0.1\%\) test on the same data is off by a factor of two and a half; a Sharpe ratio's standard error is right to three decimals under normality and forty percent too narrow under a realistic skew; a variance estimator is consistent for a number fourteen times smaller than the one the formula needs; the sign of a converging estimate never converges at all. In each case the arithmetic completes, the output is a plausible number, and the theorem being invoked is true. What the theorem was asked to certify is something it never claimed, and the diagnostics that would catch the gap — a standard error read against the calendar, a simulated size read against a nominal one, a statistic recomputed on nested subsamples — are cheap enough that their absence is the point.

Topics

Topic Focus
The Weak Law of Large Numbers Chebyshev applied to an average and a limit taken, convergence in probability as a claim about each \(n\) separately, a bound five times too wide and a rate needing a century to halve, the same theorem read on the empirical distribution as the bootstrap's licence, a Cauchy mean whose spread is identical at ten draws and ten thousand, and two published fits that disagree about whether the theorem applies
The Strong Law of Large Numbers Almost-sure convergence as limsup and liminf coinciding, Borel–Cantelli and the one word — summable — that separates the two laws, a weaker hypothesis buying a stronger conclusion, a sequence that converges in probability and on no path whatever, the path-wise question costing twice the ensemble one at twenty-five years, and a zero-edge strategy touching Sharpe \(2.0\) one time in twenty-three
The Central Limit Theorem The statement for sums with the mean as a corollary, a Taylor expansion inside a transform and exactly where finite variance is spent, skewness dying like \(1/\sqrt{k}\) under convolution while kurtosis dies like \(1/k\), a \(5\%\) level exact at six thousand observations while the \(0.1\%\) level is not, Berry–Esseen bounds that are vacuous or absent precisely where they are needed, and Cramér–Wold making every multivariate failure directional
The Delta Method A first-order Taylor expansion with an \(o_p(1)\) remainder discarded, the relative error \(1/\sqrt{2n}\) that makes volatility estimable, transforming endpoints rather than estimates and the coverage it buys back, the Sharpe ratio's standard error derived and the published \(\pm0.20\) reproduced to three decimals, skew and kurtosis terms whose omission costs forty percent of an error bar, and a vanishing slope returning a standard error of exactly zero
Slutsky's Theorem Two modes of convergence combined and why one limit must be constant, an \(o_p(1)\) term added to a converging sequence without changing it, the t-statistic as the theorem's one indispensable use, a consistent estimator converging to a constant that is not the one the formula needs, HAC as a large improvement and not a fix, and a frozen denominator whose calibration error is identical at \(252\) observations and \(25{,}200\)
Continuous Mapping Theorem One hypothesis carrying three theorems, the deterministic limit law applied one history at a time, consistency passing through a square root where unbiasedness cannot, a discontinuity at the limit point failing outright rather than approximately, a continuity set that only has to catch the limit and the second condition routinely conflated with it, and the two questions that separate a plug-in's two failure modes