Skip to content

Part XI — Parameter Estimation

Part X built the objects an inference is made out of — a population, a sample, the law of a statistic computed from it, a model as a set of candidate laws, the summaries that may stand in for the data, and the split of an error into a constant and a wobble — and stopped deliberately short of building anything that turns them into a number. This part is that construction, and it takes the closing sentence of the previous one literally: the question is not what a sample says but what a rule built from it can be guaranteed to do. The answer, page after page, is that the guarantees are real, provable and conditional. Each of them is a theorem about a description of the world, each description was asserted before the calculation rather than checked after it, and no estimator in this part has any mechanism for noticing when the assertion was wrong. What the part adds to the previous one is the vocabulary of a decision — a loss, a risk, a bound, a coverage — and what it discovers under that vocabulary is that the best available rule is almost never the unbiased one, that the cheapest property is the one everybody quotes, and that the interval making the most honest statement about a parameter can be the one that covers it least often.

The dependencies run in file order and the first page is load-bearing for all seven that follow, because it fixes what a loss, a risk and a standard error are, and every later page is a different way of restricting a question it shows has no answer. Properties of Estimators supplies the information bound that Maximum Likelihood Estimation attains and Method of Moments forfeits, so those three read as a chain; and 03 and 04 are a pair, fitted to the same tail index with opposite outcomes, so reading either alone misses the comparison that is the point of both. Bayesian Estimation and Maximum A Posteriori Estimation are a pair in the other direction — one integrates the posterior and one maximizes it, and the second is the first minus an integral — and both need the loss functions of the first page rather than anything in Part XVI. Confidence Intervals and Bootstrap Confidence Intervals are a pair and read in order, since the second is an accuracy ordering measured against the pivots the first inverts. Where this part stops is worth stating too: no hypothesis is tested and no critical region drawn, so the duality between an interval and a test is spent here and proved in Part XII; nothing is regressed on anything, which is Part XIII; no model is ranked and no complexity penalized, which is Part XIV; no maximum is corrected for the size of the search that found it, which is Part XV; no prior is constructed, no conjugacy derived and no belief updated, which is Part XVI; no optimizer is run and no expectation maximized, which is Part XVII; no resampling scheme is built from scratch, which is Bootstrap Methods; and no limit theorem is proved anywhere, because Part VII proved them and this part spends them.

One failure runs through the part and it has a single shape: the guarantee is a theorem about a model and it is read as a fact about the estimate, and the gap between the two never appears in the output. A plug-in optimizer whose every input converges — the drift's error falling from \(20.455\) annualized percentage points to \(3.213\), the covariance matrix's from \(0.15387\) to \(0.02454\) — still returns a median out-of-sample Sharpe of \(0.289\) against equal weighting's \(0.480\) after twenty years, and loses money \(38.5\%\) of the time. An estimator that is consistent, asymptotically normal and superefficient at a point carries \(n\,\mathrm{MSE}=158.9119\) against the sample mean's flat \(1.0017\), with the ratio climbing without bound in \(n\). A tail index is pinned to \(2.662\pm0.107\) by a likelihood that would have converged just as tightly on the wrong family, and the published evidence for that family sits below the fifth percentile of what the family itself produces. The same tail index estimated by matching a kurtosis converges on \(4.127\) with a standard deviation of \(0.112\) — thirteen of its own standard deviations from a truth of \(2.65\), and improving. A shrinkage verdict of "all the way" is correct at \(0.00325\) against James–Stein's \(0.00402\) and wrong at \(0.01064\) against \(0.00970\), on the same fifty variants with a different cross-sectional spread. One flat prior returns volatilities of \(18.828\%\), \(19.271\%\) and \(19.747\%\) depending only on which coordinate its mode was taken in, while the posterior median is identical in all three to machine precision. A binomial interval covers \(0.9496\) at six thousand observations and \(0.8714\) at five hundred, so more data made it worse, and a \(\chi^{2}\) volatility interval covers \(0.4592\) under a realistic tail while a robust one holds \(0.8630\) at two and a third times the width. And the bootstrap construction that theory ranks first covers \(0.905\) when it divides by the only standard error the industry has for a Sharpe, against \(0.940\) for the identical construction divided by a jackknife. In each case the arithmetic is correct, the estimator does exactly what its derivation promised, and the promise was made about a world nobody re-examined after the code shipped.

Topics

Topic Focus
Point Estimation An estimand as a functional of a law, an estimator as a rule with a distribution and an estimate as a number without one, the plug-in principle charging curvature times variance so that \(\mathbb{E}[1/\hat s^{2}]\sigma^{2}=(n-1)/(n-3)\) overstates a Kelly fraction by \(11\%\) on a month of data while its sign is wrong \(34.874\%\) of the time on a year, a loss function as the only thing that orders two rules, risk functions crossing at exactly one standard error so the rule that ignores the data wins below \(0.03945\) and the optimal shrinkage at the course's own \(0.075\) is \(0.217\), and a plug-in portfolio whose every input converges while its out-of-sample Sharpe stays at \(0.289\) against \(1/N\)'s \(0.480\)
Properties of Estimators Unbiasedness as a fence around a class rather than a certificate on a member, Rao–Blackwell as the law of total variance applied to a conditional expectation with completeness buying uniqueness, the Cramér–Rao bound as one Cauchy–Schwarz on the score with a regularity condition the uniform beats by a whole order, a drift's information \(T/\sigma^{2}\) making the course's \(\pm0.039\) a floor rather than a measurement while the efficient estimator is algebraically identical to one reading two prices, Hodges' estimator superefficient at zero and carrying \(n\,\mathrm{MSE}=158.9119\) against the sample mean's \(1.0017\), and a sample mean \(1.000\) efficient under the normal and \(0.627\) under the tails the course fits while one row moves it by \(0.630\) in every law
Maximum Likelihood Estimation The likelihood as a density read backwards and the log as the only representation in which it exists at scale, score equations as the first-order condition with the observed information delivering a standard error the optimizer already computed, invariance holding for the argmax while every likelihood level shifts by \(n\log 100=29519.1\) under a change of units, a tail index at \(2.662\pm0.107\) on the course's sample against \(2.731\pm0.605\) on one year with a log-likelihood gap whose fifth percentile of \(1116.6\) exceeds the published \(939.3\), and a misspecified fit converging on the Kullback–Leibler projection with a free standard error \(0.113\) of the truth and an implied one-percent loss threshold of \(-2.326\) against \(-2.657\)
Method of Moments Moment matching as a system of equations rather than an optimization, consistency and asymptotic normality from the delta method alone with the \(k\)th moment charging the existence of the \(2k\)th, a Yule–Walker autocorrelation of \(0.7436\) at sixty observations turning a \(3.49\)-day half-life into \(2.49\) and understating the horizon four times in five, a kurtosis-matched degrees of freedom converging on \(4.127\pm0.112\) when the truth is \(2.65\) so the error is thirteen of its own standard deviations and shrinking, and overidentification buying the only self-diagnosis in the part at a size of \(0.0463\) and a power of \(1.0000\)
Bayesian Estimation A posterior as an inference and an estimate as a decision about it, the mean minimizing quadratic risk and the median absolute risk while a \(10{:}1\) cost asymmetry walks the optimum to the \(90.9\)th percentile and charges \(131.96\%\) for using the mean, every Bayes estimator as a precision-weighted average whose weight is the shrinkage the earlier theorem refused to identify, a prior worth a factor of \(11.6\) in risk where it is right and \(29.5\) where it is not while one wide enough to be uncontroversial does nothing, and James–Stein overtaking full shrinkage at \(0.00970\) against \(0.01064\) once cross-sectional dispersion reaches half a standard error
Maximum A Posteriori Estimation The posterior mode as the one summary needing no integral, MAP as penalized likelihood with a Gaussian prior giving ridge at \(\lambda=\sigma^{2}/\tau^{2}\) and a Laplace prior lasso, the course's skeptical prior translating to \(\lambda=14781.3\) with a shrinkage weight of \(0.7059\) fixed before any data is loaded, one posterior read in three coordinates returning \(18.828\%\), \(19.271\%\) and \(19.747\%\) while the median is identical in all three to machine precision, and lasso-MAP zeroing \(47.57\) of fifty coefficients where the mean of the identical posterior zeroes none
Confidence Intervals Coverage as an unconditional property of a procedure that the sample can contradict while the stated level stays valid, every interval as an inverted pivot with the Wald interval the single exception every package returns by default, exact binomial coverage of \(0.9496\) at \(n=6{,}158\) and \(0.8714\) at \(n=500\) where more data lowers it, the scale an interval is built on deciding its coverage so a \(\chi^{2}\) volatility interval reads \(0.4592\) where a robust log interval holds \(0.8630\) at \(2.3\) times the width, and a credible interval carrying a correct probability statement about the parameter and a frequentist coverage of exactly \(0.0000\)
Bootstrap Confidence Intervals Six constructions ordered by the power of \(n\) in their coverage error rather than by taste, the studentized interval second-order accurate only in the standard error it divides by so Lo's formula drops it to \(0.905\) where a jackknife holds it at \(0.940\), BCa's two constants estimating the transformation the analyst would otherwise have to choose, a GARCH series destroying the independent bootstrap's interval for a standard deviation at \(0.537\) while leaving its interval for a mean at \(0.952\), and \(B\) cutting Monte Carlo error to \(0.0406\) against a sampling error of \(0.9354\) that does not move