Skip to content

Part VI — Multivariate Probability

Part V named laws one variable at a time and never let two of them into the same formula; this part stacks them into a vector. The trade the objects here offer is a compression: a joint law on \(\mathbb{R}^{N}\) — a function, with no finite description — becomes a mean vector and a covariance matrix, \(N(N+3)/2\) numbers from which portfolio variance, hedge ratios, effective bet counts and conditional forecasts all follow by arithmetic. The bill arrives twice. Once in estimation, where the matrix has more free parameters than the data has independent observations and the failure is a spectrum rather than an entry; and once in the tail, where the compression discards exactly the joint behaviour a risk system was built to describe.

The dependencies run in file order, with two things worth knowing. Multivariate Random Variables introduces the vector notation that Parts III and IV deliberately refused, so it reads first and everything after it assumes the conventions in its table; and Linear Transformations needs only Covariance Matrices, so it can be read before Correlation Matrices — though 03 derives the floor that 04 then shows a projection attains exactly, and the pair is better read in order. Multivariate Gaussian Distribution needs 04's sandwich formula together with The Gaussian Distribution, and Conditional Gaussian Distributions reads only after 05, because its central proof is one line of 05 applied twice. Where this part stops is worth stating too: the matrix algebra itself — quadratic forms, eigenvalues, the condition number, Cholesky, projection — is developed in Basic Linear Algebra Review and assumed rather than repeated; no limit in \(N\) or \(T\) is taken, so every convergence claim is Part VII; nothing is indexed by time, so a vector-valued process is Part VIII; no sampling distribution is derived for anything estimated, which is Part XI; the regression that page 06 shows is a conditional expectation is fitted in Part XIII; and the dependence structure a correlation matrix compresses is parameterised in Copulas.

One failure runs through the part and it has a single shape. Every object here is meaningful only as a whole, and every one of them is routinely assembled, estimated, or repaired one entry at a time: covariances computed pair by pair on whatever sample each pair happens to have, correlation matrices stitched from two vendors and an expert view, an estimator that is unbiased in every entry and badly biased in every eigenvalue, a joint law inferred from margins that cannot determine it. In each case the arithmetic completes, the output is a plausible number, and nothing announces the defect. The diagnostics that do announce it are matrix-level — a spectrum, a condition number, a ratio of observations to assets, the sign of the smallest eigenvalue — and they are the numbers this part argues belong beside every risk figure that depends on them.

Topics

Topic Focus
Multivariate Random Variables One map into \(\mathbb{R}^{n}\) rather than \(n\) maps into the line, a mean vector that needs no hypothesis at all, what a risk system stores against what a joint law contains, independence that pairwise checks cannot certify, the sample size a joint density would actually need, and the intersection of calendars nobody chose
Covariance Matrices One expectation of an outer product, the quadratic form that is the whole of positive semi-definiteness, the observation the estimator spends on the mean, an unbiased matrix whose eigenvalues are not, shrinkage that buys conditioning and cannot buy a mean, and the three ways back onto the cone
Correlation Matrices The standardization and the constraint it adds, a spectrum that is a budget of exactly \(N\), four diversification diagnostics reading that budget and disagreeing, the floor at \(-1/(N-1)\) that no equicorrelated universe can breach, a repair problem strictly harder than the covariance one, and nine tickers holding one bet
Linear Transformations The mean carrying the intercept and the covariance not, the sandwich behind every risk number in the book, demeaning as the projection that manufactures \(-1/(N-1)\), whitening defined only up to a rotation, the rank a singular map destroys permanently, and the basis nobody signed off on
Multivariate Gaussian Distribution The density as one quadratic form under an exponential, the definition by linear combinations that a portfolio actually needs, normal margins that are not a normal joint, ellipsoids whose axis-alignment is exactly independence, a Mahalanobis distance that is chi-square, and no tail dependence at any correlation below one
Conditional Gaussian Distributions The two partitioned formulas, a conditional mean that is affine and therefore the best predictor outright, the Schur complement as the variance a regression removes, zeros in the precision matrix as conditional independences, sequential updating that equals updating at once, and a conditional covariance that does not know what was observed