medium · Covariance matrices and portfolio variance
Let assets each have variance , with every distinct pair having correlation exactly . Write for the covariance matrix and for the all-ones vector.
Since has eigenvalue on and on the orthogonal complement , the eigenvalues of are
Both must be nonnegative, so
The lower limit is the interesting one: assets cannot all be strongly negatively correlated with one another. At the vector is in the kernel, meaning the equally weighted portfolio is exactly riskless.
With , the quadratic form picks out the eigenvalue:
Only the idiosyncratic part diversifies away, at rate . The common part does not shrink no matter how many names you add.
For fully invested portfolios the minimizer of subject to is
from the Lagrangian , whose stationarity condition is . Here is an eigenvector of , so
provided so that is invertible and positive definite, which also makes the stationary point a genuine minimum. The minimum variance is therefore the quantity computed in part 2:
Note. The practical reading is that the floor on portfolio risk is set by correlation, not by breadth: with and , no number of equally risky names gets annualized volatility below . The tempting error is to treat as the governing rate; it governs only the slice.
medium · EM, the ELBO, and what monotonicity does not buy
Let be observed, latent, and a joint model. For a distribution over define
does not depend on , so it equals its own expectation under any . Insert the definition of conditional probability and split the logarithm:
The first term is and the second is , which proves the identity. Since a Kullback-Leibler divergence is nonnegative and vanishes only when its arguments agree almost everywhere, with equality exactly at . Note the identity is exact for every : the bound's slack is not merely bounded by the KL term, it is the KL term.
The first inequality is part 1 applied at . The second is the M-step, which maximizes over and so cannot do worse than the incumbent. The final step is an equality, and it is the one that carries the argument: the E-step chose to be the posterior at , so the bound is tight there.
Without that tightness you would only be comparing a bound to a bound, which proves nothing about the likelihood itself. This is the step a rushed proof skips.
It does not give a global maximum. What follows is only that the sequence is nondecreasing, hence convergent whenever the likelihood is bounded above.
Three gaps separate that from what one wants. The likelihood value can converge while the parameters do not; the limit point, if the parameters do converge, is in general only a stationary point, so a local maximum or even a saddle is possible; and establishing convergence of at all requires regularity conditions on the model rather than following from the iteration (this is the content of Wu's 1983 analysis, which corrected the original claim in Dempster, Laird and Rubin).
For a Gaussian mixture the likelihood is not concave, is invariant under relabeling the components — so every maximum comes with copies — and is unbounded above as a component variance goes to zero with its mean pinned to a data point. Monotone ascent in that landscape guarantees you stop going down, nothing more, which is why EM is run from several random starts and the best run kept.
Two new problems every morning at 8am · every day so far