medium · Lagrange multipliers and portfolio optimization
You may allocate wealth across three risky assets whose returns are mutually uncorrelated. Asset has expected return and return variance :
| Asset | ||
|---|---|---|
| 1 | ||
| 2 | ||
| 3 |
A portfolio is a weight vector with ; short positions are permitted, so the may be any real numbers. Its expected return is and, because the assets are uncorrelated, its variance is .
For a general set of uncorrelated assets, use Lagrange multipliers to find the portfolio of minimum variance subject to a fixed budget and a target expected return . Express the optimal weights in closed form in terms of the scalars , , .
Using the table, compute the weights of the minimum-variance portfolio with target return , and its variance.
Show that the minimum attainable variance is a quadratic function of the target return , give that function for the table above, and identify the global minimum-variance portfolio (the one with no return constraint imposed).
Minimize (the is cosmetic) subject to the two linear constraints and . The objective is strictly convex in and the constraints are affine, so a stationary point of the Lagrangian is the global minimizer — Lagrange multipliers give a necessary and sufficient condition here.
The step a candidate skips is using two multipliers, one per constraint:
Stationarity gives
so the optimal weight is
Each weight is an affine function of scaled by the precision — this is the two-fund structure. Now impose the constraints. Writing , , :
Solve the system. With determinant (positive by Cauchy–Schwarz unless all are equal),
Hence the closed form:
The precisions are , so
For :
Then :
These sum to and reproduce , as they must. The variance is
i.e. a return standard deviation of about .
At the optimum the variance can be read off without recomputing directly. Since ,
Substituting the solved multipliers,
The minimum-variance frontier is therefore a parabola in the variance–return plane (a hyperbola in the –return plane). For the table,
Check: gives . Consistent.
The global minimum-variance (GMV) portfolio minimizes over : setting the derivative gives
At the return multiplier vanishes (), so its weights are simply proportional to the precisions:
The clean route to part 3 is the identity , which uses the constraints themselves rather than squaring the weights — a common time-saver that also generalizes verbatim to correlated assets by replacing with , , . The tempting error is to impose only the budget constraint and then be surprised the target return is not met; both constraints must carry their own multiplier.
hard · Overparameterized least squares and double descent
Fix a signal with . Draw a design matrix whose rows are i.i.d. , and set Fit least squares, taking the minimum--norm solution whenever it is non-unique: with the Moore–Penrose pseudoinverse (so when and when , both a.s.). Define the excess risk the expectation taken over and .
You may use the inverse-Wishart mean: if with and , then .
First show that for an independent test point with , , the prediction risk equals ; hence is the object of interest. Then, in the underparameterized regime , compute in closed form and describe its behaviour as .
In the overparameterized regime , derive the exact bias–variance decomposition of and give it in closed form. Identify precisely which step introduces a nonzero bias and why it is absent in Part 1.
Treating as a continuous variable on , find the minimizer of the overparameterized risk, state the exact condition on under which an interior minimum exists, and compute in closed form. Interpret the result by comparing with the risk of the null predictor .
Throughout write .
For an independent test point, Condition on (a function of the training data, independent of ). Since and , The cross term vanishes and , so the prediction risk is . The isotropy of is what turns the test risk into a plain squared parameter error; this is the reason the problem is exactly solvable.
Here , so is a.s. invertible and . The model is well specified, so There is no bias: because . Conditioning on , Now is Wishart with , . For the inverse-Wishart mean gives , hence As the denominator and : the risk diverges at the interpolation threshold. This is the first descent's collapse.
Now , the rows of span an -dimensional subspace a.s., and . Substitute and let be the orthogonal projector onto the row space of (an -dimensional subspace of ). Then Because , the cross term drops when we take expectations, leaving
Bias. This is the step a candidate misses: the minimum-norm interpolator can only fit the component of lying in the row space; the orthogonal component is irrecoverable. Since is idempotent, . The row space of a Gaussian is a uniformly random -dimensional subspace (the standard Gaussian is rotationally invariant), so and . Hence In Part 1 the row space had dimension and equalled all of the column-relevant space, so : no bias.
Variance. . With , so the variance is . Now is Wishart with , ; for ,
Therefore As the variance (the peak again); as the variance while the bias . The interpolator never forgets that it has thrown away a fraction of the signal.
On differentiate: Setting gives . Both sides positive, so take positive roots: , i.e. , whence Since exactly when , an interior minimizer with exists iff (signal norm exceeds the noise level). When , for all large and decreases monotonically to its infimum : the best overparameterized model merely matches the null predictor.
Evaluate . Writing , one has , and Hence that is
Interpretation. The null predictor has excess risk . The optimally overparameterized interpolator beats it by exactly . So heavy overparameterization is genuinely useful, but only when ; the benefit is the squared gap between signal and noise scales, discounted by .
Two new problems every morning at 8am · every day so far