Saturday 5 September 2026

Quant interview

The minimum-variance portfolio for uncorrelated assets

medium · Lagrange multipliers and portfolio optimization

You may allocate wealth across three risky assets whose returns are mutually uncorrelated. Asset i has expected return μi and return variance σi2:

Asset μi σi2
1 0.08 0.01
2 0.12 0.02
3 0.20 0.04

A portfolio is a weight vector w=(w1,w2,w3) with ∑iwi=1; short positions are permitted, so the wi may be any real numbers. Its expected return is ∑iwiμi and, because the assets are uncorrelated, its variance is ∑iwi2σi2.

  1. For a general set of n uncorrelated assets, use Lagrange multipliers to find the portfolio of minimum variance subject to a fixed budget ∑iwi=1 and a target expected return ∑iwiμi=m. Express the optimal weights in closed form in terms of the scalars A=∑i1/σi2, B=∑iμi/σi2, C=∑iμi2/σi2.

  2. Using the table, compute the weights of the minimum-variance portfolio with target return m=0.12, and its variance.

  3. Show that the minimum attainable variance is a quadratic function of the target return m, give that function for the table above, and identify the global minimum-variance portfolio (the one with no return constraint imposed).

Solution

Setting up the Lagrangian

Minimize 12∑iwi2σi2 (the 12 is cosmetic) subject to the two linear constraints ∑iwi=1 and ∑iwiμi=m. The objective is strictly convex in w and the constraints are affine, so a stationary point of the Lagrangian is the global minimizer — Lagrange multipliers give a necessary and sufficient condition here.

The step a candidate skips is using two multipliers, one per constraint:

ℒ(w,λ,γ)=12∑iwi2σi2−λ(∑iwi−1)−γ(∑iwiμi−m).

Stationarity ∂ℒ/∂wi=0 gives

wiσi2−λ−γμi=0,

so the optimal weight is

wi=λ+γμiσi2.

Each weight is an affine function of μi scaled by the precision 1/σi2 — this is the two-fund structure. Now impose the constraints. Writing A=∑i1/σi2, B=∑iμi/σi2, C=∑iμi2/σi2:

∑iwi=λA+γB=1,

∑iwiμi=λB+γC=m.

Solve the 2×2 system. With determinant D=AC−B2 (positive by Cauchy–Schwarz unless all μi are equal),

λ=C−BmD,γ=Am−BD.

Hence the closed form:

wi=1σi2·(C−Bm)+(Am−B)μiAC−B2.

Plugging in the numbers

The precisions are 1/σi2=(100,50,25), so

A=100+50+25=175.

B=100(0.08)+50(0.12)+25(0.20)=8+6+5=19.

C=100(0.08)2+50(0.12)2+25(0.20)2=0.64+0.72+1.00=2.36.

D=AC−B2=175(2.36)−192=413−361=52.

For m=0.12:

λ=2.36−19(0.12)52=0.0852,γ=175(0.12)−1952=252.

Then wi=(1/σi2)(λ+γμi):

w1=100(0.0852+252(0.08))=613,

w2=50(0.0852+252(0.12))=413,

w3=25(0.0852+252(0.20))=313.

These sum to 1 and reproduce ∑wiμi=0.12, as they must. The variance is

∑iwi2σi2=(613)2(0.01)+(413)2(0.02)+(313)2(0.04)=2325≈0.006154,

i.e. a return standard deviation of about 7.84%.

The efficient frontier and the GMV portfolio

At the optimum the variance can be read off without recomputing ∑wi2σi2 directly. Since wiσi2=λ+γμi,

σ2(m)=∑iwi2σi2=∑iwi(λ+γμi)=λ∑iwi⏟=1+γ∑iwiμi⏟=m=λ+γm.

Substituting the solved multipliers,

σ2(m)=C−BmD+Am−BDm=Am2−2Bm+CAC−B2.

The minimum-variance frontier is therefore a parabola in the variance–return plane (a hyperbola in the σ–return plane). For the table,

σ2(m)=175m2−38m+2.3652.

Check: m=0.12 gives (175·0.0144−38·0.12+2.36)/52=(2.52−4.56+2.36)/52=0.32/52=2/325. Consistent.

The global minimum-variance (GMV) portfolio minimizes σ2(m) over m: setting the derivative 2Am−2B=0 gives

m⋆=BA=19175≈0.1086,σmin2=1A=1175≈0.005714.

At m=m⋆ the return multiplier vanishes (γ=0), so its weights are simply proportional to the precisions:

wiGMV=1/σi2A=(100175,50175,25175)=(47,27,17).

Closing note

The clean route to part 3 is the identity σ2=λ+γm, which uses the constraints themselves rather than squaring the weights — a common time-saver that also generalizes verbatim to correlated assets by replacing A,B,C with 1⊤Σ−11, 1⊤Σ−1μ, μ⊤Σ−1μ. The tempting error is to impose only the budget constraint and then be surprised the target return is not met; both constraints must carry their own multiplier.

Statistics in machine learning

The second descent of the minimum-norm interpolator

hard · Overparameterized least squares and double descent

Fix a signal β\*∈ℝp with ‖β\*‖22=r2. Draw a design matrix X∈ℝn×p whose rows are i.i.d. N(0,Ip), and set y=Xβ\*+ϵ,ϵ~N(0,σ2In) independent of X. Fit least squares, taking the minimum-ℓ2-norm solution whenever it is non-unique: β^=X+y, with X+ the Moore–Penrose pseudoinverse (so β^=(X⊤X)−1X⊤y when p<n and β^=X⊤(XX⊤)−1y when p>n, both a.s.). Define the excess risk R(p)=𝔼\|β^−β\*\|22, the expectation taken over X and ϵ.

You may use the inverse-Wishart mean: if W=∑i=1mzizi⊤ with zi~iidN(0,Ik) and m>k+1, then 𝔼[W−1]=1m−k−1Ik.

  1. First show that for an independent test point (x0,y0) with x0~N(0,Ip), y0=x0⊤β\*+ϵ0, the prediction risk equals R(p)+σ2; hence R(p) is the object of interest. Then, in the underparameterized regime p≤n−2, compute R(p) in closed form and describe its behaviour as p↑n.

  2. In the overparameterized regime p≥n+2, derive the exact bias–variance decomposition of R(p) and give it in closed form. Identify precisely which step introduces a nonzero bias and why it is absent in Part 1.

  3. Treating p as a continuous variable on (n+1,∞), find the minimizer p\* of the overparameterized risk, state the exact condition on (r,σ) under which an interior minimum exists, and compute R(p\*) in closed form. Interpret the result by comparing R(p\*) with the risk of the null predictor β^≡0.

Solution

Throughout write r2=‖β\*‖22.

The risk functional

For an independent test point, 𝔼[(x0⊤β^−y0)2]=𝔼[(x0⊤(β^−β\*)−ϵ0)2]. Condition on β^ (a function of the training data, independent of (x0,ϵ0)). Since 𝔼[ϵ0]=0 and Cov(x0)=Ip, 𝔼[(x0⊤v)2]=v⊤𝔼[x0x0⊤]v=‖v‖22,v=β^−β\*. The cross term vanishes and 𝔼[ϵ02]=σ2, so the prediction risk is 𝔼‖β^−β\*‖22+σ2=R(p)+σ2. The isotropy of x0 is what turns the test risk into a plain squared parameter error; this is the reason the problem is exactly solvable.

Part 1 — underparameterized regime

Here p<n, so X⊤X is a.s. invertible and β^=(X⊤X)−1X⊤y. The model is well specified, so β^−β\*=(X⊤X)−1X⊤ϵ. There is no bias: 𝔼[β^∣X]=β\* because X⊤y=X⊤Xβ\*+X⊤ϵ. Conditioning on X, 𝔼[‖β^−β\*‖22∣X]=σ2tr((X⊤X)−1X⊤X(X⊤X)−1)=σ2tr((X⊤X)−1). Now X⊤X=∑i=1nxixi⊤ is Wishart with k=p, m=n. For n>p+1 the inverse-Wishart mean gives 𝔼[(X⊤X)−1]=1n−p−1Ip, hence R(p)=σ2pn−p−1,p≤n−2. As p↑n the denominator →0 and R(p)→∞: the risk diverges at the interpolation threshold. This is the first descent's collapse.

Part 2 — overparameterized regime

Now p>n, the rows of X span an n-dimensional subspace a.s., and β^=X+y=X⊤(XX⊤)−1y. Substitute y=Xβ\*+ϵ and let P:=X+X=X⊤(XX⊤)−1X be the orthogonal projector onto the row space of X (an n-dimensional subspace of ℝp). Then β^=Pβ\*+X+ϵ,β^−β\*=−(I−P)β\*+X+ϵ. Because 𝔼[ϵ∣X]=0, the cross term drops when we take expectations, leaving R(p)=𝔼\|(I−P)β\*\|22⏟bias+𝔼\|X+ϵ\|22⏟variance.

Bias. This is the step a candidate misses: the minimum-norm interpolator can only fit the component of β\* lying in the row space; the orthogonal component (I−P)β\* is irrecoverable. Since I−P is idempotent, ‖(I−P)β\*‖22=β\*⊤(I−P)β\*. The row space of a Gaussian X is a uniformly random n-dimensional subspace (the standard Gaussian is rotationally invariant), so 𝔼[P]=npIp and 𝔼[I−P]=p−npIp. Hence bias=p−npr2. In Part 1 the row space had dimension p and equalled all of the column-relevant space, so I−P=0: no bias.

Variance. 𝔼‖X+ϵ‖22=σ2tr((X+)⊤X+). With X+=X⊤(XX⊤)−1, (X+)⊤X+=(XX⊤)−1XX⊤(XX⊤)−1=(XX⊤)−1, so the variance is σ2𝔼tr((XX⊤)−1). Now XX⊤=∑j=1p(colj)(colj)⊤ is Wishart with k=n, m=p; for p>n+1, 𝔼[(XX⊤)−1]=1p−n−1In,variance=σ2np−n−1.

Therefore R(p)=(1−np)r2+σ2np−n−1,p≥n+2. As p↓n+1 the variance →∞ (the peak again); as p→∞ the variance →0 while the bias ↑r2. The interpolator never forgets that it has thrown away a p−np fraction of the signal.

Part 3 — the second descent

On (n+1,∞) differentiate: R′(p)=nr2p2−σ2n(p−n−1)2. Setting R′(p)=0 gives r2p2=σ2(p−n−1)2. Both sides positive, so take positive roots: rp=σp−n−1, i.e. r(p−n−1)=σp, whence p\*=(n+1)rr−σ. Since rr−σ>1 exactly when r>σ>0, an interior minimizer with p\*>n+1 exists iff r>σ (signal norm exceeds the noise level). When r≤σ, R′(p)<0 for all large p and R decreases monotonically to its infimum r2: the best overparameterized model merely matches the null predictor.

Evaluate R(p\*). Writing a=n+1, one has p\*−a=aσr−σ, and nr2p\*=nr(r−σ)a,σ2np\*−a=σn(r−σ)a. Hence R(p\*)=r2−nr(r−σ)a+σn(r−σ)a=r2−n(r−σ)2a, that is R(p\*)=r2−nn+1(r−σ)2,r>σ.

Interpretation. The null predictor β^≡0 has excess risk ‖β\*‖22=r2. The optimally overparameterized interpolator beats it by exactly nn+1(r−σ)2>0. So heavy overparameterization is genuinely useful, but only when r>σ; the benefit is the squared gap between signal and noise scales, discounted by nn+1.

Closing notes


Two new problems every morning at 8am · every day so far