8 6 2 Exact Rank Recovery

We consider the multivariate linear model (8.2):

(1)
\begin{align} Y = \mathbf{X} A^* + E . \end{align}

We denote by $r^*$ the rank of $A^*$, by $\sigma_1(M) \geq \sigma_2(M) \geq \ldots$ the singular values of $M$ ranked in decreasing order, and we consider the selection procedure (8.13):

(2)
\begin{align} \widehat{r} \in \underset{r}{\operatorname{argmin}}\left\{\left\|Y-\mathbf{X} \widehat{A}_r\right\|_F^2+\lambda r\right\} \end{align}

where the estimators $\widehat{A}_r$ are defined by (8.6):

(3)
\begin{align} \widehat{A}_r \in \underset{\operatorname{rank}(A) \leq r}{\operatorname{argmin}}\|Y-\mathbf{X} A\|_F^2 \end{align}

and with the choice:

(4)
\begin{align} \lambda = K \left( \sqrt{T} + \sqrt{q} \right)^2 \sigma^2 \end{align}

We recall that we denote

(5)
\begin{align} P = \mathbf{X} \left( \mathbf{X}^\top \mathbf{X} \right)^{+} \mathbf{X}^\top \end{align}

for the projection onto the range of $\mathbf{X}$, with $\left( \mathbf{X}^\top \mathbf{X} \right)^{+}$ the Moore-Penrose pseudo-inverse of $\mathbf{X}^\top \mathbf{X}$.

1. We admit the following result from Exercise 8.6.1, Question 2:

(6)
\begin{align} \widehat{r} = \max \left\{ r : \sigma_r (PY) \geq \sqrt{\lambda} \right\} \end{align}

which is equivalent to the characterisation:

(7)
\begin{align} \widehat{r} = r \iff \left( \sigma_r (PY) \geq \sqrt{\lambda} \quad \text{and} \quad \sigma_{r+1} (PY) < \sqrt{\lambda} \right) \end{align}

Applying this for $r=r^*$ and taking the complement we get:

(8)
\begin{aligned} \mathbb{P} \left( \widehat{r} \neq r^* \right) &= \mathbb{P} \left( \sigma_{r^* + 1} (P Y) \geq \sqrt{\lambda} \quad \text { or } \quad \sigma_{r^*} (P Y) < \sqrt{\lambda} \right) \end{aligned}

2. Since $\operatorname{rank} (\mathbf{X} A^*) \leq r^*$, by the Weyl inequality (Theorem C.6 in Appendix C):

(9)
\begin{align} \sigma_{r^*} ( \mathbf{X} A^* ) - \sigma_{r^*} (P Y) \leq \sigma_1 (P E) \quad \text{and} \quad \sigma_{r^* + 1} (P Y) \leq \sigma_1 (P E) \end{align}

Therefore, the event $\left( \sigma_{r^* + 1} (P Y) \geq \sqrt{\lambda} \quad \text { or } \quad \sigma_{r^*} (P Y) < \sqrt{\lambda} \right)$ is included in $\left( \sigma_1 (P E) \geq \sqrt{\lambda} \quad \text { or } \quad \sigma_1 (P E) \geq \sigma_{r^*} ( \mathbf{X} A^* ) - \sqrt{\lambda} \right)$.

Hence, combined with the previous question:

(10)
\begin{aligned} \mathbb{P} \left( \widehat{r} \neq r^* \right) &\leq \mathbb{P} \left( \sigma_1 (P E) \geq \min \left( \sqrt{\lambda}, \sigma_{r^*} \left( \mathbf{X} A^* \right) - \sqrt{\lambda} \right) \right) \end{aligned}

3. Let us assume that:

(11)
\begin{align} \sigma_{r^*} \left(\mathbf{X} A^*\right) \geq 2 \sqrt{\lambda} \end{align}

Then, by the previous question:

(12)
\begin{aligned} \mathbb{P} \left( \widehat{r} \neq r^* \right) &\leq \mathbb{P} \left( \sigma_1 (P E) \geq \sqrt{\lambda} \right) \end{aligned}

By Lemma 8.3, $\mathbb{E} [ \sigma_1 (PE) ] \leq \left( \sqrt{T} + \sqrt{q} \right) \sigma = \frac{\sqrt{\lambda}}{\sqrt{K}}$. Since the map $\sigma_1 (P \cdot)$ is $1$-Lipschitz with respect to the Frobenius norm, the Gaussian concentration Inequality (B.2), page 301, states that $\sigma_1 (P E) - \mathbb{E} [ \sigma_1 (PE) ]$ is subgaussian $(\sigma^2)$.

Therefore, in this case, the probability to recover the exact rank $r^*$ is lower-bounded by:

(13)
\begin{aligned} \mathbb{P} \left( \widehat{r} = r^* \right) &\geq 1 - \mathbb{P} \left( \sigma_1 (P E) \geq \sqrt{\lambda} \right) \\ &\geq 1 - \mathbb{P} \left( \sigma_1 (P E) - \mathbb{E} [ \sigma_1 (PE) ] \geq (1 - \sqrt{K}^{-1}) \sqrt{\lambda} \right) \\ &\geq 1 - \exp \left(- \frac{( \sqrt{K} - 1 )^2 \lambda}{2 K \sigma^2} \right) \\ \mathbb{P} \left( \widehat{r} = r^* \right) &\geq 1 - \exp \left(- \frac{( \sqrt{K} - 1 )^2}{2} ( \sqrt{T} + \sqrt{q} )^2 \right) \end{aligned}
Unless otherwise stated, the content of this page is licensed under Creative Commons Attribution-ShareAlike 3.0 License