Under construction. This chapter is still being developed and may change as the analysis and experiments are refined.
The usual objection is that Householder tridiagonalization costs O(n³) and Jacobi rotations immediately create fill-in. This chapter asks a different question: whether the transformed starting matrix can nevertheless reduce later Jacobi work. It analyzes energy redistribution and parallel pivot structure, then tests six symmetric matrix families with NVIDIA cuSOLVER.
Tridiagonalization does not give Jacobi a permanently sparse matrix, but it changes where the matrix energy is located before the Jacobi iterations begin.
The chapter therefore reaches a conditional conclusion: Householder preprocessing can substantially help a parallel Jacobi eigensolver, but not for every matrix.
Because the possible benefit is not the permanent preservation of tridiagonal zeros. Tridiagonalization changes the starting distribution of diagonal and off-diagonal energy and concentrates all off-diagonal energy in the first band. That can give the standard first parallel Jacobi set a much stronger collection of pivots, although Householder similarity can also move energy away from the diagonal, so the effect is matrix-dependent.
Experiments with NVIDIA cuSOLVER on six families of dense symmetric matrices showed a consistent separation between five ordinary test families and one deliberately adverse unequal variance-covariance family.
In the eigenvalue-only experiments, the complete Householder preprocessing pipeline had lower mean runtime than standalone Jacobi for all five ordinary families from n = 32 onward on the RTX 3060 and RTX 5060 Ti, and from n = 128 on the A100. Averaged across the five ordinary families, the complete-pipeline runtime reduction ranged from 7.5% to 10.0% at n = 128, and from 20.8% to 31.6% at n = 1024. Repeating the RTX 3060 experiment with five different base random seeds produced the same qualitative family-level behavior.
A separate experiment requested eigenvectors in both paths. With the explicit-Q reconstruction used here, the same qualitative separation remained, but the onset of sustained positive mean runtime reduction moved to larger matrix sizes for several families.
The unequal variance-covariance family behaved in the opposite direction on all three GPUs. Its large initial diagonal spread allowed substantial unfavorable transfer of energy away from the diagonal, followed by more Jacobi work and longer total runtime.
A separate schedule experiment kept the standard first Jacobi set but replaced only the second merry-go-round set. This deterministic substitution produced a modest reduction in complete sweep count across the tested matrix sizes and families.
The chapter is available as a PDF. Page links below are best-effort: most browsers support them but some viewers may ignore the page hint.