References and further reading
Books, papers, software, and worked examples
Look up a concept or explore it further. For tensors, start with the Kolda–Bader survey. Then choose the paper for the method you need.
Background: Chapter 2 of Deep Learning. See the pre-work before the session.
Linear algebra
From an introduction to numerical methods.
- Strang, G. — Introduction to Linear Algebra — an introduction.
Author pages: Gilbert Strang - Trefethen, L. N. & Bau, D. — Numerical Linear Algebra — numerical methods and SVD.
Author pages: Nick Trefethen - Golub, G. H. & Van Loan, C. F. — Matrix Computations — matrix algorithms.
Author pages: Gene Golub · Charles Van Loan
Tensors
Start with the survey. Then choose a method.
- Kolda, T. G. & Bader, B. W. (2009). Tensor Decompositions and Applications, SIAM Review 51(3), 455–500 — the survey; start here.
Author pages: Tamara Kolda · Brett Bader - Ballard, G. & Kolda, T. G. (2025). Tensor Decompositions for Data Science, Cambridge University Press — the textbook the survey grew into; the authors keep a full draft free to read.
Author pages: Grey Ballard - Hong, D., Kolda, T. G. & Duersch, J. A. (2020). Generalized Canonical Polyadic Tensor Decomposition, SIAM Review 62(1), 133–163 — read this one when squared error is the wrong question: counts, binary data, anything whose noise is not Gaussian. Table 1 is the list of losses. Kolda also gives the talk.
Author pages: David Hong - Kruskal, J. B. (1977). Three-way arrays: rank and uniqueness of trilinear decompositions, with application to arithmetic complexity and statistics, Linear Algebra and its Applications 18(2), 95–138 — the condition under which a CP decomposition is essentially unique, which is what section 11 claims and this is where it comes from.
- Kolda, T. G. (2021). Monkey BMI Tensor Dataset — the 43 × 200 × 88 neural tensor deep dive 14 decomposes, with the preprocessing that produced it. Cite it if you use it.
- Tucker, L. R. (1966). Some mathematical notes on three-mode factor analysis, Psychometrika 31, 279–311 — Tucker decomposition.
- Carroll, J. D. & Chang, J.-J. (1970). Analysis of individual differences in multidimensional scaling via an N-way generalization of “Eckart-Young” decomposition, Psychometrika 35, 283–319 — CP decomposition.
- Harshman, R. A. (1970). Foundations of the PARAFAC procedure, UCLA Working Papers in Phonetics 16, 1–84 — PARAFAC, also known as CP.
Author pages: Richard Harshman - Eckart, C. & Young, G. (1936). The approximation of one matrix by another of lower rank, Psychometrika 1, 211–218 — optimal low-rank matrix approximation.
Software
Implementations for your own projects.
- Harris, C. R., Millman, K. J., van der Walt, S. J. et al. (2020). Array programming with NumPy, Nature 585, 357–362 — the foundational NumPy paper: array programming and the scientific Python ecosystem.
tensorly— implementations of Tucker and CP.
Author pages: Jean Kossaifipyttb— the Tensor Toolbox in Python, from the authors of the survey. Use it for the things tensorly has no equivalent of: gcp_opt fits CP under a loss you choose.
Posts from the ML blog
One idea per post, with examples.
Every post lives on The ML blog.
- Why so many matrix factorizations?
- Factorizations as optimization
- Rotate, stretch, rotate again
- The directions a matrix refuses to turn
- Can you invert a recursive function?
- Tensor factorizations and inverses
- What a tensor factorization buys you
- Tensor inverses in practice
- Tensor inverses, worked through
- Sparse tensors
- Attention as two contractions
Related resources
- Prerequisites — prepare for the session.
- Handbook — theory, exercises, and solutions.
- Companion — NotebookLM summaries and practice questions.