Tensors for Machine Learning — from real data to modern models. Ravi Kalia and Sebastian Laverde Chunza.

Predict first: same numbers — still a voice?

Play the voice as built, then switch the layout and play it again: the same 263,169 numbers, only in a different order.

Open the voice widget ↗

The workshop idea: four cards show scalar, vector, matrix and tensor as 0D, 1D, 2D and ND arrays, each defined by its shape. The guiding question of the day: before operating, what does each axis represent?

Four ideas, one object. The whole day in four cards. Card one: a tensor generalizes a matrix. A scalar has no axes, a vector one, a matrix two, and a tensor as many as the data needs. Card two: it holds the data. Images, video frames, taxi trips and prices are all one box of numbers, and every axis stands for something. Card three: its axes can be moved. Transpose permutes them. Reshape re-reads the same flat numbers, so the same shape can carry a different meaning. Card four: it factors. Numbers into primes, quadratics into roots, matrices into LU, QR and SVD, and tensors too. Where an inverse does not exist, the pseudoinverse answers. Carry this all day: before you operate on a tensor, name its axes.

Real data, real problems. Four real datasets used throughout the workshop: handwritten digits, stained histology, California housing, and NYC taxi trips. Each carries the role it plays later.

Agenda — 210 minutes

Start Time Duration (min) Part Segment Name
00:00 5 — Setup and welcome
00:05 20 I What a tensor is
00:25 20 II Thinking in N dimensions
00:45 30 III Indexing & broadcasting · Reshape & transpose
01:15 10 🎯 Kahoot 1 + break
01:25 15 III Video pipeline design (group)
01:40 15 IV Contraction with einsum
01:55 5 — Break
02:00 15 IV Inverses and the pseudoinverse
02:15 5 🎯 Kahoot 2
02:20 25 IV Recursion · Matrix factorizations
02:45 5 — Break
02:50 15 IV Tucker decomposition
03:05 5 🎯 Kahoot 3
03:10 15 IV Tensor factorizations
03:25 5 — Wrap-up and take-homes

00 · Setup and welcome

—

5 min

Practise today

Name possible axes and predict a slice in the entry check.

Explore later

After the entry check, revisit data-quality examples. Complete runtime setup before class.

Colab ↗

01 · What a tensor is

I

20 min

Practise today

Name image axes and distinguish shape, order, and element count.

Explore later

Extract slices and fibers, and unfold images into matrices.

Colab ↗

📖 NumPy to JAX ↗

Map of factorizations: six matrix and tensor factorizations arranged on a ladder from LU through Cholesky, QR, SVD, Tucker and CP, each with one phrase describing what it buys.

Section 09 walks this map on real data: every factorization written as a constrained optimization, the cost of each one derived and then measured, and SVD as the best rank-k approximation there is. The “which one, and what does it cost me” question this ladder raises is answered two hours later, in the room.

📖 Why so many matrix factorizations? ↗

What does a factorization give you? A four-step diagram — number, polynomial, matrix, tensor — shows that factorizing reveals useful structure rather than just rewriting an expression.

NumPy: one buffer, different views

An ndarray pairs values in memory with metadata that tells us how to read them.

import numpy as np
T = np.arange(24, dtype=np.float32).reshape(2, 3, 4)
T_t = T.transpose(2, 0, 1)
np.shares_memory(T, T_t)  # True
Array Shape Strides · bytes
T (2, 3, 4) (48, 16, 4)
T_t (4, 2, 3) (4, 48, 16)

Same data. Reordered axes.
Writing through the view also changes the original.

96 bytes

24 values × 4 bytes per float32
Both arrays share this data buffer.

Strides
Byte steps along each axis. Transpose reorders these steps.
Vectorization
Compiled numeric loops reduce Python overhead.
Broadcasting
Align shapes from the right: sizes match or one is 1. No tiled inputs.

Travis Oliphant · 2005
Combined Numeric and Numarray to create NumPy.

Numeric 1995  →  Numarray 2001
NumPy 2005  →  NumPy 1.0 2006

Explore NumPy ↗
Scientific Python’s array foundation

Hardware for tensor operations

The workload determines which architecture helps.

D = AB + C
Matrix multiply–accumulate

Architecture How it computes Best suited to Memory examples
CPUGeneral purpose Multiple cores + SIMD
Flexible control flow
General tasks
Low latency
Caches + DDR DRAM
GPUNVIDIA SIMT + Tensor Cores
CUDA programming platform
Parallel contractions
Matrix multiply–accumulate
GDDR or HBM
TPUGoogle · custom ASIC Systolic matrix units (MXUs)
Operands pass between neighbors
Dense matrix products
Reuse data across operations
On-chip buffers + HBM
Other ML ASICs Specialized dataflow Target workloads SRAM, DRAM or HBM

Memory and supported precision vary by device generation.

Moving data costs time.Reuse data on the device. Small operations and repeated transfers can erase the speedup.

Numerical libraries & ML frameworks

Familiar tensor operations sit above optimized numerical routines.

Inside a NumPy operation

A @ BMatrix productBLAS
np.linalg.svd(A)Singular value decompositionLAPACK
BLAS · standard operations at three levels
1 · Vector 2 · Matrix–vector 3 · Matrix–matrix
Dot product y = Ax C = AB

LAPACK builds on BLAS for linear systems, SVD and eigenproblems.

Accelerator libraries
cuBLAS · GPU BLAS
cuTENSOR · tensor contractions
oneDNN · neural network primitives

Frameworks add more

Automatic differentiation · device placement · compilation

PyTorch
Dynamic autograd
torch.compile
JAX
Composable transformations
jit · vmap · grad
TensorFlow
Eager execution + graphs
tf.function
Keras 3
Model API over
TensorFlow / JAX / PyTorch
Haiku / Flax
Neural network modules on JAX
Name the axes. Check the shapes.Those habits transfer across frameworks.

02 · Thinking in N dimensions

II

20 min

Practise today

Keep images and labels paired when shuffling; explain why shuffling time changes a sequence.

Explore later

Build padded video batches and interpret additional experimental axes.

Colab ↗

Batch is not time. Sixteen shuffled handwritten digits stay recognisable, while shuffled video frames turn the storm sequence meaningless. Reordering a batch axis is safe. Reordering a time axis is not.

Colab ↗

03 · Indexing and broadcasting real data

III

15 min

Practise today

Predict broadcasting shapes and standardize pixel columns safely when variance is zero.

Explore later

Select observations with named features, fancy indexing, and Boolean masks.

Colab ↗

Indexing in action: the same four indexing, masking, satellite-cropping and broadcasting panels, with a summary strip naming each technique.

Colab ↗

Predict first: what shape comes out?

Before you press anything, guess the result shape — then try two shapes that do not line up.

Open the broadcasting widget ↗

04 · Reshape and transpose real images

III

15 min

Practise today

Convert HWC to CHW and test that pixel values keep their meaning.

Explore later

Build image batches, compare NHWC with NCHW, and inspect memory layout.

Colab ↗

Transpose is not reshape: the histology transpose-versus-reshape comparison, plus a real NHWC-to-NCHW batch of a histology slide, an astronaut photo and a cat photo, each correctly reordered.

Colab ↗

📖 A tensor in pure Python ↗

Predict first: does a reshape survive this photo?

The cold open’s bug, on a photo this time: reshape instead of transpose on the same batch.

Open the image tensor widget ↗

Kahoot 1 — Tensor Vocabulary & Shapes

Active break: Kahoot 1, Tensor Vocabulary & Shapes. 6 questions covering axes, shapes, indexing and transpose. Join at kahoot.it; answer with intuition first, then justify with shapes.

Kahoot ↗

05 · Video pipeline design

III

15 min

Practise today

Trace sampled frames to source indices and defend a sampling strategy for a brief event.

Explore later

Compare complete pipelines, padding costs, and synchronized-camera layouts.

Colab ↗

Myth or fact?

Back from the break: three claims from sections 02–04. Hands up for myth.

  1. Shuffling four frames leaves their mean unchanged.
    Fact. 1.5 before and after — which is why the mean cannot tell you the order survived.
  2. NaNs after standardizing mean the formula has a bug.
    Myth. A column that never varies has a standard deviation of 0; three digit pixels never vary.
  3. If a transpose and a reshape both give shape (3, 2, 2), they hold the same values in the same places.
    Myth. The same entry, [1, 0, 0], is 1 after the transpose and 4 after the reshape.

File to frames to tensor to model: a four-stage pipeline showing how a video becomes a masked tensor, with the validity-mask formula spelled out.

Colab ↗

06 · Contraction with einsum

IV

15 min

Practise today

Contract colour with einsum, name surviving axes, and verify a weighted pixel sum.

Explore later

Express matrix operations with einsum and compare digit similarity measures.

Colab ↗

Anatomy of einsum: the expression np.einsum(‘ij,jk->ik’, A, B) is broken into its input labels, its shared index, and the indices that survive, with a diagram of the matrices each letter indexes.

NumPy to einsum: a table maps inner product, outer product, matmul, transpose, trace, diagonal, batched matmul and axis-sum to their einsum equivalents.

Colab ↗

07 · Inverses and the pseudoinverse

IV

15 min

Practise today

Show how duplicate columns give identical predictions from different coefficients; explain the minimum-norm choice of pinv.

Explore later

Verify the four Moore–Penrose identities, fit housing data, and apply pinv to unfolded images.

Colab ↗

Myth or fact?

Back from the break: three claims from sections 05–06. Hands up for myth.

  1. clip[1] is the second thing that happened in the video.
    Myth. clip keeps every 45th frame of 720, so clip[1] is source frame 45.
  2. np.einsum('ii->', A) adds up the whole matrix.
    Myth. A repeated letter keeps only the diagonal: 8 for [[5, 2], [7, 3]], whose total is 17.
  3. np.einsum('ij->', A) adds up the whole matrix.
    Fact. Two different letters, both summed away: 17.

Same predictions, different coefficients

For duplicated columns A = [x, x], predictions depend on the sum of the coefficients.

Coefficients Prediction Squared norm
(2, 0) 2*x 4
(0, 2) 2*x 4
(1, 1) 2*x 2

pinv(A) @ (2*x) selects (1, 1), the minimum-norm solution.

What remains unknown? The separate effects of the duplicated features.

Colab ↗

Predict first: how far does the answer move?

Nudge the target by 2% and watch what happens to the coefficients when two columns say almost the same thing.

Open the collinearity widget ↗

Explore later

California Housing and the pseudoinverse: a real 20,433-by-7 least-squares fit — predicted versus true housing values, and the residual histogram — computed with the Moore-Penrose pseudoinverse.

Colab ↗

📖 Rotate, stretch, rotate again ↗

Explore later

What about tensors? Three cards. First: no tensor inverse is in common use. That is a fair question with an honest answer, not a gap in your reading. Second: definitions do exist, built on the Einstein product and on the t-product for order-3 tensors, and they are active research. Third: in practice you unfold the tensor into a matrix, apply the matrix pseudoinverse, and fold the result back. A 4 by 3 by 5 tensor unfolds to 4 by 15, whose pseudoinverse is 15 by 4. M A-plus M returns M exactly, because unfolding loses nothing.

This is also where the second half’s through-line starts: the pseudoinverse in section 07, Tucker in section 10 and Richardson-Lucy deconvolution in take-home 13 are three instances of one idea — when a problem has no exact answer and no true inverse, find the best stable approximation instead.

Colab ↗

📖 Tensor inverses in practice ↗

Kahoot 2 — Einsum, Distance & the Pseudoinverse

Active break: Kahoot 2, Einsum, Distance & the Pseudoinverse. 6 questions covering einsum, similarity and the pseudoinverse. Join at kahoot.it; decide the operation first, then justify with indices and shapes.

Kahoot ↗

08 · Recursion with matrices and vectors

IV

10 min

Practise today

Explain a matrix state update and connect repeated updates with a matrix power.

Explore later

Use power iteration and evaluate recursive forecasts on airline traffic.

Colab ↗

Gradient descent is also recursion: the scalar, vector and matrix forms of gradient descent are shown side by side. The same update rule scales from a number to vectors and matrices.

Colab ↗

📖 The directions a matrix refuses to turn ↗ 📖 Can you invert a recursive function? ↗

09 · Matrix factorizations

IV

15 min

Practise today

Compare solver residuals and coefficient sensitivity; trace reduced coordinates through reconstruction.

Explore later

Benchmark other factorizations, compress images with SVD, and inspect NMF factors.

Colab ↗

Factor once, solve many: the same real least-squares problem three ways. Normal equations are fastest to write, and they square the condition number. QR is the stable default, and never forms A-transpose-A. SVD is the most expensive and the most informative, because Eckart-Young gives the truncation error without building the truncation.

Colab ↗

📖 Eigenvectors or singular vectors? ↗

10 · Tucker decomposition on real data

IV

15 min

Practise today

Explain core and factor shapes, measure storage and error, and choose Tucker ranks against an error limit.

Explore later

Implement HOSVD and inspect unfoldings, core interactions, and rank sweeps.

Colab ↗

Myth or fact?

Back from the break: three claims from sections 07–09. Hands up for myth.

  1. If pinv returns an answer without complaint, the matrix was invertible.
    Myth. pinv answers for every matrix. With a duplicated column it answers for a rank-2 matrix, and P @ A is not the identity.
  2. A forecast 1% off one step ahead can be over 3% off twelve steps ahead.
    Fact. At w = 1.1, the error grows to 1% × 1.1¹² = 3.14%.
  3. Two fits can both be almost exact and still disagree on their coefficients.
    Fact. Their predictions differ by 0.000001 and their coefficients by 1.41.

Table to tensor to HOSVD to reconstruction: a four-step pipeline turns a trip table into an order-3 tensor, computes its HOSVD core and factors, and reconstructs it with visible error.

Colab ↗

📖 What a tensor factorization buys you ↗

Compression golf: hole 1

Store the taxi tensor in as few numbers as you can, with relative error under 7%. Fewest numbers wins.

  • Tee off at ranks (2, 2, 3): 102 numbers, 6.70% error.
  • The rank explorer prints your score under its heatmaps.
  • Post one line: ranks · numbers · error.

Predict first: which hour does the hour factor pick?

480 taxi counts become a small core and three factor matrices. The notebook printed one hour; find it in the picture.

Open the Tucker widget ↗

Kahoot 3 — Convolution & Tensor Decompositions

Active break: Kahoot 3, Convolution & Tensor Decompositions. 6 questions covering convolution, correlation, Tucker and CP. Join at kahoot.it; recognise the structure first, then choose the operator.

Kahoot ↗

11 · Tensor factorizations

IV

15 min

Practise today

Compare CP and Tucker on the same tensor using parameter counts, reconstruction errors, and the budget gap.

Explore later

Study Tensor Train and t-SVD, investigate CP uniqueness, and compress neural-network layers.

Colab ↗

Explore later

Four decompositions, four bargains. CP stores a sum of rank-1 components, R times I plus J plus K numbers, and is the choice when the components have to be read one by one. Tucker stores one subspace per mode plus a core, and is the choice when each mode needs a rank of its own. Tensor Train stores a chain of small cores, growing only linearly with order, and is the choice when a dense core would explode. And t-SVD takes an FFT along mode 3, matrix SVDs, and an inverse FFT, and is the choice when the third mode carries a meaning of its own.

Colab ↗

📖 CP or Tucker ↗

Compression golf: hole 2

Any model now. Relative error under 2%; fewest numbers wins.

  • golf("cp", 4) or golf("tucker", (4, 4, 3)) prints one scorecard line.
  • Hole 1 was won with 60 numbers. This bar costs more.
  • Post your best line.

Predict first: same budget — who wins?

The same parameter budget, spent as CP and spent as Tucker — once the pairs have their own numbers.

Open the budget widget ↗

One idea connects the workshop

Approximate useful structure when the exact solution is not enough.

Pseudoinverse

Section 07

House prices against the pseudoinverse's predictions for all 20,433 districts: a cloud of points around a dashed line of perfect prediction.

X.shape = (20433, 7): 20,433 equations, 7 unknowns, no exact solution. pinv(X) @ y is the least-squares fit; the dashed line is a perfect prediction.

Tucker

Section 10

Taxi trips per hour of the day as counted, drawn as bars, and as rebuilt from 60 numbers, drawn as a line. Both peak at hour 18, which is marked.

480 taxi counts kept as 60 numbers, ranks (3, 3, 1), at 4.69% error. Trips per hour, counted and rebuilt: both peak at 18:00.

Deconvolution

Take-home 13

The camera-operator photograph, the same crop blurred and noisy, and the Richardson-Lucy recovery, which is sharper but not exact.

A known 9 × 9 blur, undone approximately: Richardson–Lucy, 20 iterations, 18.9% → 13.1% error. Plausible, not exact.

hard problem → structure → useful approximation

Key idea: we are not always looking for algebraic exactness; we are looking for a representation that preserves what matters.

Colab ↗

12 · Wrap-up and take-homes

—

5 min

Practise today

Transfer axis reasoning to new data and explain why Tucker restores shape while losing information.

Explore later

Choose take-home work on PCA, attention, CP, Cholesky, audio, or deep dives 13–16.

Colab ↗

📖 Fourier finally clicked ↗