Interactive
No install and no account. Each of these runs in your browser, in English or Spanish, and answers a question the pages can only describe: move a control and the shapes, the strides and the numbers move with it. Every one is also linked from the section it belongs to, in the handbook and in the notebooks.
🧮 Broadcasting simulator
Type two shapes. Watch NumPy line them up from the right, pad the short one, and stretch every axis of size 1 — or refuse, and say which axis. Section 03.
🧊 The image tensor
Three real photos as one 4-D tensor, a cube per byte, up to 128 × 128. Transpose it and watch the cubes follow the new shape while not one byte of the buffer moves; reshape it and watch the picture break; rewrite the buffer C, Fortran or channels_last and watch only the memory move. Every click prints the NumPy and PyTorch that does it. Section 04.
Straight to transpose, reshape or the memory layout.
👁️ Attention, from words to weights
One four-word sentence, “I know you know”, followed through a single attention head with every number small enough to add up by hand: the words cut into tokens, the tokens coded as ids, the ids fetching rows of an embedding table, those rows projected into queries, keys and values, the scores, why they are divided by √d_k, softmax turning them into weights that sum to one, and the weighted average that is the output. Then more heads, and the batch a real model is handed. Every picture carries the NumPy that computes it. Sections 04 and 06, and Appendix B.
Straight to why √d_k, the scores, softmax or every head gets its own axis.
📐 Projection & SVD stage
Least squares is a projection: scroll, and watch a target vector’s shadow land on the plane its predictors span, with the residual meeting it at a right angle. Then the SVD portal, where a unit circle becomes an ellipse and A v = σ u stops being notation, and on to the eigenvectors, two nearly collinear columns and float32 dividing by exactly zero. One problem, sections 07, 08 and 09.
Straight to the SVD portal, the eigenvectors or float32 dividing by zero.
🎙️ The audio tensor
Air becomes numbers: 48 000 a second, each rounded to 16 bits, written down as an array you can listen to. Then one window is cut out of it, the transform is asked which frequencies are in that window, and the window hops along to build a frequency-by-time matrix, 513 × 465 — and you can hear what a transpose does to it. It ends on the rank-4 batch a model is handed, three recordings stacked. Sections 00, 02, 04 and 09, and Appendix E.
Straight to one window, the transform, the best rank-k there is or the batch.
🚕 Tucker and CP
One real tensor, factorised two ways: 6,383 New York taxi trips, pickup borough by dropoff borough by hour of day, as a cube of voxels you can turn and click — three quarters of it one Manhattan route. Watch the cube lay itself flat into each unfolding, see what the SVD of one finds, and reassemble a small core plus three factor matrices: 480 numbers down to 102, at 6.7% error, one core entry carrying 99% of the fit. Then build the best rank-1 term there is, see what several of them look like, play alternating least squares one solve at a time from five starts that do not all agree, and spend one parameter budget both ways. Sections 10 and 11, and Appendix C.
Straight to the cube laid flat, the core and its three factors, one rank-1 term, ALS, solve by solve or the same budget spent two ways.