The purpose of today's post is to gather the essential linear algebra prerequisites that support deep learning. It is meant as a summary. Our topic is the fundamental building blocks of linear algebra: scalars, vectors, and matrices.
A scalar k∈R is a quantity that can be represented by a real number but has no direction. A vector, on the other hand, is a geometric object that has both a numerical magnitude and a direction. The conventional notation v=x1...xn, x1,...,xn∈R is in the form of a column matrix. The number of rows of a vector indicates the dimension of the space it belongs to, such that v∈Rn is an element of the n-dimensional real space.
x1,...,xn are called coordinate axes. Each coordinate axis measures the length of the vector along that axis. The origin O=(0,...,0)∈Rn is the starting point for all vectors. Therefore, every point in space represents a vector whose starting point is the origin.
Figure 1: An example of vector and point correspondence in two-dimensional space.
Unfortunately, no one can be told what the Matrix is. You have to see it for yourself.
— Morpheus
A matrix is the numerical representation of a linear transformation that keeps the origin fixed, preserves lines as lines, and does not break the parallelism of parallel lines. Formally, for a transformation T to be linear, it must satisfy the following two conditions:
T(u+v)=T(u)+T(v),T(kv)=kT(v)
This definition covers operations such as rotation, reflection, scaling, shearing, and projection.
Let A∈Rm×n. Then A is a matrix of size m×n, where m is the number of rows and n is the number of columns.
The intuition behind defining matrix multiplication this way is as follows: matrix multiplication is the composition of transformations. The expression AB means applying transformation B first and then transformation A. This intuition is the key to understanding why successive layers in deep learning are expressed through matrix multiplication. Matrix multiplication is not commutative: in general, AB=BA. Changing the order of transformations yields a different result.
Multiplying a matrix by a vector applies the linear transformation represented by that matrix to the vector.
[acbd][xy]=x[ac]+y[bd]=[ax+bycx+dy]
Let's work through an example:
A=[2113]∈R2×2 be a matrix and v=[11]∈R2 be a vector.
v′=Av=[2113][11]=[2⋅1+1⋅11⋅1+3⋅1]=[34]
The new vector resulting from the matrix multiplication v′=[34] is the linear transformation of v. That is, the vector has been transformed — its direction and magnitude have changed.
Figure 2: The orange vector represents v, and the blue vector represents v′.
Linear algebra is one of the best tools we have for modeling the spaces we scientifically explore. Fitting all those spaces onto paper is no small feat. If the reader wishes to dive deeper, they can explore the resources below.