Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed

Matrix Methods for Data Analysis & ML · Lecture 1 of 36 · 7:04

Course Introduction to Matrix Methods for Data Analysis

Course Introduction of 18.065 by Professor Strang on YouTube

Study guide

What this lecture covers

Professor Gilbert Strang opens 18.065 by explaining why matrix methods sit at the center of modern data analysis and machine learning. Rather than diving into proofs, he sketches the shape of the whole course: which topics matter, why they connect, and where the course is headed.

You come away knowing what to expect from the term: a course built around the best matrices in linear algebra, the mechanics of deep learning, the optimization that trains it, and the statistics that keeps the numbers well behaved.

Key ideas

  • Symmetric and orthogonal matrices: Strang calls these the stars of linear algebra, and factoring a matrix into combinations of them is central to the course.
  • Singular value decomposition (SVD): a factorization of a matrix into orthogonal times diagonal times orthogonal, described as critical but often skipped in standard linear algebra courses.
  • Deep learning as function construction: a learning function takes input data (an image, handwriting, speech) and produces an output (a label, a zip code digit, a meaning).
  • Matrix multiplication plus nonlinearity: the learning function alternates matrix multiplications with a simple nonlinear step, F(x) = x for positive x and F(x) = 0 for negative x, since a purely linear function would fail.
  • Optimization: training means finding the matrix entries that minimize error, a multivariable calculus problem with hundreds of thousands of variables.
  • Statistics: keeping the mean and variance of values under control as matrices multiply repeatedly, since products can otherwise explode or vanish.

Before you watch

  • No prior background is assumed for this introduction; it previews the course rather than teaching technical content.
  • Familiarity with basic matrix multiplication is helpful for following the later descriptions.

Check your understanding

  1. Why does Strang say linear algebra alone cannot build a working learning function?
  2. What role does the nonlinear function F(x) play between matrix multiplications?
  3. Why does deep learning training require ideas from statistics as well as optimization?
  4. What makes the singular value decomposition an important factorization for this course?

Vocabulary

matrix (noun)
A rectangular grid of numbers arranged in rows and columns.
A matrix can represent the weights inside a neural network layer.
pillar (noun)
One of the main supporting parts of a larger structure or plan.
The course is built around four main pillars.
sketch (verb)
To describe something roughly, without full detail.
Strang sketches the shape of the whole course in the first lecture.
symmetric matrix (noun)
A matrix that looks the same if you flip it across its diagonal.
Symmetric matrices are one of the stars of linear algebra.
orthogonal (adjective)
At a right angle to something else, or independent in direction.
Orthogonal vectors point in completely independent directions.
factorization (noun)
Breaking a matrix into a product of simpler matrices.
The singular value decomposition is one key factorization.
singular value decomposition (noun)
A way to break any matrix into three simpler matrices that reveal its structure.
The singular value decomposition works for any matrix, even a rectangular one.
diagonal (adjective)
Describing a matrix with nonzero values only along the line from top-left to bottom-right.
The middle matrix in the SVD is diagonal.
nonlinear (adjective)
Not following a straight-line relationship between input and output.
A nonlinear step lets the network learn more than straight lines.
minimize (verb)
To make something as small as possible.
Optimization finds the matrix entries that minimize the error.
multivariable calculus (noun)
A branch of math dealing with functions of many variables at once.
Training a network is a multivariable calculus problem.
variance (noun)
A measure of how spread out a set of values is.
Statistics keeps the variance of values under control during training.
explode (verb)
To grow extremely large very quickly, out of control.
Repeated matrix products can explode if not controlled.
vanish (verb)
To shrink toward zero and disappear.
Values can vanish after many multiplications if not managed carefully.
term (noun)
A single semester-long period of study at a school.
This is what to expect from the term ahead.
proof (noun)
A step-by-step argument showing a mathematical statement must be true.
The course sketches ideas rather than diving into every proof.
handwriting (noun)
Text written by hand, used here as an example of image data.
A learning function can take handwriting as its input.
entries (noun)
The individual numbers that make up a matrix.
Training finds the matrix entries that minimize error.
well behaved (adjective)
Describing values that stay stable and don't grow or shrink out of control.
Statistics keeps numbers well behaved as matrices multiply.
shape of the term (phrase)
A general preview of what an upcoming period of study will look like.
Strang sketches the shape of the term to come.

Chapters

From the YouTube description

MIT 18.065 Matrix Methods in Data Analysis, Signal Processing, and Machine Learning, Spring 2018
Instructor: Gilbert Strang
View the complete course: https://ocw.mit.edu/18-065S18
YouTube Playlist: https://www.youtube.com/playlist?list=PLUl4u3cNGP63oMNUHXqIUcrkS2PivhN3k

Professor Strang describes the four topics of the course: Linear Algebra, Deep Learning, Optimization, Statistics. He provides examples of how Linear algebra concepts are key for understanding & creating machine learning algorithms.

License: Creative Commons BY-NC-SA
More information at https://ocw.mit.edu/terms
More courses at https://ocw.mit.edu

Interview: Gilbert Strang on Teaching Matrix Methods →