Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed

Machine Learning · Lecture 13 of 21 · 1:18:54

Lecture 12: Debugging ML Models and Error Analysis

Lecture 12 - Debugging ML Models and Error Analysis | Stanford CS229: Machine Learning (Autumn 2018) on YouTube

Study guide

What this lecture covers

Most learning algorithms do not work the first time, and the lecture's question is what to do next once yours does not either. Rather than guessing which fix to try, Andrew Ng walks through a set of diagnostics that turn debugging from a matter of intuition into a repeatable process: deciding whether a poor result comes from bias or variance, whether an optimizer is failing to converge or is optimizing the wrong objective, and which part of a multi-stage pipeline is actually responsible for your errors.

This lecture sits after the course has covered the core supervised learning algorithms, so it assumes you can already build and train a classifier; the focus here is entirely on what happens when that classifier underperforms. After watching, you should be able to read a learning curve to tell bias from variance, design a diagnostic to compare an optimization algorithm against its objective function, and run error analysis or ablative analysis on a pipeline to decide where to spend your time.

Key ideas

  • Bias vs variance diagnostic: plot training and dev/test error against training set size; a large gap between them signals high variance, while both errors staying high and close together signals high bias.
  • High variance fixes: more training data, fewer features, or a simpler model tend to help when the gap between training and dev error is large.
  • High bias fixes: more features, a more expressive model, or a different algorithm help when training error itself is too high.
  • Optimization algorithm vs optimization objective: if a competing model achieves a higher value of your own cost function than your trained model does, your optimizer failed to converge; if it achieves a lower value of your cost function but performs better in practice, you are optimizing the wrong objective.
  • Quick and dirty first pass: for an unfamiliar application, build a simple implementation first and let diagnostics tell you what to improve, rather than guessing in advance.
  • Error analysis: in a multi-stage pipeline, manually substitute ground-truth output at each stage (in dev-set order) and track how accuracy improves, which shows how much each stage is currently costing you.
  • Ablative analysis: for a system built by adding features or components, remove them one at a time and track how accuracy drops, which shows how much each one is contributing.
  • Judgment still matters: these diagnostics inform prioritization, not a formula; you still weigh them against how hard each fix would be.

Walkthrough

Framing debugging as a systematic workflow (4:42)

The lecture opens with a spam classifier example: logistic regression with regularization gets 20% test error, and a team typically has a long list of things it could try, from collecting more data to switching algorithms entirely. Ng argues that brainstorming the full list of options already puts you ahead of teams that pick one idea at random, but the real gain comes from having a diagnostic that tells you which option is likely to help before you spend days or weeks on it.

Reading bias and variance from learning curves (9:40)

The core diagnostic plots dev/test error and training error against the number of training examples. Training error typically starts near zero on tiny datasets and rises as the dataset grows, while dev error falls. A high-variance signature shows a wide, persistent gap between the two curves, meaning more data can help close it. A high-bias signature shows both curves converging early to an error level above your target, meaning more data alone will not fix the problem; you need more features, a more expressive model, or a different algorithm. The lecture ties this back to the four candidate fixes from the spam example, showing which ones address variance (more data, fewer features) and which address bias (more features, more expressive features).

Diagnosing the optimization algorithm vs the optimization objective (29:46)

A second diagnostic addresses a different failure mode: when an SVM outperforms your Bayesian logistic regression on the metric you actually care about, is the problem that gradient ascent has not converged, or that you are maximizing the wrong cost function? The test compares the value of your model's own objective function J evaluated at both models' parameters. If the competing model's parameters achieve a higher J, your optimizer failed to find the optimum, and the fix is to run it longer or switch to a better optimization method such as Newton's method. If your model achieves a higher J but the competing model still performs better on the metric you actually care about, J itself is the wrong thing to be maximizing, and the fix is to change the objective, for example by adjusting the regularization weight or switching cost functions entirely.

Applying the same diagnostic to a helicopter control problem (41:56)

Ng extends the framework to a three-part pipeline for autonomous helicopter flight: a simulator, a reinforcement learning algorithm that minimizes a cost function in simulation, and the cost function itself. When the learned controller flies worse than a human pilot, the diagnostic first checks whether the controller flies well in simulation; if it does but fails in real life, the simulator is the likely problem. If it also flies poorly in simulation, comparing the cost achieved by a human pilot against the cost achieved by the learned controller separates a bad optimizer (the human achieves a lower cost) from a bad cost function (the human achieves a higher cost yet flies better). The lecture notes this process can shift focus between the three components over successive rounds as each bottleneck is cleared.

Error analysis on a multi-stage pipeline (35:55)

Using a face-recognition pipeline as the example (background removal, face detection, eye/nose/mouth segmentation, then classification), the lecture describes replacing each stage's output with the ground truth, one stage at a time in pipeline order, and recording overall accuracy after each substitution. The stage whose perfect output produces the largest accuracy jump is the one currently limiting the system most, and is the best candidate to prioritize. Ng warns that this is a prioritization tool, not a hard rule, and recommends running the substitutions in both cumulative and non-cumulative order if the choice of order is likely to change the conclusion.

Ablative analysis for explaining what mattered (1:14:35)

Ablative analysis runs the opposite direction: starting from a system built up with several added features or components, remove them one at a time and measure how much accuracy drops. Applied to a spam classifier built from simple logistic regression plus features like spelling correction, sender host information, and text parsing, the component whose removal causes the biggest drop is the one that mattered most. This is useful both for deciding where further engineering effort should go and for explaining, in a report or paper, which parts of a system actually drove its performance.

Before you watch

  • Be comfortable with logistic regression, regularization, and the basic idea of gradient-based optimization, since the diagnostics are built directly on top of these.
  • Recall the general notion of bias and variance (underfitting vs overfitting) from earlier in the course; this lecture assumes you already know the definitions and focuses on diagnosing them in practice.
  • Some familiarity with SVMs is helpful, since one running example compares an SVM against logistic regression.

Check your understanding

  1. Given a plot of training and dev error against training set size, how do you decide whether a model has a high-bias or a high-variance problem?
  2. If a competing model achieves a lower value of your objective function J but performs better on the metric you care about, what does that tell you about your objective function?
  3. In the helicopter example, what evidence would point to the simulator being the bottleneck rather than the reinforcement learning algorithm or the cost function?
  4. How does error analysis on a pipeline differ from ablative analysis, and when would you use each?
  5. Why does Ng recommend building a quick, simple implementation first rather than a more sophisticated one, for an unfamiliar application?

Vocabulary

diagnostic (noun)
A test or procedure used to figure out what is causing a problem.
A diagnostic tells you whether a model has high bias or high variance.
learning curve (noun)
A plot showing how error changes as the amount of training data increases.
A learning curve reveals whether more data would help.
high variance (noun)
A problem where a model performs much better on training data than new data.
A wide gap between training and dev error signals high variance.
high bias (noun)
A problem where a model performs poorly on both training and new data.
High bias shows up as both errors staying high together.
convergence (noun)
The point where an optimization process settles at a stable result.
Failure to reach convergence means the optimizer stopped too early.
cost function (noun)
A formula that measures how well a model's predictions match reality.
Comparing which model achieves a lower cost function reveals the problem.
objective (noun)
The specific goal a learning algorithm is trying to optimize.
The wrong objective can lead to poor real-world performance despite good optimization.
pipeline (noun)
A sequence of processing stages that data passes through to reach a final result.
A face-recognition pipeline has several stages before final classification.
ground truth (noun)
The correct, verified answer used as a reference for checking accuracy.
Error analysis substitutes ground truth at each pipeline stage.
error analysis (noun)
Examining where a system's mistakes come from to guide improvement.
Error analysis shows which pipeline stage limits performance most.
ablative analysis (noun)
Removing parts of a system one at a time to measure their individual contribution.
Ablative analysis reveals which feature mattered most to accuracy.
bottleneck (noun)
The part of a system that limits its overall performance the most.
The simulator turned out to be the bottleneck in the helicopter pipeline.
cumulative (adjective)
Building up gradually by adding each step on top of the previous ones.
The substitutions can be tested in cumulative or non-cumulative order.
prioritize (verb)
To decide which tasks or fixes deserve attention first.
Diagnostics help prioritize which fix is worth trying first.
regularization weight (noun)
A setting controlling how strongly a model is penalized for complexity.
Adjusting the regularization weight can fix the wrong objective problem.
brainstorm (verb)
To generate a wide range of possible ideas or solutions.
Brainstorming the full list of fixes helps before committing to one.
substitute (verb)
To put one thing in place of another for a test or comparison.
Error analysis substitutes ground truth output at each stage.
segmentation (noun)
The process of dividing an image or data into meaningful parts.
The pipeline includes eye and mouth segmentation before classification.
controller (noun)
A system or program that decides what actions to take to control a machine.
The learned controller flies the helicopter using its trained policy.
simulator (noun)
A program that imitates a real system's behavior for testing without real-world risk.
The helicopter's simulator may be the bottleneck instead of the algorithm.

Chapters

From the YouTube description

For more information about Stanford’s Artificial Intelligence professional and graduate programs, visit: https://stanford.io/ai

Andrew Ng
Adjunct Professor of Computer Science
https://www.andrewng.org/

To follow along with the course schedule and syllabus, visit:
http://cs229.stanford.edu/syllabus-autumn2018.html

← Lecture 11: Backprop and Improving Neural Networks · Lecture 13: Expectation-Maximization Algorithms →