Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed
Computer Security · Lecture 8 of 22 · 1:22:08
Lecture 9: Securing Web Applications
Study guide
What this lecture covers
Continuing from the previous lecture on the same-origin policy, this lecture goes deeper into how real web applications fail and how they can be defended. It opens with live demonstrations of the Shellshock bug and a cross-site scripting (XSS) attack, then works through defenses: browser XSS filters, HTTP-only cookies, domain-based privilege separation, content sanitization (Django templates), less expressive markup languages, and Content Security Policy (CSP). It then covers SQL injection, session management with cookies versus stateless authentication using Message Authentication Codes (as in AWS request signing), and a set of side-channel attacks that let a malicious site infer which other sites a user has visited.
After watching, you can explain why naive content sanitization fails, describe how CSP restricts where scripts and other resources can load from, contrast cookie-based and MAC-based (stateless) authentication, and recognize link-color, cache-timing, DNS-timing, and rendering-timing attacks that leak browsing history.
Key ideas
- Shellshock: a Bash vulnerability where a malformed function definition passed through an HTTP header used as an environment variable tricks Bash into executing attacker-supplied commands after the function definition.
- Reflected XSS: an attacker's script is embedded in a URL and printed directly into the response HTML; it executes only while the user has that URL open.
- Persistent XSS: malicious HTML or script saved server-side (e.g. in a comment or profile) runs for every visitor until removed.
- Content sanitization: escaping special characters (like
<,>,") before inserting untrusted input into HTML, as Django's template system does by default; it reduces but does not eliminate XSS risk because HTML/CSS/JS grammars are complex and browsers tolerate malformed markup. - Content Security Policy (CSP): an HTTP response header that restricts which origins a page's scripts, images, and other resources may load from, blocks inline
<script>blocks, and disableseval. - HTTP-only cookies: a
Set-Cookieflag telling the browser not to expose a cookie to JavaScript, which blocks cookie theft via XSS but not cross-site request forgery. - Message Authentication Code (MAC) authentication: a stateless alternative to cookies (used by AWS) where the client signs each request with a shared secret key instead of sending a persistent session token.
- History-sniffing side channels: link-color inspection, cache-timing, DNS-timing, and render-timing attacks that let a malicious page infer which other sites a visitor has previously loaded.
Walkthrough
Live demos: Shellshock and XSS (0:00)
The lecture opens with two demonstrations. First, a toy CGI server that copies HTTP request headers into Bash environment variables is shown executing an attacker-supplied command via a malformed Bash function definition - the Shellshock bug. Second, a toy CGI server that prints a query-string value directly into HTML is used to demonstrate reflected XSS: a naive <script>alert('XSS')</script> payload is blocked by the browser's built-in XSS filter, but a malformed img"""><script> string confuses the HTML parser enough to bypass the filter and execute.
Browser XSS filters and privilege separation (14:16)
Browsers like Chrome and IE use heuristics (such as detecting a script tag embedded in a URL) to block likely XSS payloads, but these filters are incomplete and can be bypassed, and they cannot stop persistent XSS stored on the server. HTTP-only cookies stop JavaScript from reading a cookie but not CSRF-style requests. Sites like Google historically served user-submitted content from a separate domain (googleusercontent.com) so that an XSS bug in that content is contained away from the main domain.
Content sanitization and its limits (24:38)
Django's template system escapes characters like < and > into HTML entities so untrusted values can't break out of their intended context. This helps but is not foolproof: because HTML, CSS, and JavaScript grammars are ambiguous and browsers must tolerate broken markup to stay usable, clever payloads can still confuse the parser. A less expressive markup language, such as Markdown compiled to HTML, is easier to sanitize completely than full HTML/CSS/JS - though allowing inline HTML inside Markdown reopens the same risk.
CSP and SQL injection (34:44)
CSP lets a server declare, for example, that scripts may only load from self or *.mydomain.com; the browser then refuses any script reference outside that allowlist, blocks inline scripts, and disables eval and related dynamic code generation. The lecture then turns to SQL injection, showing how an unsanitized WHERE UserID = <input> query lets an attacker terminate the query early and append commands like DELETE TABLE. The fix is rigorous escaping of untrusted input before it reaches the database, which frameworks like Django provide but developers sometimes bypass for performance.
Cookies and stateless MAC authentication (46:03)
Session cookies map a random session ID to server-side user state, making the cookie a high-value target for theft. As an alternative, the lecture explains stateless authentication using Message Authentication Codes, illustrated with AWS request signing: each request carries the user's ID in the clear plus a signature computed from a shared secret key over key parts of the request (verb, content hash, content type, date, resource path). This avoids a persistent session token but raises its own issues, including where the secret key is stored and how requests are protected from replay (partly addressed with expiration timestamps). DOM storage and client-side certificates are mentioned as further cookie alternatives.
GIFAR, protocol parsing bugs, and history-sniffing side channels (1:03:38)
Inconsistent URL parsing between Flash and browsers could let a script from one origin run with another origin's authority. The GIFAR attack combined a GIF (parsed top-down) with a JAR/Java applet (parsed bottom-up) into one file that passed image-upload validation but executed as code when referenced by an injected applet tag. The lecture closes with covert channels that leak browsing history: reading visited-link colors (now blocked by browsers lying about link state to JavaScript), timing how fast cached resources load, timing DNS resolution for candidate hostnames, and timing how quickly a browser renders previously visited content - each of which an attacker's page can use to infer what sites a visitor has been to.
Before you watch
- Watch the previous lecture on the same-origin policy, cookies, and CSRF, since this lecture builds directly on those concepts (especially in the cookie and MAC discussion).
- Basic familiarity with HTTP requests/responses, shell scripting, and SQL queries is assumed.
Check your understanding
- Why did the toy CGI server demo make it possible to execute an attacker's command by setting a custom HTTP header, and how does this relate to how Shellshock affected real servers using Apache?
- Why can browser XSS filters and server-side content sanitization reduce but not eliminate cross-site scripting risk?
- How does a Content Security Policy header stop an XSS payload from exfiltrating data to an attacker's domain, and why does it also disable
eval? - In the AWS-style MAC authentication scheme, what does the server need to know in order to validate a signed request, and what problem does an expiration field help mitigate?
- Name two different side-channel techniques (other than reading link colors) that let a malicious site infer a user's browsing history, and explain why each works.
Vocabulary
- vulnerability (noun)
- A weakness in software that an attacker can exploit.
Shellshock is a vulnerability in the Bash shell. - malformed (adjective)
- Badly or incorrectly formed.
A malformed function definition tricks Bash into running extra commands. - environment variable (noun)
- A named value stored by the operating system that programs can read.
The server copies HTTP headers into environment variables. - cross-site scripting (XSS) (noun)
- An attack that injects malicious script into a page viewed by other users.
The demo shows a live cross-site scripting attack. - reflected XSS (noun)
- An XSS attack where the malicious script comes from the URL and is shown back immediately in the response.
Reflected XSS only affects users who click the crafted link. - persistent XSS (noun)
- An XSS attack where the malicious script is saved on the server and runs for every visitor.
Persistent XSS keeps attacking users until the bad comment is deleted. - payload (noun)
- The actual malicious code or data an attack delivers.
The XSS payload was disguised inside a broken image tag. - bypass (verb)
- To get past a security check without being stopped.
A clever string can bypass the browser's XSS filter. - heuristic (noun)
- A rough rule used to make quick decisions, not always exact.
Browsers use heuristics to guess whether a URL contains an attack. - privilege separation (noun)
- Splitting a system so different parts have different levels of access.
Serving user content from a separate domain is a form of privilege separation. - sanitize (verb)
- To clean untrusted input so it cannot cause harm.
Django's templates sanitize user input before showing it in HTML. - escape (verb)
- To convert special characters into a safe form so they are treated as plain text.
The template engine escapes < and > characters. - foolproof (adjective)
- So well designed that it cannot fail or be misused.
Sanitization helps but is not foolproof against every attack. - ambiguous (adjective)
- Having more than one possible meaning or interpretation.
HTML's grammar is ambiguous, which browsers must tolerate. - markup language (noun)
- A system of tags or symbols used to structure text, like HTML.
Markdown is a simpler markup language than full HTML. - Content Security Policy (CSP) (noun)
- A browser security header that limits which sources a page may load scripts and content from.
CSP blocks any script that isn't listed in the allowed sources. - allowlist (noun)
- A list of sources or items that are explicitly permitted.
CSP works by defining an allowlist of trusted origins. - inline script (noun)
- JavaScript code written directly inside an HTML page rather than in a separate file.
CSP blocks inline script blocks to reduce XSS risk. - SQL injection (noun)
- An attack that inserts malicious database commands through unsanitized input.
SQL injection can let an attacker delete an entire table. - terminate (verb)
- To end something early or unexpectedly.
The attacker's input terminates the query early. - session (noun)
- A period of interaction between a user and a server, often tracked by a cookie.
A session cookie links requests to the logged-in user's data. - stateless (adjective)
- Not keeping stored information between requests.
MAC authentication is a stateless alternative to session cookies. - Message Authentication Code (MAC) (noun)
- A short code computed from data and a secret key, used to prove a message was not changed.
AWS signs each request with a Message Authentication Code. - shared secret key (noun)
- A private value known only to the two parties that need to trust each other.
Both the client and server hold the shared secret key for signing requests. - replay attack (noun)
- Reusing a previously captured valid message to trick a system.
An expiration timestamp helps prevent a replay attack. - covert channel (noun)
- A hidden way of passing information that was not intended to be used that way.
Timing attacks act as a covert channel for browsing history. - side channel (noun)
- An indirect way of learning secret information, based on things like timing rather than the data itself.
Cache-timing is a side channel that leaks browsing history. - infer (verb)
- To work out something indirectly from clues.
The attacker can infer which sites a user has visited. - cache (noun)
- A temporary storage area that keeps recently used data for faster access.
A resource that loads instantly is probably already in the cache. - resolve (DNS) (verb)
- To look up and find the address matching a domain name.
Timing how fast a name resolves can reveal earlier visits.
From the YouTube description
MIT 6.858 Computer Systems Security, Fall 2014
View the complete course: http://ocw.mit.edu/6-858F14
Instructor: James Mickens
In this lecture, Professor Mickens continues looking at how to build secure web applications.
License: Creative Commons BY-NC-SA
More information at http://ocw.mit.edu/terms
More courses at http://ocw.mit.edu
← Lecture 8: Web Security Model · Lecture 10: Symbolic Execution →
