← all weeks
MLn CLUB · WEEK 7

In-context Learning and Induction Heads

Reading

About this week

When does a model suddenly learn to learn — and can you spot the moment on the loss curve? If a circuit is defined to do nothing but copy random text, why does the same circuit also translate French?

The sequel to the Mathematical Framework paper, this work argues that induction heads — which complete patterns by finding and copying earlier occurrences — are the mechanism behind most in-context learning, and that the bulk of a model’s in-context ability appears in a sharp phase change exactly when these heads form.

Join us at CASI for discussion at 8 pm, (optional) quiet reading from 7 pm.

Handwritten discussion notes — 'Is most of intelligence really induction heads?', the [A][B]…[A] copy pattern, a User/AI prompt, and a student/teacher curve
Week 7 — induction heads and in-context copying.