MLn CLUB · WEEK 7
In-context Learning and Induction Heads
Reading
About this week
When does a model suddenly learn to learn — and can you spot the moment on the loss curve? If a circuit is defined to do nothing but copy random text, why does the same circuit also translate French?
The sequel to the Mathematical Framework paper, this work argues that induction heads — which complete patterns by finding and copying earlier occurrences — are the mechanism behind most in-context learning, and that the bulk of a model’s in-context ability appears in a sharp phase change exactly when these heads form.
Join us at CASI for discussion at 8 pm, (optional) quiet reading from 7 pm.