A weekly ML ’n related topics reading group.
We meet weekly on Mondays in person to talk about new ML papers! As always, optional quiet reading from 7 pm, discussion starts 8 pm. Come and join the discussion! Click any week to register and grab the reading list, notes & resources for running the discussion yourself.
Mondays · 7–9pm quiet reading 7pm · discussion 8pm follow on Luma ↗
What's this?
What's this?
- A super warm group of folks discussing their favorite topics!
- In the first half, we host an optional quiet reading space
- In the second half, we have a discussion where people can talk about what they found interesting about the reading and ask questions about things they didn't understand
When/Where:
- CMU AI Safety Initiative's Office, 201 Craig Street, right across the PNC bank. Look for the open door up the stairs.
- 8pm discussion, 7pm optional quiet reading time.
Here's how it usually goes:
- 7:00 PM
- arrival and settling in
- 8:00 PM
- introductions
- 8:10 PM
- discussion time
- 9:00 PM
- wrap up then open discussion
Who's it for?
People who've been wanting to read up on the latest papers in ML and other fields but just haven't been able to find the time/motivation.
Why:
- We've been procrastinating too much on our readings, even though we have so much fun doing them. We know we're not alone in this and want to keep others accountable for learning more about what they're passionate about!
- We've also met a ton of really fun friends by discussing what we care about!
Rules/guidelines on how to act:
- Act like a host, include people in conversations, talk to people even if they're strangers, offer to explain what you know, and keep an open mind! come to read stuff and find super fun friends :)
- Bring snacks if you're feeling kind!
-
upcoming WEEK 13 Sep 7Attention Residuals: Rethinking Information Flow in LLMs
Standard residual connections with PreNorm accumulate every layer’s output with fixed unit weights, letting the hidden state grow uncontrollably with depth and diluting each layer’s unique contribution — what if each layer could instead learn which earlier representations to read from?
open week → - reading now
upcoming WEEK 12 Aug 31Manifold-Constrained Hyper-Connections
How does constraining residual connections to a geometric manifold restore identity mappings and enable stable scaling in large language models? Can richer cross-layer connectivity improve reasoning and representation quality without sacrificing optimization stability?
open week → -
past WEEK 11 Jul 27Fundamental Limitations of Single-Vector Embeddings
Why does something as simple as “finding people who like apples” break our best models? What do alternatives like multi-vector models or cross-encoders mean for real products?
open week → -
past WEEK 10 Jul 20A global workspace in language models
How can we identify and understand the internal “workspace” where large language models integrate information, reason, and broadcast concepts across their networks? Can mechanistic interpretability reveal whether LLMs develop functional equivalents of global information sharing found in cognitive architectures?
open week → - 🎉 symposium week
past WEEK 9 Jul 13Symposium Week
A little different this week — no single paper. Bring something that excites you and present it: great creative work is historically done together, across backgrounds and disciplines, in small, high-trust groups.
open week → -
past WEEK 8 Jul 6Nested Learning: The Illusion of Deep Learning Architectures
How might nested optimization structures allow models to generalize more efficiently by reusing global representations while adapting locally to new tasks? Could the hierarchical nature of nested learning make large models inherently more resilient to poisoning or backdoor attempts by isolating adversarial influence within inner-loop adaptations?
open week → -
past WEEK 7 Jun 29In-context Learning and Induction Heads
When does a model suddenly learn to learn — and can you spot the moment on the loss curve? If a circuit is defined to do nothing but copy random text, why does the same circuit also translate French?
open week → -
past WEEK 6 Jun 22Self-Distillation Enables Continual Learning
Why does on-policy learning dramatically reduce catastrophic forgetting? Does SDFT outperform SFT, RLHF-style pipelines, and continual pre-training?
open week → -
past WEEK 5 Jun 15A Mathematical Framework for Transformer Circuits
What does it mean to “fully understand” a model — and is a one-layer transformer just a lookup table? If the residual stream is only a communication channel, where does the computation actually live?
open week → -
past WEEK 4 Jun 8Touching the Elephant — TPUs
What if the real moat in AI isn’t the model, but the chip? Should AI run on custom silicon or general-purpose chips?
open week → -
past
WEEK 3 Jun 1Tiny Recursive Models
Using just a single 2-layer network with 7M parameters, TRM outperforms much larger language models by recursively refining both its current answer and internal reasoning state.
open week → -
past WEEK 2 May 25LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics
LeJEPA pushes in the opposite direction: instead of stabilizing training with heuristics, it builds a system where the objective itself prevents collapse.
open week → -
past WEEK 1 May 18ML Reading Group
Ilya Sutskever: “If you really learn [these 30 papers], you’ll know 90% of what matters today.”
open week →