what i'm learning and notes i take from it.
Comparing optimizers — order, preconditioners, momentum, EMA, and curvature.
Notes from Stanford's Language Modeling from Scratch course.
Using TMA in the Hopper architecture to load data from global to shared memory.
interesting blogs or papers and my notes from them.
Weight decay tied to the square of the learning rate
zines!
A visual explainer on MoE in transformers.