longer technical writeups
memory coalescing, shared mem, warp shuffling, and vectorized loading
worklog: can we learn a policy for pretraining data selection
a few grpo limitations
training a model to play wordle with grpo on modal for ~$5-6
intro to neuromorphic computing
stuff im learning, reading, paper dumps, etc.
Comparing optimizers — order, preconditioners, momentum, EMA, and curvature.
Notes from Stanford's Language Modeling from Scratch course.
Using TMA in the Hopper architecture to load data from global to shared memory.
Weight decay tied to the square of the learning rate
some saved links
learning something new everyday