home

blogs

Worklogs and write-ups documenting my learning.

Optimizing Layer Normalization Kernel with CUDA

memory coalescing, shared mem, warp shuffling, and vectorized loading

GRPO's Limitations for Reasoning Tasks

a few grpo limitations

Train Qwen3-1.7B to play Wordle with GRPO

training a model to play wordle with grpo on modal for ~$5-6

How to Build a Brain (without losing yours)

intro to neuromorphic computing