Articles
FlashDance: Flash Attention vs Vanilla Attention
How much faster fused attention really is
Dope: Comparing Position Encodings
Four position encodings on longer sequences
Marcello
Making a style-transfer model that could not style-transfer
Automatic Downlink
Failure to build a VLM for satellite imagery
VocabVacation: Does Vocab Size Matter?
Training tiny transformers with different vocab sizes to find the sweet spot
LoRAdor: LoRA from Scratch
Implementing Low-Rank Adaptation in pure PyTorch
TokTok: A Spanish Tokenizer
Training a Spanish-only tokenizer and comparing it to GPT-4
Ultra Low Power Encoder
How to build an ultra low power encoder for LLMs