What's the strongest AI model you can train on a laptop in five minutes?
TL;DR Highlight
Training a GPT-style transformer in just 5 minutes on a MacBook Pro — exploring optimal model size, dataset, and training configuration with measured results.
Who Should Read
Developers who want to train small language models hands-on or iterate quickly on ML experiments in a local environment. ML engineers trying PyTorch/MLX-based model training for the first time or looking for optimization sweet spots.
Core Mechanics
- Final result: ~1.8M parameter GPT-style transformer trained on ~20M tokens from TinyStories, achieving ~9.6 perplexity on held-out split in 5 minutes.
- Under a 5-minute constraint, a sweet spot exists between model size and training tokens — too large wastes time on initialization, too small hits capacity limits.
- On Apple Silicon, MPS-only activation is the key — torch.compile and float16 conversion can actually hurt due to launch overhead being the real bottleneck.
- Domain-specific fine-tuning on tiny datasets can produce surprisingly coherent outputs even at 1.8M parameters.
Evidence
- GPT-2 speedrun project (modded-nanogpt) techniques were suggested for further improvement: Muon optimizer, better weight initialization, learning rate tuning could achieve lower perplexity in the same time.
- Apple Silicon quirk: standard GPU optimization techniques (torch.compile, float16) actually degrade performance due to MPS backend's launch overhead characteristics.
- ~9.6 perplexity on TinyStories held-out split — generates grammatically correct simple stories
How to Apply
- For quick local training experiments: don't blindly apply 'standard optimizations' like torch.compile or float16. On Apple Silicon, just activate MPS and measure a baseline with a simple training loop first — complex optimizations can backfire due to launch overhead.
- For small domain-specific models: use a curated dataset like TinyStories as a template — quality and domain match matter more than dataset size at small scale.
- Use this as a learning exercise: 5 minutes to a working transformer helps build intuition about model size/data/training dynamics tradeoffs.
Terminology
Related Papers
Is One Layer Enough? A Single Transformer Layer Matches Full-Parameter RL Train
LLM의 RL 후처리 학습(post-training)에서 성능 향상의 대부분이 중간 레이어 소수에 집중되며, 단 하나의 레이어만 학습해도 전체 파라미터 학습과 비슷하거나 더 나은 결과를 낼 수 있다는 연구 결과. 이는 RL 학습 비용을 대폭 줄일 수 있는 가능성을 시사한다.
Knowledge Distillation of Black-Box Large Language Models (2024)
GPT-4 같은 내부 구조에 접근할 수 없는 독점 LLM에서 작은 모델로 지식을 효과적으로 전달하는 Proxy-KD 기법을 소개하는 논문으로, 전통적인 White-Box 방식보다 성능이 높다는 점에서 주목할 만하다.
Show HN: NanoEuler – GPT-2 scale model in pure C/CUDA from scratch
PyTorch나 autograd 없이 C와 CUDA만으로 GPT-2 수준의 LLM을 처음부터 구현한 교육용 프로젝트로, 역전파·BPE 토크나이저·FlashAttention까지 직접 손으로 작성했다.
Show HN: Neural Particle Automata
고정된 격자 대신 움직이는 파티클 위에서 동작하는 Neural Cellular Automata의 확장 버전으로, 형태 생성·포인트 클라우드 분류·텍스처 합성 등 다양한 작업에서 자기조직화 동작을 학습할 수 있다.
The annotated PyTorch training loop
PyTorch 학습 루프의 각 코드 줄이 왜 그 위치에 있어야 하는지, 순서를 바꾸거나 빠뜨렸을 때 어떤 문제가 생기는지를 단계별로 설명한 심층 가이드다.
When Good Verifiers Go Bad: Self-Improving VLMs Can Regress on New Tasks
VLM 자가학습 루프에서 verifier가 특정 태스크에 맞지 않으면 학습할수록 오히려 성능이 떨어지는데, DPO 손실값은 멀쩡히 내려가서 눈치채기도 어렵다.