NanoChat – The best ChatGPT that $100 can buy
TL;DR Highlight
Andrej Karpathy's LLM training framework: train a GPT-2-level model from scratch in ~4 hours ($100 or less) on 8xH100 GPUs and chat with it via a ChatGPT-style web UI.
Who Should Read
ML engineers who want to understand LLM internals by building one hands-on, or developers wanting to train small domain-specific models from scratch.
Core Mechanics
- nanochat is an experimental framework covering the full LLM pipeline — tokenizer → pretraining → fine-tuning → evaluation → inference → chat UI — in a single codebase. Minimal code designed for easy hacking.
- Core philosophy: turn a single '--depth' knob (transformer layer count) and all other hyperparameters (width, heads, learning rate, weight decay) auto-calculate to be compute-optimal. GPT-2 level is roughly depth 26.
- In 2019, training GPT-2 cost ~$43,000. With nanochat on 8xH100, it takes ~2 hours and $48. Spot instances can bring it to $15.
- Lineage: Karpathy's nanoGPT → Keller Jordan's modded-nanoGPT (extreme training speed optimization) → nanochat. Influenced by Muon optimizer (for linear layers, replacing AdamW) from modded-nanoGPT.
- Runs a GPT-2 Speedrun Leaderboard for the community to compete on how fast they can train GPT-2-level models. Evaluation metric: DCLM CORE score.
- Karpathy himself disclosed writing nearly 100% of the code manually — he tried Claude/Codex a few times but found them unhelpful because 'the code is too far from existing data distributions.'
Evidence
- The '$100' title was called misleading — it actually means $100 for 8xH100 cloud node rental, not local execution. Some were disappointed expecting local capability.
- Karpathy's disclosure about AI coding tools being unhelpful went viral — cited as evidence that AI coding agents still struggle with original code far from training data distributions.
- A user ran training live, sharing W&B links in real-time, promising to release the model 4 hours later. Community actively participating in experiments.
- Someone wanting to train on personal CPU even if it takes 3 months, but consensus was that meaningful results without GPUs are unrealistic.
How to Apply
- To experience the full LLM training pipeline end-to-end, nanochat's speedrun.sh script runs everything from pretraining to chat UI. Rent 8xH100 spot instances from Lambda Labs or RunPod for $15-48.
- For in-house small domain-specific model experiments, vary nanochat's '--depth' parameter to quickly compare models of different sizes. Lower depth dramatically cuts cost, ideal for prototyping.
- If researching LLM training optimization, join the GPT-2 Speedrun Leaderboard to experiment with training speed improvement techniques (Muon optimizer, custom schedulers) and share with the community.
Code Example
# nanochat quick start (on 8xH100 node)
git clone https://github.com/karpathy/nanochat.git
cd nanochat
bash runs/speedrun.sh # pre-training → inference → chat UI all at once
# adjust model size with depth parameter (GPT-2 scale = depth 26)
python nanochat/train.py --depth 26Terminology
Related Papers
Is One Layer Enough? A Single Transformer Layer Matches Full-Parameter RL Train
LLM의 RL 후처리 학습(post-training)에서 성능 향상의 대부분이 중간 레이어 소수에 집중되며, 단 하나의 레이어만 학습해도 전체 파라미터 학습과 비슷하거나 더 나은 결과를 낼 수 있다는 연구 결과. 이는 RL 학습 비용을 대폭 줄일 수 있는 가능성을 시사한다.
Knowledge Distillation of Black-Box Large Language Models (2024)
GPT-4 같은 내부 구조에 접근할 수 없는 독점 LLM에서 작은 모델로 지식을 효과적으로 전달하는 Proxy-KD 기법을 소개하는 논문으로, 전통적인 White-Box 방식보다 성능이 높다는 점에서 주목할 만하다.
Show HN: NanoEuler – GPT-2 scale model in pure C/CUDA from scratch
PyTorch나 autograd 없이 C와 CUDA만으로 GPT-2 수준의 LLM을 처음부터 구현한 교육용 프로젝트로, 역전파·BPE 토크나이저·FlashAttention까지 직접 손으로 작성했다.
Show HN: Neural Particle Automata
고정된 격자 대신 움직이는 파티클 위에서 동작하는 Neural Cellular Automata의 확장 버전으로, 형태 생성·포인트 클라우드 분류·텍스처 합성 등 다양한 작업에서 자기조직화 동작을 학습할 수 있다.
The annotated PyTorch training loop
PyTorch 학습 루프의 각 코드 줄이 왜 그 위치에 있어야 하는지, 순서를 바꾸거나 빠뜨렸을 때 어떤 문제가 생기는지를 단계별로 설명한 심층 가이드다.
When Good Verifiers Go Bad: Self-Improving VLMs Can Regress on New Tasks
VLM 자가학습 루프에서 verifier가 특정 태스크에 맞지 않으면 학습할수록 오히려 성능이 떨어지는데, DPO 손실값은 멀쩡히 내려가서 눈치채기도 어렵다.