Neural Networks: Zero to Hero
TL;DR Highlight
Andrej Karpathy teaches everything from backprop to GPT by building it in code — hands-on lectures for engineers who learn best by implementing.
Who Should Read
Software engineers and ML practitioners who want deep intuition for how neural networks and LLMs actually work, not just how to use APIs.
Core Mechanics
- Karpathy's lecture series (Neural Networks: Zero to Hero) covers the full stack from basic backpropagation through modern GPT architecture, all built from scratch in Python/PyTorch.
- The pedagogical approach: implement everything yourself rather than using libraries as black boxes — you understand autograd by building it, not by reading about it.
- The series covers: micrograd (backprop from scratch), makemore (bigram → MLP → attention), and nanoGPT (a minimal but complete GPT implementation).
- The content is targeted at engineers with Python knowledge but limited ML background — accessible without requiring a deep math background.
- nanoGPT became a widely used reference implementation because it's readable, not just functional.
- The lectures are freely available on YouTube and have become a standard self-study resource for engineers entering ML.
Evidence
- The series has millions of views and is frequently cited as the best free resource for engineers learning ML fundamentals.
- HN discussions of the series are consistently positive, with experienced ML engineers recommending it even to practitioners with existing backgrounds.
- nanoGPT's GitHub repo has tens of thousands of stars and is regularly forked for research experiments — evidence of practical utility beyond just education.
- Several professional ML engineers noted that working through the series filled gaps in their understanding that years of using high-level frameworks hadn't addressed.
How to Apply
- Work through the series sequentially — don't skip micrograd, even if you already use autograd. The implementation details matter for debugging mental models.
- After each lecture, try to extend the implementation yourself before looking at solutions — the struggle is where the learning happens.
- Use nanoGPT as a starting point for research experiments: it's small enough to fit in your head and modify confidently.
- After completing the series, you'll have the foundation to read ML papers directly rather than relying on blog post summaries.
Terminology
Related Papers
Is One Layer Enough? A Single Transformer Layer Matches Full-Parameter RL Train
LLM의 RL 후처리 학습(post-training)에서 성능 향상의 대부분이 중간 레이어 소수에 집중되며, 단 하나의 레이어만 학습해도 전체 파라미터 학습과 비슷하거나 더 나은 결과를 낼 수 있다는 연구 결과. 이는 RL 학습 비용을 대폭 줄일 수 있는 가능성을 시사한다.
Knowledge Distillation of Black-Box Large Language Models (2024)
GPT-4 같은 내부 구조에 접근할 수 없는 독점 LLM에서 작은 모델로 지식을 효과적으로 전달하는 Proxy-KD 기법을 소개하는 논문으로, 전통적인 White-Box 방식보다 성능이 높다는 점에서 주목할 만하다.
Show HN: NanoEuler – GPT-2 scale model in pure C/CUDA from scratch
PyTorch나 autograd 없이 C와 CUDA만으로 GPT-2 수준의 LLM을 처음부터 구현한 교육용 프로젝트로, 역전파·BPE 토크나이저·FlashAttention까지 직접 손으로 작성했다.
Show HN: Neural Particle Automata
고정된 격자 대신 움직이는 파티클 위에서 동작하는 Neural Cellular Automata의 확장 버전으로, 형태 생성·포인트 클라우드 분류·텍스처 합성 등 다양한 작업에서 자기조직화 동작을 학습할 수 있다.
The annotated PyTorch training loop
PyTorch 학습 루프의 각 코드 줄이 왜 그 위치에 있어야 하는지, 순서를 바꾸거나 빠뜨렸을 때 어떤 문제가 생기는지를 단계별로 설명한 심층 가이드다.
When Good Verifiers Go Bad: Self-Improving VLMs Can Regress on New Tasks
VLM 자가학습 루프에서 verifier가 특정 태스크에 맞지 않으면 학습할수록 오히려 성능이 떨어지는데, DPO 손실값은 멀쩡히 내려가서 눈치채기도 어렵다.