Claude Haiku 4.5
TL;DR Highlight
Anthropic launched Claude Haiku 4.5, a small model delivering Sonnet 4-level coding performance at 1/3 the price and 2x+ the speed. A cost-effective option for developers needing agentic coding and real-time responses.
Who Should Read
Backend/full-stack developers running coding agents or chatbots on Claude API who want to reduce cost and latency. AI engineers looking for sub-agent models in multi-agent architectures.
Core Mechanics
- Claude Haiku 4.5 matches the coding performance of Sonnet 4 (a frontier model from 5 months ago) at 1/3 the price ($1/M input, $5/M output) and 2x+ speed.
- Achieved 90% of Sonnet 4.5 performance on Augment's agentic coding benchmark. Even outperformed Sonnet 4 on Computer Use tasks.
- Particularly useful in multi-agent setups. Anthropic directly proposed an orchestration pattern: Sonnet 4.5 decomposes complex problems and plans, while multiple Haiku 4.5 instances handle subtasks in parallel.
- Recorded statistically significantly lower misalignment behavior rates than Sonnet 4.5 and Opus 4.1 in safety evaluations — evaluated as Anthropic's safest model. Released at ASL-2 safety level.
- Available immediately in Claude Code and via API with model ID `claude-haiku-4-5`.
Evidence
- Early tests showed Haiku 4.5 avoids touching unrelated code compared to GPT-5, potentially making actual usage costs lower than the token price difference suggests. However, the 4x price increase from Haiku 3.5 ($0.25→$1/M input) was noted as a concern.
- A direct comparison found Haiku 4.5 hallucinated function outputs giving wrong answers while Sonnet was accurate — the small model hallucination limitation persists.
- On NYT Connections benchmark, Haiku 4.5 scored 20.0 (2x Haiku 3.5's 10.0) but still trails Sonnet 4.0 (26.6) and Sonnet 4.5 (46.1).
- A freelancer said '3x faster responses outweigh slight quality loss for productivity' and planned switching daily driver from Sonnet 4.5 to Haiku 4.5.
How to Apply
- In multi-agent systems, use Sonnet 4.5 as the main agent for planning/decomposition and Haiku 4.5 as parallel sub-agents to significantly cut costs while boosting throughput.
- For quick prototyping or simple code fixes in Claude Code, switching to Haiku 4.5 cuts response wait time by half or more.
- For chatbots or customer service agents where latency matters, Haiku 4.5 instead of Sonnet delivers 1/3 cost savings and improved responsiveness simultaneously.
- However, for tasks requiring accurate fact lookup or code documentation reference, maintain Sonnet due to hallucination risk, and route only simple generation/transformation tasks to Haiku.
Terminology
Related Papers
Claude Code sends 33k tokens before reading the prompt; OpenCode sends 7k
동일한 모델과 작업 환경에서 Claude Code와 OpenCode의 실제 토큰 사용량을 API 레벨에서 측정한 결과, Claude Code가 시스템 프롬프트 오버헤드만으로 OpenCode 대비 4.7배 더 많은 토큰을 소비한다는 것을 확인했다.
Mesh LLM: distributed AI computing on iroh
사무실, 집, 클라우드에 흩어진 GPU들을 하나의 OpenAI 호환 API로 묶어주는 분산 LLM 실행 시스템으로, 비싼 API 비용 없이 큰 모델을 직접 운영할 수 있다.
Show HN: Frugon – Find which LLM calls a cheaper model could handle (local, MIT)
내 LLM API 비용이 어디서 새는지 로컬에서 분석해주는 오픈소스 CLI 도구로, 비싼 모델 대신 저렴한 모델로 전환 가능한 호출을 골라낸다.
Jamesob's guide to running SOTA LLMs locally
2천 달러짜리 RTX 3090 한 장부터 4만 달러짜리 RTX PRO 6000 4장 셋업까지, 로컬에서 최신 LLM을 직접 돌리는 방법을 하드웨어 선택·구성·실행 설정까지 통째로 정리한 실전 가이드다.
Faster embeddings: how we rebuilt the ONNX path in Manticore
Manticore Search가 기존 SentenceTransformers/Candle 백엔드를 ONNX Runtime으로 교체해 텍스트 임베딩 생성 속도를 평균 14배 향상시켰다. 별도 모델 서비스 없이 DB 내부에서 직접 임베딩을 처리하는 구조에서 INSERT 속도가 곧 임베딩 속도이기 때문에 이 개선은 실질적인 ingest 처리량 향상으로 직결된다.
Asymmetric Quantization: Near-Lossless Retrieval with 97% Storage Reduction
멀티벡터 검색 모델의 문서 벡터를 1비트 이진값으로 압축하고 쿼리 벡터만 int8로 유지하는 비대칭 양자화 기법으로, 스토리지를 97% 줄이면서 검색 품질 손실을 0.61점(NDCG@10 기준)에 그치게 만든 실제 프로덕션 적용 사례다.