1M context is now generally available for Opus 4.6 and Sonnet 4.6
TL;DR Highlight
Anthropic rolled out 1M token context windows for Opus 4.6 and Sonnet 4.6 — this changes what's practical for long-context tasks.
Who Should Read
Developers building applications with large documents, long conversation histories, or codebases that need to be processed as a single context, and ML engineers benchmarking long-context performance.
Core Mechanics
- Claude Opus 4.6 and Sonnet 4.6 now support 1 million token context windows — enough for entire medium-sized codebases, very long books, or months of conversation history.
- 1M tokens is approximately 750,000 words or roughly 3,000 pages of text.
- This makes certain use cases that previously required RAG or chunking feasible as direct in-context tasks: analyzing a full codebase, processing large legal document sets, or maintaining very long agent memory.
- The key question is whether the model's attention quality degrades in the middle of a 1M token context (the 'lost in the middle' problem) — early reports suggest Anthropic has made improvements here.
- Pricing at this scale becomes a significant consideration: 1M tokens of input is expensive relative to a well-tuned RAG retrieval that only brings in the relevant 10k tokens.
Evidence
- Anthropic announced the 1M context expansion with API availability, confirmed through the API documentation.
- HN commenters ran their own tests — feeding in full codebases and asking questions across the entire codebase. Results were generally positive for code navigation tasks.
- Some found that retrieval quality degrades for context items in the 'middle' of a very long context, consistent with the known 'lost in the middle' problem in long-context models.
- Cost comparisons showed that for high-recall tasks, 1M context could actually be cheaper than complex RAG pipelines with re-retrieval and reranking.
How to Apply
- For codebase analysis and navigation tasks, try loading the entire relevant codebase into a single 1M context request before investing in RAG-based code search.
- For long-document processing (legal, research, financial), test whether direct 1M context gives better answers than chunked RAG — quality may outweigh cost for high-value queries.
- Structure prompts for 1M contexts carefully: put the most important content at the beginning or end, not the middle — attention tends to be strongest there.
- Monitor costs carefully: 1M token inputs at current pricing can be expensive at scale — model a full-context approach vs. RAG cost/quality tradeoff before committing to architecture.
Terminology
Related Papers
Claude Code sends 33k tokens before reading the prompt; OpenCode sends 7k
동일한 모델과 작업 환경에서 Claude Code와 OpenCode의 실제 토큰 사용량을 API 레벨에서 측정한 결과, Claude Code가 시스템 프롬프트 오버헤드만으로 OpenCode 대비 4.7배 더 많은 토큰을 소비한다는 것을 확인했다.
Mesh LLM: distributed AI computing on iroh
사무실, 집, 클라우드에 흩어진 GPU들을 하나의 OpenAI 호환 API로 묶어주는 분산 LLM 실행 시스템으로, 비싼 API 비용 없이 큰 모델을 직접 운영할 수 있다.
Show HN: Frugon – Find which LLM calls a cheaper model could handle (local, MIT)
내 LLM API 비용이 어디서 새는지 로컬에서 분석해주는 오픈소스 CLI 도구로, 비싼 모델 대신 저렴한 모델로 전환 가능한 호출을 골라낸다.
Jamesob's guide to running SOTA LLMs locally
2천 달러짜리 RTX 3090 한 장부터 4만 달러짜리 RTX PRO 6000 4장 셋업까지, 로컬에서 최신 LLM을 직접 돌리는 방법을 하드웨어 선택·구성·실행 설정까지 통째로 정리한 실전 가이드다.
Faster embeddings: how we rebuilt the ONNX path in Manticore
Manticore Search가 기존 SentenceTransformers/Candle 백엔드를 ONNX Runtime으로 교체해 텍스트 임베딩 생성 속도를 평균 14배 향상시켰다. 별도 모델 서비스 없이 DB 내부에서 직접 임베딩을 처리하는 구조에서 INSERT 속도가 곧 임베딩 속도이기 때문에 이 개선은 실질적인 ingest 처리량 향상으로 직결된다.
Asymmetric Quantization: Near-Lossless Retrieval with 97% Storage Reduction
멀티벡터 검색 모델의 문서 벡터를 1비트 이진값으로 압축하고 쿼리 벡터만 int8로 유지하는 비대칭 양자화 기법으로, 스토리지를 97% 줄이면서 검색 품질 손실을 0.61점(NDCG@10 기준)에 그치게 만든 실제 프로덕션 적용 사례다.