GPT-5: Key characteristics, pricing and system card
TL;DR Highlight
OpenAI launched the GPT-5 model family (regular, mini, nano) focused on stable, less error-prone practical improvements rather than revolutionary leaps, with aggressively competitive pricing.
Who Should Read
Backend/fullstack developers using or evaluating OpenAI API, or team leads focused on LLM model selection and cost optimization.
Core Mechanics
- In ChatGPT, GPT-5 is a hybrid system with a router auto-switching between fast and deep reasoning models. In the API, it's 3 variants (regular/mini/nano) × 4 reasoning levels (minimal/low/medium/high).
- 272K input tokens, 128K output tokens (including reasoning tokens), supports text+image+audio+video multimodal input.
- Input pricing halved vs. GPT-4o, with 90% token caching discount — very aggressive cost positioning.
- Reduced hallucination highlighted by Simon Willison, though community debate about whether Claude Sonnet/Opus remains more reliable for daily coding.
Evidence
- GPT-5 being 'incremental' rather than 'revolutionary' was seen as evidence of diminishing returns from pure scaling — the shift to router optimization and sub-model composition itself signals the old approach's limits.
- Simon's reduced hallucination claims were debated — some argued Claude 4 Sonnet/Opus was more reliable for daily coding tasks.
- Token caching with 90% discount makes repeated similar requests significantly cheaper — beneficial for chatbot and agent use cases.
How to Apply
- If using GPT-4o or o3 via API, just swap to GPT-5 for halved input costs with equal or better quality. Start testing with reasoning level 'minimal' to control reasoning token costs.
- For chat-based services, leverage the 90% token caching discount by structuring conversation history as a long system prompt to maximize cache hits.
- Use the reasoning level parameter to optimize cost per use case: 'minimal' for simple Q&A, 'high' for complex math/code tasks.
Terminology
Related Papers
Claude Code sends 33k tokens before reading the prompt; OpenCode sends 7k
동일한 모델과 작업 환경에서 Claude Code와 OpenCode의 실제 토큰 사용량을 API 레벨에서 측정한 결과, Claude Code가 시스템 프롬프트 오버헤드만으로 OpenCode 대비 4.7배 더 많은 토큰을 소비한다는 것을 확인했다.
Mesh LLM: distributed AI computing on iroh
사무실, 집, 클라우드에 흩어진 GPU들을 하나의 OpenAI 호환 API로 묶어주는 분산 LLM 실행 시스템으로, 비싼 API 비용 없이 큰 모델을 직접 운영할 수 있다.
Show HN: Frugon – Find which LLM calls a cheaper model could handle (local, MIT)
내 LLM API 비용이 어디서 새는지 로컬에서 분석해주는 오픈소스 CLI 도구로, 비싼 모델 대신 저렴한 모델로 전환 가능한 호출을 골라낸다.
Jamesob's guide to running SOTA LLMs locally
2천 달러짜리 RTX 3090 한 장부터 4만 달러짜리 RTX PRO 6000 4장 셋업까지, 로컬에서 최신 LLM을 직접 돌리는 방법을 하드웨어 선택·구성·실행 설정까지 통째로 정리한 실전 가이드다.
Faster embeddings: how we rebuilt the ONNX path in Manticore
Manticore Search가 기존 SentenceTransformers/Candle 백엔드를 ONNX Runtime으로 교체해 텍스트 임베딩 생성 속도를 평균 14배 향상시켰다. 별도 모델 서비스 없이 DB 내부에서 직접 임베딩을 처리하는 구조에서 INSERT 속도가 곧 임베딩 속도이기 때문에 이 개선은 실질적인 ingest 처리량 향상으로 직결된다.
Asymmetric Quantization: Near-Lossless Retrieval with 97% Storage Reduction
멀티벡터 검색 모델의 문서 벡터를 1비트 이진값으로 압축하고 쿼리 벡터만 int8로 유지하는 비대칭 양자화 기법으로, 스토리지를 97% 줄이면서 검색 품질 손실을 0.61점(NDCG@10 기준)에 그치게 만든 실제 프로덕션 적용 사례다.