Show HN: GoModel – an open-source AI gateway in Go
TL;DR Highlight
GoModel unifies access to OpenAI, Anthropic, Gemini, and other AI providers through a single, OpenAI-compatible API, offering a compiled-language alternative to LiteLLM.
Who Should Read
Backend developers simultaneously using multiple LLM providers, or those interested in the performance, supply chain security, and Go ecosystem integration benefits over LiteLLM.
Core Mechanics
- GoModel is a Go-written AI gateway that integrates various providers—OpenAI, Anthropic, Gemini, xAI, Groq, OpenRouter, Z.ai, Azure OpenAI, Oracle, and Ollama—into a single OpenAI-compatible API.
- It can be launched with a single Docker command, requiring only the API keys for the desired providers as environment variables; at least one provider key is needed for operation.
- Positioned as an alternative to LiteLLM, it natively supports observability (monitoring), guardrails (safety filters), and streaming (streaming responses).
- Its use of the Go compiled language is highlighted as a strength, offering greater security against runtime supply chain attacks compared to Python-based LiteLLM due to fixed dependencies at compile time.
- It supports Prometheus metric integration and includes separate configuration files (prometheus.yml) and a docker-compose.yaml for easy monitoring environment setup.
- A semantic caching layer appears to be present, with the gateway embedding requests and using vector similarity search to determine cache hits.
- A Helm chart is included, enabling deployment in Kubernetes environments.
- Currently, it has 319 stars and 20 forks on GitHub and is actively being committed, indicating an early-stage project.
Evidence
- "In response to a question about the importance of being written in Go, a comment pointed out that Go compiled binaries have a significantly smaller runtime supply chain attack surface than Python-based tools, a point also made by the developer of a similar Go gateway (sbproxy.dev). An experienced AI proxy maintainer noted that the most challenging aspect is adapting to changing input/output structures with each model/provider release, emphasizing that integration within 24 hours of a new model launch is crucial for a well-managed project. Concerns were raised about the maintenance burden of keeping up with provider updates due to the lack of a robust Go SDK compared to JavaScript and Python, a challenge the author acknowledges. A vllm user inquired about Ollama integration, and requests were made for cost tracking per model/route, particularly for mixed free/paid model usage. Questions were also raised about potential open-source rug pulls, and the need for the unified API to abstract provider-specific parameters like temperature, reasoning effort, and tool choice mode."
How to Apply
- "If you're using multiple LLM providers and want to avoid modifying client code with each model switch, deploy GoModel as an intermediary gateway and route all requests to its OpenAI-compatible endpoint at `http://localhost:8080`. Provider switching is then managed through environment variables. If you're running LiteLLM and concerned about Python runtime supply chain security or memory/performance overhead, consider switching to GoModel. Its compiled binary has no runtime dependencies and the Docker image is lightweight. For centralized management of AI traffic in Kubernetes, leverage the included Helm chart to deploy GoModel to your cluster and integrate it with Prometheus to monitor model response times and error rates. If your team manages AI provider keys individually, use GoModel as an internal gateway, directing team members to its endpoint to centralize key management."
Code Example
# Minimal execution (using only OpenAI)
docker run --rm -p 8080:8080 \
-e OPENAI_API_KEY="your-openai-key" \
enterpilot/gomodel
# Using multiple providers simultaneously
docker run --rm -p 8080:8080 \
-e OPENAI_API_KEY="your-openai-key" \
-e ANTHROPIC_API_KEY="your-anthropic-key" \
-e GEMINI_API_KEY="your-gemini-key" \
-e GROQ_API_KEY="your-groq-key" \
-e OPENROUTER_API_KEY="your-openrouter-key" \
-e XAI_API_KEY="your-xai-key" \
-e AZURE_API_KEY="your-azure-key" \
-e AZURE_BASE_URL="https://your-resource.openai.azure.com/openai/deployments/your-deployment" \
-e AZURE_API_VERSION="2024-10-21" \
enterpilot/gomodel
# Then, in the client, only change the base_url
# openai.OpenAI(base_url="http://localhost:8080", api_key="any-value")Terminology
Related Papers
Claude Code sends 33k tokens before reading the prompt; OpenCode sends 7k
동일한 모델과 작업 환경에서 Claude Code와 OpenCode의 실제 토큰 사용량을 API 레벨에서 측정한 결과, Claude Code가 시스템 프롬프트 오버헤드만으로 OpenCode 대비 4.7배 더 많은 토큰을 소비한다는 것을 확인했다.
Mesh LLM: distributed AI computing on iroh
사무실, 집, 클라우드에 흩어진 GPU들을 하나의 OpenAI 호환 API로 묶어주는 분산 LLM 실행 시스템으로, 비싼 API 비용 없이 큰 모델을 직접 운영할 수 있다.
Show HN: Frugon – Find which LLM calls a cheaper model could handle (local, MIT)
내 LLM API 비용이 어디서 새는지 로컬에서 분석해주는 오픈소스 CLI 도구로, 비싼 모델 대신 저렴한 모델로 전환 가능한 호출을 골라낸다.
Jamesob's guide to running SOTA LLMs locally
2천 달러짜리 RTX 3090 한 장부터 4만 달러짜리 RTX PRO 6000 4장 셋업까지, 로컬에서 최신 LLM을 직접 돌리는 방법을 하드웨어 선택·구성·실행 설정까지 통째로 정리한 실전 가이드다.
Faster embeddings: how we rebuilt the ONNX path in Manticore
Manticore Search가 기존 SentenceTransformers/Candle 백엔드를 ONNX Runtime으로 교체해 텍스트 임베딩 생성 속도를 평균 14배 향상시켰다. 별도 모델 서비스 없이 DB 내부에서 직접 임베딩을 처리하는 구조에서 INSERT 속도가 곧 임베딩 속도이기 때문에 이 개선은 실질적인 ingest 처리량 향상으로 직결된다.
Asymmetric Quantization: Near-Lossless Retrieval with 97% Storage Reduction
멀티벡터 검색 모델의 문서 벡터를 1비트 이진값으로 압축하고 쿼리 벡터만 int8로 유지하는 비대칭 양자화 기법으로, 스토리지를 97% 줄이면서 검색 품질 손실을 0.61점(NDCG@10 기준)에 그치게 만든 실제 프로덕션 적용 사례다.