Executing programs inside transformers with exponentially faster inference
TL;DR Highlight
A new approach runs programs directly inside Transformer weights without external tool calls — the LLM itself acts as the compute substrate.
Who Should Read
ML researchers interested in in-context computation, and engineers exploring alternatives to tool-calling architectures for agent reasoning.
Core Mechanics
- The key idea: instead of having an LLM call external tools to run code, the paper proposes a method for encoding and executing programs directly within the Transformer's weight space.
- This is distinct from in-context learning — it's not giving the model examples. It's more like compiling programs into the model's activation patterns.
- The approach could eliminate round-trip latency for tool calls and reduce dependence on external compute infrastructure for certain types of computation.
- Current limitations: the types of programs that can be efficiently encoded this way are constrained — not general Turing-complete programs but specific computation patterns.
- This is an early-stage research result, not a practical engineering approach yet, but points toward a future where the line between 'model weights' and 'program' blurs.
Evidence
- The paper demonstrated the approach on a set of algorithmic tasks, showing programs executing correctly within the forward pass of a Transformer.
- HN commenters with ML theory backgrounds engaged with the theoretical implications — asking whether this relates to the 'mechanistic interpretability' view of Transformers as general computers.
- Skeptics questioned whether this is meaningfully different from a very capable model doing multi-step reasoning — the distinction between 'executing a program' and 'reasoning through a program' is philosophically tricky.
- Others connected this to older work on neural Turing machines and differentiable programming.
How to Apply
- This is research-stage work — don't plan your architecture around it yet. Monitor for follow-up work on scaling these results to more general computation.
- For tool-calling architectures: the latency problem this addresses is real — even if this specific approach isn't practical, keep an eye on alternatives to round-trip tool calls for computationally simple tasks.
- For ML researchers: the connection to mechanistic interpretability is worth exploring — if programs can be encoded in weights, understanding those weight patterns is a new lens on model internals.
Terminology
Related Papers
Claude Code sends 33k tokens before reading the prompt; OpenCode sends 7k
동일한 모델과 작업 환경에서 Claude Code와 OpenCode의 실제 토큰 사용량을 API 레벨에서 측정한 결과, Claude Code가 시스템 프롬프트 오버헤드만으로 OpenCode 대비 4.7배 더 많은 토큰을 소비한다는 것을 확인했다.
Mesh LLM: distributed AI computing on iroh
사무실, 집, 클라우드에 흩어진 GPU들을 하나의 OpenAI 호환 API로 묶어주는 분산 LLM 실행 시스템으로, 비싼 API 비용 없이 큰 모델을 직접 운영할 수 있다.
Show HN: Frugon – Find which LLM calls a cheaper model could handle (local, MIT)
내 LLM API 비용이 어디서 새는지 로컬에서 분석해주는 오픈소스 CLI 도구로, 비싼 모델 대신 저렴한 모델로 전환 가능한 호출을 골라낸다.
Jamesob's guide to running SOTA LLMs locally
2천 달러짜리 RTX 3090 한 장부터 4만 달러짜리 RTX PRO 6000 4장 셋업까지, 로컬에서 최신 LLM을 직접 돌리는 방법을 하드웨어 선택·구성·실행 설정까지 통째로 정리한 실전 가이드다.
Faster embeddings: how we rebuilt the ONNX path in Manticore
Manticore Search가 기존 SentenceTransformers/Candle 백엔드를 ONNX Runtime으로 교체해 텍스트 임베딩 생성 속도를 평균 14배 향상시켰다. 별도 모델 서비스 없이 DB 내부에서 직접 임베딩을 처리하는 구조에서 INSERT 속도가 곧 임베딩 속도이기 때문에 이 개선은 실질적인 ingest 처리량 향상으로 직결된다.
Asymmetric Quantization: Near-Lossless Retrieval with 97% Storage Reduction
멀티벡터 검색 모델의 문서 벡터를 1비트 이진값으로 압축하고 쿼리 벡터만 int8로 유지하는 비대칭 양자화 기법으로, 스토리지를 97% 줄이면서 검색 품질 손실을 0.61점(NDCG@10 기준)에 그치게 만든 실제 프로덕션 적용 사례다.