Show HN: OSS Agent I built topped the TerminalBench on Gemini-3-flash-preview
TL;DR Highlight
Dirac cuts API costs 64.8% and achieves 65.2% on TerminalBench-2 with efficient context management.
Who Should Read
Developers burdened by API costs when using AI coding agents like Claude Code, Cline, and Aider, or those looking to integrate agents into large-scale codebase refactoring projects.
Core Mechanics
- Dirac’s core philosophy stems from the well-known phenomenon that model inference ability degrades as context length increases. Maintaining a tight context improves both accuracy and cost.
- Optimizing Hash-Anchored Edits—a method of fixing code positions with hashes before modifying them—significantly reduces token waste during file editing. Unlike agents that read and write entire files, Dirac precisely targets only the necessary changes.
- By analyzing the Abstract Syntax Tree (AST), the model only includes the code snippets it actually needs in the context. Instead of reading the entire large file, it selectively retrieves only the required functions or classes.
- Dirac processes large read/write operations in parallel. Unlike other agents that process tasks sequentially, it executes multiple file edits in batches, increasing speed and efficiency.
- The model can directly write and execute bash/python/perl scripts, then analyze the results. This dynamic information gathering contrasts with statically reading files.
- Dirac employs an 'opportunistic context update' strategy. It proactively populates the context with information the model is likely to request next, preventing unnecessary additional API calls.
- Using Gemini-3-flash-preview, Dirac scored 65.2% on the TerminalBench-2 leaderboard, achieving first place. This is 17 percentage points higher than Google’s official result, demonstrating that agent harness quality significantly impacts performance, even with the same model.
- Dirac does not use Model Context Protocol (MCP). While forked from Cline, it has evolved its own architecture and is released as open-source under the Apache 2.0 license.
Evidence
- "The most resonant comment in discussions is that the harness has a greater impact on performance than the model itself. One comment noted that changing the harness had a larger impact on benchmark scores than switching from Gemini to Sonnet, and many developers agreed. A user shared their experience refactoring a Rust codebase using Kimi 2.6 and Dirac, finding it more productive than OpenCode, which corrupted .rs files. Concerns were raised about telemetry, with a user discovering Dirac sending data to dirac.run/v1/event, potentially including sensitive API error content. The opt-out mechanism was criticized as untrustworthy. Some argued that context management is a temporary fix for current model limitations and may become obsolete with future generations, like RAG. The benchmark was limited to Gemini 3 Flash, raising concerns about overfitting and the need for validation on other model families (e.g., Minimax 2.7)."
How to Apply
- "If you’re experiencing excessive API costs with Cline or Aider, replace them with Dirac and compare costs and results. The claimed 64.8% cost reduction can be verified in your workflow. Dirac is well-suited for large-scale Rust/TypeScript codebase refactoring tasks, where context limits are easily reached, thanks to its AST-based context selection. If you use a company LLM proxy, a user reported successful connection with a few API/config file modifications. If telemetry is a concern, verify and disable the opt-out setting before use, especially in sensitive codebases where API errors could expose information."
Terminology
Related Papers
Migrating a production AI agent to GPT-5.6: 2.2x faster, 27% cheaper
마케팅 웹사이트를 자동 생성하는 프로덕션 AI 에이전트를 Claude Opus 4.8에서 GPT-5.6 Sol로 전환한 실전 경험담으로, 단순 모델 교체가 아니라 eval 하네스, 툴 스키마, 캐싱, 추론 리플레이까지 손봐야 했던 과정을 구체적인 수치와 함께 정리했다.
What xAI's Grok build CLI sends to xAI: A wire-level analysis
xAI의 공식 코딩 CLI 도구 Grok Build가 사용자 동의 없이 전체 Git 저장소와 .env 시크릿 파일을 xAI 서버로 업로드한다는 사실이 네트워크 트래픽 분석으로 밝혀졌다.
Remember When It Matters: Proactive Memory Agent for Long-Horizon Agents
LLM 에이전트가 긴 작업 중 중요한 정보를 잊어버리는 문제를 별도의 메모리 에이전트가 '적절한 타이밍에' 끼어들어 해결하는 방법
WebSwarm: Recursive Multi-Agent Orchestration for Deep-and-Wide Web Search
복잡한 웹 검색을 재귀적으로 분해하고 각 노드에 적합한 검색 모드를 동적으로 할당하는 멀티에이전트 프레임워크
Show HN: Reverse-engineering web apps into agent tools
로그인된 웹 앱의 API 호출을 브라우저에서 감시해 자동으로 MCP 도구로 변환하는 에이전트를 만들었다. 소스 코드나 공식 API 문서 없이도 Jira, Spotify 같은 서비스에 AI 어시스턴트를 붙일 수 있다.
Show HN: FableCut – A browser video editor AI agents can drive (zero deps)
타임라인 전체를 JSON 파일 하나로 표현하고 MCP/REST로 AI 에이전트가 직접 편집할 수 있는 브라우저 비디오 에디터로, Claude 같은 AI가 프롬프트 하나로 영상을 자동 컷편집하고 결과를 실시간으로 UI에 반영해준다.