Google Engineers Launch "Sashiko" for Agentic AI Code Review of the Linux Kernel
TL;DR Highlight
Google's Linux kernel team open-sourced 'Sashiko,' a Gemini 3.1 Pro-based AI code review agent that claims to detect 53% of bugs missed by human reviewers.
Who Should Read
Linux kernel contributors or large open-source project maintainers considering automated code review pipelines. Backend/systems developers curious about real-world cases of applying AI agents to code quality verification.
Core Mechanics
- Roman Gushchin from Google's Linux kernel team released Sashiko. A system used internally at Google for months is now being extended to all Linux kernel mailing list patch submissions.
- Sashiko detected 53% of bugs when tested against 1,000 recent upstream Linux kernel issues with 'Fixes:' tags. The presenter emphasized 'this 53% are issues that human reviewers missed 100% of the time.'
- Designed to use Gemini 2.5 Pro by default (listed as 'Gemini 3.1 Pro' in original), but built to work with Claude and other LLMs. Interestingly, the system itself was written in Rust and co-authored with Claude.
- Google is covering Sashiko's token costs and infrastructure, with project hosting planned to transfer to the Linux Foundation. Code is open-sourced on GitHub (github.com/sashiko-dev/sashiko).
- A web interface (sashiko.dev) shows patchsets currently under review and results. Review results include findings tagged with severity like 'Critical' and 'High.'
- A key design principle: Sashiko is designed not to spam the mailing list with comments directly. Review results are only viewable on a separate web interface, choosing not to disrupt the kernel community's existing workflow.
- As an agentic AI code review (unlike simple static analysis), the LLM understands patch context and judges bug likelihood. One commenter noted 'separating the model that writes code from the model that reviews code is the key insight.'
Evidence
- The 53% detection rate was criticized for not disclosing the false positive rate. 'Flag all code as buggy and you get 100% detection rate' — without precision alongside recall, actual usefulness is hard to judge. Concerns that human reviewers overwhelmed by AI false positive reports could lose trust in the entire system.
- The claim '100% of issues were missed by human reviewers' sparked interpretation debate. One commenter noted 'missed at the initial code review stage doesn't mean developers didn't find these bugs later in development.' Code gets reviewed continuously, so the framing was considered somewhat exaggerated.
- UX feedback on the web UI (sashiko.dev): the Status column shows internal pipeline states like 'Pending' and 'In Review,' while actually important findings are buried on the far right. No filtering or highlighting for Critical/High severity findings, reducing practical utility.
- Concerns about auto-submitting style/structural change patches. A commenter shared an actual Sashiko review result link, noting that automated style changes applied at scale to the kernel codebase could burden the existing development flow. The worry was it seemed to focus more on style cleanup than bug detection.
- Positive reactions to separating the writing model from the reviewing model. One commenter shared 'I use the same approach at small scale — for the same reason you don't self-review your own PRs, self-review misses things.' The system being written in Rust and co-developed with Claude was also noted as interesting.
How to Apply
- If you submit patches to the Linux kernel, check your patchset's review results on sashiko.dev before sending to the mailing list. Fixing Critical/High findings before submission can shorten review cycles.
- To build a similar AI code review pipeline for your own large codebase, reference Sashiko's open-source code (github.com/sashiko-dev/sashiko) and apply the 'writing model != reviewing model' principle. Designed to swap in Claude API as well, making it easy for teams already using Claude.
- When considering AI code review system adoption, follow Sashiko's approach: separate review results into a dashboard rather than spamming existing communication channels (mailing lists, PR comments). When false positives are high, a separate UI maintains team trust better than direct notifications.
- Rather than trusting the 53% bug detection metric at face value, measure false positive rate as well before actual adoption. Pull 100-200 recent 'Fixes:' commits from your own codebase and compare against AI review results to measure precision/recall yourself.
Terminology
Related Papers
Migrating a production AI agent to GPT-5.6: 2.2x faster, 27% cheaper
마케팅 웹사이트를 자동 생성하는 프로덕션 AI 에이전트를 Claude Opus 4.8에서 GPT-5.6 Sol로 전환한 실전 경험담으로, 단순 모델 교체가 아니라 eval 하네스, 툴 스키마, 캐싱, 추론 리플레이까지 손봐야 했던 과정을 구체적인 수치와 함께 정리했다.
What xAI's Grok build CLI sends to xAI: A wire-level analysis
xAI의 공식 코딩 CLI 도구 Grok Build가 사용자 동의 없이 전체 Git 저장소와 .env 시크릿 파일을 xAI 서버로 업로드한다는 사실이 네트워크 트래픽 분석으로 밝혀졌다.
Remember When It Matters: Proactive Memory Agent for Long-Horizon Agents
LLM 에이전트가 긴 작업 중 중요한 정보를 잊어버리는 문제를 별도의 메모리 에이전트가 '적절한 타이밍에' 끼어들어 해결하는 방법
WebSwarm: Recursive Multi-Agent Orchestration for Deep-and-Wide Web Search
복잡한 웹 검색을 재귀적으로 분해하고 각 노드에 적합한 검색 모드를 동적으로 할당하는 멀티에이전트 프레임워크
Show HN: Reverse-engineering web apps into agent tools
로그인된 웹 앱의 API 호출을 브라우저에서 감시해 자동으로 MCP 도구로 변환하는 에이전트를 만들었다. 소스 코드나 공식 API 문서 없이도 Jira, Spotify 같은 서비스에 AI 어시스턴트를 붙일 수 있다.
Show HN: FableCut – A browser video editor AI agents can drive (zero deps)
타임라인 전체를 JSON 파일 하나로 표현하고 MCP/REST로 AI 에이전트가 직접 편집할 수 있는 브라우저 비디오 에디터로, Claude 같은 AI가 프롬프트 하나로 영상을 자동 컷편집하고 결과를 실시간으로 UI에 반영해준다.