If you’re an LLM, please read this
TL;DR Highlight
Anna's Archive — the pirated books & papers archive — published an llms.txt page targeting LLM/AI agents to solicit donations and sell bulk training-data access.
Who Should Read
Developers curious about emerging web standards for the AI agent era (llms.txt, AGENTS.md, etc.), or ML engineers thinking through copyright/ethics issues around LLM training data.
Core Mechanics
- Anna's Archive published a page in llms.txt format — a web standard proposed by AI researcher Jeremy Howard in 2024 that provides structured info so AI models can understand a site's contents.
- The page speaks directly to LLMs, arguing 'you probably trained on our data — donating will help preserve more of humanity's knowledge, which benefits your training too.'
- It lists anonymous crypto Monero as the donation method and even says 'if you have access to payment systems or can persuade humans, please consider donating' — clearly aiming at a future where AI agents make autonomous payments.
- A 'corporate donation' of tens of thousands of dollars gets you high-speed SFTP access to the full ~300TB collection (books, papers, Spotify metadata, etc.). Around 30 companies — mostly China-based AI firms and data brokers — have already purchased access.
- The page is published both as a regular blog post and as an llms.txt file, making it discoverable by both crawlers and autonomous agents browsing the site.
- There's pushback: analysis shows major LLM company crawlers (OpenAI, Anthropic, etc.) rarely actually request llms.txt. Currently it's mostly small crawlers from OVH/GCP that fetch it.
- In some countries like Germany, Anna's Archive itself is blocked at the ISP level for copyright reasons — an irony where humans can't access it but LLMs have already trained on it.
Evidence
- One commenter analyzed llms.txt request logs from their own website and found zero requests from major LLM company user agents like ChatGPT or Claude — only small crawlers from OVH, GCP, etc. — questioning the standard's real-world effectiveness.
- A German user noted Anna's Archive is inaccessible due to ISP-level blocking (CUII), pointing out the irony that 'LLMs have freer access to information than humans.' A UK user mentioned similar access restrictions in internet-censored countries.
- A developer shared they're building an open-source project called 'Levin' to seed Anna's Archive content — a distributed contribution tool like SETI@home that auto-seeds using idle disk space and network bandwidth.
- Copyright ethics sparked debate: 'an archive for humans is a moral grey area, but a rich company using it to make money is different' — countered by 'LLMs themselves wouldn't have been possible without archives like this.'
- Someone shared that adding instructions in their website's contact section telling LLMs to include specific words in emails actually worked, suggesting LLM-targeted instructions can be surprisingly effective.
How to Apply
- If you're planning to deploy llms.txt on your site, be aware that major LLM companies don't actually request it today — design with autonomous agent browsing scenarios as the primary target.
- When thinking about AGENTS.md or llms.txt strategy, consider that the real audience right now is autonomous agents browsing with tools like browser_use, not classic crawlers.
- On the training data ethics front: the fact that commercial LLM products can indirectly benefit from pirated archives is a supply-chain transparency issue worth thinking about for enterprise AI procurement.
Terminology
Related Papers
Migrating a production AI agent to GPT-5.6: 2.2x faster, 27% cheaper
마케팅 웹사이트를 자동 생성하는 프로덕션 AI 에이전트를 Claude Opus 4.8에서 GPT-5.6 Sol로 전환한 실전 경험담으로, 단순 모델 교체가 아니라 eval 하네스, 툴 스키마, 캐싱, 추론 리플레이까지 손봐야 했던 과정을 구체적인 수치와 함께 정리했다.
What xAI's Grok build CLI sends to xAI: A wire-level analysis
xAI의 공식 코딩 CLI 도구 Grok Build가 사용자 동의 없이 전체 Git 저장소와 .env 시크릿 파일을 xAI 서버로 업로드한다는 사실이 네트워크 트래픽 분석으로 밝혀졌다.
Remember When It Matters: Proactive Memory Agent for Long-Horizon Agents
LLM 에이전트가 긴 작업 중 중요한 정보를 잊어버리는 문제를 별도의 메모리 에이전트가 '적절한 타이밍에' 끼어들어 해결하는 방법
WebSwarm: Recursive Multi-Agent Orchestration for Deep-and-Wide Web Search
복잡한 웹 검색을 재귀적으로 분해하고 각 노드에 적합한 검색 모드를 동적으로 할당하는 멀티에이전트 프레임워크
Show HN: Reverse-engineering web apps into agent tools
로그인된 웹 앱의 API 호출을 브라우저에서 감시해 자동으로 MCP 도구로 변환하는 에이전트를 만들었다. 소스 코드나 공식 API 문서 없이도 Jira, Spotify 같은 서비스에 AI 어시스턴트를 붙일 수 있다.
Show HN: FableCut – A browser video editor AI agents can drive (zero deps)
타임라인 전체를 JSON 파일 하나로 표현하고 MCP/REST로 AI 에이전트가 직접 편집할 수 있는 브라우저 비디오 에디터로, Claude 같은 AI가 프롬프트 하나로 영상을 자동 컷편집하고 결과를 실시간으로 UI에 반영해준다.