Google Titans architecture, helping AI have long-term memory
TL;DR Highlight
Google's new architecture handles 2M token contexts without Transformer's quadratic complexity — using 'Titans' memory blocks instead of attention.
Who Should Read
ML researchers working on long-context models, and engineers building RAG or document processing systems who care about context length scaling.
Core Mechanics
- The architecture replaces standard attention with 'Titans' — memory units that store compressed representations of past context and can be retrieved efficiently, avoiding the quadratic cost of full attention over 2M tokens.
- Two types of memory: short-term (sliding window attention for recent tokens) and long-term (Titans persistent memory blocks trained to compress and retrieve key information).
- The long-term memory is a small neural network that learns to compress context — it's trained end-to-end with the main model, not a fixed external retrieval system.
- On benchmarks requiring long-range recall (e.g., needle-in-haystack at 2M tokens), performance is competitive with or better than full attention models at much lower compute cost.
- The architecture has implications for very long document understanding, multi-session memory, and any task where full-context attention is currently prohibitively expensive.
Evidence
- Google published benchmark results showing the Titans architecture matches or beats standard Transformers on long-context tasks while being significantly more efficient.
- HN commenters noted this is in the lineage of SSM/Mamba-style architectures but with a different approach to learned memory — some skepticism about whether benchmark gains hold in practice.
- Several researchers flagged that 'compressing' context into fixed-size memory inherently loses information — the question is whether the compression is smart enough for real tasks.
- The approach was compared favorably to RAG for some use cases, since the memory is learned end-to-end rather than relying on a separate retrieval index.
How to Apply
- If you're hitting context window limits on document processing tasks, watch this architecture closely — it could enable processing entire codebases or document repositories in a single forward pass.
- For multi-turn agents that need long memory without growing context costs, Titans-style architectures may be worth experimenting with as they become available in open-source form.
- RAG architects should benchmark against learned memory approaches for tasks where chunking and retrieval introduce too much information loss.
Code Example
# Core logic of Surprise-based memory update in Titans architecture (conceptual code)
import torch
import torch.nn as nn
class NeuralMemory(nn.Module):
def __init__(self, dim, memory_lr=0.01):
super().__init__()
# Long-term memory = small MLP (key-value associative memory)
self.memory_mlp = nn.Sequential(
nn.Linear(dim, dim * 2),
nn.SiLU(),
nn.Linear(dim * 2, dim)
)
self.memory_lr = memory_lr # Memory update rate
def compute_surprise(self, query, target):
"""Surprise = magnitude of prediction error (gradient norm)"""
pred = self.memory_mlp(query)
loss = nn.functional.mse_loss(pred, target)
grad = torch.autograd.grad(loss, self.memory_mlp.parameters())
surprise = sum(g.norm() for g in grad) # More surprising = stronger memorization
return surprise, loss
def update_memory(self, key, value, forget_gate):
"""Write surprising information into memory (test-time update)"""
surprise, loss = self.compute_surprise(key, value)
# Forgetting Gate: decay old memories + write new information
effective_lr = self.memory_lr * surprise.item() * forget_gate
for param in self.memory_mlp.parameters():
if param.grad is not None:
param.data -= effective_lr * param.grad
def recall(self, query):
"""Retrieve information from long-term memory using a query"""
return self.memory_mlp(query)
# Example usage with MAC (Memory as Context) variant
# memory_output = neural_memory.recall(current_query)
# context = torch.cat([memory_output, short_term_kv_cache], dim=1)
# output = attention(query, context) # Integration of long-term + short-term memoryTerminology
Related Papers
Claude Code sends 33k tokens before reading the prompt; OpenCode sends 7k
동일한 모델과 작업 환경에서 Claude Code와 OpenCode의 실제 토큰 사용량을 API 레벨에서 측정한 결과, Claude Code가 시스템 프롬프트 오버헤드만으로 OpenCode 대비 4.7배 더 많은 토큰을 소비한다는 것을 확인했다.
Mesh LLM: distributed AI computing on iroh
사무실, 집, 클라우드에 흩어진 GPU들을 하나의 OpenAI 호환 API로 묶어주는 분산 LLM 실행 시스템으로, 비싼 API 비용 없이 큰 모델을 직접 운영할 수 있다.
Show HN: Frugon – Find which LLM calls a cheaper model could handle (local, MIT)
내 LLM API 비용이 어디서 새는지 로컬에서 분석해주는 오픈소스 CLI 도구로, 비싼 모델 대신 저렴한 모델로 전환 가능한 호출을 골라낸다.
Jamesob's guide to running SOTA LLMs locally
2천 달러짜리 RTX 3090 한 장부터 4만 달러짜리 RTX PRO 6000 4장 셋업까지, 로컬에서 최신 LLM을 직접 돌리는 방법을 하드웨어 선택·구성·실행 설정까지 통째로 정리한 실전 가이드다.
Faster embeddings: how we rebuilt the ONNX path in Manticore
Manticore Search가 기존 SentenceTransformers/Candle 백엔드를 ONNX Runtime으로 교체해 텍스트 임베딩 생성 속도를 평균 14배 향상시켰다. 별도 모델 서비스 없이 DB 내부에서 직접 임베딩을 처리하는 구조에서 INSERT 속도가 곧 임베딩 속도이기 때문에 이 개선은 실질적인 ingest 처리량 향상으로 직결된다.
Asymmetric Quantization: Near-Lossless Retrieval with 97% Storage Reduction
멀티벡터 검색 모델의 문서 벡터를 1비트 이진값으로 압축하고 쿼리 벡터만 int8로 유지하는 비대칭 양자화 기법으로, 스토리지를 97% 줄이면서 검색 품질 손실을 0.61점(NDCG@10 기준)에 그치게 만든 실제 프로덕션 적용 사례다.