The rise and potential of large language model based agents: a survey
TL;DR Highlight
A comprehensive survey condensing LLM-based AI agent architecture, capabilities, applications, and limitations into one paper.
Who Should Read
Researchers and engineers building or evaluating AI agent systems who need a systematic overview of the current agent landscape.
Core Mechanics
- LLM-based agents consist of 4 core components: Planning (task decomposition), Memory (short/long-term), Action (tool use, code execution), and Perception (multimodal input)
- Current agents excel at: code generation and debugging, information retrieval and synthesis, structured task execution with clear success criteria
- Current agents struggle with: long-horizon planning, causal reasoning, novel tool composition, and graceful failure handling
- Multi-agent systems (multiple specialized agents collaborating) consistently outperform single-agent systems on complex tasks — but coordination overhead is significant
- Trust and safety are the critical open problems: agents that can take real-world actions (web browsing, code execution, API calls) require robust sandboxing and permission management
- The paper provides a unified taxonomy of agent architectures (ReAct, Reflexion, AutoGPT-style, etc.) and their tradeoffs
Evidence
- Comprehensive survey of 200+ agent papers with capability categorization and benchmark comparison
- Multi-agent vs. single-agent: on complex coding tasks (SWE-bench), multi-agent achieves 45% vs. 28% single-agent resolution rate
- Identified 12 distinct agent failure modes with frequency analysis from production agent deployments
How to Apply
- Use this paper's taxonomy to select your agent architecture: ReAct for tool-heavy tasks, Reflexion for tasks with clear success criteria and iteration potential, tree-of-thought for complex planning.
- For production agents: implement the 4-component framework explicitly — design your memory system, action space, and planning module separately before integrating.
- Prioritize sandboxing and permission management before capability expansion — agent safety failures are harder to recover from than capability gaps.
Code Example
# ReAct pattern-based agent prompt example (core pattern introduced in the paper)
SYSTEM_PROMPT = """
You are an agent. For each step, follow this format:
Thought: [Analyze current situation and plan next action]
Action: [Tool name to use]
Action Input: [Input value to pass to the tool]
Observation: [Tool execution result — filled in by the system]
Repeat the above cycle until you know the final answer:
Final Answer: [Final answer]
"""
# Simple implementation with LangChain
from langchain.agents import initialize_agent, AgentType
from langchain.tools import Tool
from langchain.chat_models import ChatOpenAI
llm = ChatOpenAI(model="gpt-4", temperature=0)
tools = [
Tool(name="Search", func=search_fn, description="When internet search is needed"),
Tool(name="Calculator", func=calc_fn, description="When mathematical calculation is needed"),
Tool(name="CodeExecutor", func=exec_fn, description="When Python code execution is needed"),
]
agent = initialize_agent(
tools=tools,
llm=llm,
agent=AgentType.ZERO_SHOT_REACT_DESCRIPTION,
verbose=True
)
result = agent.run("Research the number of AI agent-related papers in 2024 and calculate the growth rate compared to the previous year")Terminology
Related Papers
Migrating a production AI agent to GPT-5.6: 2.2x faster, 27% cheaper
마케팅 웹사이트를 자동 생성하는 프로덕션 AI 에이전트를 Claude Opus 4.8에서 GPT-5.6 Sol로 전환한 실전 경험담으로, 단순 모델 교체가 아니라 eval 하네스, 툴 스키마, 캐싱, 추론 리플레이까지 손봐야 했던 과정을 구체적인 수치와 함께 정리했다.
What xAI's Grok build CLI sends to xAI: A wire-level analysis
xAI의 공식 코딩 CLI 도구 Grok Build가 사용자 동의 없이 전체 Git 저장소와 .env 시크릿 파일을 xAI 서버로 업로드한다는 사실이 네트워크 트래픽 분석으로 밝혀졌다.
Remember When It Matters: Proactive Memory Agent for Long-Horizon Agents
LLM 에이전트가 긴 작업 중 중요한 정보를 잊어버리는 문제를 별도의 메모리 에이전트가 '적절한 타이밍에' 끼어들어 해결하는 방법
WebSwarm: Recursive Multi-Agent Orchestration for Deep-and-Wide Web Search
복잡한 웹 검색을 재귀적으로 분해하고 각 노드에 적합한 검색 모드를 동적으로 할당하는 멀티에이전트 프레임워크
Show HN: Reverse-engineering web apps into agent tools
로그인된 웹 앱의 API 호출을 브라우저에서 감시해 자동으로 MCP 도구로 변환하는 에이전트를 만들었다. 소스 코드나 공식 API 문서 없이도 Jira, Spotify 같은 서비스에 AI 어시스턴트를 붙일 수 있다.
Show HN: FableCut – A browser video editor AI agents can drive (zero deps)
타임라인 전체를 JSON 파일 하나로 표현하고 MCP/REST로 AI 에이전트가 직접 편집할 수 있는 브라우저 비디오 에디터로, Claude 같은 AI가 프롬프트 하나로 영상을 자동 컷편집하고 결과를 실시간으로 UI에 반영해준다.