This repository curates high-signal resources on Coding Agent technology, emphasizing technical principles, evaluation, and research (not API docs).
| Rank | Agent | Company/Team | Summary | CLI | IDE Plugin | Native IDE | App | OSS Repo | Blog/News |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Claude Code | Anthropic | AI coding agent for terminal and IDE workflows. | ✓ | ✓ | ✓ | Docs-only: https://github.com/anthropics/claude-code | https://www.claude.com/blog | |
| 2 | Codex | OpenAI | OpenAI coding agent for IDE, CLI, and cloud tasks. | ✓ | ✓ | ✓ | https://github.com/openai/codex | https://openai.com/index/introducing-codex/ | |
| 3 | GitHub Copilot | GitHub | AI coding assistant integrated across developer tools. | ✓ | ✓ | Closed-source | https://devblogs.microsoft.com/visualstudio/tag/github-copilot/ | ||
| 4 | Cursor | Anysphere | AI-first code editor with agent features. | ✓ | ✓ | ✓ | Closed-source | https://cursor.com/blog | |
| 5 | Windsurf | Cognition | AI IDE focused on agentic workflows. | ✓ | ✓ | Closed-source | https://windsurf.com/blog | ||
| 6 | Amp | Amp Inc. | Agentic coding tool for CLI and IDEs. | ✓ | ✓ | Closed-source | https://ampcode.com/news | ||
| 7 | Moltbot | Open-source personal AI assistant. | ✓ | https://github.com/gstarwd/clawbot | |||||
| 8 | OpenCode | OpenCode | Open-source coding agent for terminal and desktop. | ✓ | ✓ | ✓ | https://github.com/opencode-ai/opencode | https://opencode.ai/changelog | |
| 9 | Gemini CLI | Open-source Gemini terminal agent. | ✓ | https://github.com/google-gemini/gemini-cli | https://google-gemini.github.io/gemini-cli/docs/releases.html | ||||
| 10 | OpenHands | All Hands AI | Open-source platform for autonomous coding agents. | ✓ | ✓ | https://github.com/All-Hands-AI/OpenHands | https://openhands.dev/blog | ||
| 11 | Cline | Cline | Open-source IDE agent focused on automation. | ✓ | https://github.com/cline/cline | https://cline.bot/blog | |||
| 12 | Continue | Continue | Open-source coding assistant for IDEs and CLI. | ✓ | ✓ | https://github.com/continuedev/continue | https://blog.continue.dev | ||
| 13 | Goose | Block | Open-source local agent by Block. | ✓ | ✓ | https://github.com/block/goose | https://block.github.io/goose/blog | ||
| 14 | Crush | Charmbracelet | Open-source terminal coding agent. | ✓ | https://github.com/charmbracelet/crush | https://charm.land/blog | |||
| 15 | Qwen Code | Alibaba | Open-source Qwen CLI coding agent. | ✓ | ✓ | https://github.com/QwenLM/qwen-code | https://qwenlm.github.io/blog/ | ||
| 16 | Roo Code | Roo Code | Open-source VS Code coding agent. | ✓ | https://github.com/RooCodeInc/Roo-Code | https://docs.roocode.com/update-notes | |||
| 17 | Kilo Code | Kilo | Open-source agent for VS Code and JetBrains. | ✓ | ✓ | ✓ | https://github.com/Kilo-Org/kilocode | https://blog.kilo.ai | |
| 18 | Kimi Code CLI | Moonshot AI | Open-source Kimi coding agent for the terminal. | ✓ | https://github.com/MoonshotAI/kimi-cli | https://www.kimi.com/blog/kimi-k2-5.html | |||
| 19 | Augment | Augment | Agentic coding platform focused on context quality. | ✓ | Docs-only (Auggie repo): https://github.com/augmentcode/auggie | https://www.augmentcode.com/blog | |||
| 20 | CodeBuddy | Tencent Cloud | Tencent Cloud AI coding assistant. | ✓ | ✓ | Closed-source | |||
| 21 | Mux | Coder | Parallel coding agents with isolated workspaces. | ✓ | ✓ | https://github.com/coder/mux | https://coder.com/blog | ||
| 22 | Kode | ShareAI Lab | Open-source terminal agent CLI. | ✓ | https://github.com/shareAI-lab/Kode-cli | ||||
| 23 | Pi | badlogic | Open-source CLI coding agent in the Pi monorepo. | ✓ | https://github.com/badlogic/pi-mono | ||||
| 24 | Mistral Vibe | Mistral AI | Open-source Mistral terminal agent. | ✓ | https://github.com/mistralai/mistral-vibe | https://mistral.ai/news | |||
| 25 | Junie | JetBrains | JetBrains AI coding agent for IDEs. | ✓ | Closed-source | https://blog.jetbrains.com/junie/ | |||
| 26 | Kiro CLI | AWS | Spec-driven CLI and agentic IDE tooling. | ✓ | Docs-only: https://github.com/kirodotdev/Kiro | https://kiro.dev/blog | |||
| 27 | Antigravity | Agent-first IDE platform for autonomous execution. | ✓ | Closed-source | |||||
| 28 | MCPJam | MCPJam | Open-source MCP inspector and local agent tools. | ✓ | ✓ | https://github.com/MCPJam/inspector | https://www.mcpjam.com/blog | ||
| 29 | Neovate | Neovate | Open-source coding agent and tooling. | ✓ | https://github.com/neovateai/neovate-code | https://neovateai.dev/blog | |||
| 30 | Qoder | Alibaba | Agentic coding platform focused on real software tasks. | ✓ | ✓ | ✓ | Closed-source | ||
| 31 | Replit | Replit | Cloud development platform with AI agents. | ✓ | Closed-source | https://blog.replit.com/ | |||
| 32 | Trae | ByteDance | AI-native IDE from ByteDance. | ✓ | Open-source module (Trae Agent CLI; not full end-user CLI): https://github.com/bytedance/trae-agent | ||||
| 33 | Pochi | TabbyML | Open-source VS Code agent by TabbyML. | ✓ | ✓ | https://github.com/TabbyML/pochi | https://tabby.tabbyml.com/blog/ | ||
| 34 | Command Code | Langbase | CLI coding agent that learns your coding taste. | ✓ | Closed-source | https://commandcode.ai/launch | |||
| 35 | Droid | Factory | AI coding agent for software development. | ✓ | ✓ | ✓ | Closed-source | https://factory.ai/news | |
| 36 | Trae CN | ByteDance | China site for the Trae AI IDE. | ✓ | Open-source module (Trae Agent CLI; not full end-user CLI): https://github.com/bytedance/trae-agent | https://www.trae.cn/announce/ | |||
| 37 | Zencoder | Zencoder | AI coding agent and IDE plugin. | ✓ | Configs-only (not OSS): https://github.com/zencoderai/zenagents-library | https://zencoder.ai/blog | |||
| 38 | AdaL | SylphAI | AdaL CLI coding agent by SylphAI. | ✓ | Docs-only: https://github.com/SylphAI-Inc/adal-cli | ||||
| 39 | Aider | Aider-AI | AI pair programming in your terminal. | ✓ | https://github.com/Aider-AI/aider | https://aider.chat/ | |||
| 40 | Devin | Cognition | AI software engineer by Cognition. | ✓ | Closed-source | https://cognition.ai/introducing-devin | |||
| 41 | Deep Agents CLI | LangChain | CLI for deep agent workflows and developer task automation. | ✓ | Docs-only: https://docs.langchain.com/oss/python/deepagents/cli/overview |
Legend: ✓ indicates official availability. App refers to a standalone web/desktop/mobile app (non-IDE).
- 2025-05-16 - Introducing Codex - Research preview and core Codex design. Key technical takeaways: parallel tasks in isolated cloud containers; codex-1 RL on real-world coding tasks; no-internet execution with verifiable logs/tests.
- 2025-05-16 - Addendum to OpenAI o3 and o4-mini system card: Codex - Safety addendum covering codex-1 deployment and safeguards. Key technical takeaways: codex-1 RL on real-world tasks; cloud containers with no internet; citations of terminal logs/files for verification.
- 2025-09-15 - Introducing upgrades to Codex - Codex product upgrades and GPT-5-Codex rollout. Key technical takeaways: GPT-5-Codex optimized for agentic coding across interactive and long tasks; unified experience across terminal/IDE/web/phone; revamped CLI/IDE and GitHub workflows.
- 2025-09-15 - Addendum to GPT-5 system card: GPT-5-Codex - Safety addendum for GPT-5-Codex. Key technical takeaways: GPT-5-Codex trained via RL on real-world coding tasks; safety mitigations include model-level harmful-task and prompt-injection training plus product-level sandboxing and network controls.
- 2025-10-06 - Codex is now generally available - GA announcement covering Codex availability plus SDK/admin features. Key technical takeaways: Slack integration for task delegation; Codex SDK to embed the agent with structured outputs and context management; admin controls for environment management, monitoring, and analytics.
- 2025-11-13 - Introducing GPT-5.1 for developers - Developer-focused model update for agentic workflows. Key technical takeaways: adds apply_patch and shell tools for structured edits/execution; expands prompt caching for longer-lived sessions; improves dynamic reasoning for coding tasks.
- 2025-11-19 - Building more with GPT-5.1-Codex-Max - Long-horizon coding capabilities and compaction. Key technical takeaways: compaction across multiple context windows enables long-running tasks; auto compaction preserves context for refactors and multi-hour loops.
- 2025-11-19 - GPT-5.1-Codex-Max System Card - Safety system card for GPT-5.1-Codex-Max. Key technical takeaways: first model natively trained for multi-context-window compaction; trained on agentic tasks across domains; safety mitigations and preparedness evaluation.
- 2025-12-12 - How we used Codex to build Sora for Android in 28 days - Case study on Codex-driven mobile delivery. Key technical takeaways: parallel Codex sessions for throughput; AGENTS.md as a living spec/guardrail; plan-first and review/feedback loops to keep changes aligned.
- 2025-12-18 - Introducing GPT-5.2-Codex - Model update focused on real-world coding and security. Key technical takeaways: improvements in long-horizon work via compaction; stronger performance on refactors/migrations and Windows; stronger cybersecurity capabilities and tool calling.
- 2025-12-18 - Addendum to GPT-5.2 System Card: GPT-5.2-Codex - Safety addendum for GPT-5.2-Codex. Key technical takeaways: compaction and project-scale improvements plus Windows and cybersecurity boosts; safety mitigations with model-level training and sandbox/network controls.
- 2026-01-09 - Datadog - System-level code review case study. Key technical takeaways: Codex reviews land before CI to flag risks; incident replay harness to validate fixes; emphasizes cross-module reasoning over lint checks.
- 2026-01-20 - Cisco - Enterprise-scale Codex deployment. Key technical takeaways: enterprise-scale workflows and collaboration; accelerates large codebase iteration; focuses on complex development tasks.
- 2026-01-23 - Unrolling the Codex agent loop - Deep dive into the Codex harness and tool loop. Key technical takeaways: harness orchestrates prompt -> inference -> tool calls; tool-call iterations with context window management; prompt built from roles/tools/inputs via the Responses API.
- 2026-01-29 - Inside our in-house data agent - Internal agent architecture writeup. Key technical takeaways: internal data agent for employee questions; multi-step agentic workflows beyond single-turn chat; integrates with internal data sources.
- 2026-02-02 - Introducing the Codex app - Product launch for a dedicated Codex app experience. Key technical takeaways: app-level workflow for coding tasks with Codex and broader access pattern across devices.
- 2026-02-05 - Introducing GPT-5.3-Codex - Model update focused on coding quality and agent performance. Key technical takeaways: new Codex model generation for stronger real-world coding and agentic execution.
- 2026-02-05 - GPT-5.3-Codex system card - Safety/system-card documentation for GPT-5.3-Codex. Key technical takeaways: evaluation and mitigation details for the GPT-5.3-Codex deployment.
- 2025-11-04 - Code execution with MCP: Building more efficient agents - Engineering walkthrough of code execution via MCP. Key technical takeaways: shifts execution from shell wrappers to typed MCP tool calls; simplifies orchestration and reduces coordination overhead in agent loops.
- 2025-11-24 - Introducing advanced tool use on the Claude Developer Platform - Production lessons for higher-fidelity tool invocation. Key technical takeaways: tool-use upgrades improve multi-step execution reliability and expand structured interaction patterns.
- 2025-11-26 - Effective harnesses for long-running agents - Harness design for multi-hour agent runs. Key technical takeaways: harness architecture, retries, and state management matter as much as model quality for long-horizon coding tasks.
- 2026-01-09 - Demystifying evals for AI agents - Practical guidance for measuring agent quality. Key technical takeaways: evaluator design, task realism, and variance control are critical for trustworthy agent benchmarks.
- 2026-02-05 - Building a C compiler with a team of parallel Claudes - Experiment on parallel-agent software construction. Key technical takeaways: parallel specialized agents can accelerate compiler implementation while keeping integration/test loops explicit.
- Codex Open Source - Official index of open-source Codex components (CLI, SDK, app server, skills) and collaboration entry points. Key technical takeaways: component map plus extension points for agent building.
- Building effective agents - Anthropic research on agent/workflow distinctions, composable patterns, and when to use agents. Key technical takeaways: decision boundary between workflows vs agents and reusable pattern library.
- Designing technical evaluations to be AI-resistant - Anthropic engineering guidance on interview/evaluation design in the agent era. Key technical takeaways: anti-leakage and anti-outsourcing evaluation patterns for measuring true human engineering skill.
- Claude Code Sub-Agents - Official description of sub-agent concepts, configuration scopes, and delegation triggers. Key technical takeaways: subagent definitions, scope boundaries, and delegation rules.
- Fast mode (research preview) - Official Claude API documentation for Fast mode. Key technical takeaways: lower-latency execution mode, tradeoffs versus standard quality profiles, and integration guidance for agent workflows.
- Extending Claude's capabilities with skills and MCP servers - Explains the separation of tool connectivity (MCP) and workflow logic (skills). Key technical takeaways: MCP handles connectivity while Skills encode workflow logic and orchestration.
- Cowork: Claude Code for the rest of your work - Research preview describing file-scoped autonomy, planning, and safety boundaries for agentic work. Key technical takeaways: file-scoped autonomy and plan-and-execute loops with safety constraints.
- Anthropic Cookbook - Official recipes and patterns for tool use, sub-agents, and evaluation loops (beyond API docs). Key technical takeaways: practical tool-use patterns, subagent recipes, and evaluation scaffolds.
- knowledge-work-plugins - Official Anthropic repository of plugins for Claude Cowork knowledge-work workflows. Key technical takeaways: reusable plugin patterns and integration surfaces for practical non-trivial workflows.
- claude-agent-sdk-demos - Official Anthropic demos for the Claude Agent SDK. Key technical takeaways: runnable examples for SDK-based agent orchestration and tooling patterns.
- Agent Trace: Capturing the Context Graph of Code - Cognition proposal for an open tracing format for coding agents. Key technical takeaways: context-graph traces for model/tool/file actions and a shared observability/evaluation substrate.
- Agent Trace (spec) - Specification site for the Agent Trace format. Key technical takeaways: interoperable schema for recording and analyzing agent execution traces.
- Devin's MCP Marketplace - Cognition post on MCP ecosystem integration in Devin. Key technical takeaways: curated server discovery, one-click MCP deployment, and governance-oriented integration workflows.
- Kimi K2.5: Visual Agentic Intelligence - Moonshot AI technical blog describing K2.5 and Agent Swarm. Key technical takeaways: self-directed swarm up to 100 sub-agents and 1,500 tool calls; K2.5 availability via Kimi.com/App/API/Kimi Code; Agent Swarm beta on Kimi.com.
- Qwen3-Coder-Next Technical Report (PDF) - Official Qwen3-Coder-Next technical report from QwenLM. Key technical takeaways: consolidated reference for model design, training/evaluation setup, and coding-agent capabilities.
- 2025-06-18 - Remote MCP support in Claude Code - Remote MCP server support and secure auth flow. Key technical takeaways: remote MCP connectivity and secure authentication for tooling.
- 2025-07-24 - How Anthropic teams use Claude Code - Internal adoption patterns and workflows. Key technical takeaways: internal rollout patterns and workflow guardrails.
- 2025-08-06 - Automate security reviews with Claude Code - Security review workflows and GitHub Action support. Key technical takeaways: automated security checks embedded in CI workflows.
- 2025-10-06 - Optimize code performance quickly - Performance optimization workflow with Claude Code. Key technical takeaways: profiling-to-optimization loop for performance work.
- 2025-10-08 - Beyond permission prompts: making Claude Code more secure and autonomous - Sandboxing + approval improvements for autonomy. Key technical takeaways: sandboxing and permission tiers for safer autonomy.
- 2025-10-09 - Customize Claude Code with plugins - Plugin system and extensibility model. Key technical takeaways: plugin architecture and extension points.
- 2025-10-10 - Build responsive web layouts - Responsive UI workflow and refactoring guidance. Key technical takeaways: responsive layout iteration and refactor patterns.
- 2025-10-15 - How to scale agentic coding across your engineering organization - Adoption playbook and team rollout guidance. Key technical takeaways: org-level rollout, governance, and adoption playbook.
- 2025-10-16 - Introducing Agent Skills - Official Skills announcement and core properties. Key technical takeaways: Skills as reusable tool-augmented behaviors.
- 2025-10-16 - Equipping agents for the real world with Agent Skills - Real-world agent skills architecture. Key technical takeaways: productionization of skills and real-world connectors.
- 2025-10-27 - How to integrate APIs seamlessly - API integration workflow guidance. Key technical takeaways: API integration patterns for agentic coding.
- 2025-10-28 - Fix software bugs faster with Claude - Debugging workflow guidance with Claude Code. Key technical takeaways: bug localization -> patch -> test loop.
- 2025-10-30 - Introduction to agentic coding - Defines agentic coding and Claude Code's workflow model. Key technical takeaways: agentic coding definition and loop model.
- 2025-10-31 - What is Model Context Protocol? Connect AI to your world - Official MCP overview and connector model. Key technical takeaways: MCP as a standard for tool and data connectivity.
- 2025-10-31 - Claude Code power user customization: How to configure hooks - Hooks for automation, guardrails, and context injection. Key technical takeaways: hook-based automation, guardrails, and context injection.
- 2025-11-10 - Best practices for prompt engineering - Prompting guidance grounded in real Claude usage patterns. Key technical takeaways: structured prompting patterns and constraint setting.
- 2025-11-12 - Improving frontend design through Skills - Frontend design skill patterns and prompting. Key technical takeaways: Skills templates for UI design improvements.
- 2025-11-13 - Skills explained: How Skills compares to prompts, Projects, MCP, and subagents - Taxonomy and tradeoffs across Skills and other mechanisms. Key technical takeaways: boundaries between Skills, prompts, Projects, MCP, and subagents.
- 2025-11-17 - How three YC startups built their companies with Claude Code - Startup case studies and workflows. Key technical takeaways: startup-scale agentic workflows and adoption lessons.
- 2025-11-19 - How to create Skills: Key steps, limitations, and examples - Practical skill creation workflow. Key technical takeaways: skill definition, testing, and limitations.
- 2025-11-25 - Using CLAUDE.md files: Customizing Claude Code for your codebase - Project-specific instruction patterns and context control. Key technical takeaways: repo-scoped instructions and context management.
- 2025-12-01 - The key benefits of transitioning to agentic coding - Practical and org-level benefits of agentic coding adoption. Key technical takeaways: productivity and quality gains plus org alignment.
- 2026-01-21 - Eight trends defining how software gets built in 2026 - Agentic coding trends and collaboration patterns. Key technical takeaways: emerging agentic workflows and collaboration trends.
- 2026-01-22 - Building agents with Skills: Equipping agents for specialized work - Skills ecosystem and layered agent architecture (loop/runtime/MCP/skills). Key technical takeaways: layered agent architecture and Skills ecosystem design.
- 2026-01-23 - Building multi-agent systems: when and how to use them - When to use multi-agent architectures and orchestration patterns. Key technical takeaways: selection criteria and orchestration patterns for multi-agent systems.
- 2026-01-29 - A complete guide to building skills for Claude - End-to-end guidance for skill design, testing, and distribution (including MCP + Skills). Key technical takeaways: full skill lifecycle (design -> test -> distribute) and MCP + Skills composition.
- 2026-01-29 - Building Skills for Claude Code - Skill structure and best practices in Claude Code. Key technical takeaways: skill structure and operational best practices for Claude Code.
- 2026-01-29 - Understand Claude Code's impact with contribution metrics - Introduces contribution metrics for PR/code impact measurement. Key technical takeaways: contribution metrics for agentic code impact and PR attribution.
- 2026-02-05 - Claude Opus 4.6 - Model release update from Anthropic. Key technical takeaways: newest Opus-series capability update relevant to coding quality and advanced agent workflows.
- OpenAI Podcast - Episode 6: The future of coding with AI - Greg Brockman and Codex lead Thibault Sottiaux on harnesses, agentic coding, and GPT-5 Codex evolution. Key technical takeaways: harness design principles, agent loop evolution, and safety constraints.
- Dev Interrupted: Scaffolding Is Coping, Not Scaling - Codex lead Thibault Sottiaux on codex-cli, agent reliability, and practical constraints. Key technical takeaways: reliability constraints and real-world boundaries for codex-cli.
- Agent Factory Recap: Deep Dive into Gemini CLI with Taylor Mullen - Gemini CLI creator on origin story, design philosophy, and roadmap. Key technical takeaways: Gemini CLI architecture choices and product roadmap.
- Decoder: Why tech is racing to adopt AI coding - Michael Truell on Cursor product choices and the long arc of agentic IDEs. Key technical takeaways: Cursor product strategy and agentic IDE trajectory.
- The a16z Show: How Cursor Builds at the Speed of AI - Constraints, editor-first strategy, and hiring philosophy. Key technical takeaways: editor-first strategy and constraints that shape agent behavior.
- Lenny's Podcast: The rise of Cursor - "After code" vision, custom model strategy, and product focus. Key technical takeaways: custom model strategy and product focus for agentic IDEs.
- YC Startup Podcast: Cursor CEO - Going Beyond Code - Fireside chat on "beyond code" trajectory and taste. Key technical takeaways: product taste and "beyond code" trajectory for agentic tools.
- We're All Addicted To Claude Code (Y Combinator) - YC conversation on real-world Claude Code usage patterns. Key technical takeaways: high-frequency coding workflows, collaboration habits, and practical tradeoffs in agent-assisted development.
- Peak XV: Aman Sanger on 0-$100M in 12 Months - Origin story, shipping culture, and future of programming. Key technical takeaways: shipping culture and operational focus for scaling agentic products.
- 2026-01-27 - Securely indexing large codebases - Securely reusing teammate indexes to cut time-to-first-query on very large repos.
- 2026-01-15 - Building a better Bugbot - Custom AI-driven metric used to systematically improve Bugbot quality.
- 2026-01-14 - Scaling long-running autonomous coding - Long-running agent orchestration patterns and execution lessons.
- 2026-01-06 - Dynamic context discovery - Let agents fetch context on demand rather than upfront.
- 2025-11-06 - Improving agent with semantic search - Semantic search as a measurable boost for agent performance.
- 2025-11-11 - The productivity impact of coding agents - Study on how coding agents change PR throughput.
- 2025-10-29 - Composer: Building a fast frontier model with RL - RL training notes for a fast code-editing model.
- 2025-09-12 - Improving Cursor Tab with online RL - Online RL to improve Tab completion acceptance.
- 2025-08-29 - 1.5x faster MoE training with custom MXFP8 kernels - Custom kernels that speed up MoE training on modern GPUs.
- 2024-09-01 - Iterating with shadow workspaces - Shadow workspaces for safe, isolated AI iterations.
- 2024-05-14 - Editing Files at 1000 Tokens per Second - High-speed full-file edit inference for code.
- 2024-05-25 - More problems - Updated list of research/engineering problems.
- 2023-10-12 - Our problems - Early list of open problems for AI coding.
- 2023-07-20 - Inference characteristics of Llama - Practical inference and latency characteristics of Llama.
- 2023-06-11 - Prompt design - Early prompt design lessons from production usage.
- 2026-01-09 - Best practices for coding with agents - Practical workflow guidance on planning, context, and review.
- 2025-12-22 - Hooks for security and platform teams - Hooks integrations for security and platform workflows.
- 2025-12-11 - A visual editor for the Cursor Browser - Visual editing flow for the Cursor Browser prototype.
- 2025-12-10 - Introducing Debug Mode: Agents with runtime logs - Debugging agents with runtime logs.
- 2025-12-04 - Improving Cursor’s agent for OpenAI Codex models - Agent harness updates optimized for Codex models.
- 2025-10-31 - Introducing Cursor for Enterprise - Enterprise controls, security, and deployment workflow.
- 2025-10-30 - Cloud Agents - Remote long-running agents and background workflows.
- 2025-10-29 - Introducing Cursor 2.0 and Composer - New interface and multi-agent experience.
- 2025-10-07 - Introducing Plan Mode - Planning workflows for longer agent runs.
- 2025-10-01 - Improving Java support in Cursor - Java LSP performance and ecosystem improvements.
- 2025-08-21 - Bringing the Cursor Agent to Linear - Linear integration for background agent tasks.
- 2025-08-07 - Cursor Agent CLI - CLI/headless agent usage for workflows.
- 2025-07-24 - Bugbot is out of beta - Bugbot release and PR review automation.
- 2025-06-30 - Cursor on web and mobile - Web and mobile access to the Cursor Agent.
- 2025-01-13 - A new Tab model - Tab model upgrade for code completion.
- 2025-01-06 - Character Prefix Conditioning - Code completion sampling technique for better prefix control.
- AB Method - Spec-driven workflow that turns large problems into incremental missions using Claude Code subagents. Key technical takeaways: spec-driven decomposition into incremental missions with subagents.
- Agentic AI Systems (workflow patterns) - Workflow and orchestration patterns with diagrams and taxonomy. Key technical takeaways: workflow taxonomy and orchestration patterns.
- Claude Code Handbook - Best practices, fundamentals, tips, and workflow guidance. Key technical takeaways: practical Claude Code workflow fundamentals and guardrails.
- claude-code-guide - Continuously updated Claude Code CLI guide. Key technical takeaways: broad command/workflow coverage and practical operating patterns.
- Claude Code Infrastructure Showcase - Skill auto-activation patterns using hooks and agents. Key technical takeaways: hook-driven auto-activation infrastructure.
- Claude Flow Skills: Complete Introduction Tutorial (Issue #821) - Community tutorial issue covering Claude Flow skill builder and flow skills usage. Key technical takeaways: setup flow for skills, builder-oriented workflow, and practical skill composition patterns.
- claude-code-best-practice - Community best-practice collection for Claude Code usage. Key technical takeaways: pragmatic usage patterns and guardrails distilled from hands-on practice.
- Claude Code System Prompts - System prompt segments, tool descriptions, and subagent prompts for studying agent design. Key technical takeaways: system prompts and tool descriptions that shape agent behavior.
- Claude Code PM (ccpm) - Spec-driven project management workflow using GitHub Issues and worktrees for parallel agent execution. Key technical takeaways: spec-driven PM with worktrees for parallelism.
- Claude Code Spec Workflow - Structured pipelines for new features (Requirements -> Design -> Tasks -> Implementation) and bug fixes (Report -> Analyze -> Fix -> Verify). Key technical takeaways: structured pipelines for features and bug fixes.
- Claude CodePro - Spec-driven workflow with TDD enforcement, modular rules, and persistent memory for production-quality work. Key technical takeaways: TDD enforcement, modular rules, and persistent memory.
- ContextKit - Context engineering and planning system to reduce micromanagement and improve first-pass quality. Key technical takeaways: context planning system to improve first-pass quality.
- claude-mem - Claude Code plugin that captures session activity, compresses it with Claude, and injects relevant memory into future sessions. Key technical takeaways: automated memory capture/compression and cross-session context reinjection for continuity.
- claude-howto - Visual, example-driven Claude Code guide with practical templates. Key technical takeaways: stepwise examples from basics to advanced agent workflows.
- oh-my-opencode - Community harness toolkit around opencode agent workflows. Key technical takeaways: harness-oriented automation patterns for improving agent execution flow.
- SuperClaude Framework - Configuration framework with specialized commands, cognitive personas, and development methodologies. Key technical takeaways: command framework with personas and structured methods.
- Simone for Claude Code - Project management framework for Claude Code with task planning and MCP-aware workflows. Key technical takeaways: task planning with MCP-aware workflows.
- Field Notes From Shipping Real Code With Claude - Practical workflow lessons on shipping production code with Claude, focusing on context control and guardrails. Key technical takeaways: context control and guardrails for production work.
- Context Priming - Systematic project priming commands for robust context loading. Key technical takeaways: repeatable context priming commands.
- Design Review Workflow - Agentic UI/UX review workflow with subagents and quality criteria. Key technical takeaways: subagent-driven UI/UX review with quality criteria.
- Project Bootstrapping and Task Management - Structured project bootstrap and task management command set. Key technical takeaways: bootstrap and task management command set.
- Project Management, Implementation, Planning, and Release - End-to-end SDLC command suite for planning through release. Key technical takeaways: SDLC command suite for planning -> release.
- Project Workflow System - Comprehensive workflow commands embedded in a real-world dotfiles system. Key technical takeaways: end-to-end workflow commands in a production dotfiles setup.
- RIPER Workflow for Claude Code - Research/Innovate/Plan/Execute/Review phase separation with subagents and memory bank for context efficiency. Key technical takeaways: phased loop with subagents and memory bank.
- How OpenAI Codex Works Behind-the-Scenes (and How It Compares to Claude Code) - System-level breakdown of Codex CLI internals with a side-by-side comparison to Claude Code. Key technical takeaways: harness architecture comparison and loop design differences.
- Claude Code: Behind-the-scenes of the master agent loop - Deep dive into the control loop and execution harness in Claude Code. Key technical takeaways: master loop orchestration and tool-call sequencing.
- Watch my AI Engineering talk: How Claude Code Works - Talk distilling Claude Code design signals from real usage and system behavior. Key technical takeaways: observable design signals from behavior and logs.
- Decoding Claude Code - Behavioral analysis based on usage/logs; emphasizes a simple control loop, small-model usage, tool design, and LLM-first search. Key technical takeaways: simple loop design, tool-first execution, and LLM-first search strategy.
- Decoding Claude Code (Chinese repost) - Chinese repost of the MinusX analysis. Key technical takeaways: Chinese repost of the core loop and tool design analysis.
- claude-code-reverse - Runtime/API log interception and visualization to study prompts, tools, and workflow structure. Key technical takeaways: prompt/tool tracing via runtime logs.
- analysis_claude_code - Obfuscated-code reverse engineering repo; documents architectural hypotheses (steering, multi-agent, context management). Key technical takeaways: deobfuscation-driven architecture hypotheses.
- claude-code-source-code-deobfuscation - Deobfuscation notes and artifacts for Claude Code internals. Key technical takeaways: source-level deobfuscation artifacts and notes.
- system-prompts-and-models-of-ai-tools - Large collection of system prompts, internal tools, and model configs across many AI tools/agents.
- agentic-system-prompts - Curated library of system prompts and tool definitions from production coding agents, with source attribution and comparisons.
- System-Prompt-Agent-Prompts - Large collection of system/agent prompts across AI models and dev tools with source attribution.
- awesome-ai-system-prompts - Curated system prompt collection for popular AI tools, plus prompt-pattern analysis.
- system_prompts_leaks - Collection of extracted system prompts from popular chatbots (ChatGPT, Claude, Gemini).
- What we can learn from Anthropic's system prompt updates - Extracts design patterns from Claude system prompt revisions across 2024-2025. Key technical takeaways: prompt-design patterns and governance signals.
- Claude Code has changed how we do engineering - Org-level workflow and prioritization shifts after adopting Claude Code. Key technical takeaways: org workflow shifts and prioritization changes.
- The Agentic System Design Interview - Interview framework for evaluating agentic system design skills. Key technical takeaways: evaluation rubric for agentic system design.
- 2025 State of AI Engineering Survey - Survey results on tooling, model churn, evals, and org practices. Key technical takeaways: tooling adoption, model churn, and eval practices.
- Prompt Engineering with Anthropic Claude - Practical prompting tips from Anthropic's "Prompt Doctor". Key technical takeaways: practical prompting patterns and guardrails.
- Michael Truell on X: “We built a browser with GPT-5.2 in Cursor” - CEO announcement of the week-long agent run and 3M+ LOC browser experiment. Key technical takeaways: long-horizon agent run and large-scale output claim.
- Scaling long-running autonomous coding - Cursor’s research write-up on planner/worker/judge roles and the browser-from-scratch experiment. Key technical takeaways: planner/worker/judge orchestration for long-running builds.
- FastRender (wilsonzlin/fastrender) - Public repo for the experimental browser engine referenced in the blog post. Key technical takeaways: reference implementation and artifact for the experiment.
- The Register: “Cursor used agents to write a browser…” - Critical coverage, including debate over code quality and reliance on open-source components (e.g., Servo). Key technical takeaways: external critique on quality and OSS dependence.
- GIGAZINE coverage of Cursor’s scaling-agents post - Summary report with links to the official blog and repo. Key technical takeaways: summary and link aggregation to primary sources.
- [zh] Zhihu: “Cursor一夜翻车,AI 300万代码写浏览器被打假”】【新智元】 - Chinese commentary discussing build failures and claims about heavy use of existing open-source components. Key technical takeaways: Chinese commentary on build reliability and OSS reuse claims.
- The Register forums discussion - Comment thread reacting to the experiment and its claimed results. Key technical takeaways: community reactions and critiques.
- AGENTS.md - Open format for agent-facing project instructions that complements README.md. Key technical takeaways: dedicated agent instruction layer separate from README.
- agentsmd/agents.md - Reference repository and site implementation for the standard. Key technical takeaways: canonical spec and implementation reference.
- OpenAI agents.md repo - OpenAI reference repository and examples for AGENTS.md. Key technical takeaways: OpenAI-backed examples and patterns.
- AGENTS.md examples on GitHub - Code search for real-world adoption. Key technical takeaways: real-world adoption inventory.
- agents-md - Compose AGENTS.md from fragments and auto-regenerate via pre-commit hooks. Key technical takeaways: templated composition and regeneration workflow.
- agentsmd - Templates and git automation to keep AGENTS.md updated locally. Key technical takeaways: templates and Git automation for maintenance.
- SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering - Shows how specialized agent-computer interfaces improve repository navigation and tool use.
- AgentCoder: Multi-Agent-based Code Generation with Iterative Testing and Optimisation - Multi-agent generation with programmer/tester roles and feedback loops.
- CodeDelegator: Mitigating Context Pollution via Role Separation in Code-as-Action Agents - Separates planning and implementation to reduce context pollution.
- Meta Context Engineering via Agentic Skill Evolution - Introduces autonomous skill evolution loops to improve context engineering and downstream agent performance.
- MegaFlow: LLM Orchestration for Agentic AI - Distributed orchestration system for large-scale agent training and evaluation.
- SWE-Gym: Training Software Engineering Agents and Verifiers - Training environment with real repos and executable tests.
- Agentic Refactoring: An Empirical Study of AI Coding Agents - Large-scale analysis of agent-generated refactorings and their characteristics.
- How do Agents Refactor: An Empirical Study - Compares agent vs human refactoring patterns and quality effects.
- Beyond Accuracy: Behavioral Dynamics of Agentic Multi-Hunk Repair - Studies multi-hunk bug repair trajectories and failure modes.
- On the Use of Agentic Coding: An Empirical Study of Pull Requests on GitHub - PR acceptance and task distribution in agent-generated contributions.
- On the Use of Agentic Coding Manifests: An Empirical Study of Claude Code - Empirical analysis of Claude.md manifests and common structure patterns.
- SWE-bench - Real GitHub issue benchmark for software engineering agents.
- SWE-bench Verified Technical Report (Verdent) - Reports pass@1/pass@3 resolved rates for Verdent and compares pass@1 results for multiple code agents based on official leaderboard data.
- SWE-Dev: Evaluating and Training Autonomous Feature-Driven Software Development - Feature-driven dataset with runnable environments.
- SWE-Dev: Building Software Engineering Agents with Training and Inference Scaling - Pipeline for scaling agent trajectories and test synthesis.
- Saving SWE-Bench: A Benchmark Mutation Approach for Realistic Agent Evaluation - Mutates benchmarks to more realistic user queries.
- Does SWE-Bench-Verified Test Agent Ability or Model Memory? - Tests for possible contamination and localization bias.
- The SWE-Bench Illusion: When State-of-the-Art LLMs Remember Instead of Reason - Evidence of memorization effects.
- Are "Solved Issues" in SWE-bench Really Solved Correctly? - Patch validity analysis and overestimation risk.
- AI Agentic Programming: A Survey of Techniques, Challenges, and Opportunities - Taxonomy of agentic programming techniques and open challenges.
- Large Language Model-Based Agents for Software Engineering: A Survey - Survey of agent methods in SE.
- Survey on Evaluation of LLM-based Agents - Evaluation methods and benchmark landscape.
- Model Context Protocol (MCP) - Official overview of the protocol.
- Augment MCP overview - Augment documentation on MCP architecture and integration model. Key technical takeaways: MCP connectivity model in Augment context services and practical setup concepts.
- MCP Servers (directory) - Community directory of MCP servers.
- MCP Market (Chinese) - Chinese MCP marketplace and listings.
- MCP Registry - Alternative server index.
- MCP Serve - MCP server directory.
- claude-code-mcp - Claude Code exposed as a one-shot MCP server for agent-in-agent workflows. Key technical takeaways: packaging Claude Code capabilities behind MCP for composable orchestration.
- MCP Safety Audit - Security risks and auditing tool for MCP servers.
- Model Context Protocol at First Glance - Large-scale empirical study of MCP server security and maintainability.
- Beyond the Protocol: Unveiling Attack Vectors in the MCP Ecosystem - Systematic analysis of malicious MCP servers.
- "MCP Does Not Stand for Misuse Cryptography Protocol" - Cryptographic misuse at scale in MCP servers.
- MCP Guardian - Security-first layer for MCP-based systems.
- SoK: Prompt Injection Attacks in Agentic LLM-based Assistants - Systematizes prompt-injection threat models and defenses for agentic tools.
- Evaluating and mitigating growing risk of zero-day jailbreaks for LLM safety safeguards - Anthropic red-team analysis of adaptive jailbreak attacks. Key technical takeaways: zero-day attack framing, transferability concerns, and layered mitigation strategy for model and product defenses.
- AI Agent, AI Spy (39C3) - Signal's Whittaker and Tiwari analyze OS-level agent integration as surveillance risk and outline mitigation principles (YouTube mirror: https://youtu.be/0ANECpNdt-4).
- Agent Skills standard - Open standard for skills packaging and interoperability.
- skills.sh - Community catalog of skills.
- Agent Skills Index - Skills directory and categories.
- Agent-Skills-for-Context-Engineering - Community repository of reusable agent skills for context engineering. Key technical takeaways: practical skill modules and context-oriented prompting patterns.
- awesome-agent-skills - Curated collection of reusable skills for coding and agent workflows. Key technical takeaways: ecosystem map of skill packs and implementation references.
- Claudeception - Claude Code skill for autonomous skill extraction and continuous learning. Key technical takeaways: self-improving skill loop and automatic knowledge capture from ongoing work.
- Context Engineering Kit - Plugin marketplace focused on improving agent result quality across multiple coding agents.
- Trail of Bits Skills - Security-focused Claude Code skills for audit and vulnerability workflows.
- everything-claude-code - Configuration corpus spanning agents, skills, hooks, commands, rules, and MCPs.
- open-claude-cowork - Open-source Cowork-style desktop agent with tool integrations.
- openwork - Open-source alternative to Claude Cowork powered by opencode. Key technical takeaways: open implementation of Cowork-style autonomous workflows.
- Auto-Claude - Autonomous multi-session coding app with parallel agent terminals, worktree isolation, self-validating QA, and memory across sessions.
- Continuous-Claude-v3 - Context-management and orchestration framework for Claude Code with hooks and handoffs. Key technical takeaways: ledger-based state continuity and isolated context windows for long-running loops.
- Ralph for Claude Code - Autonomous development loop with intelligent exit detection (Ralph Wiggum technique).
- ralph-orchestrator - Expanded Ralph Wiggum orchestration system for autonomous agent runs.
- Ralph Playbook - Practical guide to running autonomous Ralph loops with theory and guardrails.
- awesome-claude-code - Workflows, skills, and tooling for Claude Code.
- Awesome-Agent4SE - Agents in software engineering resources.
- Awesome-Code-LLM - Code LLM papers, benchmarks, and leaderboards.
- awesome-llm-agents - LLM agent frameworks and tools.
- awesome-devins - Devin-style agent projects.
- LLM-Agent-Survey - Survey repo and paper list on LLM agents.
- learn-claude-code - Minimal, step-by-step agent implementation path.
- MCP protocol evolution: SSE to Streamable HTTP - Transport-layer changes and architectural implications.
- What is MCP and why it matters - Conceptual overview and design rationale.
- Zhihu series: Coding Agent topics - Series entry point aggregating Coding Agent/Claude Code discussions.
- Are we evaluating coding agents in the wrong direction? - Chinese discussion on benchmark alignment and real-task validity for coding-agent evaluation.
- From beginner to practical Agent Skills: MCP + Context Engineering - Chinese walkthrough on combining Agent Skills with MCP and context-engineering workflows.
- Claude Code safety and plugin chain analysis - Discussion of plugin/MCP control flow risks.
- Claude Code: deep experience and practice - Workflow-focused experience report.
- Claude Code vs Gemini CLI (comparative review) - Comparative evaluation of Claude Code and Gemini CLI.
- GPT-5-Codex agent landscape - Survey-style overview of coding agents around GPT-5-Codex.
- OpenAI Codex open-source code analysis - Zhihu analysis of the open-source Codex release and code structure.
- GPT-5-Codex in-depth podcast summary - Zhihu recap referencing OpenAI Podcast Episode 6 with Codex details.
- Zhihu: Claude system prompt analysis - Community analysis of Claude system prompts.
- WeChat Official Account MCP service listing - MCP-based integration overview for public accounts.
- OpenAI Podcast recap: Codex past, present, future - Chinese summary of the OpenAI podcast interview on Codex.
- Interview: OpenAI Codex product lead - Chinese interview-style writeup on the origins and roadmap of Codex.
- Gemini CLI interview notes (Agent Factory podcast) - Chinese summary of the Gemini CLI interview with Taylor Mullen.
- Anthropic official talk: Claude Code principles and scenarios - Principles and internal best practices.
- Deep dive: How the Claude Code team built an AI-native product - Interview-style breakdown of Claude Code product and engineering choices.
- OpenAI Codex research preview (Greg Brockman talk) - Talk covering Codex research preview and product framing.
- Reverse engineering Claude Code workflow - Reverse analysis of Claude Code core logic.
- Claude Code best practices (19 min) - Explains Anthropic official best practices and usage tips.
- Claude Code external behavior framework - Framework-level analysis from observed behavior.
- Claude Code task/sub-agent mechanism - Task and sub-agent behavior walkthrough.
- OpenAI Codex team discussion - Interview-style conversation on Codex capabilities and product framing.
- Claude system prompt analysis (video) - Community walkthrough of Claude system prompt content.
- Claude 3.7 system prompt (video) - Focused walkthrough of Claude 3.7 system prompt content.
- Agent principles: ReAct and Plan-and-Execute walkthrough - Short-form breakdown of agent concepts and loops.
- OpenAI Codex quick intro - Short-form overview of Codex capabilities and positioning.
- Xiaohongshu AI agent mini-app docs - Platform guidance for agent-oriented mini-apps.