GPT-5.6 (Sol / Terra / Luna) is now evaluated on TrustVector โ€” with day-1 independent verification, incl. METR's benchmark-cheating findings.

Read the evaluation
Evaluation record ยท cursor-agent

Cursor Agent

v3.x

Anysphere (SpaceX acquisition announced 2026-06-16, expected to close Q3 2026)

Agentideagentic-codingparallel-agentsproprietary
72
Adequate
About This Agent

Agent mode of Cursor, Anysphere's AI-native IDE. The product's primary surface is now agentic: parallel local agents, cloud/background agents running in isolated VMs, and the in-house Composer model line (Composer 2.5) alongside Claude, GPT, and Gemini. Anysphere IPO'd on Nasdaq in June 2026 and days later agreed to a $60B all-stock acquisition by SpaceX (announced 2026-06-16, expected to close Q3 2026 under its xAI subsidiary).

Last Evaluated: July 9, 2026
Official Website

Trust Vector Analysis

Dimension Breakdown

๐Ÿš€Performance & Reliability
+
task completion accuracy

Assessment of agentic edit and task completion quality across Composer and frontier model options, drawing on vendor benchmarks and user reports

Evidence
Cursor Blog - Composer 2 โ€” Composer 2 (2026-03-19) improved agentic coding accuracy and speed over Composer 1, which shipped with Cursor 2.0 in October 2025
mediumVerified: 2026-07-09
tool use reliability

Review of agent tool loop reliability across edit, search, terminal, and MCP integrations

Evidence
Cursor Documentation - Agent โ€” Agent reliably uses codebase search, terminal commands, file edits, lints, and MCP tools within the IDE loop
highVerified: 2026-07-09
multi step planning

Evaluation of plan construction and multi-file refactoring on complex tasks in the agent-first interface

Evidence
InfoQ - Cursor 3 Agent-First Interface โ€” Cursor 3 (April 2026) redesigned the IDE around agent plans and task orchestration as the primary workflow
mediumVerified: 2026-07-09
memory persistence

Review of rules, memories, and codebase indexing persistence

Evidence
Cursor Documentation - Rules and Memories โ€” Project rules, user rules, and Memories persist conventions and learned context across sessions
mediumVerified: 2026-07-09
error recovery

Assessment of automatic iteration on lints, build errors, and failing tests

Evidence
Cursor Documentation - Agent โ€” Agent reads linter and test failures and iterates automatically; loop detection prevents most runaway retries
mediumVerified: 2026-07-09
agent collaboration

Review of parallel and background agent orchestration capabilities

Evidence
InfoQ - Cursor 3 Agent-First Interface โ€” Parallel agents run concurrently on separate tasks, locally via worktrees or in cloud VMs as background agents
highVerified: 2026-07-09
๐Ÿ›ก๏ธSecurity
+
tool sandboxing

Security review of cloud VM isolation versus local execution approval model

Evidence
Cursor Documentation - Background Agents โ€” Cloud/background agents execute in isolated VMs; local agents run terminal commands on the user's machine with configurable approval and sandbox settings
mediumVerified: 2026-07-09
access control

Review of org-level controls, SSO, and repository scoping

Evidence
Cursor Security โ€” Teams/Enterprise offer SSO, admin controls, and centrally enforced privacy mode; repository access scoped via GitHub app for cloud agents
mediumVerified: 2026-07-09
prompt injection defense

Threat surface analysis of untrusted content ingestion against documented mitigations

Evidence
Cursor Security โ€” Agents process untrusted repo content, web results, and MCP outputs; command allowlists and approvals mitigate but injection hardening is not publicly detailed
lowVerified: 2026-07-09
data isolation

Review of privacy mode guarantees and tenant isolation for cloud agents

Evidence
Cursor Security โ€” Privacy mode guarantees code is not stored or trained on; SOC 2 attested infrastructure with per-tenant separation for cloud agents
mediumVerified: 2026-07-09
open source transparency

Source availability assessment

Evidence
Cursor โ€” Proprietary VS Code fork; Composer models and agent harness are closed source
highVerified: 2026-07-09
๐Ÿ”’Privacy & Compliance
+
data retention

Review of privacy mode retention guarantees and default settings

Evidence
Cursor Privacy โ€” Privacy mode (default for Teams) prevents code storage and training use; without it, telemetry and code data may be retained
mediumVerified: 2026-07-09
gdpr compliance

Compliance documentation assessment

Evidence
Cursor Security โ€” SOC 2 Type II attestation, DPA availability, and GDPR-aligned processing terms for business customers
mediumVerified: 2026-07-09
third party data sharing

Data flow analysis across multi-provider model routing

Evidence
Cursor Documentation - Models โ€” Prompts and code context are routed to the selected model provider (Anysphere Composer, Anthropic, OpenAI, Google), expanding the data-processing surface
mediumVerified: 2026-07-09
local deployment option

Deployment options assessment

Evidence
Cursor Documentation โ€” IDE runs locally but all model inference is cloud-based; no on-premises or fully offline mode
highVerified: 2026-07-09
๐Ÿ‘๏ธTrust & Transparency
+
documentation quality

Documentation completeness review

Evidence
Cursor Documentation โ€” Thorough docs covering agent mode, background agents, rules, MCP, models, and enterprise administration
highVerified: 2026-07-09
execution traceability

Review of action visibility, diff review workflow, and agent logs

Evidence
Cursor Documentation - Agent โ€” Every agent action is visible: diffs reviewable before apply, terminal output streamed, and per-agent activity logs in the agent-first UI
highVerified: 2026-07-09
decision explainability

Assessment of plan narration and change rationale quality

Evidence
InfoQ - Cursor 3 Agent-First Interface โ€” Agents narrate plans and rationale in the task view, though depth of explanation varies by selected model
mediumVerified: 2026-07-09
open source code

Open source assessment

Evidence
Cursor โ€” Closed-source editor fork and proprietary Composer models; no published weights or harness internals
highVerified: 2026-07-09
community activity

Community engagement and release cadence analysis

Evidence
Cursor Community Forum โ€” Very large active user community, busy forum, and rapid release cadence (Cursor 2.0 Oct 2025, Composer 2 Mar 2026, Cursor 3 Apr 2026, Composer 2.5 mid-2026)
highVerified: 2026-07-09
โš™๏ธOperational Excellence
+
ease of integration

Onboarding and integration friction assessment

Evidence
Cursor Documentation โ€” VS Code fork imports existing extensions, settings, and keybindings; developers are productive in minutes
highVerified: 2026-07-09
scalability

Scalability assessment of parallel and cloud agent execution

Evidence
Cursor Documentation - Background Agents โ€” Cloud agents offload long tasks to isolated VMs and run in parallel, decoupling agent throughput from local hardware
mediumVerified: 2026-07-09
cost predictability

Pricing model analysis including the June 2026 Teams restructure; flat tiers are clear but usage-based overages reduce predictability for heavy agent users

Evidence
Cursor Pricing โ€” Free / Pro $20 / Pro+ $60 / Ultra $200 per month; Teams $40 per seat. Tiers include usage allowances with overage billing for heavy frontier-model use
Cursor Blog - Improvements to Teams Pricing โ€” June 2026 Teams overhaul (effective 2026-07-01 for renewals): Standard seat $40/mo ($32 annual) with split usage pools for Composer/Auto vs third-party API models; new Premium seat at $120/mo ($96 annual) with 5x usage; Cursor estimates lower costs for ~90% of teams
highVerified: 2026-07-09
monitoring capabilities

Monitoring and admin analytics features assessment

Evidence
Cursor Documentation - Teams โ€” Team usage dashboards and admin analytics; deep observability of agent behavior requires external tooling
mediumVerified: 2026-07-09
production readiness

Product maturity and adoption assessment; score held at 85 as strong adoption and revenue are offset by pending ownership change

Evidence
InfoQ - Cursor 3 Agent-First Interface โ€” Mature, widely adopted product with massive enterprise and individual user base and sustained rapid iteration through Cursor 3
TechCrunch - SpaceX to acquire Cursor for $60B โ€” Annualized revenue reached ~$4B by early June 2026; Nasdaq IPO June 2026 followed by $60B all-stock SpaceX acquisition agreement (close expected Q3 2026), introducing ownership-transition uncertainty alongside strong commercial traction
highVerified: 2026-07-09
Strengths
  • +Agent-first IDE redesign (Cursor 3, April 2026) makes agent orchestration the primary workflow
  • +In-house Composer model line (latest Composer 2.5) delivers very fast agentic edits alongside Claude/GPT/Gemini choice
  • +Parallel agents and cloud background agents in isolated VMs scale work beyond one task at a time
  • +Human-in-the-loop by design: reviewable diffs, command approvals, and visible terminal output
  • +Seamless VS Code compatibility for extensions, themes, and keybindings
  • +Privacy mode with no-storage/no-training guarantee, enforceable org-wide
Limitations
  • !Closed-source editor and models limit independent security and behavior auditing
  • !No offline or on-premises inference; all model calls go to cloud providers
  • !Usage-based overages above tier allowances make heavy agent usage costs less predictable
  • !Local agent terminal execution depends on user-configured approvals; misconfiguration widens risk
  • !Multi-provider model routing complicates data governance reviews
  • !Rapid release cadence occasionally introduces regressions and workflow changes
  • !Pending SpaceX/xAI acquisition (announced 2026-06-16, closing Q3 2026) creates governance, roadmap, and data-stewardship uncertainty for enterprise buyers
Metadata
license: Proprietary
supported models
0: Composer 1/2/2.5 (Anysphere in-house)
1: Anthropic Claude
2: OpenAI GPT
3: Google Gemini
programming languages
0: All languages supported by VS Code ecosystem
deployment type: Local IDE (VS Code fork) + cloud agents in isolated VMs
tool support
0: Codebase search and indexing
1: Terminal execution
2: MCP servers
3: Browser/web context
4: GitHub integration for cloud agents
first release: 2023 (IDE); Cursor 2.0 with Composer Oct 2025; Cursor 3 April 2026
pricing: Free / Pro $20 / Pro+ $60 / Ultra $200 per month; Teams (June 2026): Standard $40/mo ($32 annual), Premium $120/mo ($96 annual) per seat
company: Anysphere (Nasdaq IPO June 2026; ~$4B ARR; $60B all-stock SpaceX acquisition announced 2026-06-16, expected close Q3 2026 under xAI subsidiary)

Use Case Ratings

code generation

Best-in-class agentic IDE experience; parallel agents, fast in-house Composer models, and frontier model choice

data analysis

Strong for building and iterating on analysis code and notebooks, though not an analytics product itself

research assistant

Useful for technical research within codebases and docs; general research is outside its design focus

education

Visible agent reasoning and diffs help learners understand changes; generous free tier lowers the barrier