Summary
Cursor built an LLM-powered security review system to examine every pull request, identify and validate vulnerabilities, and provide feedback directly in source control. The system combines specialized asynchronous review agents, analytical triage agents, fast-model deduplication, precomputed code facts, developer feedback, and a serverless MCP-based tool gateway. After an initial GitHub Action became too slow at scale, the architecture was rebuilt around parallel processing and selective agent execution. The resulting system became a hard merge gate for unacknowledged security findings, while remaining integrated with developers’ coding agents and continuously adapting to threat-model decisions and feedback. The company reports improved coverage and actionable findings, but the system’s effectiveness depends on model validation, prompt quality, human oversight, reliable infrastructure, and careful control of cost and latency.
In Production
Four LLMOps case studies like this one, in your inbox every Tuesday and Thursday.
On this page
Overview
Cursor developed a production security agent intended to review every pull request for vulnerabilities without becoming a bottleneck for a very high-velocity engineering organization. Rather than relying exclusively on conventional static-analysis rules, the system uses LLM agents to reason about potential attack paths, inspect surrounding code, validate findings, and return actionable comments in the pull request. It eventually became a hard security gate: code containing an unacknowledged active security issue cannot be merged unless an incident is referenced or the responsible manager is paged.
The system is best understood as a hybrid LLMOps control plane rather than a single prompt wrapped around a model. It combines deterministic CI and source-control integration with parallel specialist reviewers, separate triage agents, fast-model deduplication, an MCP-based tool interface, persistent state, precomputed code facts, and feedback loops driven by developer responses. The reported outcome is broad review coverage and findings that developers’ own coding agents can often remediate automatically. However, the supplied results are primarily internal operational observations rather than an independently measured security benchmark, so claims about detection quality, precision, and prevented vulnerabilities should be treated as directional.
Problem and Requirements
When Cursor’s security program was being expanded, dependency scanning surfaced approximately 3,500 vulnerabilities. The immediate challenge was not simply finding more issues; it was establishing a process for vulnerability management and reducing noise while the engineering organization continued to ship code at an unusually fast rate. Traditional security review introduces a substantial cognitive-switching cost because security staff must decide which changes deserve manual inspection and then repeatedly move between development and security workflows.
The security-agent project was designed around several operational requirements. It needed to be fast enough that security review would not become the slowest part of CI, scalable enough to handle a rapidly increasing pull-request volume, and dependable enough to support a blocking merge policy. It also had to present findings where developers were already working, particularly in GitHub and source control, rather than requiring a separate security application. Finally, the system needed to evolve as the company’s threat model and paved-road engineering practices changed. A fixed list of rules was considered insufficient for that purpose.
Initial Deployment and Iteration
The first implementation was a GitHub Action triggered by pull requests. It used a CLI wrapper and a simple instruction to find vulnerabilities. This produced useful discoveries, but the early findings were not initially considered reliable or actionable enough to expose broadly to developers. The implementation was refined for several months, with additional instructions requiring the agent to validate its claims instead of reporting speculative problems.
The validation approach emphasized proving the attack chain: the agent had to establish how input moved through the code, whether it was attacker-controlled, what happened downstream, and what the practical impact would be. The system was explicitly pushed to validate findings line by line and file by file. This is an important production safeguard because a language model can produce plausible but unsupported paths, and a security gate cannot treat an unverified suggestion as equivalent to a demonstrated vulnerability.
Cursor Automations later replaced much of the custom trigger logic. Automations could run an agent in response to configured events and provide connected tools, allowing the security work to focus more on reviewer quality rather than maintaining all of the surrounding GitHub-action plumbing. The system then expanded from a general security reviewer into specialized agents for infrastructure-as-code issues, company-specific conventions, and other “paved road” or product-specific security requirements.
Architecture and Agent Roles
The eventual architecture runs from Buildkite and uses an internal package, described as ASR local, against the latest main branch or a pinned revision for testing. For each pull-request diff, the system determines the affected file paths and selects the applicable reviewers. Reviewers execute asynchronously and in parallel rather than serially, which is central to reducing wall-clock latency.
The design separates broad discovery from focused judgment. A review agent is creative and fast: it examines the relevant diff and looks for possible security issues with limited initial validation. A triage agent receives one candidate finding at a time and has the narrower responsibility of deciding whether the finding meets the reporting criteria. Giving these personas different prompts and responsibilities prevents every agent from performing the full investigative workload and helps balance completion times. Different vulnerability domains have their own review and triage prompts, so an infrastructure-as-code reviewer does not have to use the same context or criteria as a reviewer for application code or organization-specific conventions.
When a reviewer produces a candidate, it is placed into a queue. A fast LLM performs an initial deduplication pass so that equivalent findings are not repeatedly sent through triage. After triage, another deduplication pass handles cases where separate reviewers identified the same underlying issue or where triage determined that two findings should be merged. This layered approach is intended to reduce both developer noise and wasted model calls.
The agents interact with an MCP interface that acts as a lightweight API gateway over serverless components, including Lambda and DynamoDB. Through this boundary, agents can retrieve relevant comments and state, propose or record facts, and set gate status without receiving broad direct access to GitHub credentials. The gateway also provides a persistence layer and a controlled interface to external systems. This illustrates a practical LLMOps pattern: tools are used to define explicit capabilities and authorization boundaries, while deterministic services retain responsibility for state transitions and CI integration.
Context Management and Precomputed Facts
A major optimization is to avoid making every reviewer rediscover the same codebase facts. The system runs an automation several times per day to harvest new facts and update existing facts when code movement changes their locations. Examples include whether an input is attacker-controlled, whether it is sanitized, and what downstream operation receives it. A fuzzy matching process selects facts relevant to the current code location and supplies only those facts to the agents that need them.
This design has several benefits. It reduces repeated investigation, limits context size, and avoids asking a single agent to reason about every fact in the repository. It also improves consistency across reviewers. At the same time, precomputed facts create an operational dependency: stale or incorrectly matched facts could bias later decisions. Periodic refresh and location updates are therefore important, but the material does not provide a formal freshness SLA or an independent accuracy measurement for this fact store.
Gating, Remediation, and Feedback
The system posts findings as pull-request comments and exposes commit statuses. Developers can respond to comments, and their coding agents can monitor CI, make changes, and iterate until the security findings are resolved. In this workflow, many developers do not need to read every finding manually; the coding agent consumes the feedback and attempts remediation. A finding is marked fixed only after an LLM validates that the issue was resolved in a later commit, creating a record of the claimed remediation rather than simply relying on a changed line count.
The hard gate is deliberately designed to create enough friction that the issue receives attention. An informational comment can be ignored, but a failed merge status gives the developer’s agent a concrete exit condition. If a finding is legitimate but intentionally accepted, the exception requires an active incident reference or manager escalation. This is a strong policy choice: it can improve accountability and reduce silent risk acceptance, but it also makes availability and precision critical. A system that fails closed, blocks on false positives, or cannot complete reliably can directly disrupt delivery.
Developer feedback is incorporated into the operating loop. A finding may be technically real but outside the company’s threat model, such as an accepted risk or an issue that the organization has decided not to prioritize. The system distinguishes this kind of noise from a model hallucination or an incorrectly inferred attack path. Comments and responses are harvested to identify recurring patterns. Repeated negative feedback on a class of finding can trigger review of its rule, prompt, fact, or reporting threshold rather than requiring the entire reviewer to be removed. Candidate facts proposed by agents are themselves reviewed by another recurring agent before becoming triage knowledge.
Scaling and Cost Tradeoffs
The first implementation eventually exceeded eight minutes of runtime as more reviews were added and execution became congested. The architecture was rebuilt around parallel reviewers, selective routing, queues, deduplication, and the separation of review from triage. This was necessary because the system’s value depends not only on detection quality but also on predictable pull-request latency.
Model selection and token cost were significant operational concerns. An experiment that upgraded all components to a more capable model produced a reported token spend of about $100,000 in one day, after which the configuration was dialed back to a less expensive medium or fast setting. This demonstrates why model quality cannot be optimized in isolation. A production security system needs per-stage model selection, constrained context, routing, and usage monitoring. Fast models are useful for deduplication and potentially lightweight screening, while more capable reasoning may be reserved for cases that survive initial filtering.
The system also benefits from specializing work, both to reduce token usage and to avoid degrading agent performance by giving one prompt too many unrelated responsibilities. Nevertheless, no complete cost-per-pull-request, false-positive rate, recall estimate, or service-level objective is provided. Those measurements would be needed to evaluate whether the design remains economically sustainable as repository size and pull-request volume grow.
Results and Tradeoffs
Cursor reports that developers reacted positively once the findings became sufficiently actionable, and that the system reached a level of internal adoption where teams wanted additional specialized reviewers. It also reports that the gate is faster than the intended CI threshold after the redesign, that reviewers operate across all pull requests, and that many findings are automatically fixed through the coding-agent feedback loop. The recorded fixed status provides a way to quantify completed remediation events, although it does not by itself prove that every vulnerability was correctly detected or fully eliminated.
The principal strength of this approach is coverage: agents can inspect changes that a small security team would otherwise triage away. The principal risk is that an LLM-based security gate can produce unsupported findings, miss subtle vulnerabilities, become stale as code and threats change, or fail operationally at exactly the point where it is enforcing a merge decision. The architecture addresses some of these risks with attack-chain validation, specialist prompts, multiple triage stages, human feedback, fact review, controlled tools, and asynchronous execution. It does not eliminate the need for conventional deterministic checks, incident response, threat modeling, or human security judgment.
Overall, the case demonstrates how LLMOps for security review extends beyond prompt engineering. Reliable production use requires explicit agent roles, controlled tool access, state management, queueing, deduplication, latency and cost controls, feedback-driven policy updates, and a clearly defined relationship between model output and a consequential CI gate. The reported implementation is a promising internal security automation pattern, but its claims should be validated with longer-term metrics for precision, recall, remediation correctness, availability, cost, and the number and severity of vulnerabilities prevented.