Summary
Replit redesigned its Agent harness so the core language model, rather than a fixed router or prescribed workflow, decides how much reasoning to use, whether to delegate, which specialist to invoke, and whether to reuse an existing subagent. The production system combines domain-aware subagents, model and effort tiers, reusable worker context, and mid-turn effort adjustment. In Replit’s reported evaluations, the Astra-based Agent scored 72% on DeepSWE v1.1 at $2.11 per task and 49% on Terminal-Bench 4.0 at $2.53, outperforming the company’s single-worker sidekick architecture by 11 and 16 percentage points respectively, while generally offering a better cost-quality tradeoff than Astra alone. These results are promising but are based on Replit’s own runs and comparisons with published baselines, so they should not be treated as an independent, universally representative measurement of production quality.
In Production
Four LLMOps case studies like this one, in your inbox every Tuesday and Thursday.
- Industry
- Tech
On this page
Overview
Replit operates an AI development platform in which users describe software tasks in natural language and Replit Agent works on applications, websites, and codebases. The case study concerns the production design of that agent, not merely an offline prompt experiment. Replit argues that conventional model routers have a structural limitation: a router that examines a request and selects an LLM cannot be more capable than the model making the selection. Instead, Replit gives the main agent—or core loop—control over delegation and reasoning decisions while the harness supplies available tools, specialist types, model tiers, and safety or quality guardrails.
The resulting architecture is adaptive. At each step, the core loop can decide how hard to think, whether to perform work itself or delegate it, which domain-aware subagent to use, what model tier and effort level to assign, and whether to return to a previously briefed subagent. Replit reports that this approach achieved better cost-quality tradeoffs than both Astra running alone in mini-swe-agent and a simpler architecture with one persistent sidekick worker. The evidence is useful for LLMOps practitioners because it describes a production control plane for frontier models, but the results remain vendor-reported and depend on selected models, benchmarks, exclusions, and comparison configurations.
Problem and design rationale
Replit’s underlying problem is that coding, UI design, slide creation, exploration, testing, and writing are different workloads. A single strongest model is not necessarily the best model for every domain, while using an expensive frontier model for every operation can make an agent uneconomical. A fixed orchestration plan also risks encoding assumptions that become obsolete as models improve. Replit therefore treats the harness as a composable set of capabilities rather than a rigid workflow.
The system is intentionally less prescriptive than a traditional router. The harness determines which kinds of specialists exist and what tools they can use, but the core model chooses when to invoke them. A small mechanical change may require no subagent. Searching a codebase may justify a read-only explorer. A difficult bug may justify a larger worker at high effort. Independent tasks can be dispatched in parallel, while a tester or reviewer can validate the result. This allows orchestration to vary with the trajectory of the task instead of being decided once from the initial request.
Production architecture
The core loop is built around four stated primitives. Domain-aware subagents include a general worker as well as read-only explorers, browser testers, reviewers, and a design-oriented subagent for UI and slides. Each specialist has its own model and tooling. The source says that the set of specialist types is still chosen by the harness; the model controls their use during execution. This is an important distinction: the system is not unrestricted self-modification, but model-directed selection within an engineered action space.
Subagents are available in small, standard, and large tiers, with an effort setting for each tier. Tier affects the cost and capability envelope, while effort controls how much reasoning is applied. The model can therefore send routine work to a smaller, lower-effort configuration and reserve a larger, higher-effort configuration for unresolved or complex work. Replit describes this as a per-dispatch decision rather than a global setting for the entire user request.
The architecture also supports reusable subagents. Rather than maintaining one always-on sidekick for the whole session, the core loop can return to a subagent that has already been briefed. Multiple subagents can remain warm across types and tiers, and the model selects which one to reactivate. Reuse preserves relevant context and can avoid repeating an initial briefing, while the absence of a single permanently active worker limits the requirement to keep one fixed decomposition alive throughout a task.
Dynamic effort tuning is another important control. Replit reports that GPT-6 Astra and some other models can change effort mid-turn without a cache miss. Its escalation system checks the trajectory at each step and adjusts effort to the apparent difficulty of the work. On models where effort changes or model switches rebuild the cache, the same operation has a higher latency and cost implication. Thus, cache behavior is not an incidental implementation detail; it directly affects the economic viability of adaptive reasoning and model switching.
Observed production behavior
Replit provides one-week production observations for Fable 5 in August 2026 and Fable 5.1 and GPT-6 Astra in September 2026, all at medium reasoning effort. Fable 5 dispatched a subagent on 32% of turns, Fable 5.1 on 21%, and Astra on 36%. The percentage of turns handing work to a general worker was 0.9%, 2.3%, and 20%, respectively. Among dispatches, returns to an existing subagent were reported at 17% for Fable 5, 29% for Fable 5.1, and 42% for Astra.
These figures suggest that model behavior matters substantially even when the surrounding harness is unchanged. Replit reports that the Fable models more often use explorers and reviewers while retaining implementation work, whereas Astra more frequently delegates implementation to general workers and returns to workers it has already briefed. The figures are descriptive rather than a controlled causal analysis: they cover short production windows, one effort setting, and particular model versions. They nevertheless illustrate an operational requirement for instrumentation. A production agent needs traces showing dispatches, worker identity, task type, model tier, effort, cache behavior, and returns in order to understand cost and quality.
Evaluation and reported results
Replit evaluated Agent in Max mode with Astra as the core loop on DeepSWE v1.1 and Terminal-Bench 4.0. The company compared it with Astra alone in the published mini-swe-agent configurations and with a sidekick architecture that replaced the subagent primitives with one long-lived worker. Replit states that its own results are means of four repetitions with 95% intervals calculated as mean plus or minus 1.96 standard errors.
On DeepSWE v1.1, covering long-horizon changes to active open-source repositories, Replit Agent scored 72% at $2.11 per task. The cited Astra-alone results were 67% at $1.60 with low effort and 74% at $4.43 with xhigh effort. The sidekick architecture scored 61% at $1.34. On Terminal-Bench 4.0, which evaluates multi-step shell work, Replit Agent reached 49% at $2.53 per task. Astra-alone results were 42% at $2.25 at low effort and 60% at $5.86 at xhigh effort, while the sidekick scored 33% at $1.84.
On the reported plots, Replit characterizes its system as Pareto-efficient against Astra alone: the highest-scoring Astra settings cost substantially more, while lower-cost Astra settings score less. It also reports an 11-point improvement over the sidekick on DeepSWE and a 16-point improvement on Terminal-Bench. These are meaningful signals for architecture selection, especially because the sidekick costs less but sacrifices score. However, the comparison is not a fully independent head-to-head study. Astra’s comparison numbers come from published mini-swe-agent baselines, the Replit runs use four repetitions, three GPU tasks are excluded from the Terminal-Bench runs, and the benchmarks may not represent all customer workloads. The source also does not provide production customer success rates, latency distributions, failure recovery rates, or human-review burden.
LLMOps implications and tradeoffs
The case demonstrates that LLMOps for agentic systems includes runtime policy, not just model hosting and prompt versioning. Replit’s policy surface includes delegation thresholds, specialist availability, tier selection, effort escalation, context reuse, cache lifetime, and validation roles. These decisions influence quality, token consumption, latency, and failure modes simultaneously. A trace-driven control loop is therefore necessary to evaluate whether an agent is making economically sensible decisions rather than merely producing plausible outputs.
The main benefit of model-directed delegation is adaptability. Stronger models may discover useful decompositions without requiring engineers to hard-code every workflow, and the system can avoid spawning workers for trivial requests. Reusing already briefed workers may improve context continuity, while parallel dispatch can reduce the need for one model to serialize independent investigations. Domain-specific workers can also constrain tools and responsibilities, making exploration, testing, reviewing, and implementation more separable.
The tradeoff is reduced predictability. Letting the model decide when and how to delegate can lead to unnecessary workers, missed validation, excessive reasoning, or inconsistent behavior across model versions. Replit explicitly notes that each frontier model delegates differently, which means a model upgrade can alter cost and execution patterns even without a harness change. The production system consequently needs regression evaluation not only for final task scores, but also for delegation rates, worker reuse, token and cache costs, latency, tool errors, and unsafe or out-of-scope actions.
The results also do not establish that less scaffolding is always better. Replit’s harness still supplies the specialist inventory, tools, tiers, effort controls, and guardrails. Its success may depend on those carefully designed primitives and on the capabilities of the selected frontier models. The practical lesson is more limited and more actionable: expose a bounded set of composable actions, measure their outcomes in production, and allow the core model to select among them when the model is demonstrably capable of doing so. This should be validated against fixed workflows and simpler single-agent baselines rather than assumed from the general claim that newer models will delegate well.
Overall assessment
Replit presents a credible production pattern for cost-aware multi-agent coding: a capable core model dynamically allocates work among specialists and adjusts reasoning effort as task difficulty changes. Its reported benchmark results favor the adaptive architecture over a single persistent sidekick and indicate a useful quality-cost frontier relative to Astra alone. The strongest evidence is the combination of production traces and benchmark comparisons, while the main limitations are vendor-controlled evaluation, reliance on published external baselines, a small number of repetitions, benchmark exclusions, and limited disclosure of operational reliability metrics. For teams adopting a similar design, the central LLMOps requirement is continuous, model-version-aware evaluation of both outcomes and orchestration behavior.