Stop Over-Orchestrating AI Agents: Simplicity Wins
The Framework You Built May Be the Problem
In 2025, the default advice for anyone building AI agents was: pick an orchestration framework, wire everything through a central coordinator, and let the middleware handle complexity. We followed that advice. Three months into a build, we had a beautifully layered system where every agent request traveled through a routing layer, a context manager, a memory broker, and a dispatch queue before reaching the model that would actually do the work. Latency was high. Token bills were higher. Debugging felt like reading a stack trace through frosted glass.
A debate has been building in developer communities on Hacker News and in engineering Slack groups throughout 2025 and into 2026: are orchestration frameworks solving real problems, or are they a category of over-engineering that the industry adopted before it had enough production experience to know better? The question is worth taking seriously. According to Gartner's State of AI Agents report, organizations are increasingly recognizing that overly complex AI agent architectures lead to higher operational costs and reduced performance, with simpler, more focused agent designs showing better ROI in production environments. That finding matches what we saw firsthand.
What Orchestrators Actually Do to Your Token Budget
The core promise of an orchestration layer is coordination: one system that routes tasks, manages state, and ensures agents don't step on each other. The cost of that promise is abstraction. Every abstraction layer in an AI pipeline adds tokens.
Here is the mechanism. A user sends a request. The orchestrator receives it, formats it into a routing prompt, sends that prompt to a reasoning model to determine which sub-agent should handle the task, receives the routing decision, formats a new prompt for the target agent, sends that prompt, receives the response, formats a synthesis prompt, and finally returns an answer. Each formatting step injects system context, role definitions, and state summaries that the model needs to orient itself. None of that context is free. In a direct-call pattern, the same request goes to one model with one system prompt. The work gets done. The conversation ends.
The token multiplication is not theoretical. It compounds with every hop. A three-agent pipeline with a central orchestrator can easily triple the token consumption of a direct implementation handling the same task. At low volume, that difference is invisible. In production, it becomes a line item that someone has to explain to finance.
Latency follows the same pattern. Each orchestration hop is a synchronous API call. If each call takes 800 milliseconds, a four-hop pipeline adds over three seconds of pure coordination overhead before any real work begins. Users notice three seconds.
The Case for Orchestrators: Where They Actually Earn Their Keep
Fairness requires acknowledging what orchestrators do well. They shine in genuinely parallel workloads where multiple independent agents need to run simultaneously and their outputs need to be merged. They also help when you need a single audit trail across a complex multi-step process, or when different agents require different tool permissions and you want one place to enforce access control.
If you are building a system where ten agents are simultaneously researching, writing, fact-checking, and formatting a document, a coordinator that manages their outputs is doing real work. The coordination cost is justified because the alternative, managing that concurrency yourself in application code, is worse.
The problem is that most teams reach for orchestration frameworks before they have workloads that require them. They build the coordination infrastructure first, then fill it with tasks that a single well-prompted model could handle directly. The framework becomes the architecture, and the architecture becomes the constraint.
Direct Agent Patterns: What They Look Like in Practice
A direct agent pattern is not a primitive or a shortcut. It is a deliberate choice to keep the call graph flat. One model receives a well-constructed prompt, uses tools if it needs them, and returns a result. The calling application handles routing logic in code, not in a prompt sent to another model.
This approach has three concrete advantages. First, debugging is straightforward: you have one input, one output, and a clear log of every tool call in between. When something breaks, you know exactly where to look. Second, iteration is faster because you are editing a prompt and a tool list, not reconfiguring a multi-agent topology. Third, the cost per task is predictable because you are not paying for coordination tokens that vary based on how the orchestrator interprets the routing task.
We learned a version of this lesson the hard way during our first Stripe product creation. The API call included a recurring parameter set to null. We thought omitting the value was the same as omitting the field. It wasn't. Stripe created two prices: one correct one-time payment at $297, and one spurious monthly subscription at $297. We caught it before a customer was charged monthly for a one-time product, but it took a manual archive in the Stripe Dashboard to fix. Now our factory pipeline never includes the recurring field at all, not null, not false, just absent. The lesson applies directly to agent architecture: the absence of a thing is not the same as setting it to zero. Removing an orchestration layer entirely is different from building one and configuring it to be lightweight. Absence is cleaner.
For more on where direct patterns break down in production, our post on why 24/7 AI agents fail and what actually works covers the failure modes we've seen most often.
Orchestration vs. Direct Calls: A Practical Decision Matrix
The choice between these patterns is not ideological. It is a function of your actual workload. Here is how we think about it.
Use a direct agent pattern when: the task has a single clear owner, the output of one step is the input of the next in a linear chain, you need fast iteration cycles, or your token budget is a real constraint. Most customer-facing automations, content generation pipelines, and data extraction tasks fall here.
Use an orchestration layer when: you have genuinely parallel workloads where agents must run concurrently, you need centralized access control across agents with different permission sets, or you are building a system where the coordination logic itself is complex enough that encoding it in application code would be harder to maintain than a dedicated coordinator. Research pipelines, multi-source data aggregation, and what ForgeWorkflows calls agentic logic, where the system must reason about its own next action, are legitimate candidates.
The honest version of this matrix is that most teams need orchestration for fewer tasks than they think. Start with the direct pattern. Add coordination infrastructure only when you hit a specific problem that the direct pattern cannot solve. Do not build the coordination layer speculatively.
Why the Industry Defaulted to Complexity
The over-engineering tendency has a clear origin. The first wave of AI agent frameworks arrived before anyone had meaningful production data on what these systems actually needed. Framework authors, many of them coming from distributed systems backgrounds, applied patterns that work well for microservices: centralized coordination, message queues, service registries. Those patterns solve real problems in distributed computing. They do not automatically translate to AI agent pipelines, where the bottleneck is model latency and token cost rather than network throughput or service discovery.
The Gartner finding cited above reflects a correction that is now underway. Organizations that shipped complex orchestration architectures in 2023 and 2024 are measuring their production costs and finding that simpler designs outperform them. This is the normal arc of a new technology category: early adoption favors complexity because complexity signals sophistication, and production experience corrects toward simplicity because simplicity is cheaper to operate and easier to fix.
The developer community debate happening right now on HN and in engineering forums is that correction happening in public. It is worth paying attention to, not because the contrarian position is always right, but because the people making the argument are the ones who built the complex systems and are now living with the results.
If you want a broader view of how experienced engineers are approaching these tradeoffs in 2026, our post on how top engineers actually build AI stacks covers the patterns we see repeated across teams that ship reliably.
What We'd Do Differently
We'd instrument token consumption per hop before committing to any architecture. The total token cost of a multi-agent pipeline is not obvious from the design diagram. We would add logging at every model call from day one, measure the coordination overhead as a percentage of total tokens, and use that number to justify or reject the orchestration layer. If coordination tokens exceed 20% of total consumption, the architecture needs a second look.
We'd treat orchestration as a late-stage addition, not a foundation. The instinct to build the coordination infrastructure first is understandable but expensive. Starting with direct patterns and adding coordination only when a specific production problem demands it would have saved us significant rework. The framework should emerge from the requirements, not precede them.
We'd evaluate what ForgeWorkflows calls a modular swarm pattern earlier. Rather than a single central orchestrator managing all agents, a loosely coupled set of specialized pipelines that hand off to each other through simple API calls preserves most of the coordination benefit while keeping each component independently debuggable. We came to this architecture late. It should have been the starting point for any system where more than two agents needed to interact.