Dev.to AI 🤖 Ai 👁 0 📖 7 min read

Sequential pipelines are killing your agent throughput, here's the concurrent alternative that 3x'd ours

I watched a user abandon a research report at second 41. The pipeline was still running — 45 seconds of sequential agent work, every step dependent on the last, every agent waiting its turn. By the time the output landed

I watched a user abandon a research report at second 41. The pipeline was still running — 45 seconds of sequential agent work, every step dependent on the last, every agent waiting its turn. By the time the output landed, the tab was closed. That single abandoned session cost us more in trust than the entire GPU bill for the week.

That was the day I stopped treating sequential orchestration as a reasonable default and started treating it as a latency tax I'd been paying without noticing.

The Tax I Couldn't See

The pipeline looked reasonable on the whiteboard. A planner decomposed the query into sub-questions. A retriever fetched sources for each. An analyzer produced findings. A verifier checked claims. A synthesizer merged everything into a final answer. Five stages, each one logically dependent on the last.

In production, it was a serial queue of LLM calls. The retriever couldn't start until the planner finished. The analyzer couldn't start until every retrieval completed. The verifier waited on the analyzer. The synthesizer waited on everything. Every agent was idle while the previous one finished work that had no actual dependency on it.

Microsoft's own orchestration documentation states the arithmetic plainly. If market analysis takes 12 seconds and risk assessment takes 10 seconds, sequential execution takes 22 seconds. Parallel execution completes in 12 — the maximum of the two, not the sum. For workflows with independent tasks, parallel execution cuts latency proportionally to the number of concurrent agents. Four agents running in parallel complete in roughly the time of the slowest agent, rather than the sum of all agents.

I had four agents that could have run concurrently. I was paying their sum instead of their maximum.

What the Research Already Knew

The 2026 benchmark literature had already quantified what I was feeling. A NYU study compared four orchestration architectures — sequential pipeline, parallel fan-out with merge, hierarchical supervisor-worker, and reflexive self-correcting loop — across 10,000 SEC filings and five frontier models. The finding that stopped me cold: parallel fan-out with merge is latency-optimal because latency is dominated by the slowest parallel branch plus merge overhead, not the sum of all stages.

The financial document benchmark confirmed it with a second data point: sequential pipelines are the cheapest and most stable at large scale, especially above 100,000 documents per day — but parallel fan-out wins decisively on latency at higher token cost. The architecture is a choice, not a verdict. I had chosen the wrong trade-off for a user-facing product where responsiveness determines whether anyone ever sees the output.

The LangGraph parallel execution POC made the math visceral. Four nodes in a serial graph: 4.97 seconds. The same four nodes with parallel edges: 2.29 seconds. A 53.9% reduction in wall-clock time, achieved by changing three lines of graph wiring and nothing else.

I had been tuning prompts for weeks. The fix was the topology.

The Fan-Out That Fixed It

The pattern is called fan-out and fan-in, and it's simpler than the name suggests. Fan-out spawns multiple agent tasks simultaneously. Fan-in waits for all (or some threshold of) agents to complete before proceeding.

I rebuilt the pipeline around three changes.

The planner still runs first — because it genuinely must. Decomposition is a true sequential dependency. You cannot retrieve without knowing what to retrieve. But the planner produces a list of independent sub-questions, and that list is where the serial chain ends.

Retrieval, analysis, and verification fan out per sub-question. Instead of one retrieval step that processes five sub-questions serially, five retrieval agents run concurrently. Instead of one analysis step, five analyses run concurrently. The wall-clock time for each stage drops from the sum of all sub-questions to the slowest single sub-question.

The synthesizer runs after the fan-in. It's the only stage that genuinely needs everything upstream. It waits for all parallel branches to complete, reconciles their outputs, and produces the final answer.

The architectural shift is that the dependency graph is explicit, not implied by code order. Three of my five stages were never actually dependent on each other. The chain had been a habit, not a requirement.

The Numbers

Before the rebuild: 45 seconds end-to-end. After: 14 seconds. A 3.2× reduction.

The math holds up. If the sequential pipeline spent roughly 12 seconds on retrieval (five sub-questions × 2.4 seconds each), 15 seconds on analysis, and 8 seconds on verification, the parallel version runs each stage in the time of its slowest branch. Retrieval drops to ~3 seconds. Analysis to ~4. Verification to ~2. The synthesizer, which genuinely needs everything, stays at ~5 seconds. Total: ~14 seconds.

The Amdahl's Law bound is real and it matters. A 2026 study on LLM teams as distributed systems confirmed that speedup remains significantly below the theoretical Amdahl bound even in highly parallel conditions, because the serial fraction — planning, synthesis, merge overhead — is irreducible. My 3.2× wasn't a failure of the pattern. It was the pattern working as designed: the serial fraction bounded the gain, and the parallel fraction delivered the rest.

What Production Teams Are Running

Microsoft's Contoso Capital reference architecture runs risk assessment and market analysis agents in parallel, reducing report generation from 90 seconds to 25 seconds — a 3.6× improvement — because the agents operate on independent data sources with no shared state.

The LangGraph multi-agent supervisor benchmark, reproducible in six seconds with no API key, shows a 1.99× speedup on deterministic latency and 1.31× on live HotpotQA with real TTFT jitter. The 1.99× is the theoretical ceiling: four sequential stages of 0.5 seconds each, collapsed to the slowest single stage. The live run shows what real-world variance does — per-question speedup ranges from 1.05× to 1.45× because each layer's slowest specialist becomes the bottleneck.

iFood's Rosie support agent parallelizes independent workflows and reduced average resolution time from 38 minutes to 10, with 70% of cases fully automated. The parallelization isn't the only factor, but it's the architectural enabler that made the latency reduction possible at scale.

Scalable Inference Architectures for Compound AI Systems, a production deployment study, reports over 50% reduction in tail latency (P95) and up to 3.9× throughput improvement compared to prior static deployments, with specific analysis of multi-model fan-out overhead and cascading cold-start propagation.

The Guardrails That Made It Survivable

Fan-out isn't free. It introduces failure modes that a sequential chain never had, and I hit every one of them in the first week.

The barrier bottleneck. If one branch takes 10× longer than the others, the entire fan-in waits for it. The GitHub benchmark's live HotpotQA results show this precisely: per-question speedup ranges from 1.05× to 1.45× because the slowest specialist becomes the bottleneck. The fix is per-branch timeouts with fallback values, so a slow branch degrades to a partial result rather than blocking the entire merge.

Fan-in reconciliation. When five analyses return concurrently, something has to decide what "merged" means. I hit the exact trap described in the concurrency research: two branches returned conflicting numbers for the same metric, and the synthesis agent quietly picked one. The fix is structured claims with provenance — each branch emits a schema-tagged result with source and timestamp, and the fan-in refuses to blend conflicting values on the same invariant key. Halt and surface the diff. A loud failure is cheaper than a plausible hallucination.

Durable fan-out, not Promise.all. The naive implementation uses Promise.all — spawn everything, await everything. It collapses latency, but a crash mid-batch re-runs everything on recovery, and a retry of one branch blasts the same side effect twice. Resonate's durable fan-out makes every spawn and every await a durable promise. If one branch fails and retries, the others stay checkpointed and don't re-execute. Crash recovery only re-runs whatever was in flight when the worker died.

Read-only branches. Fan-out branches that write to shared state introduce race conditions that look like logic bugs. One branch reads a value, another branch writes it, the first branch commits a stale result, and nothing in the trace shows the race. Scoping each branch to its own state slice and merging explicitly at the fan-in eliminates the entire class of bug.

When Not to Fan Out

The financial document benchmark's finding is the honest counterweight: sequential pipelines are the cheapest and most stable at large scale, especially above 100,000 documents per day. Fan-out adds token cost because every branch carries its own context, and merge overhead grows with the number of branches.

Stay sequential when your task has genuine serial dependencies. If stage B needs the complete output of stage A, don't parallelize what isn't parallel.

Stay sequential when latency isn't the binding constraint. A nightly batch job that finishes in six minutes doesn't need fan-out. A user-facing report that takes 45 seconds does.

Stay sequential when your branches would write to shared state. Fan-out on write-heavy workloads is a concurrency bug waiting to happen. Read-only fan-out is where the pattern pays.

Stay sequential when your merge logic is more complex than your computation. If reconciling five branch outputs takes longer than producing them, you've moved the bottleneck rather than removed it.

What I'd Tell My Past Self

The sequential pipeline wasn't wrong because I used the wrong framework. It was wrong because I encoded dependencies that didn't exist. A chain says "B cannot begin until A completes." That's a claim about the world. When it's true, a chain is correct and efficient. When it's false, you've serialized work that should have been concurrent, and you'll never notice because the system still produces correct output — just slowly.

The teams that are getting this right aren't choosing between sequential and parallel as an ideology. They're drawing the actual dependency graph instead of the one they assumed, and parallelizing only the stages that measurement proves are independent and expensive.

So here's my question: If you drew the real dependency graph for your agent pipeline, how many of those sequential steps would survive?

I'd love to hear where you've landed. Sequential by choice, fan-out after a painful migration, or a hybrid you discovered the hard way — and what finally made you look at the graph instead of the chain?

📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.