The CI bottleneck in the AI era: what Linear changed
Your agent writes a PR in 4 minutes and CI takes 25 to tell you if it's any good. The math doesn't work. Linear published a post about exactly this: once AI joined the team's coding workflow, the CI bottleneck showed up, and they had to rebuild the pipeline to keep pace.
The post is worth reading in full for the details. Here I want to talk about what it means for people who aren't Linear. You, me, and the five-person team with a monorepo that grew without asking permission.
What happened at Linear
The thesis is simple. For years, writing code was the slowest step in the cycle. CI ran while you grabbed a coffee, and nobody complained much.
Then coding agents arrived. PR volume went up. PRs got smaller and more frequent. Some devs open three or four branches in parallel, each with an agent working on it. And every push triggers the whole pipeline.
The result is predictable: the CI queue became the place where work waits. Linear looked at this and treated CI as a product, not as plumbing. They changed how tests are selected, how work is parallelized, and how much gets reused between runs.
Why AI turned CI into a bottleneck
Think of the cycle as a conveyor belt with three stages: write, validate, review. If one stage gets ten times faster, the slowest one sets the pace. That's Amdahl's law applied to your day.
Before AI, a dev opened maybe two or three PRs a day. Today, with an agent, that number can easily triple. If each run costs 20 minutes of runner time, you don't just get more waiting. You get more cost, more queueing, and more context switching.
And there's a detail few people mention: the agent depends on CI more than you do. You know you only changed some CSS. The agent doesn't have that instinct, or at least you shouldn't trust it to. It needs the green or red signal to know it's done. Slow CI makes the agent slow too.
When writing code gets cheap, verifying code becomes the expensive work.
What changes in the pipeline in practice
You can't copy Linear's infrastructure, but the principles work for any project. These are the ones I'd apply first, in order of effort.
Cancel runs nobody will read
If you pushed three times in two minutes, only the last result matters. In GitHub Actions this is one block of config, and plenty of projects still don't have it:
concurrency:
group: ci-${{ github.ref }}
cancel-in-progress: true
With an agent pushing after every tweak, this alone cuts a good chunk of the queue.
Run only what was affected
In a monorepo, running every test for every app on every PR is burning money. Turborepo and Nx already know how to figure out what changed from the dependency graph:
bunx turbo run test lint type-check --filter='...[origin/main]'
Only the marketing app changed? There's no reason to test the payments API. This is usually the biggest win of all, and the cost is having a well-declared dependency graph. If yours isn't, now is the time to fix it.
Parallelize with shards
A big suite you can't trim, you can split. Vitest and Playwright support sharding natively:
strategy:
matrix:
shard: [1, 2, 3, 4]
steps:
- run: bunx vitest run --shard=${{ matrix.shard }}/4
Four 5-minute runners give you an answer well before one 20-minute runner. Total minutes stay about the same, but time to feedback drops a lot. And time to feedback is what blocks the team.
Real caching
Everyone has dependency caching. What makes the difference is task result caching: if the @repo/utils package didn't change, yesterday's test result for it is still valid. Turborepo or Nx remote cache handles this across machines, including between your laptop and the runner.
What about quality? Doesn't it drop?
This is the objection I hear most. "If you run fewer tests, you'll let bugs through."
It depends on where you set the bar. What works is splitting it into two moments:
- On the PR: fast feedback, only what's affected, fail early. The goal is an answer in a few minutes.
- Before merge or on main: full suite, slow tests, heavy E2E. A merge queue handles this well, validating the combined result instead of each branch in isolation.
This way you don't trade safety for speed. You just change when you pay each cost. The PR stays fast for the human and the agent to iterate on, and main stays protected.
Another honest point: flaky tests got a lot more expensive. Before, an unstable test annoyed you once a day. With ten times more runs, it annoys you ten times more and also confuses the agent, which will try to "fix" code that had nothing wrong with it. I've seen an agent rewrite an entire function because of a random timeout in an E2E test. Quarantining flaky tests is no longer a luxury.
How to measure if your CI is already the bottleneck
Before you start optimizing, look at the numbers. Three questions are enough:
- On average, how long does it take from push to CI result? If it's over 10 minutes, you have a problem.
- How many runs per day get canceled or replaced by a newer push? If it's a lot, you're missing
cancel-in-progress. - What percentage of jobs run on code the PR didn't touch? In a monorepo with no filter, the answer is usually scary.
The GitHub API gives you all of this with no extra tooling:
gh run list --limit 100 --json conclusion,createdAt,updatedAt,headBranch
Drop that into a spreadsheet and you already have a better picture than many "we need to improve CI" meetings.
There's also the money side. On a free plan, whether it's GitHub Actions or Vercel deploys, the quota runs out fast when every push turns into five jobs and three preview deploys. I've had to add folder filters to my own projects because the agent pushed so much it ate the daily limit before lunch.
My take
What I liked most about Linear's post is that it doesn't treat AI as magic. It sped up one part of the work and exposed the rest. That's normal engineering: you fix one bottleneck and the next one shows up.
The part of this whole debate that bothers me is the claim that "AI will make the team ship 10x more". It'll write 10x more, maybe. Shipping depends on validating, reviewing, and getting to production, and those steps didn't speed up on their own. If your pipeline was designed for one human opening two PRs a day, it won't hold up under a team of agents. There's no point owning a Ferrari if your garage has a turnstile for a gate.
My recommendation is to start with the cheap stuff: cancel-in-progress, affected-only filtering, and flaky test quarantine. That solves most of the problem in an afternoon. Shards, remote cache, and merge queues come later, when the numbers call for them.
If you want to see how I set up pipelines in real monorepos, take a look at my projects.
LinkedIn summary
My agent writes a PR in 4 minutes. CI takes 25 to tell me if it's any good. Linear shared that once AI joined the team's workflow, the bottleneck stopped being writing code. It became validating it. When writing code gets cheap, verifying it becomes the expensive work. And slow CI makes the agent slow too. What I'd do first, in one afternoon: cancel-in-progress, run only what was affected, and quarantine flaky tests. AI might make the team write 10x more. Shipping 10x more depends on the pipeline keeping up. I wrote on the blog about how to apply this to any monorepo. Link in the comments. #DevOps #CICD #AI #Monorepo #SoftwareEngineering