EngineeringA CI Refactor of Waiting PRs, Duplicated Config, and Coverage Gates
· AlgoSu
- #ci-cd
- #github-actions
- #ai-dev
- #refactoring
From Reference to Practice
The first time I read Channel.io's post “Backend CI Refactoring” (Channel.io is a Korean SaaS company), my immediate reaction was: "I could do this for AlgoSu." Their story of cutting a 36.6-minute CI run down to 15 minutes and 38 seconds. The method was concrete, the principles clear.
But I couldn't follow it directly. AlgoSu is different from Channel.io. It's one developer directing a team of AI agents, every build runs inside Docker, and a lead agent hands tasks out to the other agents. I could borrow the underlying principles, but I still had to translate those principles into AlgoSu's context.
This isn't a post introducing a finished CI architecture. AlgoSu has run more than a hundred sprints (short work cycles), and this is a record of five of them, S102–S106, spent reading a reference, experimenting, and sometimes dropping the plan. The focus is on two pieces of work that ended with zero lines of implementation, and why those weren't failures but wins.
Problem
There were plenty of great CI references, but none fit a pipeline run by one developer and AI agents as-is. 30 pending PRs, repeated setup steps, and a coverage check that blocked nothing had piled up.
Decision
Take only the principles from the reference and apply them myself over a roadmap planned as four sprints (one more sprint was added for the leftovers), from auto-merge to setting measurement rules.
Result
Pending PRs now merge automatically, a composite action (a reusable bundle of steps) removed 67 duplicated lines, and the coverage gate sits at 70%. Two of the last sprint's three tasks closed with zero lines after a pre-check confirmed they were already handled.
Why I Needed a Reference
An honest snapshot of where things stood at the start:
Before: Three Problems
- Dependabot, the GitHub bot that opens dependency-update PRs, kept adding PRs every week that had to be merged by hand: 30 open PRs
- The Node setup steps (
setup-node + cache + npm ci) were copied 3–4 times across job × service combinations - Test coverage was measured, but falling below the threshold didn't block PRs (measured but not gated)
30
Pending Dependabot PRs
Generated weekly, piled up as manual squash-merge burden
3–4×
Duplicate Setup Steps
setup-node + cache + npm ci repeated across each job × service
None
Coverage Gate
Measured but didn't block PR merges
These three problems looked independent, but they shared a common cause: the momentum of "good enough if it works right now." If CI passed, we deployed. Cleaning up repetitive code always got pushed to later.
Three principles I borrowed from the Channel.io post:
01
Small pilot → expand
Don't apply to all services at once. Validate on the simplest service first, then roll out.
02
Workflow + repository settings go together
A GitHub Actions file alone is half the picture. Repository settings have to back it up.
03
Removing duplication is a result, not a goal
You don't build shared modules just to remove duplication; a good abstraction removes it naturally.
There were also principles I deliberately didn't borrow. Two of Channel.io's techniques (uploading prepared setup to S3 and polling it so initialization overlaps, and distributing tests through a dynamic queue) would have been overkill for AlgoSu. GitHub Actions (GHA) artifacts were enough, and with so few tests, the cost of splitting them up outweighed the gain. Choosing what not to carry over is part of translating a reference.
One more AlgoSu-specific wrinkle: because the lead agent hands out the work, I could apply "pilot then expand" sprint by sprint, but I also had to re-check which agent should own the work before each sprint started. In the first sprint I wrongly gave the CI work to the auth & security agent, then moved it to the infrastructure agent. When you work with many agents, checking role boundaries matters as much as writing the code.
The 4-Sprint Roadmap
The roadmap I drew up after reading the Channel.io post covered four sprints (S102–S105); in practice, S106 was added for leftover items. Here's what each sprint aimed for and what it actually delivered:
S102 · Operations Automation
Dependabot grouping + Auto-merge + Branch Protection (merge rules for main)
S103 · Composite Pilot
setup-node-service action + github-worker pilot + Coverage Gate 60%
S104 · Rollout
Composite expanded to all Node services (67 lines deleted) + AI Coverage integration
S105 · Setting Measurement Rules
rebuild_all (force-rebuild every service) runbook + github-worker benchmarks + dynamic commitlint scope
S106 · Deferred Items
Coverage 70% achieved (real implementation) + L2 cache & build optimization stopped early by pre-checks
Each sprint stood on its own, yet each one's lessons fed the next design. The finding that "a PR that changes CI can't measure its own effect" became the usage rules for rebuild_all, an option that forces every service to rebuild. And the habit of a pre-check (asking another agent whether something is needed before building it) made the final sprint's zero-line conclusion possible.
S102: Auto-merge, Branch Rules
The first problem was 30 open Dependabot PRs. Reviewing and merging weekly patch/minor updates by hand had piled up into a backlog.
The solution had two steps: group dependabot.yml to reduce individual PR count, and an auto-merge workflow that merges patch/minor PRs automatically once checks pass.
Dependabot Grouping
# .github/dependabot.yml (excerpt)
updates:
- package-ecosystem: npm
directory: /services/gateway
schedule:
interval: weekly
groups:
gateway-minor-patch:
update-types:
- minor
- patchI added {service}-minor-patch groups across all 8 ecosystems (5 Node services + Frontend + Blog + Python). Two Docker image updates were excluded from groups due to potential security impact.
Auto-merge Workflow
# .github/workflows/dependabot-automerge.yml (key excerpt)
on:
pull_request_target:
types: [opened, synchronize, reopened, ready_for_review]
permissions:
contents: write
pull-requests: write
jobs:
auto-merge:
if: github.actor == 'dependabot[bot]'
runs-on: ubuntu-latest
steps:
- name: Fetch Dependabot metadata
id: metadata
uses: dependabot/fetch-metadata@v2
- name: Auto-merge patch and minor
if: |
steps.metadata.outputs.update-type == 'version-update:semver-patch' ||
steps.metadata.outputs.update-type == 'version-update:semver-minor'
run: gh pr merge --auto --squash "$PR_URL"
env:
PR_URL: ${{ github.event.pull_request.html_url }}
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
- name: Skip major updates
if: steps.metadata.outputs.update-type == 'version-update:semver-major'
run: |
echo "Major update detected — skipping auto-merge."
exit 0One important decision here: I chose pull_request_target instead of pull_request as the trigger. Dependabot PRs carry fork-like context and need pull_request_target to access secrets. To prevent code injection, I removed actions/checkout entirely. No step executing PR code means no injection path.
Three layers of defense: job-level if: github.actor == 'dependabot[bot]', step-level metadata type check, and complete removal of actions/checkout.
Repository Settings to Match
After creating the workflow, auto-merge still didn't work. The reason turned out to be repository settings.
GitHub's auto-merge requires two preconditions:
- Repository setting
allow_auto_merge: true - Branch Protection's required status checks must pass
The lead agent applied the settings directly via the gh API. Branch Protection is the GitHub setting that defines what must pass before anything merges into main.
# Repository settings (applied via gh api)
allow_auto_merge: true
delete_branch_on_merge: true
# main Branch Protection
strict: true
required_checks: ["Secret & Env Scan", "Detect Changed Services"]
allow_force_pushes: false
allow_deletions: false
required_conversation_resolution: trueRequired checks were kept minimal for a reason. Jobs like quality-nestjs and test-node get skipped when no relevant services changed. Registering a skipped job as a required check means that PR can never merge when those jobs don't run. Only the always-running Secret & Env Scan and Detect Changed Services were registered.
Result: Dependabot pending PRs 30 → 2. 28 were reorganized into 7 group PRs and queued for auto-merge. An unexpected bonus: Dependabot detected the grouping, automatically closed the existing individual PRs, and recreated them as group PRs. No manual closing needed.
Verification of the work PR:
| Check | Result |
|---|---|
| Work PR, full CI | ✅ 26 success / 10 skipped / 0 failure (counting jobs expanded by the matrix) |
| Auto-merge workflow | ✅ 7/7 success (7 group PRs) |
| Actual auto-merge confirmed | ✅ github-worker group PR (3 updates), merged by app/github-actions |
| Branch Protection in effect | ✅ Direct push to main blocked, strict mode requires up-to-date base |
S103–104: Pilot, Then Rollout
The second problem was the repeated setup-node + cache + npm ci pattern. The fix was a composite action, a GitHub Actions feature that bundles several steps into one reusable step. Three jobs, quality-nestjs, audit-npm, and test-node, each ran the same setup steps across a 5-service matrix.
I applied the Channel.io principle of "validate on the smallest service first, then expand" at the sprint level. S103: pilot on github-worker only. S104: roll out to every service.
S103: Composite Action Pilot
I chose github-worker as the pilot service for a clear reason: of the five Node services, it has the simplest structure (pure Node.js, no NestJS), so side effects from pattern changes could be verified in the smallest possible blast radius.
# .github/actions/setup-node-service/action.yml
name: 'Setup Node Service'
description: 'setup-node + lockfile cache + conditional install for a single service'
inputs:
service-path:
description: 'Relative path to the service directory'
required: true
node-version:
description: 'Node.js version'
default: '20'
install-command:
description: 'Install command to run'
default: 'npm ci'
runs:
using: composite
steps:
- name: Setup Node.js
uses: actions/setup-node@v6
with:
node-version: ${{ inputs.node-version }}
- name: Cache node_modules
id: cache
uses: actions/cache@v5
with:
path: ${{ inputs.service-path }}/node_modules
key: node-${{ inputs.node-version }}-${{ hashFiles(format('{0}/package-lock.json', inputs.service-path)) }}
restore-keys: |
node-${{ inputs.node-version }}-
- name: Install dependencies
if: steps.cache.outputs.cache-hit != 'true'
shell: bash
run: ${{ inputs.install-command }}
working-directory: ${{ inputs.service-path }}One key design decision: actions/checkout was deliberately excluded from the composite. Checkout is the same first step in every job and doesn't vary by service path. The composite extracts only what's service-specific: setup-node + cache + install.
This kept the composite general-purpose. Each job can freely decide how to checkout (e.g., sparse-checkout), and the composite just uses whatever's already there. For the audit-npm job, I passed install-command: 'npm ci --ignore-scripts' to preserve the security scan policy.
S103: Coverage Gate
Coverage was already being measured, but falling below the threshold didn't block PRs. I wrote scripts/check-coverage.mjs to close that gap.
// scripts/check-coverage.mjs (core logic excerpt)
import { readdirSync, readFileSync, existsSync } from 'fs';
import { join } from 'path';
function parseLcov(content) {
let lh = 0, lf = 0, brh = 0, brf = 0;
for (const line of content.split('\n')) {
if (line.startsWith('LH:')) lh += parseInt(line.slice(3));
else if (line.startsWith('LF:')) lf += parseInt(line.slice(3));
else if (line.startsWith('BRH:')) brh += parseInt(line.slice(4));
else if (line.startsWith('BRF:')) brf += parseInt(line.slice(4));
}
return { lh, lf, brh, brf };
}
// Guard for missing coverage directory
if (!existsSync(coverageDir)) {
process.stdout.write('No coverage artifacts found. Skipping gate.\n');
process.exit(0);
}
// Recursive scan — no script changes needed when adding new services
function findLcovFiles(dir) { /* ... */ }Written with zero external npm dependencies. No supply chain risk, and simple logic: recursively scan all lcov.info files → sum the "covered / total" counts for lines (LH/LF) and branches (BRH/BRF) → validate lines AND branches simultaneously → exit 1 if below threshold.
Started with a global 60% threshold. Since each service's line-coverage thresholds (Node 97–98%, Python 98%, Frontend 83%) already far exceed 60%, the global gate was designed as a floor guard for newly added services.
S104: Full Rollout
The rollout was straightforward. Remove the matrix.service != 'github-worker' condition and route all services through the composite.
# ci.yml (after rollout — this pattern applied to quality-nestjs / audit-npm / test-node)
- name: Setup Node service
uses: ./.github/actions/setup-node-service
with:
service-path: services/${{ matrix.service }}
# For the audit-npm job:
# install-command: 'npm ci --ignore-scripts'Removing the 3-step inline (Setup Node + Cache + Install) × 3 jobs = 67 lines deleted from ci.yml, roughly a 25% reduction. A maintainability win you feel before you even measure performance.
S105: Measurement Automation
S105 was the closing sprint of the four-sprint roadmap. Three tasks bundled together.
[A] Operating Rules for rebuild_all
The fix for that pitfall (a PR that changes CI can't measure itself) was already in ci.yml: workflow_dispatch.inputs.rebuild_all=true, an option on manual runs that forces every service to rebuild. It had existed since the pilot, but there were no rules for when, by whom, and how to use it.
The fix required zero new code, just documenting the operating rules. I created docs/runbook/ci-rebuild-all.md and added a checkbox to .github/pull_request_template.md.
Three trigger conditions were formalized:
- PRs that only change
.github/workflows/*.yml - PRs that change
.github/actions/**composites - PRs that change CI utility scripts like
scripts/check-coverage.mjs
A runbook only matters if it's rehearsed right after it's written. The [A] runbook was actually used in the [B] benchmark run within two hours of being merged.
[B] github-worker Benchmark: A Pre-Check Yields N=1
First came the question of how to run the benchmark. The original plan called for 5 after-change samples. Before running anything, I asked another agent to review the statistics of this plan. This is the step this post calls a pre-check. From here on, that agent (the analysis agent) handled every pre-check.
MDE in the answer below is the minimum detectable effect: the smallest time difference these samples could reliably call a change. The analysis agent's key finding: "Under the Welch-Satterthwaite formula, Pre n=4 locks degrees of freedom (df) at 4. Increasing Post n from 2 to 6 only improves MDE by 0.8s. The original N=5 plan is overengineering. N=1 is sufficient." In plain terms: with only four before-change samples, adding more after-change samples barely improves what the test can detect.
That stopped wasted CI time (runner-minutes) before it happened. I added a dummy comment in one PR so detect-changes would pick up github-worker, and combined the 2 resulting runs with 1 run forced by rebuild_all=true, for post n=3.
Results:
| Job | Before (n=4) | After (n=3) | Delta |
|---|---|---|---|
Quality — github-worker | 22.2s (σ 5.8s) | 22.3s (σ 2.5s) | +0.1s (+0.4%) |
Audit — github-worker | 19.8s (σ 3.7s) | 18.0s (σ 3.0s) | −1.8s (−8.9%) |
| Test GitHub Worker | 19.2s (σ 1.9s) | 20.0s (σ 1.0s) | +0.8s (+3.9%) |
All three jobs within the ±10% practical threshold. Welch t-test: |t_obs| < 0.7 (t_crit=2.776). Composite action rollout had no statistically detectable effect on github-worker per-job runtimes.
Thanks to the pre-check, I reached a conclusion with fewer runs than planned. For the first time, I felt that "approve plan → execute immediately" isn't always the optimal path.
[C] Dynamic commitlint scope-enum
CI had once failed on a scope error from commitlint, the tool that checks commit-message format. ci(actions) and ci(coverage) look like obvious scopes, but neither was in the allowed list (scope-enum) in commitlint.config.mjs, so the error only surfaced in the PR CI run. Without a local pre-commit hook, scope errors only show up after you push.
Two problems solved at once: move validation earlier with husky, a git-hook tool, eliminate manual maintenance with dynamic scope-enum generation.
I added husky at the root package.json and set up a commit-msg hook.
// package.json (root — new file)
{
"devDependencies": {
"@commitlint/cli": "^19.0.0",
"@commitlint/config-conventional": "^19.0.0",
"husky": "^9.0.0"
},
"scripts": {
"prepare": "husky"
}
}# .husky/commit-msg
npx --no -- commitlint --edit "$1"Adding a root package.json doesn't affect existing CI jobs. Each job either specifies working-directory to a service directory or goes through a composite action. Every CI job in that PR came back SUCCESS.
// commitlint.config.mjs
import { readdirSync } from 'fs';
// Auto-scan services/ — new services register automatically
const dynamicScopes = readdirSync('./services', { withFileTypes: true })
.filter((d) => d.isDirectory())
.map((d) => d.name);
const staticScopes = [
'ci', 'docs', 'blog', 'frontend', 'infra',
'deps', 'security', 'adr', 'e2e', 'runbook',
];
export default {
extends: ['@commitlint/config-conventional'],
rules: {
'scope-enum': [
2,
'always',
[...new Set([...dynamicScopes, ...staticScopes])].sort(),
],
},
};Now adding a directory under services/ automatically registers the scope. The recurring "remember to update scope-enum when adding a new directory" feedback loop was resolved structurally. Human-dependent feedback promoted to system automation.
S106: The Zero-Line Decision
The last sprint handled the three items carried over until then. The table below calls each task a track. Ultimately, one track was actually implemented, while two were closed without any code changes.
| Track | Task | Result |
|---|---|---|
| [A] Coverage 70% | Frontend branches 69.55% → 71%+ + global gate 70% | ✅ Implemented (2 PRs) |
| [B] L2 Cache Layer | Target: 40% Docker build reduction | ❌ Not introduced (stopped at pre-check) |
| [C] Frontend Build Optimization | swcMinify · optimizePackageImports · sourceMaps | ❌ Not introduced (stopped at pre-check) |
[A] Coverage 70%: Frontend Branches Was the Single Bottleneck
To raise the global coverage gate from 60% to 70%, I first had to find the bottleneck. The weighted aggregate branches across all services was already around 82%, well above 70%. So why couldn't the gate be raised?
The pre-check revealed the structural reason. check-coverage.mjs only aggregates lcov.info for services detected by the path-filter. For frontend-only PRs, only frontend branches are aggregated. And frontend branches sat at 69.55%.
The intuition that "all-service aggregate branches = 82%, global gate 70% → passes" was wrong. A frontend-only PR would aggregate at 69.55% and fail the 70% gate. In a path-filter-based CI design, the coverage gate needs to be explicit about which lcov set it operates on per PR scope.
Fixing the bottleneck meant writing new tests to bring frontend branches up to 71%.
The analysis agent's gap analysis:
| Target | Branches Hit Needed | Current | Gap |
|---|---|---|---|
| 71.0% (this sprint’s target) | 1,330 | 1,302 | +28 |
| 72.0% (safety buffer) | 1,348 | 1,302 | +46 |
Achievable with ~120–190 LOC of new tests across three recommended files (lib/feedback.ts, components/ui/CodeBlock.tsx, components/providers/EventTracker.tsx).
Actual result: 77 tests added in the test PR, frontend branches 69.55% → 76.42% (+6.87pp, exceeding the 71% target by +5.42pp). A conservatively designed scenario.
69.55%
Frontend Branches (Before)
1302 / 1872 branches hit
76.42%
Frontend Branches (After)
+5.42pp beyond the 71% target
A separate gate PR then changed the ci.yml coverage threshold from 60 → 70. Order mattered here: the test PR had to pass CI and merge first, and only then the gate PR. Merging the gate PR first would immediately make any frontend-only PR fail the coverage gate. I wrote the analysis agent's warning into the gate PR's description to enforce the order.
[B] L2 Cache: Docker Buildkit Was Already L2
The plan was an extra cache layer (an "L2 cache") that would store build output in the GHA cache a second time. Again, I asked the analysis agent before building.
What it found was striking. docker/build-push-action already had --cache-from=type=gha,mode=max set, and mode=max stores all intermediate layers from builder stages in GHA cache.
# NestJS build layers — mode=max cache coverage
Layer 1: FROM node:22-alpine AS builder → cached
Layer 2: COPY package*.json ./ → HIT when package.json unchanged
Layer 3: RUN npm ci → HIT when package.json unchanged
Layer 4: COPY . . → MISS on source changes
Layer 5: RUN npm run build ← generates dist/ → re-runs on Layer 4 MISSRUN npm run build, the dist/ generation step, was already being saved as a GHA cache layer. Adding an external GHA filesystem cache would have been 100% redundant.
An additional finding: ci.yml already had a Frontend .next/cache GHA cache step, and it was dead code that never did anything. The build-frontend job does a Docker build only, with no host-side npm run build, so .next/cache is never generated on the host filesystem. Restoring it restores nothing; saving it saves an empty directory. A leftover artifact from before the Docker-only pipeline migration.
The deeper finding: not a single build job ran npm run build on the host filesystem. All TypeScript/Next.js compilation happens inside Docker containers. GHA filesystem caching only applies to the host filesystem, so the current Docker-only architecture has no path to benefit from it at all.
Conclusion: L2 cache not introduced. Zero code changes. The dead step was removed in a separate PR.
[C] Frontend Optimization: All Three Were Already Handled
I'd planned to apply three Next.js optimization options. [C] also included recording how long the build step took, but every build ran only inside Docker, so there was nothing separate to time. The analysis agent checked each one directly in node_modules/:
| Option | Plan | Verified Result | Verdict |
|---|---|---|---|
swcMinify: true | Explicit setting | Completely removed from Next.js 15.5.15 config-schema.js (z.strictObject violation) | HARD BLOCK |
optimizePackageImports | Add @radix-ui/react-*, lucide-react | No wildcard support + lucide-react already in default list | Excluded |
productionBrowserSourceMaps: false | Explicit setting | Default is already false in config-schema.js | Excluded |
swcMinify was a valid option in Next.js 14.x when the plan was written. In 15.5.15 it was completely removed from the config schema. Adding it breaks the build via z.strictObject() validation. When a library version advances after a plan is drafted, official docs alone aren't enough. Directly reading node_modules/ source is the only reliable verification method.
Conclusion: all three options inapplicable or already the default. Zero code changes.
The pre-check as an implementation gate
Four Principles in Hindsight
Looking back across five sprints, four principles kept showing up.
(i) Pilot → Expand Is the Basic Unit of Hypothesis Testing
Pilot on github-worker, then roll out to every service the next sprint. Straightforward in hindsight, but this structure is what surfaced the "checkout excluded" design decision and the "infra PR can't benchmark itself" problem within a small blast radius. Applying it to all services at once would have meant encountering the same problems at a much larger scale.
When picking a pilot, "simplest service first" is the right call. github-worker being pure Node.js without NestJS minimized side effects. If the pilot succeeds, expand. If it fails, fix and re-pilot. Sprint-level separation naturally enforced this cycle.
(ii) A Workflow Only Works When Paired with Repository Settings
The auto-merge work taught me this the hard way. It's easy to think "the workflow file is done, so automation is complete." But without allow_auto_merge: true and required status checks, the workflow can execute without auto-merge ever being scheduled.
This principle applies across CI automation. Required checks without Branch Protection mean nothing; a coverage gate only blocks PR merges when registered in Branch Protection. Between "I built a workflow" and "automation works" there's always a settings layer.
(iii) Infra PRs Can't Benchmark Themselves
The pitfall from the composite-action pilot and rollout. A PR that modifies the CI pipeline can't generate benchmark data for that pipeline. If the path-filter doesn't detect service code changes, all service jobs are skipped.
The fix for this structural gap was rebuild_all=true workflow_dispatch. Writing down when to use it was the core of S105 [A]. A single runbook, with zero code changes, closed a measurement gap that had built up over two sprints.
(iv) The Pre-Check Is a "Necessity Gate": Zero-Line Implementation Is Still a Conclusion
The Story Isn't Over
Five sprints are done, but one structural constraint remains unresolved. Every build job in AlgoSu is Docker-only. Not a single job runs npm run build on the host filesystem.
This constraint blocked both the L2 cache ([B]) and build-timing measurement ([C]) in the last sprint. Finding that both share the same root cause is the key input for what comes next.
The path forward:
| Approach | Description | Expected Gain |
|---|---|---|
| Blog host-side SSG build | CI builds out/ on host → GHA cache → Docker does COPY only | 40–60% reduction when there is no cache (cache miss) |
| Frontend host-side build | .next/standalone on host → GHA cache → Docker COPY only | 40–60% reduction when there is no cache (cache miss) |
| Per-service independent coverage gate | check-coverage.mjs per-service threshold, resolves path-filter misunderstanding | Better visibility |
The host-side build migration isn't a simple optimization PR. It's an architectural decision that requires simultaneous changes to both Dockerfile and ci.yml. Just as Channel.io separated the prepare phase into S3, AlgoSu is moving toward separating the build phase to the host side. But rather than rushing that decision, clearly understanding the current architecture's limits is the real contribution of these five sprints.
Five sprints in summary:
30→ 2
Dependabot PRs
S102: grouping + auto-merge
−67 lines
CI Duplicate Code
S103–S104: Composite action
69.55→ 76.42%
Frontend Branches
S106: 77 tests added
2
Zero-Line Decisions
S106: stopped early by pre-checks
Reading Channel.io's post “Backend CI Refactoring” again, I realize it didn't give me answers. It gave me questions. "What can AlgoSu's pipeline fix?" Carrying those questions into AlgoSu's context and walking the path: that was these five sprints.
The reference pointed the direction. The rest, I walked myself.