RetrospectiveA Retrospective on 67 Sprints with AI
· AlgoSu
- #retrospective
- #solo-dev
- #growth
Looking Back at Sprint 67
It started with curiosity about multi-agent systems. "If several AI agents worked together, could I build a whole service on my own?" That one question started a project that, 67 sprints (short work cycles) later, had become a platform.
- 67 sprints
- 2,432 tests across the codebase, with 99.51% branch coverage in the user & auth service
- 6 microservices + 1 frontend
- 34 API endpoints in the user & auth service (Identity)
- 12 AI agents
The numbers look clean. Behind them is a loop of decisions, mistakes, and fixes.
What You Only See After Doing It
Key points across 67 sprints
Decisions Made Early
Splitting the database: decided in Sprint 12, paid off in Sprint 51. When there were only three services, I gave each its own database (identity_db, problem_db, submission_db). Honestly, it felt like overkill.
The databases were split, but one shortcut remained: the Gateway (the service that receives every request) still read the user database directly. Sprint 51 brought a big refactor to cut that off completely. Because the databases were already separate, the migration went far more smoothly. An "overengineered" call from 40 sprints earlier saved me.
Sprint 12
Split into 3 databases
Sprint 51
Gateway's direct DB access removed
40 sprints later, overengineering saved the day
httpOnly cookies: decided in Sprint 8, paid off in Sprint 65. When I added social login (OAuth), I chose to keep the login token (JWT) in an httpOnly cookie instead of browser storage (localStorage). Page scripts can't read an httpOnly cookie, so a token is much harder to steal through injected script (XSS).
Because I had moved to cookies early, the 11 leftover localStorage token functions were no longer used by the real code. When I removed them in Sprint 65 (+42 lines / −440 lines), I only had to delete them; nothing else needed changing.
Sprint 8
Adopted httpOnly cookies
Sprint 65
Removed 11 localStorage functions
57 sprints later, I only had to delete
Test obsession: 2,432 tests in total, 99.51% branch coverage in the user & auth service. The Sprint 51 refactor meant creating 34 new APIs and rewriting 19 files outright. I could do it because the tests were a safety net. I could change things boldly, without worrying "will this break something over there?"
Decisions Deferred
Not automating type synchronization. With the system split into services (MSA), adding a single enum value meant hand-editing at least five files: the user service's entity → the Gateway's types → the frontend's api.ts → the UI components.
Once I added an enum value, FEEDBACK_RESOLVED, only to the user service and never propagated it. That gap took two more sprints to clean up. A code generator or a shared package, introduced early, would have spared me.
Connecting monitoring alerts late. Metrics collection (Prometheus) and alert routing (Alertmanager) were set up long before. But the Discord channel that actually reaches a human wasn't connected until Sprint 63. Having a tool and having it actually work are different things.
Lessons from the Details
Why an obvious bug went unnoticed. The admin page's feedback filter ran in the browser (client side). Meanwhile, the server sent the list 20 items per page. What happens? Only the 20 items on screen get filtered, and everything on other pages stays invisible.
An obvious bug, and I didn't see it while building. I had to move the filter to the server and include total counts in the response. Eight commits, 17 files changed.
The Discord notification trap. The code was done, the Secret was created, and the deployment manifest in the source repo was written. Yet no notifications arrived.
The cause was the repository that holds only deployment configuration (aether-gitops). ArgoCD, the tool that applies configuration to the running cluster, syncs from that deployment-config repository. I had updated the source repo and forgotten the environment variables (env) in the deployment-config repo.
The alert problem and the enum propagation gap above were separate issues, but they landed at the same time while I was fixing this. In the end, I had to resync all three layers (user service → Gateway → frontend) in Sprint 66.
Both episodes tell the same story: each part looked right on its own, but something was missing in the whole.
AlgoSu by the Numbers
67
Total Sprints
2,432
Total Tests
User & auth service: 99.51% branch coverage
15
CI Jobs
Automated checks run on every push and PR
6
Microservices
Gateway · Identity · Submission · Problem · GitHub-Worker · AI-Analysis
12
AI Agents
The lead agent, Oracle, + 11
34
User & Auth APIs
0 → 34 in Sprint 51
From Coder to Builder
Looking back, every sprint was a first encounter with something new. The overwhelm of meeting the Saga pattern (running a multi-service operation step by step and rolling it back on failure) for the first time. The late nights digging through logs when messages in the queue (RabbitMQ) weren't being consumed. The relief when the first ArgoCD sync succeeded. These moments repeated 67 times, and somewhere along the way I changed without noticing.
At first, "how do I implement this feature?" was everything. I was a coder, someone who writes code. As sprints piled up, the questions shifted. "Will this decision still hold 20 sprints from now?" "If something fails in this architecture, where do I catch it?" "How do I prove this automation actually works?" I was becoming a builder, someone who builds systems, not just code.
Some things you only get by pouring in tokens and time.
Sprint 68 is still blank. But the sprints will go on. A project born from curiosity became a platform, and a coder became a builder. I'll feel lost in front of new technology in the next sprint too, but 67 sprints of experience tell me: do it, and you'll figure it out.