EngineeringRuntime Contracts and Infrastructure Ownership Boundaries for Deployment
· PinLog
- #GitOps
- #Kubernetes
- #DevOps
My job on PinLog was deploying each team's services to Kubernetes, connecting their dependencies, observing their health, and handling recovery.
Before development started, the team agreed in writing on structure, conventions, and who was responsible for what. This article is about how that agreement took shape in deployment.
An image lets you start a container. Its address alone does not tell you which port to connect to, what counts as healthy, or which configuration and external services it needs. A team deployment required that information as well.
Problem
An image address alone did not say which port to connect to, what counted as healthy, or which configuration and external services were needed. When information was missing, infrastructure could have filled the gap by editing another owner's code.
Decision
Application owners would hand over the image together with its runtime contract, and infrastructure would own only deployment, connectivity, observability, and recovery. Missing information meant holding the deployment and reporting the blocker and its owner.
Result
The role boundary was confirmed on July 29, 2026. Passing checks and reaching production were also kept apart: a change that had only passed its checks, like the July 27 pgvector transition, was not called deployed.
What "Deployable" Required
We agreed that the service owner would hand over a runnable image together with its runtime contract. Infrastructure would deploy and connect it according to that contract.
| Handoff information | What infrastructure needed to establish |
|---|---|
| Image commit SHA and digest | Which build artifact are we deploying? |
| Container port | Where should the Service send traffic? |
| Health, liveness, and readiness paths | Which endpoints report status? |
| Environment variables and credential keys | What configuration must be delivered securely? |
| External dependencies, including the DB | Which services and network paths are required? |
For example, the backend contract specified internal port 8080, the /api/core context path, and /api/core/actuator/health. These were not values for infrastructure to guess. The application owner defined them; infrastructure connected the service accordingly.
Credentials followed the same principle. We distinguished the required key schema from the secret values themselves. Production secrets would be injected through Secrets, not stored in the image or Git.
Not Fixing Others' Code
When I worked alone on AlgoSu, my personal algorithm study platform, I could change application code if I found a problem during deployment. That approach did not carry over unchanged to a team.
The boundary confirmed on July 29 was explicit. Infrastructure owned deployment, connectivity, observability, and recovery. It did not write or modify another owner's application code. Reviewing that code, deciding API internals, or determining database migration files and their order were also outside its scope.
| Application owner | Infrastructure owner |
|---|---|
| Deliver a runnable image and runtime contract | Deploy that image to Kubernetes |
| Own application configuration and probe internals | Configure Services, Ingress, and health checks |
| Own database migration files and ordering | Provide database connectivity and deployment/recovery procedures |
| Define credential keys and external dependencies | Provide Secret delivery and network configuration |
We also defined what to do when information was missing. Rather than filling the gap by editing application code, infrastructure would hold the deployment and report the blocker and its owner.
The same boundary applied to Pico, the AI agent I worked with on infrastructure. An agent's ability to modify code was separate from its authority to do so.
From Image to Production
We defined a Kubernetes and Argo CD GitOps flow rather than running images directly on the server. In GitOps, deployment configuration lives in Git, and changes there are applied to the cluster. Argo CD is the tool that keeps the cluster in line with Git.
Image references would not rely on a mutable tag alone. They would pin both the commit SHA and the digest (an identifier derived from the image content).
Image and contract
The application owner delivers the required information
Validate the handoff
If incomplete: hold deployment → identify blocker and owner → request an updated handoff
Infra PR and checks
Update image references and deployment configuration; validate and obtain approval
Argo CD and verification
Apply changes, then inspect rollout, health, and logs
This is the deployment flow I set. The backend deployment below has not gone through every stage yet.
A change in Git was also distinct from a healthy running service. After applying it, we still needed to inspect Pod health, events, logs, and metrics. Infrastructure's responsibility did not end with updating an image reference.
Passing Checks Is Not Deploying
The July 27 pgvector transition review makes this distinction concrete.
pgvector is a PostgreSQL extension that adds vector columns. The backend's Flyway migration (Flyway applies database schema changes in order) created the vector extension, while the AI migrations used vector columns. The production PostgreSQL image therefore needed the extension binaries. This was where an application requirement had to be matched by its operating environment.
At that point, the infrastructure PR had passed its checks but remained open. Production PostgreSQL still used its existing image, and the vector extension was not installed. The backend image-publishing workflow changes existed only in a local commit, with no remote PR yet.
| Item checked on July 27 | Status on July 27 |
|---|---|
| Checks on the pgvector infrastructure PR | Passed |
| PR merge | Not completed |
| Production DB image change and restart | Not performed |
Creation of the vector extension | Not performed |
| Backend image-publishing workflow changes | Local commit only |
So I don't call this change "deployed" yet. The deployment contract has been reviewed and the changes checked, but rolling it out to production is still pending.
The database transition also had recovery conditions. Before and after creating the extension, returning to the previous image meant different things. Once the extension existed, we would not simply revert to the old image without the required shared library. Restarting a single-replica database also required separate maintenance approval.
Nor did the existence of a backup file mean recovery was proven. The database backup archive passed its format check, but the database was still an initial database with no user tables. So I didn't count it as a restore test.
What Infrastructure Owns
Looking back, what I needed as the infrastructure owner was not authority to fix every piece of code. I needed to understand the conditions under which the artifact I received would run, and meet those conditions in the operating environment.
The application owner supplied a runnable image and a contract. Infrastructure deployed it, connected it, observed it, and handled recovery. An incomplete handoff meant stopping at that boundary and checking again.
This was how the structure, conventions, and responsibilities the team agreed on before development took shape in deployment. Receiving an image was the beginning. We also had to define what came with it, how far my responsibility extended, and which states we could not yet call complete.