Every release meant rebuilding production by hand, and one missed step took the whole API down.
- Client
- A multi-tenant B2B SaaS company
- Scale
- 150+ tenant databases, one per tenant
- Engagement
- Platform build, production migration, then day-to-day operations
The outcome in six lines
| Measure | Before | After | Evidence |
|---|---|---|---|
| A production release1 | new servers and a database fork, built and repointed by hand | one GitHub release, a 5 to 10 minute deploy with automatic rollback, then tenant migrations drain for 1 to 1.5 hours | Measured |
| The database during cutover2 | forked to a new cluster at every release | never moved or paused | Described |
| Server errors3 | not measured, one full API outage after a release | 0.024% of 20.7 million requests over 60 days | Measured |
| Logs and metrics4 | on each server, no dashboards | one place for every environment, 7 dashboards kept as code | Described |
| Production hosting, list price5 | $220 a month | $244 a month, plus $108 for a shared management cluster | Described |
| Downtime per release6 | not measured | about 2 minutes of HTTP 503 | Measured |
Three lines are measured and three describe the platform or its price list. Two lines got worse, and they are shown as they are. The small number beside each line says how it is known.
The problem, in the client's terms
Page speed and load were not the complaint. Releases were. Every release meant building a new set of three servers (web, worker, Redis) and a forked copy of the managed database, then repointing the configuration and copying the application's encryption key across by hand. There was no written runbook.
That process failed the way hand-run processes do. On one release the new servers sat on a different private network from the new database and could not reach it. Every API request timed out, and nobody could log in or load data.
The heavy work ran in the background. After each release, schema migrations ran across 150+ tenant databases, alongside imports and media conversion. Workers ran out of memory during big tenant migrations, the jobs went back on the queue, and a drain that should take about an hour took about 4.5 hours on one release.
Logs never left the server that wrote them and there were no dashboards, so diagnosing a fault meant logging into servers one at a time.
For the client, a release became something to plan around and dread rather than something to ship. Each one carried the risk of an outage their customers would see first.
- ConstraintStay with the existing hosting provider and keep the existing managed MySQL, one database per tenant.
- ConstraintA small infrastructure team on our side, so the setup had to be one a few people can run.
- ConstraintThe client's customers use the platform every working day, so the move could not take a long outage.
The architecture, before and after
Switch between the two. Releases moved from hand-built servers to one pipeline, the workload moved into a cluster that shares the database's private network, and logs and metrics moved to one place. The components are as described by the team that ran it; the layout and link directions are drawn by us.
- What hurt in this state
- Built in this engagement
Three decisions, and what we turned down
The stack is the easy part to copy. The choices behind it are what made the cutover quiet, so each one names what it replaced and what it cost.
Where the clusters run
Turned down
AWS EKS or Google GKE
Both priced higher in the estimate made before the build, and the databases, file storage and team were already on DigitalOcean, so moving clouds meant moving the data too.
Turned down
OpenShift or VMware Tanzu to manage the clusters
More than a single-application platform needs.
Chosen
DigitalOcean Kubernetes, managed through Rancher
The control plane is managed, free in its basic form, with the paid high-availability option in production, and every cluster can sit next to the managed MySQL the client already had.
What it cost: Rancher and the monitoring stack need a cluster of their own, a standing cost across every environment, and there are far more moving parts to watch than three servers.
One cluster per environment, inside its database's network
Turned down
One shared cluster reaching each database over network peering
It was the first design and it did not work. The older cluster networking we were on did not route across peered networks in our setup, and a traceroute showed traffic leaving over the public internet.
Chosen
A workload cluster per environment in its database's own private network, and one management cluster as the hub
Each application pod reaches its database privately, and production could be cut over by building the new cluster next to the live database instead of moving it.
What it cost: More clusters to build, patch and pay for, and the platform phase ran longer because the first design had to be torn up.
How a release reaches production
Turned down
GitOps with ArgoCD
More machinery than a single application in four environments needs, for a small team that would have to run it.
Turned down
Keep releasing on the old servers
Every release stayed a manual sequence where one missed step could take production down.
Chosen
Push deploys from GitHub Actions with a Helm chart we wrote
Publishing a release in the API repository builds one image and runs a Helm upgrade, which rolls back on its own if the upgrade fails.
What it cost: Nothing watches the cluster for drift between releases the way a GitOps controller would, so a change made by hand stays until the next deploy overwrites it.
How it was delivered
Platform first, then the application, then one environment at a time with production last. About 8.5 months passed from the first infrastructure work to the production cutover.
The platform
The infrastructure repository, Terraform, Rancher, monitoring and a test cluster, over about 3 months. That included the network-peering dead end and the redesign to one cluster per environment.
The application in a container
One image with three roles (web, worker, scheduler), health checks, logs to standard output, a guard so scheduled tasks run once rather than once per pod, and trust in the load balancer's forwarded headers.
The lower environments
The first moved 3 days after the test cluster came up. About 5.5 weeks later the second was built fresh on Kubernetes from a copy of the production database.
Production cutover
About 4 months later, in one step. The old web server went into maintenance, the old workers stopped and the queue drained, the release deployed to Kubernetes and DNS switched to the new load balancer. Traffic was back to normal within about 25 minutes. The old servers stayed warm, on the same database, as the rollback path. It was never used.
Clean-up
From 2 weeks after cutover. The test cluster was removed and the old servers were switched off within about 9 weeks of cutover.
Who did the work
- Platform software engineer, AtyantikDesigned and built the platform.
- Infrastructure software engineer, AtyantikRan the production cutover, the audits and day-to-day operations.
- Application software engineers, AtyantikMade the code changes the containers needed.
- The clientTheir technical operations lead and one stakeholder.
Who owns what now
Atyantik runs the clusters. The client's operations team had a walkthrough of Rancher, Grafana and what to watch in the logs, a post-cutover report on the timeline, fixes and open items, and a follow-up audit on cost, clean-up and autoscaling, reviewed on the weekly call.
Results
What the move changed, including what got worse.
- Whoever ships a release publishes a GitHub release instead of forking a database and repointing servers by hand.
- Whoever is on call reads one set of dashboards and logs instead of logging into servers one by one.
- The client's operations team can see the platform's state in Rancher and Grafana themselves.
- Queue workers were not killed for running out of memory once in a 10-day window measured after cutover. Tenant migrations still take about 1 to 1.5 hours to drain after a release, and about 4.5 hours in the worst case seen.
The old servers had no metrics stack, so there is no before figure for response time, uptime or the length of a release. Time spent on operations was not measured on either side.
$220
Before, three servers
$244
After, the production cluster
The move bought a repeatable release and one place to look, not a smaller bill.
| Option | US dollars a month at list price, production only, with the same managed database on both sides. The shared management cluster adds $108 a month across all environments. |
|---|---|
| Before, three servers | $220 |
| After, the production cluster | $244 |
Source: Source note 5
What hand-run releases cost you
We do not publish what an engagement costs, and the hands-on time of the old releases was never timed. So this works from your side: what your releases cost in people's hours today.
Hands-on release cost per year
$17,280
- Hands-on release hours per year
- 288 h
Every one of those hours is also a chance for one missed step.
What went wrong, and what we would change
Single sign-on broke at cutover
Everything came up clean, then every single sign-on login was rejected. The login library picked up the container's internal port instead of the public one and put it in the sign-in URL. A one-setting fix shipped about 4 hours later and rejections dropped to zero. Behind a load balancer an application no longer knows its own public address. Next time, single sign-on is tested in a lower environment through the same load balancer before production moves.
Autoscaling that was configured and did nothing
The configuration asks for 3 to 10 web pods at 70% CPU, but the Helm chart never created the autoscaler and the metrics service it needs was not installed. Production runs a fixed 3 web pods on 2 nodes, and at weekday peaks the web pods run close to their CPU limit. Next time, the metrics service and pod autoscaling go in, and are shown to scale, before anyone says the platform autoscales. That work is now the next step.
Is Kubernetes the answer for your platform?
It fits when
- You need several identical environments run the same way.
- Releases carry manual steps, and one missed step can take production down.
- Your data already sits in managed databases that can stay where they are while the application moves.
Look elsewhere when
- A single application with a small team and modest load. A PaaS or scripted server deploys may be cheaper to run.
- Your pain is slow pages under load rather than releases. Profile the application before changing where it runs.
- Nobody, on your team or a partner's, will own the clusters day to day.
Check it yourself before your next release
If the list runs long and the scheduler output is missing, the release process is the problem, whatever it runs on.
- Write down every step of your last release, including the ones only one person knows.
- Mark every step that is typed or clicked by hand.
- Check that last night's scheduled jobs actually ran, from their output rather than the cron file. Moving this platform found a scheduler that had not been running on the old servers either.
- If you already run Kubernetes, check that your autoscaler exists and has metrics to act on.
kubectl get hpa --all-namespaces
kubectl top pods --all-namespaces The stack, by layer
- ApplicationA modular Laravel API with Horizon queue workers
- DataManaged MySQL with one database per tenant, Redis inside the cluster, object storage for files
- PlatformDigitalOcean Kubernetes, one workload cluster per environment, managed through Rancher, built with Terraform
- DeliveryGitHub Actions, a Helm chart we wrote, one container image with web, worker and scheduler roles
- EdgeCloudflare with Full (Strict) TLS and origin certificates, a load balancer, NGINX ingress
- ObservabilityPrometheus, Grafana, Loki and Alertmanager on a management cluster, Grafana Alloy in every cluster
How each figure was measured
- Release steps on the old servers, and deploy run times from GitHub Actions on the new platform.
- How production was cut over, with the new cluster built inside the database's existing private network.
- Ingress access logs over the first 60 days after cutover, server errors (5xx) against all API requests.
- The monitoring and logging in place on the old servers and on the new platform.
- Provider list prices per month for the production servers before and the production cluster after.
- Ingress logs of 5xx responses per minute, matched to pod start times, on two production deploys.
Send us your last release, step by step, and we will tell you if Kubernetes is the fix.
Write down what was done by hand and how long it took. A software engineer reads it and replies with an honest answer, including when scripted deploys or a PaaS are the better choice.