Every release meant rebuilding production by hand, and one missed step took the whole API down.

A multi-tenant B2B SaaS company ran its API on servers it rebuilt by hand for every release. We moved production to Kubernetes without moving or pausing its database, and a deploy now runs from a pipeline in 5 to 10 minutes. In the first 60 days the new platform served 20.7 million API requests with 0.024% server errors.
Client
A multi-tenant B2B SaaS company
Scale
150+ tenant databases, one per tenant
Engagement
Platform build, production migration, then day-to-day operations

The outcome in six lines

The outcome in six lines
MeasureBeforeAfterEvidence
A production release1new servers and a database fork, built and repointed by handone GitHub release, a 5 to 10 minute deploy with automatic rollback, then tenant migrations drain for 1 to 1.5 hoursMeasured
The database during cutover2forked to a new cluster at every releasenever moved or pausedDescribed
Server errors3not measured, one full API outage after a release0.024% of 20.7 million requests over 60 daysMeasured
Logs and metrics4on each server, no dashboardsone place for every environment, 7 dashboards kept as codeDescribed
Production hosting, list price5$220 a month$244 a month, plus $108 for a shared management clusterDescribed
Downtime per release6not measuredabout 2 minutes of HTTP 503Measured

Three lines are measured and three describe the platform or its price list. Two lines got worse, and they are shown as they are. The small number beside each line says how it is known.

The problem, in the client's terms

Page speed and load were not the complaint. Releases were. Every release meant building a new set of three servers (web, worker, Redis) and a forked copy of the managed database, then repointing the configuration and copying the application's encryption key across by hand. There was no written runbook.

That process failed the way hand-run processes do. On one release the new servers sat on a different private network from the new database and could not reach it. Every API request timed out, and nobody could log in or load data.

The heavy work ran in the background. After each release, schema migrations ran across 150+ tenant databases, alongside imports and media conversion. Workers ran out of memory during big tenant migrations, the jobs went back on the queue, and a drain that should take about an hour took about 4.5 hours on one release.

Logs never left the server that wrote them and there were no dashboards, so diagnosing a fault meant logging into servers one at a time.

For the client, a release became something to plan around and dread rather than something to ship. Each one carried the risk of an outage their customers would see first.

  • ConstraintStay with the existing hosting provider and keep the existing managed MySQL, one database per tenant.
  • ConstraintA small infrastructure team on our side, so the setup had to be one a few people can run.
  • ConstraintThe client's customers use the platform every working day, so the move could not take a long outage.

The architecture, before and after

Switch between the two. Releases moved from hand-built servers to one pipeline, the workload moved into a cluster that shares the database's private network, and logs and metrics moved to one place. The components are as described by the team that ran it; the layout and link directions are drawn by us.

Before. Three servers and a database fork, rebuilt and repointed by hand at every release.
The architecture, before and after: BeforeBefore: Three servers and a database fork, rebuilt and repointed by hand at every release. Components: Cloudflare, Web server, Managed MySQL, Worker server, Redis server, Hand-run release. Cloudflare connects to Web server. Web server connects to Managed MySQL. Worker server connects to Redis server. Worker server connects to Managed MySQL. Hand-run release connects to Worker server, dashed.RequestsBackground workReleasesCloudflareproxied DNS, no balancerWeb serverNGINX and PHP-FPMManaged MySQLa database per tenantWorker serverHorizon queue workersRedis serverqueues and cacheHand-run releaseno written runbookEvery release rebuiltall three servers andforked the database.Workers ran out ofmemory on big tenantmigrations and the queuestalled for hours.Logs stayed on eachserver. There were nodashboards.
After. One cluster per environment inside its database's network, one pipeline, one place to look.
The architecture, before and after: AfterAfter: One cluster per environment inside its database's network, one pipeline, one place to look. Components: Cloudflare, Load balancer, Web pods, Managed MySQL, Worker pods, Redis, GitHub Actions, Container registry, Management cluster. Cloudflare connects to Load balancer. Load balancer connects to Web pods. Web pods connects to Managed MySQL. Worker pods connects to Redis. Worker pods connects to Managed MySQL. GitHub Actions connects to Container registry. Container registry connects to Worker pods. Worker pods connects to Management cluster, dashed.RequestsBackground workReleases and opsCloudflareFull (Strict) TLSLoad balancerinto NGINX ingressWeb pods3 pods, NGINX, PHP-FPMManaged MySQLsame private networkWorker podsHorizon and schedulerRedisinside the clusterGitHub ActionsHelm upgrade, rollbackContainer registryone image, three rolesManagement clusterRancher, Grafana, LokiThe cluster was createdinside the database'sexisting network, so nodata had to move.
  • What hurt in this state
  • Built in this engagement

Three decisions, and what we turned down

The stack is the easy part to copy. The choices behind it are what made the cutover quiet, so each one names what it replaced and what it cost.

  1. Where the clusters run

    • Turned down

      AWS EKS or Google GKE

      Both priced higher in the estimate made before the build, and the databases, file storage and team were already on DigitalOcean, so moving clouds meant moving the data too.

    • Turned down

      OpenShift or VMware Tanzu to manage the clusters

      More than a single-application platform needs.

    • Chosen

      DigitalOcean Kubernetes, managed through Rancher

      The control plane is managed, free in its basic form, with the paid high-availability option in production, and every cluster can sit next to the managed MySQL the client already had.

    What it cost: Rancher and the monitoring stack need a cluster of their own, a standing cost across every environment, and there are far more moving parts to watch than three servers.

  2. One cluster per environment, inside its database's network

    • Turned down

      One shared cluster reaching each database over network peering

      It was the first design and it did not work. The older cluster networking we were on did not route across peered networks in our setup, and a traceroute showed traffic leaving over the public internet.

    • Chosen

      A workload cluster per environment in its database's own private network, and one management cluster as the hub

      Each application pod reaches its database privately, and production could be cut over by building the new cluster next to the live database instead of moving it.

    What it cost: More clusters to build, patch and pay for, and the platform phase ran longer because the first design had to be torn up.

  3. How a release reaches production

    • Turned down

      GitOps with ArgoCD

      More machinery than a single application in four environments needs, for a small team that would have to run it.

    • Turned down

      Keep releasing on the old servers

      Every release stayed a manual sequence where one missed step could take production down.

    • Chosen

      Push deploys from GitHub Actions with a Helm chart we wrote

      Publishing a release in the API repository builds one image and runs a Helm upgrade, which rolls back on its own if the upgrade fails.

    What it cost: Nothing watches the cluster for drift between releases the way a GitOps controller would, so a change made by hand stays until the next deploy overwrites it.

How it was delivered

Platform first, then the application, then one environment at a time with production last. About 8.5 months passed from the first infrastructure work to the production cutover.

  1. The platform

    The infrastructure repository, Terraform, Rancher, monitoring and a test cluster, over about 3 months. That included the network-peering dead end and the redesign to one cluster per environment.

  2. The application in a container

    One image with three roles (web, worker, scheduler), health checks, logs to standard output, a guard so scheduled tasks run once rather than once per pod, and trust in the load balancer's forwarded headers.

  3. The lower environments

    The first moved 3 days after the test cluster came up. About 5.5 weeks later the second was built fresh on Kubernetes from a copy of the production database.

  4. Production cutover

    About 4 months later, in one step. The old web server went into maintenance, the old workers stopped and the queue drained, the release deployed to Kubernetes and DNS switched to the new load balancer. Traffic was back to normal within about 25 minutes. The old servers stayed warm, on the same database, as the rollback path. It was never used.

  5. Clean-up

    From 2 weeks after cutover. The test cluster was removed and the old servers were switched off within about 9 weeks of cutover.

Who did the work

  • Platform software engineer, AtyantikDesigned and built the platform.
  • Infrastructure software engineer, AtyantikRan the production cutover, the audits and day-to-day operations.
  • Application software engineers, AtyantikMade the code changes the containers needed.
  • The clientTheir technical operations lead and one stakeholder.

Who owns what now

Atyantik runs the clusters. The client's operations team had a walkthrough of Rancher, Grafana and what to watch in the logs, a post-cutover report on the timeline, fixes and open items, and a follow-up audit on cost, clean-up and autoscaling, reviewed on the weekly call.

Results

What the move changed, including what got worse.

  • Whoever ships a release publishes a GitHub release instead of forking a database and repointing servers by hand.
  • Whoever is on call reads one set of dashboards and logs instead of logging into servers one by one.
  • The client's operations team can see the platform's state in Rancher and Grafana themselves.
  • Queue workers were not killed for running out of memory once in a 10-day window measured after cutover. Tenant migrations still take about 1 to 1.5 hours to drain after a release, and about 4.5 hours in the worst case seen.

The old servers had no metrics stack, so there is no before figure for response time, uptime or the length of a release. Time spent on operations was not measured on either side.

Production hosting per month, at list priceAbout 11% more, before the shared management cluster

$220

Before, three servers

$244

After, the production cluster

The move bought a repeatable release and one place to look, not a smaller bill.

Production hosting per month, at list price (US dollars a month at list price, production only, with the same managed database on both sides. The shared management cluster adds $108 a month across all environments.)
OptionUS dollars a month at list price, production only, with the same managed database on both sides. The shared management cluster adds $108 a month across all environments.
Before, three servers$220
After, the production cluster$244

Source: Source note 5

What hand-run releases cost you

We do not publish what an engagement costs, and the hands-on time of the old releases was never timed. So this works from your side: what your releases cost in people's hours today.

Your releases today

an assumption, change it

This is a diagnostic of what manual releases cost you today, not a saving we claim. Outage cost is left out, and so is the pipeline's own running cost.

Hands-on release cost per year

$17,280

Hands-on release hours per year
288 h

Every one of those hours is also a chance for one missed step.

What went wrong, and what we would change

Single sign-on broke at cutover

Everything came up clean, then every single sign-on login was rejected. The login library picked up the container's internal port instead of the public one and put it in the sign-in URL. A one-setting fix shipped about 4 hours later and rejections dropped to zero. Behind a load balancer an application no longer knows its own public address. Next time, single sign-on is tested in a lower environment through the same load balancer before production moves.

Autoscaling that was configured and did nothing

The configuration asks for 3 to 10 web pods at 70% CPU, but the Helm chart never created the autoscaler and the metrics service it needs was not installed. Production runs a fixed 3 web pods on 2 nodes, and at weekday peaks the web pods run close to their CPU limit. Next time, the metrics service and pod autoscaling go in, and are shown to scale, before anyone says the platform autoscales. That work is now the next step.

Is Kubernetes the answer for your platform?

It fits when

  • You need several identical environments run the same way.
  • Releases carry manual steps, and one missed step can take production down.
  • Your data already sits in managed databases that can stay where they are while the application moves.

Look elsewhere when

  • A single application with a small team and modest load. A PaaS or scripted server deploys may be cheaper to run.
  • Your pain is slow pages under load rather than releases. Profile the application before changing where it runs.
  • Nobody, on your team or a partner's, will own the clusters day to day.

Check it yourself before your next release

If the list runs long and the scheduler output is missing, the release process is the problem, whatever it runs on.

  1. Write down every step of your last release, including the ones only one person knows.
  2. Mark every step that is typed or clicked by hand.
  3. Check that last night's scheduled jobs actually ran, from their output rather than the cron file. Moving this platform found a scheduler that had not been running on the old servers either.
  4. If you already run Kubernetes, check that your autoscaler exists and has metrics to act on.
bash
kubectl get hpa --all-namespaces
kubectl top pods --all-namespaces

The stack, by layer

  • ApplicationA modular Laravel API with Horizon queue workers
  • DataManaged MySQL with one database per tenant, Redis inside the cluster, object storage for files
  • PlatformDigitalOcean Kubernetes, one workload cluster per environment, managed through Rancher, built with Terraform
  • DeliveryGitHub Actions, a Helm chart we wrote, one container image with web, worker and scheduler roles
  • EdgeCloudflare with Full (Strict) TLS and origin certificates, a load balancer, NGINX ingress
  • ObservabilityPrometheus, Grafana, Loki and Alertmanager on a management cluster, Grafana Alloy in every cluster

How each figure was measured

  1. Release steps on the old servers, and deploy run times from GitHub Actions on the new platform.
  2. How production was cut over, with the new cluster built inside the database's existing private network.
  3. Ingress access logs over the first 60 days after cutover, server errors (5xx) against all API requests.
  4. The monitoring and logging in place on the old servers and on the new platform.
  5. Provider list prices per month for the production servers before and the production cluster after.
  6. Ingress logs of 5xx responses per minute, matched to pod start times, on two production deploys.

Send us your last release, step by step, and we will tell you if Kubernetes is the fix.

Write down what was done by hand and how long it took. A software engineer reads it and replies with an honest answer, including when scripted deploys or a PaaS are the better choice.