New Year Codebase Health Check: A January Checklist for Development Teams
A practical January checklist for development teams: audit dependencies, target test coverage, prune...
Deployments rarely fail because of the code you just wrote. They fail on the small, boring things: a memory limit nobody set, a secret that only ever existed in one engineer's shell, a database migration that cannot be undone. The ten minutes before a push to production are worth more than the ten hours you might spend afterwards explaining what went wrong.
None of the checks below are clever. Every one of them is cheaper than a rollback at 2am.
Scan images in CI, not on your laptop, and make the build fail on critical vulnerabilities rather than printing a report nobody reads. An image that was clean six months ago is not clean now: new CVEs are published against packages you never changed, so rebuild and rescan on a schedule as well as on every commit.
Start from the smallest base that will run your application. Distroless or a slim Alpine build gives attackers less to work with and gives you fewer things to patch. Multi-stage builds keep compilers and package managers out of the final layer.
Then pin what you deploy. Tags are mutable, and latest tells you nothing about what is actually running. Reference the digest alongside the tag so the exact bytes you tested are the bytes that ship. Run the container as a non-root user, set readOnlyRootFilesystem where you can, and drop Linux capabilities you do not need.
Requests and limits are not the same thing, and confusing them causes most of the strange behaviour you see under load. Requests tell the scheduler what to reserve; limits tell the kernel when to step in. A pod with no requests is scheduled blind and is among the first evicted when a node runs short.
Set memory requests close to typical usage and memory limits with a sensible margin, because going over a memory limit is an OOM kill rather than a slowdown. CPU limits cause throttling, so many teams set generous CPU requests and leave limits off for latency-sensitive services. Sidecars and init containers count too, and a Horizontal Pod Autoscaler cannot work at all without requests to calculate against.
If your team keeps forgetting, put a LimitRange and a ResourceQuota on the namespace. Enforcing a default is easier than remembering a rule.
A readiness probe gates traffic; a liveness probe restarts the container. Point liveness at something local and cheap, such as an endpoint that checks the process itself. If liveness depends on the database, a brief database blip becomes a restart loop across the whole deployment. Use a startup probe for applications that take a while to boot, so liveness does not kill them mid-initialisation.
Once Kubernetes decides to remove a pod, the clock starts. Set terminationGracePeriodSeconds longer than your slowest in-flight request, add a short preStop sleep so endpoint removal propagates before the process stops accepting connections, and handle SIGTERM properly. If your service holds queue connections or long-lived streams, drain them explicitly. A decent test is to delete a pod under load and watch the error rate.
Base64 is an encoding, not encryption. Anything in a Kubernetes Secret that has been committed to git, pasted into a CI log or echoed into a shell is effectively public. Keep the source of truth in a secret manager and pull it into the cluster with something like the External Secrets Operator or the Secrets Store CSI driver. If you need secrets in git for GitOps, use sealed or encrypted secrets with keys held elsewhere.
Mount secrets as files rather than environment variables where the application allows it. Environment variables leak through crash dumps, subprocesses and /proc, and a surprising number of frameworks print their entire configuration at startup. Tighten RBAC so only the workloads and controllers that genuinely need to read Secrets can do so, and check that rotation does not mean editing forty manifests by hand.
Confirm the previous good image digest still exists in the registry, and that your deployment keeps enough revision history for kubectl rollout undo to mean something. Then think about state, because that is where rollbacks become painful.
Ship schema changes in expand-and-contract steps: add the new column as nullable, backfill it, deploy code that reads both shapes, and remove the old column in a later release. Never combine a destructive migration with the code that stops using it. Feature flags let you switch a feature off without a deploy, which is often faster than a rollback and easier to reason about. Rehearse the rollback in staging as an actual step rather than a line in a runbook, and give stateful components their own documented plan.
Choose the rollout shape deliberately. A rolling update with sensible maxSurge and maxUn Photo: StockSnap / Pixabay
A practical January checklist for development teams: audit dependencies, target test coverage, prune...
CSS Grid usually breaks on mobile because of sizing floors, not the grid itself. Here's why implicit...
A practical checklist for keeping personal data out of application logs, setting sensible retention...