New Year Codebase Health Check: A January Checklist for Development Teams
A practical January checklist for development teams: audit dependencies, target test coverage, prune...
Logs are usually the first thing a team adds to a new service and the last thing anyone audits. A stack trace with a customer's email buried in it. A request body written to disk because parsing it was more effort. An IP address kept for three years because deleting things feels risky. None of that looks like a data protection problem until someone asks how you would find and remove one person's data. Under UK GDPR, log files are processing like any other: they hold personal data, they are shared with third-party tools, and they need an owner.
Logs are messier than databases. They are unstructured, replicated, shipped to vendors and indexed somewhere nobody quite remembers setting up. There is rarely a schema, and often no named owner. The accountability principle expects you to demonstrate compliance rather than simply intend it, which means being able to say what is in your logs, who can read them, where they go and when they are deleted.
Databases force these questions because they require migrations, access reviews and backup policies. Logs rarely do. A new field can appear in a log line without a review, a new vendor can be added without a data protection impact assessment, and a debug flag can stay on in production for months. That is how personal data ends up in places nobody expected.
If you cannot describe what is in your logs, you cannot answer a subject access request about them.
A field does not need to contain a name to be personal data. If a value relates to someone who can be identified, on its own or combined with other information you hold, it counts.
Some of it is worse than merely identifying. If someone describes a health condition in a support form, or a profile field records ethnicity, that content can fall into special categories and deserves extra care. An error message such as "failed login for [email protected]" is also a security problem, because it puts a credential and a direct identifier into a system that may be broadly readable.
Five questions, asked while writing the code rather than six months later, prevent most of the trouble:
These questions are cheap. Retrofitting redaction across a dozen services after an incident is not.
Structured logging with an explicit allow-list of fields is the single most effective habit. You have to add a field deliberately, which means somebody has to think about it. Redact at the point of logging rather than in a downstream pipeline that can fail quietly. Drop request bodies by default; if you genuinely need them, record a length, a schema or a keyed hash instead.
A simple example: logging "created user" with an internal numeric ID is usually enough to debug the flow. Logging the full signup payload is not. The same applies to search endpoints. The fact that a query returned zero results may be useful; the exact search term, which could reveal a health concern or political view, usually is not.
Then look at what you did not write. Error trackers capture breadcrumbs. HTTP clients log full URLs. ORMs log queries with their parameter values. Authentication middleware logs headers. Debug logging is the classic offender, so check that your production flag actually disables it in every service.
Pseudonymisation helps but is not a magic trick. A keyed hash of an email address remains personal data for as long as you hold the key, so the mapping table is a separate store with its own access controls. Never log credentials, API keys, session tokens or full payment card numbers. Add a scrubbing layer, and a continuous integration test that fails the build when a known secret pattern appears in output.
The storage limitation principle asks for a defined period, not an aspiration. A workable pattern is tiered: raw logs for a short window, security-relevant events for longer, and aggregated metrics indefinitely because they no longer identify anyone. Whatever you choose, automate the deletion — index lifecycle policies, TTLs in the store, a scheduled job whose failures raise an alert. Deletion that depends on someone remembering is deletion that will not happen.
Retention also needs to survive vendor changes. If you move from one log platform to another, do not simply copy the old index "just in case". Check what the new default retention is, whether backups inherit it, and whether support access creates a separate copy. A retention policy that exists only in a wiki page is not a control.
Retention is only half the story. A log that is kept for thirty days but readable by everyone in the company is still a weak point. Treat log stores as production systems: role-based access, audit trails on queries, no shared dashboards with public links, and no exporting raw logs to a spreadsheet for a meeting. The people who need to debug an incident are not always the people who need to see every customer's email address.
Third-party log and error tools are processors. You need a data processing agreement, a clear description of what is sent, and a view on where it is stored and transferred. Ask vendors about sub-processors, retention controls, deletion APIs and whether support staff can access your data. If a tool cannot delete a single user's events or apply a custom retention period, that is a design constraint you need to know before an incident, not after.
Also think about internal sharing. Pasting a stack trace into a ticket, chat channel or AI assistant can move personal data into another system with different retention and access rules. If you need help debugging, redact first or share only the minimum viable excerpt.
If someone asks for a copy of their data, logs may be in scope. You do not have to hand over every noisy line, but you do need to be able to search for identifiers and explain what you found. That is difficult if email addresses are scattered across free-text messages or if user IDs are not indexed. Design logs so you can find records by the identifiers you actually use for data subject requests, and document which fields are searchable.
Breach detection is the other side. Logs are often how you discover unauthorised access, but they can also be the breached data. If an attacker reads a log file containing session tokens, those tokens become credentials. That is why secrets should never be logged, why raw request bodies are dangerous, and why log stores need the same patching, encryption and access controls as any other production datastore.
Have a simple plan for log-based incidents:
Before you merge a logging change, run through this short list. It is deliberately boring, which is the point.
None of this requires a legal degree. It requires treating logs as a product surface with users, risks and lifecycle rules. Teams that do this ship faster in the long run, because they can answer questions about their data without a frantic search through five systems. The alternative is a log pile that nobody owns, which is exactly where compliance problems like to hide.
Photo: Digital Buggu / Pexels
A practical January checklist for development teams: audit dependencies, target test coverage, prune...
CSS Grid usually breaks on mobile because of sizing floors, not the grid itself. Here's why implicit...
A practical comparison of AWS, Azure and Google Cloud for UK SMEs, covering UK regions and data residency,...