How PII escapes Doctrine into Monolog, Messenger failures, profiler traces and exception reports — and the allowlist pattern that closes the gap.

Engineering analysis, framed around the GDPR. Not legal advice.

A support ticket comes in: a customer wants to know why a payment failed. You open Kibana, search the order id, and there it is — not just the order id. The full request payload, pretty-printed by an exception handler: name, email, cardholder name, shipping address. The database column for that customer’s email is encrypted. The log line sitting three tabs away in your browser is plaintext, retained for ninety days, readable by every engineer with log access, and outside every access control you built for the database.

Nobody put it there on purpose. An exception listener serialized the failed command for debugging. Nothing in that listener knows the word “personal data.” It just knows how to call json_encode on whatever object it was handed.

The popular fix, and why it only closes half the hole

The usual first response is a denylist: a Monolog processor that redacts keys named email, password, ssn. It catches the obvious case — a flat array with an email key — and it is worth having. It does not catch the case that actually filled the log with PII above: a customer's full name embedded inside a free-text exception message ("Payment declined for Jane Doe, card ending 4242"), or a nested DTO where the field is called contactName, or a Messenger envelope whose serialized body contains the whole command object as one opaque string. A denylist keyed on field names cannot see into prose, into nested structures it wasn't told about, or into a blob it never parses.

The gap is structural, not a matter of adding more key names to the denylist. The moment PII is allowed to travel through logging as “whatever the object happens to contain,” a denylist is chasing an open set.

Where a Symfony application actually leaks

Personal data reaches your log storage and observability stack through more paths than most teams have inventoried:

  • Automatic serialization of commands, events and DTOs. An exception subscriber, a debug processor, or an APM integration that calls json_encode or the Symfony Serializer on "the object that was being handled" to make errors reproducible. If that object is a command carrying a customer's contact details, the serialized dump carries them too.
  • Monolog context arrays and exception messages. $logger->error('Failed to update customer', ['customer' => $customer]) looks harmless until a Monolog processor or formatter stringifies $customer. Free-text exception messages built with string interpolation are worse — the PII is inside prose, not a structured field, so no key-based filter will ever find it.
  • Symfony Messenger retries and failure transports. An asynchronously transported message is serialized whole for transport. When handling fails, Messenger can retry it; after the configured retry limit, it can send the serialized message to a failure transport, if one is configured. That failure transport may be backed by Doctrine, Redis, AMQP or any other transport, and it has its own retention and access-control characteristics — usually different from your primary database’s.
  • The Symfony Profiler and request/response logging. The Profiler can persist request data and other debugging context collected for a request — exactly what depends on which data collectors are active. In a normal Flex application it is a development tool, and Symfony explicitly warns against enabling it in production — but a “temporarily enable profiler in prod to debug an incident” is exactly the kind of decision made under pressure, and it silently reopens this surface for as long as it stays on.
  • Tracing, APM and error-reporting tools. Distributed tracing spans and error-reporting SDKs (Sentry-style breadcrumbs, APM request tags) often capture request/response bodies or exception context by default, because that is what makes them useful for debugging. That default was not written with your PII inventory in mind.
  • Web-server, reverse-proxy and gateway logs. Personal data placed in a query string, a URL path or selected headers can be captured before Symfony sees the request at all — by nginx, a load balancer, an API gateway or a CDN. No Monolog processor can redact a log line that was written upstream of the application.
  • PII embedded in filenames, cache keys, labels and metrics. A cache key built as "customer_export_{$email}" or a metric label tagged with a customer id ends up in infrastructure — cache dashboards, metrics backends — that was never scoped as a place personal data lives.

None of these paths involve a developer typing a customer’s email into a log statement on purpose. Each one is a generic, reasonable piece of infrastructure doing exactly what it was built to do: capture enough context to debug a failure. The context happens to include personal data.

Allowlist instead of denylist

The fix that actually closes the gap inverts the default. Instead of trying to enumerate every field, message, and object shape that might carry PII and strip it, define what is safe to log and treat everything else as unsafe by default.

final readonly class SafeLogContext
{
private function __construct(
private array $fields,
) {
}

public static function forOrder(
OrderId $orderId,
CustomerId $customerId,
string $status,
): self {
return new self([
'order_id' => $orderId->toString(),
'customer_id' => $customerId->toString(),
'status' => $status,
]);
}

public function toArray(): array
{
return $this->fields;
}
}

SafeLogContext has no constructor path that accepts a raw entity, a raw command, or a raw string. Every field that goes in was chosen deliberately, by name, by a developer who wrote the specific named factory method for that log site. There is no generic fromEntity() escape hatch — adding one would reopen exactly the hole this class exists to close.

The second half is a Monolog processor that acts as a backstop for the paths SafeLogContext doesn't reach — third-party bundles, framework-level log lines, anything not yet migrated:

final readonly class PersonalDataRedactionProcessor
{
private const array DENYLISTED_KEYS = ['email', 'phone', 'fullName', 'address', 'cardHolder'];

public function __invoke(LogRecord $record): LogRecord
{
$record->extra['context_redacted'] = $this->redactRecursively($record->context) !== $record->context;

return $record->with(context: $this->redactRecursively($record->context));
}

private function redactRecursively(array $context): array
{
$redacted = [];

foreach ($context as $key => $value) {
$redacted[$key] = match (true) {
in_array((string) $key, self::DENYLISTED_KEYS, strict: true) => '[REDACTED]',
is_array($value) => $this->redactRecursively($value),
is_object($value) => '[object:' . $value::class . ']',
default => $value,
};
}

return $redacted;
}
}

The order of the match arms matters. Key-based redaction runs first, so a denylisted field is removed regardless of whether its value is a scalar, an array or an object — ['email' => ['value' => 'alice@example.com']] becomes ['email' => '[REDACTED]'] instead of being traversed into a nested key the denylist has never heard of. For non-denylisted fields, arrays are traversed recursively and any object is replaced by its class name rather than serialized, which is what closes the "nested DTO with an unexpected field name" hole a pure denylist leaves open. The processor never calls json_encode or the Serializer on application objects, so there is no reflection-based path for a new property to leak through unnoticed. In Monolog 3, core record fields such as context are readonly and are replaced with LogRecord::with(), while extra stays intentionally mutable for processors — which is why the redaction marker goes into extra directly. Combined with SafeLogContext as the primary path and this processor as a backstop, a class name in a log line becomes a signal that something bypassed the allowlist — a fact worth alerting on, not silently accepting.

The same allowlist discipline applies to identifiers used outside the log context array. A fingerprint or a technical id — the order id, a hashed correlation id — belongs in a log line, filename, or cache key. The raw value that the fingerprint stands in for does not. If two log lines need to be correlated to the same customer without ever printing who that customer is, a stable pseudonymous identifier carries that correlation; article #3 in this series builds the same shape for a different purpose — a blind index that supports lookups without exposing the underlying value.

Three ways this still fails

The processor runs after the leak, not before it. A Monolog processor redacts the log record before it is written to a handler — but if a third-party bundle logs directly to a file, syslog, or an external SDK that bypasses the Symfony logger channel entirely (many APM and error-reporting integrations register their own hook alongside Monolog, not through it), the processor never sees that data. The allowlist on the primary path is the actual control; the processor is a backstop for what still flows through Monolog, not a universal filter.

Messenger’s failure transport still needs its own decision. Redacting what gets logged about a failed message does not redact the message itself sitting in the failure transport’s storage, if a failure transport is configured — that is a separate retention and access-control surface, and it needs its own answer (a bounded retention period, restricted access, or a normalizer that strips or fingerprints PII fields before a command reaches the bus). Treating “we redact our logs” as if it also covers “we retain full commands in a failure queue for 30 days” is the gap that turns an otherwise careful implementation into a false sense of safety.

The Profiler and debug tooling are an environment-configuration risk, not a code risk. No amount of SafeLogContext discipline in application code controls what an enabled Profiler and its active data collectors capture. If an emergency-debug procedure enables it in production, treat the resulting profiles as another potentially sensitive data store and verify exactly which collectors are active. This is a deployment and configuration concern — verify framework.profiler.enabled is false in the prod environment, and audit whatever emergency-debug procedure exists for turning it on, because that procedure is exactly when this control gets bypassed under pressure.

What this guarantee covers, and what it does not

An allowlist-based logging discipline gives you a testable boundary: the log context this code writes contains only fields a developer named on purpose. It does not give you:

  • A universal sanitizer for every string that reaches Monolog. A raw exception message built by string interpolation can still carry PII in prose; the fix there is discipline in how exception messages are constructed, not a filter that reads natural language.
  • Coverage of tools that bypass the Monolog channel. Third-party SDKs with their own logging or breadcrumb collection need their own PII configuration, checked against that vendor’s documentation.
  • A guarantee about historical log data. This closes the leak going forward; data already written under the old, unguarded logging calls is still sitting in your log storage and needs its own retention or purge decision.

The mental model that holds all of this together: logging is a second data store you did not design, with none of the access control, encryption, or retention policy you built for the first one — so anything that reaches it should be something you chose to put there, not something a generic serializer decided to include.

A useful, cheap regression check: seed a test fixture with a distinctive, obviously-fake marker string in a PII field — something like zz-pii-marker-do-not-log-zz — run it through the request paths that are most likely to leak (a failing command, an exception handler, a Messenger failure), and grep the resulting log output for that marker. It will not catch every leak this article describes, but it turns "we think we redact PII" into a test that fails the day someone adds a new debug processor that dumps the whole command object.

TL;DR

  • Encrypting a database column does not stop the same value from reaching logs, traces, the Messenger failure transport, or the Profiler in plaintext — those are separate storage systems with separate retention and access control.
  • A denylist of field names (email, phone) misses PII embedded in prose exception messages, nested objects with unexpected property names, and anything serialized wholesale by a debug tool.
  • Invert the default: build log context from a named allowlist (SafeLogContext) where every field is chosen deliberately, and use a Monolog processor that redacts by key and replaces any stray object with its class name as a backstop, not the primary control.
  • Use a fingerprint or technical id for correlation, never the raw PII value, in log lines, cache keys, and filenames.
  • Messenger’s failure transport, the Profiler in production, third-party APM/error-reporting SDKs, and upstream web-server, proxy and gateway logs are separate surfaces that a Monolog processor cannot reach — each needs its own explicit decision.
  • A test that seeds a distinctive PII marker and greps log output after a failing request is a cheap, durable regression check — it will not catch everything in this article, but it catches the next accidental full-object dump.
  • The goal is not a universal sanitizer. It is minimization: nothing reaches the log unless a developer decided, field by field, that it belongs there.

Where has a logging or tracing tool surprised you by capturing more than you expected — and did you find it by grepping logs, or did a compliance review find it for you?

Part of the Privacy Engineering with Symfony and Doctrine series, which develops the Privacy Architecture pillar article.

Previous posts: Doctrine Lifecycle Listeners Are a Dangerous Place for GDPR Erasure — the write-path counterpart to this post’s read/observability path. Next: how I prove a Symfony application really forgot a customer, including the same allowlist discipline applied to test assertions.