
Name the outcome before naming the cause
SMTP distinguishes transient negative replies, permanent negative replies and successful hand-off semantics in RFC 5321. A 4xx deferral may remain queued, a 5xx reply rejects that transaction, and a 2xx acceptance transfers responsibility. None of these alone proves inbox placement or a message read.
Start with four categories: rejected during SMTP, deferred and still retrying, accepted but later bounced, or accepted with uncertain placement. A spam-folder report belongs to the last category, not to an authentication rejection. Mixing them leads to duplicate retries and irrelevant DNS edits.
Write the user-visible symptom and protocol outcome separately. “Customer did not see invoice” is a symptom; “receiver returned 550 after RCPT” is evidence.
Name the exact outcome before choosing a cause. The same words 'we are not getting mail' can describe mail that was never accepted, mail accepted but filtered, mail delivered to a folder nobody checks, or mail the recipient's own rule moved out of sight. Each of those has different evidence and a different fix, so writing down the observable state first prevents you repairing the wrong layer.
Preserve evidence before queues and logs rotate
Save UTC timestamp, sender and recipient domains, redacted addresses, message or campaign ID, sending service, sending IP, complete SMTP reply, delivery-status notification and current queue state. Keep the original text because provider-specific diagnostic links and identifiers matter.
For accepted mail, capture raw headers from the recipient where available. Authentication-Results fields record what a receiver or trusted intermediary evaluated, using the format defined by RFC 8601. Keep Received lines in order and preserve DKIM signatures. Redact subject, body, local parts, internal hosts and tokens before wider sharing.
Do not click unexpected “fix” links in a bounce. Open the known provider console and match the event by ID and time. If no corresponding send exists, treat the notification as suspicious.
Scope by route, receiver and time
Build a small matrix of mailbox, campaign, invoice, website and alert routes against affected recipient providers. Mark success, failure and untested. Add the first failure time and any DNS, provider, template, list or credential change near it.
Compare exact replies rather than dashboard labels such as hard or soft bounce. One invalid address is not a domain incident; every route failing after a selector change probably is. A problem isolated to one receiver requires that receiver’s evidence, while a route-wide fault points towards sender configuration.
Stop broad resending until queue state is known. The original service may still retry deferred messages, and a second campaign can create duplicates.
build a sender inventory firstRead identities, not only pass and fail
RFC 8601 shows how receivers communicate SPF, DKIM and DMARC results. Record the SPF-evaluated domain, DKIM d= domain and selector, visible From domain, reason text and any alignment result. A DKIM pass for a provider domain may not align with the sender’s From domain.
Compare with Google sender guidelines, which requires authentication and other controls for mail to personal Gmail accounts but does not promise inbox placement. A green DNS checker proves only what it queried at that moment, not which identity a real message used.
If Authentication-Results came from an untrusted forwarded copy, identify which server added it. Prefer evidence from the final receiver or your controlled test destination.
For a Gmail authentication rejection, the Gmail 550 5.7.26 investigation shows how to trace the exact failed route.
For a refresher on the SPF, DKIM and DMARC records an incident review keeps coming back to, see what SPF, DKIM and DMARC actually do.
Separate configuration, reputation, recipient and content causes
Authentication faults include missing authorisation, selector lookup failure, body-hash change and unaligned identities. Infrastructure faults include PTR, TLS, queue and shared-IP problems. Recipient faults include a nonexistent mailbox or local policy. Placement can also reflect complaints, permission, content and sending behaviour.
Do not prescribe simulated engagement, domain rotation or indiscriminate warming. These actions can conceal poor permission and erase the baseline. Authentication supports identity but cannot guarantee placement.
Choose the owner by mechanism: DNS owner for records, sending provider for shared infrastructure and queue behaviour, list owner for invalid recipients and consent, security for compromise, and privacy staff for data-handling concerns.
Make one reversible correction
- State the leading cause. Tie it to a reply, header, queue event or provider notice.
- Save the baseline. Export current DNS values, configuration and affected counts.
- Choose the smallest correction. Repair the selector, restore an authorised return path, suppress an invalid address or pause a compromised credential.
- Name rollback authority. Record the previous value and decision window.
- Apply once. Do not combine content, domain, IP and authentication changes.
- Wait for the relevant state. Respect DNS cache and provider queue behaviour rather than repeatedly editing.
- Test through the original route. Use the same service, From identity, template class and receiver type.
Share a minimal incident packet
A header can expose addresses, internal hostnames, IPs, route identifiers, mailing-list data and authentication tokens. A bounce can include message content. Create a redacted working copy and retain the original only in a restricted case store with an operational retention period.
Give providers exact UTC times, message IDs, sending IPs, recipient domains and replies. Replace personal local parts unless the provider genuinely needs them to locate the event. Never post full headers to a public forum or generic online analyser without reviewing its data terms.
For complaint and unsubscribe incidents, involve the privacy owner before circulating recipient-level evidence. Delivery troubleshooting does not suspend data minimisation.
Verify recovery against the same failure mechanism
For rejection, obtain a new accepted SMTP transaction through the same route and compare the reply. For authentication, inspect the receiving header and confirm the intended aligned identity passes. For deferral, confirm the original queue resolved or expired before authorising a new send. For placement, use controlled recipients and provider evidence without claiming universal results.
Check side effects: replies, links, attachments, unsubscribe, transactional completion and duplicate prevention. Monitor the next ordinary volume window for the same receiver and route. One successful administrator message cannot validate an invoicing platform.
Record changed value, time, tester, evidence and residual uncertainty. Close only the scoped incident, not every deliverability risk.
Recovery is confirmed when the agreed path passes on the same route and the original symptom disappears, not when a dashboard turns green. Re-check the actual folder the recipient reported, allow for DNS propagation time where a record changed, and keep watching for a full business cycle rather than closing on one delivered message. A single successful inbox test is evidence for that route, not a warrant that everything is fixed.
Use exact stop and escalation conditions
Stop sending when credentials appear compromised, complaints or invalid addresses are not being suppressed, critical mail is repeatedly rejected, or a queue may duplicate messages. Stop DNS editing when authoritative answers conflict, the owner is unknown or the proposed change cannot be rolled back.
Escalate shared-IP and PTR issues to the provider; suspicious access and malware to security; legal or rights issues to the privacy lead; and persistent receiver-specific rejection with a complete packet to sender support or a deliverability specialist.
Resume after the cause is evidenced, the one correction is verified on the same route, queues are reconciled, affected recipients are protected and the accountable owner approves release. If evidence contradicts the theory, restore the safe baseline where appropriate and reopen diagnosis rather than stacking another change.
Build a timeline that can disprove the theory
Put all events on one UTC timeline: last known success, first user report, first matching SMTP failure, DNS publication, provider configuration, campaign launch, credential event, queue retries and corrective action. State the source of every time because mailbox display time, application time and DNS-provider audit time may use different zones. A timeline prevents a later change from being mistaken for the original cause.
Write at least two competing explanations and the evidence each predicts. If a DKIM selector is missing, messages using that selector should fail lookup across receivers while routes using another selector may pass. If a recipient mailbox is invalid, failures should cluster on that address rather than every destination. If a shared IP is blocked by one receiver, other customers or domains on the IP may show related replies. Seek a fact that distinguishes explanations before changing anything.
Use a controlled test matrix with one variable per comparison. Hold the From domain and template constant while changing receiver, or hold receiver and message type constant while comparing routes. Never use a real recipient who has opted out, and do not repeatedly probe an invalid address. Mark tests with internal identifiers that do not expose incident details in the subject.
Stop the investigation run if new tests increase duplicate or privacy risk, the queue cannot be frozen, credentials may be compromised or logs are being overwritten. Escalate for preservation and containment first. Resume testing after access is secure, queues are known and the matrix has an accountable owner. A disproved theory is progress because it prevents an unsupported production change.
For customer-facing disruption, coordinate communication with the service owner. State which messages and time window are affected, whether retries remain queued, and what recipients should do without asking them to whitelist unsafe mail. Never claim that all delayed messages will arrive until queue evidence supports it. When recovery is verified, reconcile invoices, password links or notifications that may have expired and contact only affected people through an authorised route. This closes the business harm rather than merely clearing an SMTP graph, while avoiding duplicate or unsolicited replacements.
Archive a short incident record containing the confirmed cause, affected routes, correction, rollback position and verification artefacts. Add a preventive action only where it follows from the evidence, such as selector-expiry monitoring or queue ownership. Do not turn one receiver’s local response into a universal sending rule. Review the record with the route owner after normal traffic returns.
honest email blacklist monitoringSources and further reading
- RFC 5321: Simple Mail Transfer Protocol
- RFC 8601: Authentication-Results header field
- Google: Email sender guidelines