Inbox placement testing: useful signals with honest limits

An inbox placement test sends mail to mailboxes you control and records where it appears. It is a diagnostic tool: a good result invites one more check, and no small sample can promise placement for every subscriber. Use it to compare evidence, not to certify delivery.

A row of overlapping brown padded envelopes on a white background.

Separate acceptance, placement and engagement

SMTP acceptance is a transport event: the receiving server takes responsibility for the message. RFC 5321: Simple Mail Transfer Protocol Placement asks a different question: did the observed mailbox put it in an inbox folder, a category, junk, or somewhere else? Engagement asks whether a real recipient wanted or acted on it. Keep those outcomes separate in your report.

Do not infer inbox placement from opens alone. Google says it does not track open rates, cannot verify third-party open-rate accuracy, and warns that low opens are not necessarily an accurate indicator of delivery or spam-classification problems. Google sender guidelines A pixel report cannot replace direct mailbox observation.

The three things a test can record are not the same. Acceptance means the receiver took the message at the SMTP boundary; placement means the mailbox system put it in a particular folder such as inbox or spam; engagement means what the human recipient did with it. A seed test measures the first two on the mailboxes you control. It can tell you a message was accepted and how that receiver categorised it, but it cannot tell you whether a real customer will open or reply.

Focus the test on the thing you can actually act on. If the concern is whether a message was rejected or deferred at the boundary, the SMTP reply and your platform logs answer that. If the concern is placement, the controlled test shows what that receiver did on that day. Knowing which of the two you are measuring stops a placement result being blamed for a rejection, or a rejection being used as proof of spam filtering.

Why seed results do not represent everyone

A seed list is a selected group of test accounts, not a random sample of your subscribers. Its provider mix, account histories and settings may differ from your audience. Treat any overall percentage as a description of the tested accounts, not a forecast for your next campaign.

Personal settings matter. Google explains that marking messages as not spam teaches Gmail how to handle mail addressed to that user; contacts and administrator bypass rules can also influence classification. Gmail spam preferences A long-used staff account with your address saved is therefore not equivalent to an unfamiliar subscriber’s inbox.

Authentication is another limit, not a shortcut. Google’s guidelines state that meeting requirements does not guarantee the expected delivery outcome and explain the effects of complaints and reputation. Google sender guidelines A passing test today does not demonstrate that a different sending route, later message or larger campaign will behave identically.

Seed accounts sit in a network a testing provider manages, so they warm up together with other test traffic and may be handled differently from your actual subscriber base. They also represent a small set of mailboxes and providers at one point in time. Geographies, new provider changes and the specific receiving judgement applied to your sender can all differ for real customers. A seed result is a sample with limits, not a census of your audience.

The providers and settings a seed panel can reach are a fixed set at a fixed moment, so results reflect that panel, not every customer's provider, device or app. A business that sends to a Gmail-heavy UK audience may see different behaviour from one serving subscribers across many small providers. Before treating a result as your reality, ask how closely the seed network resembles the providers, regions and mailbox clients your real recipients actually use.

Run a controlled test in order

  1. Write the question first. For example, ask whether messages from the new newsletter route reach your authorised test accounts without an authentication failure. Avoid vague objectives such as “prove our deliverability”.
  2. Choose relevant receiving services. Match the important parts of your audience as far as practical. Record which services and account types are missing. Keep consumer and organisation-managed accounts distinguishable.
  3. Document the accounts. Note contacts, safe-sender entries, forwarding, rules and prior interactions. Avoid silently changing these between tests. Do not disable security protections merely to produce a passing result.
  4. Use the real sending route. Send representative, non-sensitive content through the actual campaign system, From domain and link setup. A personal mailbox test is not a substitute for the newsletter platform.
  5. Record transport and placement separately. Save the send time, message identifier, SMTP outcome, observed arrival time and folder. Check authentication headers on received copies. Mark unresolved messages as “not observed”, not automatically “spam”.
  6. Change one suspected cause. If you correct authentication or replace a broken link, retain the earlier evidence and rerun a comparable test. Record changed settings and the observation window.

A useful test controls the variables you can control: send the same campaign version through the same route at a defined time, keep the audience and infrastructure constant, and record the baseline before the change you are evaluating. If you are testing whether a template or sending platform change moved placement, only that one variable should differ between runs. Sending two different builds through two different senders at once tells you which one changed, not which one improved.

Plan the test around a decision, not around the dashboard. Write down the question you are trying to answer, such as 'does campaign A place better than campaign B on Outlook for the same audience?', and which result would prompt action. Repeating a short test at the same point in the week, on the same route, gives you a steadier comparison than a single run, because filtering can vary by day and by sender activity. Preserve the raw result alongside the interpretation so someone can re-check it later.

Read patterns without inventing certainty

If several accounts at one provider put the message in junk, investigate that receiver’s guidance, authentication results and available reputation data. If only one account differs, inspect its rules and history before rewriting the campaign. These are diagnostic priorities, not proof of causation.

Report inbox categories separately where available. Do not label every non-primary category as spam. For missing messages, check the queue, rejection logs and search results before drawing a conclusion. SMTP delivery and a mailbox’s visible filing decision are different stages. RFC 5321: Simple Mail Transfer Protocol

Keep comparisons honest: record the accounts tested, observations collected, missing results and exclusions. Do not combine incomparable providers into a headline “deliverability score” without explaining its weighting and denominator. A small selected sample cannot establish an audience-wide percentage.

A poor placement result is a reason to investigate, not a verdict on your whole domain, and a good result is not a licence to ignore everything else. Before changing course, check whether the pattern is consistent, whether it matches the specific provider behaviour your own recipients use, and whether operational evidence such as replies, complaints and unsubscribe requests agrees. Confidence comes from several evidence channels pointing the same way.

A small number of seed mailboxes also means a single moved message can swing a percentage, so prefer the pattern across several runs and providers over one headline figure. Look for whether the outcome is stable for a route, whether it changes with the specific variable you altered, and whether it is reproducible before you redesign sending around it. Disciplined interpretation is what keeps a useful diagnostic from becoming a false alarm or a false reassurance.

free email deliverability checks can and cannot tell you

Use tests alongside operational evidence

Combine observations with provider-specific complaint trends, bounce responses and authentication monitoring. Google recommends monitoring spam rate and domain/IP reputation through Postmaster Tools. Google sender guidelines Where those reports are unavailable, state the gap rather than substituting a fabricated measurement.

Repeat a focused test after a material sending change or a documented incident, then watch real delivery outcomes. This guide describes a method, not a hands-on review of any testing product; no proprietary tools or claimed campaign results were used.

Because a seed network is a snapshot, pair it with the evidence that reflects your real traffic: how recipients interact with your mail, whether replies arrive, how complaint and unsubscribe rates move, and what your sending platform's own event logs show about deferrals and bounces. Testing tells you where to look next; operational measurement tells you whether the thing you are looking at is actually changing for your customers.

Keep a short record of each test: the date, the route, the campaign version, the seed panel, the raw placement outcome and what you decided next. Over a few months that log becomes evidence about your own sending that is more relevant than a generic best-practice number. When a real customer later reports a problem, you can check whether the controlled test already saw a warning sign and whether it has since moved.

an email deliverability tool cost guide

Sources and further reading