Email verification tests often fail for a reason that looks like timing, but is really identity. The test receives a message, finds a token, and assumes that token belongs to the current run. A busy CI inbox can quietly disprove that assumption.
This article describes a small pattern I use in Playwright work: every test must claim one message before it asserts anything about the message. The claim gives the test a boundary, a cursor, and evidence that can be inspected after a failure.
The flaky test is usually a correlation problem
Consider a signup test that creates a user and waits for a verification email. A first version might poll until the inbox contains a message with a matching subject, then click the first link it finds.
That looks reasonable, but several messages may match:
- A previous retry may have delivered its email late.
- Two workers may be testing the same subject at the same time.
- The provider may return messages in an order that is not the delivery order.
- A broad subject match may accept an email for another user.
The result is a test that passes on a laptop and fails in CI. The assertion is not necessarily wrong. The test has not proved which message it is reading.
The same risk appears when a team searches for a facebook temp email or temporary disposable mail and then reuses the resulting address across several scenarios. The address is a fixture, not an identity system. It still needs a run boundary.
Define a message claim before opening the inbox
Before sending the signup request, create a run identifier and save the time at which the request starts. Put the identifier in data that the application sends, when possible:
const runId = `pw-${test.info().parallelIndex}-${Date.now()}`;
const email = `qa+${runId}@example.test`;
If the product cannot echo a run ID, use the strongest combination available: recipient, a unique part of the subject, sender, and a message timestamp after the request began. Do not make the claim depend on the visible order of an inbox list.
A useful claim record looks like this:
type MessageClaim = {
runId: string;
recipient: string;
messageId: string;
receivedAt: string;
};
The polling operation should return a MessageClaim, not just a body string. This is a small design choice, but it save a lot of diagnosis later.
Implement the claim in a Playwright fixture
Keep inbox access behind one fixture or helper. The helper should poll with a deadline, filter candidates, and claim exactly one message. It should also report what it saw when the deadline expires.
async function waitForClaim(
inbox: InboxClient,
query: { recipient: string; runId: string; requestStartedAt: string },
timeoutMs = 30_000,
): Promise<MessageClaim> {
const deadline = Date.now() + timeoutMs;
while (Date.now() < deadline) {
const messages = await inbox.list({ recipient: query.recipient });
const match = messages
.filter((message) => message.receivedAt >= query.requestStartedAt)
.find((message) => message.subject.includes(query.runId));
if (match) {
return {
runId: query.runId,
recipient: query.recipient,
messageId: match.id,
receivedAt: match.receivedAt,
};
}
await new Promise((resolve) => setTimeout(resolve, 500));
}
throw new Error(`No verification message claimed for ${query.runId}`);
}
The timestamp must be captured before the action that sends mail. The important part is the contract: no message ID, no verification step.
Do not hide this logic inside every test. A repeated polling loop drifts quickly. It also become hard to tell whether a failure came from the application, the inbox adapter, or the test itself. A shared fixture gives the team one place to add backoff, logging, and cleanup.
For larger CI suites, it is worth reading about how to isolate email fixtures in CI. If the inbox is exposed through an API, reusable email API checks can keep transport assertions out of browser scenarios.
Keep evidence separate from the assertion
After claiming a message, save a small redacted receipt as a test attachment:
await test.info().attach("email-claim.json", {
body: JSON.stringify({
runId: claim.runId,
messageId: claim.messageId,
receivedAt: claim.receivedAt,
}),
contentType: "application/json",
});
Then fetch the claimed message by ID and assert the verification link. Avoid attaching the entire email when it can contain personal data or tokens. A receipt should help answer “which message did we use?” without becoming another secret store.
I also keep fixture labels boring and explicit. Names such as temp org mail or temp mailid may be useful while exploring, but they make failure output harder to search. A run ID that says what it is is much kinder to the next person on call.
A practical reliability checklist
Before merging a browser test that reads email, check these points:
- Does each worker get a unique recipient or correlation value?
- Is the request-start time captured before the action that sends mail?
- Does polling have a hard deadline and a useful failure message?
- Is one message claimed by ID before its body is parsed?
- Are old messages excluded without relying on list order?
- Are tokens and message bodies redacted in artifacts?
- Does cleanup happen even when the test fail?
For the last question, use a fixture teardown or try/finally. Cleanup is not just housekeeping: abandoned addresses and messages can change the next run’s matching behavior.
Questions that come up in review
Should the browser test poll the inbox directly?
Usually no. Put provider-specific behavior in an inbox adapter and let Playwright consume a small, stable contract. This makes the UI test easier to read and lets API-level tests cover polling details.
Is a unique recipient enough?
It is better than a shared address, but not always enough. Delivery retries, provider delays, and reused accounts can still produce ambiguity. Claim by a stable message ID after filtering with a run-specific value.
What should happen when email is delayed?
Fail with the run ID, recipient, deadline, and a bounded list of observed message metadata. Do not increase the timeout forever. A longer wait can hide a broken delivery path and make the whole suite slow.
Email verification becomes much less mysterious when the test treats each message like a resource that must be claimed. The browser still proves the user flow, while the claim record proves that the flow used the right evidence.
Top comments (0)