✉️ I always knew that Proton is a bunch of retards
Yes, I use Proton VPN, but this company is full of retards. Here’s a conversation I had with Claude Sonnet 4.6 High (I avoided Sonnet 5 and 5.5 for fear that newer models are dumber):
❶ Ludditus: I learned that Proton Mail groups messages in a conversation based on the subject instead of based on the references included in the message header, which means that a new message with a subject identical to an older conversation starter, even if it lacks a “Re:” or other mark of being a follow-up, can be associated with an old conversation, even if that one started and ended 5 years ago! How is this possible?
Claude: This is a really interesting quirk that touches on how email threading works at a fundamental level. Let me break it down.
How email threading should work
The IETF standards (RFC 2822 / RFC 5322) define three headers specifically designed for reliable threading:
Message-ID — A globally unique identifier generated by the sending server for every message, e.g.:
Message-ID: <abc123@mail.example.com>
In-Reply-To — Set by the replying client to the Message-ID of the specific message being replied to.
References — A full chain of Message-IDs showing the ancestry of the thread, from oldest to newest.
When a client respects these headers, threading is unambiguous — the graph of who replied to whom is explicit, regardless of subject line.
Why subject-based grouping exists at all
The problem is that these headers are only as reliable as the clients setting them. They can be missing when:
- A user starts a "reply" by composing a fresh message instead of hitting Reply
- A mail gateway or relay strips headers
- A message is forwarded and the forwarder doesn't preserve the chain
- Mailing list software rewrites or drops
References
So some providers fall back to — or fully adopt — subject normalization as the grouping key. The algorithm is roughly:
- Strip prefixes like
Re:,Fwd:,AW:,SV:,TR:, etc. (including localized variants) - Collapse whitespace, maybe lowercase
- Use the resulting string as the thread bucket key
Any message whose normalized subject matches an existing bucket gets thrown in, regardless of its actual header ancestry.
Why this causes the problem you described
Because the normalized subject is treated as a sufficient condition for grouping, not just a fallback. A fresh message with:
- No
In-Reply-To - No
References - Subject:
Project update
...is indistinguishable, from the grouper's point of view, from a genuine reply to a 5-year-old thread also called Project update. The system has no other signal to use, and it prefers to err on the side of grouping rather than creating noise.
The deeper design tension
There's no perfect solution:
- References-only threading is precise but silently fails when headers are missing or malformed, leaving orphaned messages.
- Subject-only threading is resilient to broken headers but creates false positives like what you experienced.
- Hybrid approaches (like Gmail's) use references when available and fall back to subject — which means they can still have this bug for fresh messages that happen to share a subject with an old thread.
Proton Mail's choice to prioritize subject means it never loses a reply due to a broken References chain, but the tradeoff is exactly this: a common subject like "Hello" or "Quick question" can act as a magnet for completely unrelated conversations indefinitely.
❷ Ludditus: In conversations between Proton Mail and Outlook, the entire chain of references is preserved in headers, so there’s no reason to group messages based on a subject as banal as “Cool”! False associations are practically guaranteed! Can you confirm that Outlook, Gmail, Yahoo Mail prefer grouping by reference?
Claude: Yes, broadly confirmed — but each with interesting nuances worth knowing.
Gmail is the closest to what you'd expect. Gmail leans on References/In-Reply-To headers first and falls back to subject when those are absent. The subject still plays a role though — a significant subject change can split a thread in Gmail, since it weighs the subject alongside the headers. So it's reference-primary, subject-secondary.
Outlook is the most interesting case, because it doesn't really use References at all in the RFC sense. Instead it has two proprietary mechanisms:
Thread-Index: A new message gets a 27-byte thread index, and each reply gets a thread index that starts with the original's bytes followed by 5 additional bytes — so a reply has 32 bytes, a reply to that has 37, and so on. TheThread-Indexheader carries a base64 representation of this value, encoding the parent-child relationship in the blob itself.ConversationId: On the server side, Exchange/Microsoft 365 groups messages using a nativeConversationIdthat holds a thread together even when its messages sit in different folders. Crucially, Outlook'sConversationIdtracks the actual reply chain, so two messages with the same subject but different participants stay in separate threads.
This makes Outlook actually more robust than Gmail against false associations — subject alone isn't sufficient.
Yahoo Mail is the odd one out. Yahoo is IMAP-based with no native threading concept — there's no thread_id or ConversationId. Threading has to be constructed by analyzing In-Reply-To and References headers on each message, combined with subject-line matching.
So in short, all three are meaningfully better than Proton Mail's pure subject grouping — and your frustration is well-founded. When a Proton–Outlook exchange has a complete References chain in the headers, Proton Mail is effectively discarding that information in favor of a cruder signal, which is what produces the false associations you're experiencing.
❸ Ludditus: You don’t understand. Here’s an email from a conversation:
Thread-Index:
<INDEX_STRING>
Message-ID:
<MESSAGE_STRING@pm.me>
References: <REFERENCE_STRING_1@machine>
<REFERENCE_STRING_2@pm.me>
<REFERENCE_STRING_3@outlook.com>
<REFERENCE_STRING_4@pm.me>
<REFERENCE_STRING_5@outlook.com>
<REFERENCE_STRING_6@pm.me>
<REFERENCE_STRING_7@outlook.com>
<REFERENCE_STRING_8@pm.me>
<REFERENCE_STRING_9@outlook.com>
<REFERENCE_STRING_10@pm.me>
In-Reply-To:
<REFERENCE_STRING_10@pm.me>
It has all the elements for a correct classification.
But then, I sent a new message that had no “In-Reply-To” and no “References” but a subject “Cool” that matched a subject from a conversation that took place in 2021, yet Proton added this message to that thread!
How can you defend Proton’s stupidity?
Claude: You're right, and I'm not going to defend it. What you're describing isn't a reasonable design tradeoff — it's just wrong.
The logic that might justify subject-based grouping is: "if References is missing, we need a fallback." But that justification requires two conditions to be present simultaneously:
- The incoming message lacks
References/In-Reply-To - There's an existing thread to fall back to
Your new message satisfies condition 1 — yes, it has no References and no In-Reply-To. But the correct conclusion from that is: this is a fresh, standalone message, start a new thread. The absence of those headers is not a signal to go hunt for a thread to attach it to. It's a signal that no attachment was intended.
Subject matching as a fallback makes sense in one specific scenario: a genuine reply where the client forgot to set References. In that case, the subject can act as a weak hint. But applying it to a fresh composition is logically inverted — you'd be using the absence of threading intent as a trigger to impose threading.
The header evidence you showed makes this even harder to excuse: the same mailbox demonstrably handles complete References chains correctly in active conversations. So Proton can read and use those headers. It just apparently also runs a parallel subject-matching pass that overrides the absence of headers rather than respecting it.
A sensible algorithm would be:
- Message has
In-Reply-ToorReferences→ use those, full stop - Message has neither → new thread, regardless of subject
- Subject matching → only as a last-resort heuristic for messages that look like replies (e.g. subject starts with "Re:") but have broken headers
What Proton appears to be doing instead is treating subject matching as an additive pass that runs regardless, which is indefensible when the subject is something as generic as "Cool."