Pull a day of financial news for a single large company and the first thing you notice isn’t a lack of information — it’s the repetition. One earnings release, one product recall, one merger becomes eight, thirteen, nineteen articles that all describe the same underlying event. A newswire files it, three aggregators re-syndicate it, two outlets localize it, and a handful of “markets wrap” columns mention it in passing.
If you’re building anything on top of news — a screener, an alert, a backtest — that repetition is not harmless. It double-counts. Thirteen stories about one recall look like thirteen units of pressure on the stock when there was one. Any volume-of-coverage feature you compute is really a volume-of-syndication feature, and syndication is an artifact of the distribution network, not of the market.
So the first real job isn’t sentiment or entity extraction. It’s collapsing those thirteen back into one.
Why the two obvious approaches fail
The instinct is to dedupe on exact text. It doesn’t work: each syndicator rewrites the headline, appends a disclaimer, swaps a dateline, changes “Q2” to “second quarter.” Byte-for-byte matching catches almost nothing — the stories are the same event, not the same string.
The next instinct is to go fully semantic: embed every article, and treat anything above a cosine threshold as a duplicate. This is better at recall — it surfaces the near-duplicates that text matching misses — but it’s a dangerous thing to use as the actual keep/kill decision. Two stories can sit very close in embedding space and still be genuinely different events: “Company X recalls product A” and “Company X recalls product B” are near-neighbors to an embedding model and opposites to a trader. Lean on cosine alone and you will silently merge distinct events, and you won’t notice until a real signal disappears.
The lesson we keep relearning: embeddings are excellent for recall and unreliable as a verdict. Use them to gather candidates, not to pass sentence.
The shape of the method
What actually holds up is a layered collapse — cheap and precise first, expensive and fuzzy only where needed:
- Canonicalize the event, not the article. The unit we dedupe onto isn’t the story, it’s the (who · what · when) it describes — the entities involved, the event type, and the time window. Two articles collapse when they point at the same canonical event, even if every word differs.
- Cheap structural signals before any model. Shared primary entity, same event class, overlapping timestamps, matching numbers (the recall count, the deal size) — these are fast, deterministic, and rule out most non-duplicates before a single embedding is computed.
- Semantic similarity as a candidate net, not the judge. Embeddings pull the near-duplicates that survived the cheap filters; a stricter check then decides keep-or-kill, so a high cosine score never silently merges two different events.
- Pick a canonical head, don’t just “keep one.” Which of the thirteen survives matters — it becomes the representative every downstream feature reads. We prefer the earliest substantive report over the later wraps and re-syndications, and we keep the merge reversible so a wrong collapse can be undone.
None of these steps is exotic on its own. The result is: a company’s day of news stops being a pile of thirteen and becomes a handful of distinct events, each counted once.
What it’s worth
The payoff shows up the moment you measure anything. In our own data, a single Hasbro story cluster collapsed 13 → 1; a Binance episode went 19 → 1; and the long tail of “markets wrap” columns that merely name-drop a company — without the company being the subject — drop out of that company’s event count entirely instead of inflating it.
That last part is the quiet win. Once events are deduped and attributed to the company they’re actually about (not just mentioned in), “how much did the market hear about this company today” becomes a number you can trust — and everything built on top of it, from a noise-ratio metric to an impact score, inherits that trust.
MarketIV turns raw financial news into a deduplicated, per-stock impact graph — the direction, the reason, and the noise removed. You can poke at the live version, no signup, in the playground.
