What the percentages mean

A story page shows two different percentages, and they are deliberately given two different names.

“% identical” sits beside each outlet's filing in the coverage list. It answers one narrow question: how closely does this outlet's wording resemble the filing it is listed under?

“% related” sits in “Part of a bigger story” at the foot of the page. It compares this whole story to another whole story, to point you at an ongoing situation worth following.

Both are measures of text — not of truth, quality, or bias.

The short version

A run of very high numbers — most outlets in the 90s — usually means they are working from the same wire copy. Pakistani outlets carry a great deal of agency material (APP, Reuters, AFP), and a wire story often reaches a dozen sites with the headline barely touched.

A spread of lower numbers — 60s and 70s — usually means the outlets wrote it themselves: different angles, different details led with, different emphasis.

That is the useful signal. When every number on a story is near-identical, you are reading one source repeated, not several newsrooms independently confirming something. When they vary, you are seeing genuinely separate accounts, and the differences between the headlines are worth your attention.

How it is calculated

Every article's headline and short excerpt — never the body, which we neither store nor republish — is converted into an embedding: a list of numbers representing the text's meaning, produced by a language model. Two pieces of text that say the same thing in the same way land close together; two that differ land further apart.

The percentage is the cosine similarity between two of those lists, expressed out of 100. Identical wording approaches 100%. Unrelated text sits near 0%.

Each development is scored against a reference filing: the most recent article in that group, shown at the top with no percentage of its own. Every other filing in the group is compared to it. Where a story has several developments over days, each one is scored against its own reference rather than against something from three days earlier.

Why two names. “% identical” compares one article to another article, so a high number really does mean near-identical text. “% related” compares one whole story to another whole story, averaging everything in each — a much looser claim, which is why 70% related is a useful suggestion while 70% identical means two outlets wrote it quite differently. Same underlying arithmetic, different scale, so they must not be read against each other.

What it does not tell you

  • It is not a quality or accuracy score. A low “% identical” means an outlet wrote it differently, not better or worse. A high one does not make a report wrong.
  • Carrying wire copy is normal practice, not a fault. Agencies exist to be republished, and a wire report is often the fastest accurate account available. The number tells you how many independent accounts you are actually reading — what you make of that is yours to decide.
  • It compares wording, not claims. Two outlets can word a story almost identically and still differ on a crucial number, and the score will not notice. It is a starting point for reading the headlines side by side, not a substitute for it.
  • The reference can move. It is whichever filing is newest when the page is built, so if an outlet files late, the reference changes and every percentage in that group shifts with it.
  • Short or missing excerpts skew it. An outlet publishing a bare headline has less text to compare, which tends to push its score away from the rest for reasons that have nothing to do with its journalism.

Why we show it at all

HarZaviya exists to show the same story from every angle. Part of that picture is how many genuinely distinct angles there are. A story covered by twelve outlets in twelve near-identical sentences is a different thing from one covered by twelve newsrooms that each went and looked — and until now, a list of twelve headlines made those two look the same.

More on the project in About HarZaviya, or how the crawler behaves in About the bot.