How a matched-volume contradiction pointed to a missing parameter — and what it turned out not to prove.
Key takeaways
Two cases sat side by side in the fintech corpus and refused to behave. Interactive Brokers in January 2024 and Webull in December 2025 were, on every dimension we normally check first, close to matched: comparable platforms, comparable markets, and — this is the detail that made the disagreement so hard to explain away — almost identical mention volume. 772 expressive mentions for one, 696 for the other. If mention count is doing the work we usually ask it to do, these two events should have produced roughly similar verdicts about how strongly each one moved its audience.
01
They did not. Interactive Brokers’ peak read as sharply negative by volume, −0.694 on the Balance of Expression. Webull’s read as mildly positive, +0.144. Two peaks, almost the same size, telling opposite stories. The instinct, at first, is to look for something wrong with the data — a classification error, a mismatched time window, a corpus artifact. We checked. The numbers held.
02
What forces a genuine methodological move is usually not a single anomaly — anomalies get explained away — but a matched pair that cannot be. Same order of magnitude, same general market conditions, opposite reading. If volume were the right lens, this pair should not exist. So the question became: what varies between these two events that volume, by construction, cannot see?
03
The answer, once it was framed correctly, felt almost embarrassingly simple: it is not how many people wrote something negative. It is how many people read it. A mention count treats a complaint posted by an account with twelve followers and a headline carried by a financial wire service as the same unit — one mention, one vote. But they are not remotely the same event from the audience’s side. One reaches a dozen people who already knew what they thought. The other reaches everyone deciding what to think.
04
This reframing carries a second, less comfortable implication. In most of our work, social media data plays the role of a symptom — a readout of an underlying reality, a thermometer for consumer sentiment that already exists elsewhere. But if reach determines how far a signal actually travels through an audience, then the media event is not only a symptom. It is also capable of being an independent variable in its own right — the first tile in the Signal Chain, capable of setting the rest of the sequence in motion rather than merely reporting on it. A monitoring practice built entirely on volume is, in that light, missing the one parameter that would let it distinguish a peak that merely describes public mood from a peak that is about to move it.
05
So we added reach — not as a replacement for the mention count, but as a second lens computed over the same peaks: Balance of Expression by volume, and Balance of Expression by summed audience reach within each tonal class. Applied to the matched pair, the disagreement dissolved into an explanation rather than a mystery. Interactive Brokers’ negative volume peak carried a reach balance of +0.945 — a swarm of small, low-reach complaints sitting underneath a much smaller number of high-reach items that were, on balance, positive. Webull’s mildly positive volume peak carried a reach balance of −0.968 — the mirror image. The two events were not actually contradicting each other. They were answering two different questions, and volume alone could only ever answer one of them.
January 2024 · 772 mentions
December 2025 · 696 mentions
06
The natural next move was to check whether this was a one-off rescue of an inconvenient pair or a general pattern, and it turned out to be a gradient running across the whole seven-entity corpus: platforms with large retail bases showed the same shape of disagreement — dense low-reach complaint volume diluted or reversed by a thinner layer of high-reach coverage — while the degree of reversal tracked the scale of the retail base and the nature of the triggering event. That was reassuring. It meant the hypothesis was not a coincidence built to explain two data points.
07
But the corpus had one more turn to take, and it was not the one we expected. The working assumption at that point was, roughly, that reach was simply the more honest measure and volume the naive one — that we had upgraded the instrument and could now trust the better lens across the board. Testing against downstream outcomes corrected that assumption immediately. Reach tracked market and fundamental performance closely: where financial results and equity movement could be checked, the reach-weighted balance called the direction correctly in every case, while the volume-weighted balance would have withheld or inverted that signal. So far, exactly as expected.
08
Consumer search behavior told a different story. The clearest activation of the informational-to-transactional search sequence we had specified in the Signal Chain — Coinbase, November 2024 — showed up in a peak that read as almost perfectly neutral by volume, +0.027, and strongly positive by reach, +0.781. The volume-neutral reading was not a failure of the metric. It was the exact mixed condition under which an audience goes looking for more information before forming a view — which is precisely what happened, on schedule, the following month. Reach, in that same case, would have predicted a muted reaction, because a strongly positive reading gives an audience little reason to go searching for anything. It was the least confident measure that called the most important behavioral consequence correctly.
Coinbase · November 2024
A volume-neutral, reach-positive peak — the mixed condition that sends an audience searching.
09
That was the surprise. Not that reach beats volume, and not that volume was quietly right all along — but that the two measures are not competing for the same job. Reach speaks to how an event registers with the audience that prices assets. Volume speaks to how an event registers with the audience that types into a search bar wondering what just happened. Neither one is the upgraded version of the other. Each is wired to a different downstream channel, and a monitoring practice that only ever asks ‘which number is bigger’ will misread whichever channel it wasn’t built to see.

About the author
Vadim Matyushkin
Behavioral Scientist & Sociologist · ICDS
Vadim Matyushkin is a psychologist and sociologist with close to twenty years of research into digital behavior, trust, and information dynamics, and a researcher behind the Institute of Communication and Data Science (ICDS). His work has supported organizations including Coca-Cola, PepsiCo, Mars, Danone, Nestlé, and Bayer in moving from self-reported survey data toward direct behavioral evidence.
ICDS — Institute of Communication and Data Science is an independent research institute focused on understanding how trust, behavior, and reputation are formed in digital environments. ICDS operates as an intellectual institution — not an agency, not a platform, and not an educational provider.