Skip to content
GuideEvidence & Myths12 min read

Nobody can predict which post takes off

Two studies tested how predictable a viral post is. Early signals help a little, then the future goes dark. This is the rigorous answer to why one post blew up and the identical one did not.

Aubrium ResearchEditorial ·

You posted two reels last week. Same niche. Same effort. Same time of day. One got 2,000 views. The other got 200,000. You have watched them both a dozen times. You cannot see what made the difference.

So you go looking for a pattern. Someone will sell you one. Post at this hour. Use this hook style. Get early engagement in the first thirty minutes. Hit the algorithm right and it will carry you. Miss it and you are done.

Here is what the research actually says. Early signals help. A little. Better than guessing. Nowhere near certain. Then the room goes dark. Not because we lack data. Because there is a mathematical ceiling on how predictable social outcomes can be. Even with perfect information.

Two studies show this. The first built a virality predictor on Facebook's own cascade data. The second proved a hard limit on prediction even when you know everything. One ran in 2014. The other in 2016. Both were published at WWW. That is the Web Conference. It is run by the International World Wide Web Conferences Steering Committee. Both are about the same question. Why did that one blow up. And both land in the same place. You can predict a little. Then you hit a wall.

What follows reads those two papers and one other study. Every source is linked, quoted and dated. One section describes the ceiling on what early engagement can ever do.

The first study: early signals help, but the same post can do anything

Justin Cheng, Lada Adamic, P. Alex Dow, Jon Kleinberg and Jure Leskovec published Can Cascades be Predicted? in 2014. All five were at either Facebook, Cornell or Stanford at the time. The paper does not name their employers. The arXiv page names their universities only. So we report what the page states and nothing more.

A cascade is one post and everyone who shares it. The shares spread. Like a wave. The paper tracked 150,572 photo cascades on Facebook. Each one had at least five reshares. Together they logged over nine million reshares. The data came from June 2013. The authors watched each cascade for 28 days.

They built models to predict whether a cascade would double in size. They trained those models on the first half of a cascade's growth. Then tested them on the second half. The result: 79.5% accuracy. The baseline, which is guessing, sits at 50%. So the models beat guessing by a lot. But nowhere near certainty.

The paper names what helped most. Temporal features won. That means facts about time. How fast the reshares arrived. Whether they came in a burst or a trickle. Their summary:

reshares [that] come quickly are more likely to grow significantly — Cheng et al., Can Cascades be Predicted?, 2014

Structure mattered too. A cascade's shape is a tree. One post at the root. Reshares branch off it. The paper measured breadth versus depth. A shallow, wide tree beat a deep, narrow one. Shallow and wide means many people reshared the original. Deep and narrow means people reshared the reshares. Width predicted growth better than depth.

One finding cuts through everything else in the paper. The authors took 983 clusters of identical photos. The exact same image. Uploaded at different times by different people. Ten copies per cluster, on average. They asked: can you pick the biggest one in advance? You get to see all ten after a few reshares. Pick the winner.

The models picked right 49.7% of the time. The baseline is 10%. One guess out of ten. So the models did better. But barely. The same photo, posted by different people, did wildly different things. The paper puts a number on that spread. The Gini coefficient. It measures inequality. Zero means every upload did the same. One means one upload got everything. Across the 983 clusters, the average Gini was 0.787. Hugely unequal. Same content. Completely different outcomes.

The authors state this as a finding, not a puzzle:

even the same photo uploaded at different times by different users can fare dramatically differently — Cheng et al., Can Cascades be Predicted?, 2014

And they add that the variation and the predictability sit together. You can predict better than random. And the variance stays enormous.

The paper names its own limits. It ran "entirely with Facebook data and only with photos". Photos on Facebook in 2013. Not reels on Instagram in 2026. The mechanism might be the same. The numbers are not portable.

The second study: even perfect information hits a ceiling

Travis Martin, Jake Hofman, Amit Sharma, Ashton Anderson and Duncan Watts published Exploring limits to prediction in complex social systems in 2016. All five were at Microsoft Research when the arXiv copy was posted. The paper does not state this. The arXiv page does not either. So we do not claim it as fact.

They studied Twitter. They built models to predict how big a cascade would grow. They had unprecedented access to data. User features. Content features. Past performance. The works. Their best models explained less than half the variance in final cascade size.

Less than half. With more data than anyone outside the platform could ever get.

That finding alone would be useful. But the paper goes further. It asks a harder question. Say you had perfect information. Every fact about every user. Every past action. No noise. No missing data. Could you predict cascade size then?

The answer is no. Not fully. There is a ceiling. It comes from two things. The paper calls them system homogeneity and ex-ante knowledge. Homogeneity means everyone in the system behaves the same way. Ex-ante knowledge means you know the quality of the content before anyone sees it. In the real world you have neither. People behave differently. Content quality is uncertain until it gets tested.

The paper proves this with math and then tests it with simulations. Small uncertainties in quality. Slight variation in how people respond. Both create a ceiling on predictability. And that ceiling sits well below perfect prediction. The authors' own words:

realistic bounds on predictive accuracy require both system homogeneity and perfect ex-ante knowledge — Martin, Hofman, Sharma, Anderson & Watts, WWW 2016

They did not get either. So the bound was low. And they argue it would be low for any complex social system. Not just Twitter. Any platform where people influence each other and the thing being shared is new.

What this means for early engagement

Put the two together. Cheng's 2014 study says early signals help. Faster reshares. Wider initial spread. You can see those in the first few hours. And they correlate with bigger final size. That is the evidence for caring about early engagement. It is real. It is measured. It is published.

But Martin's 2016 study says there is a ceiling. Even with perfect data. Even if you knew exactly how good the content was before anyone saw it. Even if everyone responded the same way. You still could not predict the outcome with certainty. The system is too complex. Small differences compound. Randomness stacks.

So the case for early engagement is real. It is also bounded. Early signals help. The same research that proves they help also proves there is a ceiling on how much they can ever matter. You can push the odds a little. You cannot make an outcome certain. What is left, once prediction runs out, is measurement. Instagram ships one tool for exactly that: Trial Reels publishes a reel to non-followers first and returns view metrics at 24 hours. It does not forecast an outcome. It produces one, on a sample audience, before your followers see the post.

Cheng's 79.5% accuracy is the number to hold. It is not 50%. It is also not 100%. That 20.5% error rate is not measurement noise. It is not a gap that more data would close. It is the room where randomness lives. The identical photo uploaded ten times lands all over the map. Early engagement might help it land higher. It cannot make it land in one spot.

Why the same post does different things every time

Three forces explain this. The first is the audience. Your followers do not all see the post. A platform picks who sees it. Based on when they open the app. What else is competing for their feed. Whether they usually engage with your type of content. You do not control that. So the same post gets shown to different people each time. Different people respond differently.

The second is timing. Not the hour you post. The moment you land in someone's feed. Say you post at 6pm. One follower opens the app at 6:02. Your post is at the top. Fresh. They watch it. Another opens at 6:45. By then fifty other posts have arrived. Yours is buried. They never scroll that far. Same post. Same person. Different outcome. Because of when they happened to look.

The third is the cascade itself. Once a few people engage, the platform's models read that as a signal. They might show the post to more people. Or they might not. The decision depends on how those first few people compare to the model's guess about everyone else. One person skipping after two seconds is a negative signal. It gets weighed. Against what, you do not know. The weights are not published. And even if they were, you could not control who sees it first.

Those three forces interact. The audience the platform picks affects the early signals. The early signals affect who sees it next. Who sees it next affects the cascade's shape. The shape feeds back into the model. Every step depends on the one before. And the one before had noise in it. So the noise compounds.

Salganik, Dodds and Watts proved this in 2006. They built a fake music market. Participants chose songs to download. Some participants saw what others had chosen. Some did not. The ones who could see other people's choices showed more inequality. The winners won bigger. And the winners in one world were not the winners in another. Eight parallel worlds. Same 48 songs. Eight completely different charts. More information made outcomes less equal and less predictable. Not more.

Their summary:

Increasing the strength of social influence increased both inequality and unpredictability of success. — Salganik, Dodds and Watts, Science, 2006

That paper ran in a lab. Not on a real platform. But its finding matches what Cheng found on Facebook. And what Martin found on Twitter. Outcomes spread wide. Even when the content is identical.

Where you can predict, and where you cannot

Cheng's paper shows the limits clearly. Prediction accuracy goes up as the cascade grows. It is easier to predict whether a cascade that already has 25 reshares will get another 25. Harder to predict whether one with five reshares will get to 25. Why? Because the one with 25 has a history. You can read it. The one with five barely exists yet.

The paper also shows that large cascades are less predictable than small ones. That sounds backwards. But it is not. A post that will get exactly 20 reshares is easier to predict than one that might get 20 or might get 20,000. The huge outliers are the hardest to see coming. Because their path depends on a cascade catching. And whether a cascade catches depends on who sees it, when, and whether they share it to the right people at the right time.

So you can predict the middle. You cannot predict the tails. You can predict a little better as time passes. You cannot predict well early. And you can never predict perfectly. Even with Facebook's own data. Even with Twitter's. Even with everything.

This category checks the claims creators are handed against what the research and the platform documents actually say. These six extend the argument above.

What we could not determine

Three things we could not verify from the sources we read.

The first: sample size for Martin 2016. The abstract gives no count of how many Twitter cascades they studied. The paper is available at arXiv. We read the abstract only. We could not open the full text. So we report the finding without the sample size.

The second: whether the Cheng 2014 result holds on Instagram. The paper ran on Facebook photos in 2013. Instagram reels in 2026 are a different product. Different format. Different distribution model. Different audience. The mechanism might carry over. We cannot prove it does.

The third: whether buying early engagement makes a measurable difference to final outcomes. No published experiment we could find tests this on Instagram. Not Cheng's paper. Not Martin's. Not any other. The mechanism is plausible. The proof does not exist.

How we did this

Scope: this post reads three published papers and reports what they say. No data collection of our own. No experiment.

How we read them: Cheng et al. read from the arXiv HTML page at arxiv.org/html/1403.4608v1, retrieved 18 August 2026. Martin et al. read from the arXiv abstract at arxiv.org/abs/1602.01013, retrieved 18 August 2026. Full text not opened. Salganik et al. cited from the findings in the earlier post You are a cold item, which quotes it directly. Every quotation above is verbatim. No source was paraphrased.

Limits: we could not open the full text of Martin et al. So we report only what the abstract states. We cite no sample size for it. Cheng et al. is Facebook data from 2013. We do not claim its numbers carry to Instagram in 2026.

Sources

  1. Justin Cheng, Lada Adamic, P. Alex Dow, Jon Kleinberg and Jure Leskovec, Can Cascades be Predicted?, arXiv:1403.4608v1, 2014. Presented at WWW 2014, the 23rd International World Wide Web Conference. Sample: 150,572 photo cascades on Facebook with at least 5 reshares each, totaling 9,233,300 reshares, June 2013, observed for 28 days. Method: machine learning models trained on cascade features (temporal, structural, user, content) to predict doubling in size. Read from arXiv HTML at arxiv.org/html/1403.4608v1, retrieved 18 August 2026.
  1. Travis Martin, Jake M. Hofman, Amit Sharma, Ashton Anderson and Duncan J. Watts, Exploring limits to prediction in complex social systems, arXiv:1602.01013, 2016. Presented at WWW 2016, the 25th International Conference on World Wide Web. DOI: 10.1145/2872427.2883001. Platform: Twitter cascade size prediction. Sample size not stated in abstract. Method: predictive models tested against theoretical bounds on predictability in complex social systems. Read from arXiv abstract, retrieved 18 August 2026. Full text not opened.
  1. Matthew J. Salganik, Peter S. Dodds and Duncan J. Watts, Experimental study of inequality and unpredictability in an artificial cultural market, Science 311(5762):854–856, 2006. DOI: 10.1126/science.1121066. Sample: 14,341 participants, 48 songs, eight parallel social-influence worlds. Findings cited from the earlier post You are a cold item, which quotes the paper directly with full citation.

Changelog

  • 2026-08-18: First version drafted.
viralityresearchevidenceengagementalgorithm