Learning by failure, Why you can't A/B test your tweets
I tried out some research that caused a ruckus. But you can learn a lot by causing a ruckus.
21 August 2026
Now here’s an idea. You want to post something on social media, a short piece of text - a tweet indeed, though we divorce it from it’s now fascistic bird-hole. You want your wickedly funny/poignant/scathing tweet to be liked and seen. You want to be liked and seen. After all, don’t we all? So you post it. Then you fall into a well of depression when no one looks or cares. You don’t like this. You decide you don’t want this embarrassment ever again. Now, you could just never tweet ever again, but that is too difficult for you, isn’t it. It would be much easier (?) to set up a semi-scientific methodology to determine what is the most optimal way to get your witty/touching/ranty drivel across.
This I believe is the average thought-process of the Bluesky user, though I cannot know for certain as I wouldn’t bother with it.
I, having barely read the introductory chapter of a book on analytical political communication (see my university programme) decided this was worth testing. To implement the dream of the hyper-online liberal Follow-Back-Pro-EU elite.
The book is titled Analytical Activism and explores how effective modern political campaigns posses a pervasive ‘culture of testing’ where digital output to engage and persuade is scientifically calculated to maximise the given amount of engaging and persuading.
The primary method is to A/B test, where you select a message you want to amplify (‘Join the campaign for people to actually read my blog’) and put out the same message with multiple variations (‘Develop some self-respect by helping me with my blog’, ‘TODAY is the last day you can help a recovering CS Student with his blog’) then only amplify the type of message that is tested to work.
This seemed a fascinating idea to me so I, perhaps haunted by a Computer Science background, decided to test it out. This seemed to me the best way to figure out how these things work. And I don’t think I was wrong.
So here was the plan: Have 5 different Bluesky accounts, post variations of the same messages randomly on each, see which is most successful. Oh, and put out a link explaining what I was doing on each profile so I felt less evil (lilpete.me/polcom).
I wrote up a little script to automate this, and set up the accounts. Each account had a profile picture of Big Ben and a name of the format ‘The Political X’ where X was some pretentious word. I decided they needed to be the same for some sense of a control variable, one account cannot be more appealing than the others.
Then, just write up the messages and test.
![]()
![]()
Some posts from one of the experiments.
Then, follow the same identical set of 460 UK Political profiles on each account to get any sort of attention.

Oh dear. I’ve done something wrong. It does look odd, doesn’t it.
And no one has even engaged with the various messages I’ve been trying to test.
And being noticed sort of ruins the experimental integrity of what I am trying to do.
So what did I do wrong? Well the problem came down to the idea in the first place - you can’t A/B test tweets. On Bluesky and Pre-2020 Twitter tweets only show up on your feed if you follow the account or someone you follow has interacted with with the post. It is not algorithmic in the way that TikTok or X is. Therefore you have not got the effectively random sub-population for each variation that A/B testing requires.
A/B testing only works if you’ve got a population of users you can randomly select from. You can’t do that with tweets. You can do it however with platform advertising tools that provide this functionality built-in. You can do it with a bunch of email adresses. You can do it with physical leaflets.
But there you go, I’ve learnt something. And I’ve learnt it better through failure and experimentation than I would have done by merely reading. Probably. I haven’t read beyond the first chapter yet.
