FEEDING THE BOTS: ‘Extraction without compensation’ — the AI dilemma facing SA’s digital media
A new study of 263 South African news websites has found that fewer than a third explicitly block access to AI crawlers collecting content with which to train and power AI systems – even if they might want to.
South African news publishers are confronting an increasingly uncomfortable choice: allow AI companies access to their journalism, or risk sacrificing online visibility.
This is the tension highlighted in a report, “The Protocol Gap: South Africa”, released on Tuesday by the Journalism Relay Project, the Media Leadership Think Tank (MLTT) at the Gordon Institute of Business Science, and the International Fund for Public Interest Media (IFPIM).
The report examined the robots.txt files – documents in a website’s root directory that permit or restrict access to bots – of 263 South African news websites to establish how publishers are responding to AI crawlers, the automated programmes used by companies including OpenAI, Google, Anthropic and ByteDance to harvest online content.
Although 74.1% of the websites examined had a robots.txt file, just 30.4% explicitly blocked at least one AI crawler. The figure has barely moved since December 2025, when it stood at 29.7%.
What the report suggests, however, is that the decision to block AI crawlers is possible mainly for large, well-resourced publishers – while smaller, community, vernacular and independent outlets are often still fair game for the AI crawlers.
The report attributes this partly to differences in technical expertise and resources, but also to the publishers’ differing levels of dependence on Google and other platforms to attract audiences.
The choices confronting publishers, the report says, involve trade-offs between “visibility and control, audience reach and content protection, and short-term sustainability and long-term values”.
Robots.txt files are one of the few simple and widely accessible mechanisms available to websites to try to prevent crawlers from accessing their data.
But they are by no means an impenetrable defence: the report stresses that the mechanism has no binding legal force and cannot technically prevent the AI bots from encroaching. Its effectiveness depends entirely on the companies behind the crawlers choosing to respect it.
At Tuesday’s report launch, MLTT director Michael Markovitz described robots.txt as essentially a public declaration rather than an enforcement tool.
“Robots.txt is a kind of yes/no gate,” said Markovitz. “So, in that sense, every South African publisher is exposed.”
Journalism Relay Project’s Sérgio Spagnuolo said that despite its toothlessness, the mechanism has an important potential function by stating clearly that the publisher’s attitude is: “I didn’t consent to that.”
A publisher that explicitly states that a crawler may not use its site has at least created a dated public record of its refusal. The report notes that robots.txt directives have already featured in legal complaints over unauthorised scraping in Canada, the US and UK.
5News aggregated this summary from the outlet’s public feed. The full article, with all the context, is on www.dailymaverick.co.za — the content belongs to Daily Maverick.