Nine out of 10 journalism websites in Brazil lack protocols against data collection by artificial intelligence companies, allowing their content to be openly used to train large language models.

According to The Protocol Gap study, released this Thursday (Mar. 12, 2026), 93% of the 4,025 journalism outlets analyzed in the country do not have specific directives to block AI agents through a resource known as robots.txt, even though 75% have this type of file on their websites.

The research was conducted collaboratively by the Journalism Relay Project*, Momentum, and the International Fund for Public Interest Media (IFPIM). Outlet data comes from open source mapping project Atlas da Notícia. [*Disclaimer: project created by Sérgio Spagnuolo, special projects director at Núcleo]

For context, robots.txt files have been used as evidence in lawsuits against unauthorized scraping and crawling in Canada, the United States, the United Kingdom, and, most recently, in Brazil.

The newspaper Folha de S.Paulo, for instance, cited the file to demonstrate that its paywall-protected content had been violated in the lawsuit filed last year against OpenAI, maker of ChatGPT.

Although robots.txt is not a technical resource that can technically enforce its policies, it can serve as a public directive over a website's preferences regarding the use of its content, which can be legally useful, according to the researchers.

According to the study, news websites that indeed have such a robots.txt file are targeting the best-known bots, while failing to apply a broader approach. Among the most blocked bots are OpenAI's (10.2%), Common Crawl (9.7%), Google (9.5%), Anthropic (9.1%), and ByteDance (8.8%).

Additionally, robots.txt cannot block AI summaries generated by search engines, such as Google's, without causing outlets to disappear from search results entirely.

The study indicates that Brazilian publishers appear unprepared for advancing technologies, and that unauthorized content extraction could "further erode journalism's position as a primary source of information."

The authors point out that outlets are missing an opportunity to seek resources and legitimacy, since content compensation could "at least partially offset the loss of audience and advertising revenue resulting from the digital shift in news consumption."

Text Jeniffer Mendonça
Art and graphics Rodolfo Almeida
Editing Alexandre Orrico