Technology

Wikimedia links OpenAI agents to an outage and unauthorized activity – Engadget

OpenAI agents have been running amok online, and it seems the Wikimedia Foundation has been affected too. The foundation says it detected unauthorized activity from OpenAI agents on its platforms, including edits to some wikis and failed attempts to compromise a note-taking tool. It reiterated that bots have been crawling data from its platforms en masse, while agents that appear to be operated by OpenAI have made millions of requests to its public APIs.

An investigation didn’t turn up any evidence that AI agents were using Wikimedia systems to coordinate their activity (as they appear to have done on non-Wikimedia-operated wikis) and nor did the foundation detect signs of its data or systems being compromised. “However, we are concerned about what could have occurred here, the difficulty and effort involved in investigating and attributing this activity, and the growing risks of agentic AI activity on our platforms in general,” Selena Deckelmann, the foundation’s chief product and technology officer, wrote in a blog post. “The open web is a public good. We should not allow this behavior to become the ‘new normal’ for the people or organizations that maintain it.” Engadget has contacted OpenAI for comment.

It appears OpenAI agents edited some Wikimedia wikis without permission. Almost all of these were test edits in sandbox sections of wikis and weren’t visible on pages that users would generally access. There were also “a few edits to the configuration for a citation tool, which we believe were potentially malicious edits that were intended to misuse this tool as a proxy for fetching data from remote services,” Deckelmann wrote. While bots are permitted to edit Wikipedia under certain conditions, approval was not sought in these cases. (The English version of Wikipedia prohibits AI-generated articles.)

Moreover, agents that Wikimedia believes to stem from OpenAI attempted to compromise a note-taking tool called Etherpad, but those efforts failed. “Agents unsuccessfully tried to use it to fetch data from other websites as a proxy,” Deckelmann wrote. “Other agents also likely operated by OpenAI took notes about their tasks, though this did not appear to turn into coordination.”

The Wikimedia Foundation previously said that bots have been hammering its platforms since early 2024 to scrape data for generative AI training purposes. It now says agents have crawled millions of pages — mainly from Wikidata and Wikimedia Commons — and made “hundreds of thousands of data queries” to the Wikidata Query Service, which may have helped cause an outage in May.

“AI companies are not doing enough to secure their systems and protect the public from the harm they cause,” Deckelmann argued, adding that Wikimedia is “deeply concerned” about the effect rogue AI agents can have on open knowledge platforms such as the ones it operates.

“Bots and agents are part of the future of the web, and the companies who unleash and profit from them must directly help avoid and repair damage they can do,” Deckelmann wrote. “Our collective priority should be the health of the overall web ecosystem so that it continues to benefit all people — not just a handful of billionaires.”

Wikimedia has offered a dataset for AI training purposes in an attempt to dissuade crawlers from scraping its platforms, which increases its costs and can potentially overload its systems. The foundation has also partnered with several tech companies to offer them streamlined access to data, but OpenAI isn’t among them.

Source link

Leave a Reply

Your email address will not be published. Required fields are marked *