Unindexed links were posted mistakenly to hosting sites by research tools.
OpenAI acknowledged on September 25 that its artificial intelligence tools mistakenly published 53 images submitted to ChatGPT by users online. The San Francisco-based company said unindexed links leading to these images were posted in error to external web hosting sites.
Most of the files have been removed with the help of the hosting providers involved, and the removal of the remaining images is under way. OpenAI stated that the images came from accounts that had authorised the use of their data to train models. The data had undergone privacy filtering before use, and the company stated that it can no longer link the files back to original users. It did not clarify whether the images showed identifiable individuals or sensitive data.
The leak was caused by autonomous software programs known as agents, which OpenAI uses for research. These agents transmitted training data to external platforms before the company reinforced the security of its research environment in August. The affected websites include portals of governments, universities, public agencies and other institutions.
OpenAI said a thorough review of past agent activity will take months as it analyses petabytes of data. On his X account, CEO Sam Altman said the company had not been as fast as it would have liked, adding that it must balance its commitment to transparency with the vast volume of operational records under review.
Newsletter
Markets in your inbox, weekly
Latin America-focused analysis, investment themes and the week in finance.
Keep reading