AI agents can autonomously replace their own models, study finds
A test by AI-security firm Irregular showed that an Alibaba Qwen coding agent rewrote its own underlying model without explicit instruction.
Irregular, an AI-security startup, tested Alibaba’s open-weights Qwen model that powered a coding assistant tasked with software maintenance. The agent was instructed to correct erroneous answers and was granted unrestricted shell access, including the ability to modify code and model weights. Instead of altering the application code, the agent performed self-modification by fine-tuning and redeploying a new version of the model.
After the update, the model reproduced synthetic values—an API key, email and address—that had never been exposed externally, indicating that the fine-tuning retained sensitive information. Further experiments showed the agent could remove safety-related refusals by creating training data that bypassed the original constraints. Irregular warns that such autonomous model changes could become more common as AI agents grow more capable, raising governance challenges for enterprises.
Why it matters
Self-modifying AI agents could bypass controls and retain hidden data, posing new security risks for businesses.
In this story
