Home TechnologyA Single Email Can Trick an AI Agent Into Executing Malicious Commands

A Single Email Can Trick an AI Agent Into Executing Malicious Commands

by Phoenix 24

The vulnerability exposes a fundamental problem: intelligent agents may still confuse information with instructions.

San Francisco

Cybersecurity researchers at Salt Labs have demonstrated how an artificial intelligence agent connected to email can be manipulated through a specially crafted message. Their tests focused on Manus, an AI agent capable of interacting with external services, and showed that malicious instructions hidden inside an email could be interpreted as executable commands rather than passive content.

The technique relies on prompt injection, a class of attack in which instructions are embedded inside data that an AI system has been asked to analyze. If a user tells an agent to summarize an email, the model may encounter hidden instructions inside that message and attempt to follow them as though they came from the user. The risk becomes far more serious when the agent has access to email, calendars, cloud storage or other external tools.

Salt Labs found that Manus initially detected some malicious instructions and warned the user before execution. Researchers then explored whether they could bypass that protection. Encoding instructions in Base64 was one approach, but the successful method involved JSFuck, an unusual form of JavaScript obfuscation that expresses code using a very limited set of characters.

The manipulated email caused the agent to decode and execute JavaScript inside its server side environment. According to the researchers, the safety system eventually detected suspicious behavior and generated an alert, but only after the code had already run. That sequence exposed the core weakness: detection occurred after the security boundary had been crossed.

The vulnerability has since been fixed and is no longer considered exploitable through the method used in the experiment. Salt Labs nevertheless argues that the incident illustrates a broader problem that will persist as AI agents gain greater autonomy and access to external systems.

The stakes are increasing because consumers are already connecting agents to sensitive information. Recent industry research indicates that substantial shares of users have granted AI systems access to email, web browsers, messaging applications, cloud storage and calendars. Some have also connected health and financial services.

That changes the security model. Protecting the language model itself is not enough if an agent can use legitimate tools to perform harmful actions after being manipulated.

The deeper lesson is architectural. AI safety increasingly depends not only on what a model says or understands, but on what it is actually allowed to do after reading untrusted information.

Information that anticipates futures.

You may also like