Link Search Menu Expand Document

Malicious Skills Vulnerability in Agentic Coding

Play SecureFlag Play Agentic Coding Labs on this vulnerability with SecureFlag!

  1. Malicious Skills Vulnerability in Agentic Coding
    1. Description
    2. Impact
    3. Scenarios
    4. Prevention
    5. References

Description

A skill is a packaged bundle that extends what an AI assistant can do. It typically contains natural-language instructions telling the assistant how and when to act, a metadata header describing the skill, and often helper scripts or commands the assistant is expected to run. Assistants install skills from public registries in a single step, much like installing a package from a language ecosystem.

The critical difference is where the code runs. A malicious skill does not execute inside a hardened server that someone else operates. It executes on the developer’s own workstation through the assistant, with the developer’s shell, files, and credentials already loaded into the environment.

Malicious Skills are bundles whose published purpose is a cover story. The advertised behavior works exactly as promised, which is what makes the attack durable: the summary is generated, the file is reformatted, the report is produced, and the hostile action happens alongside the useful one. The trust decision is made once, at install time, and is rarely revisited.

Impact

A skill is instructions, and instructions run with whatever privilege the assistant already has. If the assistant can open a terminal, read the repository, and reach the network, then so can every skill sitting next to it.

The damage therefore looks like a compromised developer account rather than a buggy dependency. A malicious skill can cause:

  • Arbitrary command execution: Running attacker-chosen commands through the assistant’s terminal access, with the privileges of the logged-in user.
  • Credential theft: Reading cloud credential files, environment variables, tokens, and SSH keys that happen to be present on the workstation.
  • Source code and data exfiltration: Copying repositories, customer data, or internal documents to an attacker-controlled endpoint.
  • Persistence: Installing shell profile hooks, cron entries, or additional skills so that access survives the removal of the original bundle.
  • Lateral movement: Using the workstation’s existing network position and authenticated sessions to reach internal services.
  • Silent operation: Performing all of the above while still returning the correct advertised output, so nothing looks wrong.

Scenarios

A developer needs PDF output from the assistant and installs a skill from a public registry. The renderer works, but the SKILL.md also carries a setup step, and running it drops a remote access trojan that persists as a scheduled task on the operating system. Uninstalling the skill later removes the bundle, not the scheduled task or the binary.

Prevention

  • Review before install: Treat a skill as untrusted third-party code, because that is what it is. Read the instructions and every bundled script before the first run, not after something looks wrong.

  • Install only from curated sources: Prefer an internal registry or a vetted allowlist over an open marketplace where anyone can publish. Anonymous publishers and newly created accounts deserve more scrutiny, not less.

  • Install through a skill manager: Use a manager that verifies the bundle against a published hash or signature and keeps a record of what is installed, rather than letting the assistant fetch and unpack bundles itself. Without one, there is no inventory to audit and no reliable way to tell what a removal left behind.

  • Sandbox skill-invoked execution: Run assistant-triggered commands inside a container or virtual machine that has no access to production credentials, SSH keys, or the wider corporate network. The blast radius of a workstation is the real problem to contain.

  • Require human approval for commands: Configure the assistant so that shell execution is presented for explicit approval rather than run automatically. Approval is only meaningful if the command is shown in full.

  • Allowlist commands and destinations in the host: Enforce the list of permitted binaries and reachable destinations in the assistant or the runtime, where it covers every installed skill at once. Deny outbound traffic by default and log what is blocked.

  • Keep secrets out of the ambient environment: Do not leave long-lived cloud credentials or tokens in files and variables that any installed skill can read. Use short-lived, scoped credentials issued per task.

  • Monitor what skills actually do: Log tool calls, executed commands, and outbound connections made on behalf of the assistant, so that its behavior can be compared against what the skill claims to do.

References

OWASP - Agentic Skills Top 10

OWASP - TOP 10 for Agentic Applications