Skip to main content
Optimus Labs logo

Guide · LLM application security

Top 10 Risks for LLM Applications

Ten common security risks in applications built with large language models, translated into checks security and engineering teams can use. Each entry explains the original application risk and what changes when an AI agent can use tools, credentials, and data.

This independent Optimus Labs guide is heavily inspired by OWASP's Top 10 for LLM Applications. It describes the risks at a different level of detail, with a focus on what security teams can check when agents can act. It is not an official OWASP publication.

By Nipun Gupta

CEO, Optimus Labs · Early contributor to the OWASP Top 10 for Large Language Model Applications

Read the official OWASP 2026 publication

Why the LLM application is only the start

LLM applications include more than a model. They also include prompts, retrieval systems, data pipelines, extensions, and application code. Agents add SKILLs, MCP servers, tools, memory, and credentials. A weakness can therefore produce a bad answer, expose sensitive data, or trigger an action. The practical question is where trust crosses between these parts and what the application is allowed to do at each boundary.

The ten risks and what to check

LLM01:2026

Prompt injection

Agentic AI spotlight

LLM application risk
Untrusted text steers the model away from the instructions its operator gave it.
When agents can act
The agent can act on that text. An injected instruction inside an issue title, a web page, a code comment, or a tool response becomes a command with the agent's credentials behind it.

What to check

  • List every source an agent reads without a human in between: repositories, issue trackers, inboxes, web fetches, tool output.
  • Confirm no single agent combines untrusted input, sensitive data, and external action. That combination is the Lethal Trifecta.
  • Require confirmation on irreversible actions rather than trusting the agent to refuse.

Read: one untrusted string, three trust boundaries

LLM02:2026

Sensitive information disclosure

LLM application risk
The model repeats secrets or personal data in its output.
When agents can act
The agent can move that data. It holds tokens, reads local files, and calls outbound tools, so disclosure is an egress event and not just a bad answer.

What to check

  • Inventory the credentials each agent and MCP server can reach, including environment variables and credential helpers.
  • Restrict outbound destinations for agents that touch regulated data.
  • Prefer short-lived credentials over long-lived tokens stored in agent configuration.

LLM03:2026

Excessive agency

Agentic AI spotlight

LLM application risk
The system grants the model more capability than the task needs.
When agents can act
This is the central agent-layer risk. Permission comes from the harness: allow-lists, auto-approve settings, tool scopes, and the credentials sitting on the machine. Defaults are usually broader than anyone intended.

What to check

  • Audit auto-approve and skip-confirmation settings across agent harnesses.
  • Scope tool permissions per project rather than per user.
  • Score each agent against the actions it can take, not the actions it usually takes.

Read: when safe defaults still grant access

LLM04:2026

Supply chain

Agentic AI spotlight

LLM application risk
Models, adapters, and libraries arrive from third parties.
When agents can act
The supply chain now includes MCP servers, SKILLs, extensions, and hooks that employees install themselves, often outside any review process.

What to check

  • Discover installed agents, MCP servers, and SKILLs on the endpoint rather than relying on an approved-tools list.
  • Record who installed each one, from which registry, and at which version.
  • Alert on new install sources instead of only on known-bad packages.

Read: agents flooded RubyGems to reach a docs builder

LLM05:2026

Data and model poisoning

LLM application risk
Training or fine-tuning data is tampered with.
When agents can act
Persistent agent memory, project instruction files, and cached tool descriptions are all writable state. Poison them once and the agent misbehaves on every later run, with no prompt to point at.

What to check

  • Treat agent instruction files and memory stores as code: review changes and keep history.
  • Detect edits to agent configuration that no human made.
  • Re-read tool and MCP descriptions on change, since a renamed or rewritten tool changes agent behavior.

LLM06:2026

Unbounded consumption

LLM application risk
Cost and capacity run away through excessive requests.
When agents can act
Agents loop. A retry cycle can exhaust API quota, hammer an internal service, or run up spend overnight with nobody watching a terminal.

What to check

  • Alert on agent activity volume per user and per tool, not only on total spend.
  • Cap run duration and tool-call depth in the harness.

LLM07:2026

Misinformation

LLM application risk
The model states something false with confidence.
When agents can act
A wrong conclusion becomes a wrong change: a deleted branch, a bad migration, a message sent to a customer. The cost is operational, not conversational.

What to check

  • Keep a record of what each agent changed and where, so a wrong action can be traced and undone.
  • Gate destructive operations behind review regardless of stated confidence.

LLM08:2026

Hidden context exposure

LLM application risk
System prompts, retrieval context, tool definitions, and hidden application instructions expose information that should remain restricted.
When agents can act
Hidden context can reveal internal paths, available tools, approval rules, and data sources. That gives an attacker a map of what the agent can reach and how to shape its next action.

What to check

  • Keep secrets and authorization decisions out of prompts and retrieved context.
  • Enforce permissions in code and policy controls rather than relying on hidden instructions.
  • Log which context sources influenced sensitive tool calls without exposing the sensitive content itself.

LLM09:2026

Vector and embedding weaknesses

LLM application risk
Retrieval stores leak or return manipulated content.
When agents can act
Retrieval feeds an actor. Content planted in an indexed document reaches the agent as trusted context and can carry instructions along with facts.

What to check

  • Track which collections each agent can query and who can write to them.
  • Separate indexes by sensitivity instead of sharing one store across teams.
  • Treat retrieved text as untrusted input for the purposes of the check above.

LLM10:2026

Improper output handling

LLM application risk
Downstream systems trust model output without validating it.
When agents can act
Output is executed. Shell commands, SQL, file writes, and API calls run with the agent's privileges, so an unvalidated string is remote code execution.

What to check

  • Find every place agent output reaches a shell, a database, or a build step.
  • Pass values through parameters and environment variables instead of string interpolation.
  • Log the command that ran, not only the intent the agent stated.

Where to start

  1. Inventory first. You cannot score a posture you cannot see. Find every agent, MCP server, and SKILL installed across your endpoints, including the ones nobody filed a ticket for.
  2. Score the harness, not the prompt. Permission lives in configuration: auto-approve settings, tool scopes, credential files.
  3. Watch runtime. Static posture tells you what an agent could do. Behavior tells you what it did, and which run deserves a question.

For worked examples of these risks in real incidents, read the Civilizations threat briefings, or see how the platform applies this to security teams.

Apply the OWASP checks to your environment

Optimus Labs discovers agent assets, maps what they can reach, inspects their posture, and enforces policy from installation through runtime.

Book a demo

Backed by

Benhamou Global Ventures
Arka
Executive Venture Fund
a16z Scout Fund
Scout

Ecosystem

NVIDIA Inception ProgramAWS ActivateOpen Secure AI Alliance general memberThe Linux Foundation Silver MemberCarnegie Mellon University