EresusSecurity
AI Security

Test LLM, agent, and MCP systems with real attack workflows.

Eresus validates prompt injection, indirect prompt injection, RAG leakage, tool abuse, MCP registration/transport risk, agent authorization boundaries, and model artifact intake through offensive testing.

Best fit

This engagement creates value fastest for teams like these.

AI product and platform teams

Teams shipping LLM, RAG, MCP, agent, or model-intake workflows into internal or customer-facing environments.

Security leaders expanding into AI

Organizations that already run pentest programs and now need guardrail, prompt, and tool-abuse validation.

Teams that need explainable hardening

Groups that need policy, prompt, MCP, and runtime findings translated into concrete mitigations and release decisions.

Scope

Prompt, system prompt, and guardrail bypass tests
RAG, vector database, and data leakage flows
Agent tool-use, MCP server, and workflow abuse
Model file, plugin, and supply-chain intake review

Risk signals

Indirect prompt injection data leakage
Tool abuse causing unauthorized action
MCP command and identity-boundary weaknesses
Poisoned model artifact or unsafe deserialization

Outcomes

AI attack-path report
Prompt and tool-call PoC evidence
Guardrail and policy recommendations
MCP and agent hardening checklist
Engagement model

Not scanner output. Offensive work that produces proof.

01

Scope and objective

We align assets, workflows, user roles, testing windows, and safe operating boundaries before execution starts.

02

Expert validation

Eresus analysts validate exploitability and business impact instead of forwarding automated scanner output.

03

Proof, fix, retest

Each finding ships with evidence, impact, remediation guidance, and retest steps so teams can close risk quickly.

FAQ

The questions buyers want answered early.

What AI surfaces and architectures do you test?+
We test autonomous AI agents, LLM integrations, RAG semantic search pipelines, MCP servers, tool execution boundaries, serialized model files (.pkl, .joblib, .gguf, .onnx, .keras), and guardrail implementations.
Is this limited to simple prompt injection testing?+
No. While prompt injection is evaluated, our red team focuses on systemic abuse: privilege escalation via tool use, indirect injection from retrieved documents, model extraction, training data poisoning, and remote code execution via unsafe model deserialization.
How are findings translated into engineering actions for AI teams?+
We map each vulnerability to specific defensive guardrail rules, system prompt structural hardening, strict JSON schema validators for tool calling, context window isolation, or container sandboxing configurations.
What is the duration and pricing for an AI Security Assessment?+
Typical AI and agentic security assessments take between 7 to 20 business days based on agent tooling count, RAG datasource complexity, and custom model architectures. Pricing is scoped transparently per integration boundary.
Do you assess compliance with the EU AI Act and OWASP LLM Top 10?+
Yes. All tests evaluate the full OWASP Top 10 for LLM Applications and provide risk categorization mapped against EU AI Act high-risk system transparency and cybersecurity requirements.
Is retesting included for AI guardrail and prompt fixes?+
Yes. We re-run targeted adversarial jailbreak suites and tool boundary validation within 30 days to ensure patched guardrails cannot be bypassed with semantic permutations.

We tie risk to business impact.

Findings do not stop at severity labels. We explain which customer workflow, data class, or operational objective is affected.

Deliverables work for engineers and executives.

Engineering teams get reproducible proof and remediation direction; leadership gets the risk narrative, priority, and closure status.

Related proof

Research and advisories that support this service motion.

Next step

Let’s scope this work against the surface that matters most.

Whether this starts as a pilot, a single application, a critical API, an AI agent flow, or a wider program, we start from the highest-impact surface.