Model & LLM assessments
Prompt injection, guardrail bypass, system-prompt extraction, data leakage, model extraction, and denial-of-wallet. Mapped to OWASP LLM Top 10 and MITRE ATLAS.
Private offensive security · models
A boutique atelier for AI penetration testing. We assess weights, agents, retrieval pipelines, and the infrastructure that holds them — red-team discipline applied to the model as perimeter.
The application used to be the boundary. The model is now the estate.
The house
Model Intruder exists for organisations shipping copilots, agents, and RAG systems who need more than a jailbreak screenshot. We test the model, the tools it can call, the documents it can read, the cloud it lives in, and the people who deploy it.
Engagements are scoped like a private intelligence brief: limited seats, named operators, no junior farm, no unvalidated scanner noise. Full-stack AI coverage, delivered as a house, not a platform.
Practice
Prompt injection, guardrail bypass, system-prompt extraction, data leakage, model extraction, and denial-of-wallet. Mapped to OWASP LLM Top 10 and MITRE ATLAS.
Tool-use abuse, MCP servers, multi-agent trust, goal hijacking, and excessive agency. Every function call is a privilege boundary.
Indirect injection via documents, retrieval poisoning, cross-tenant leakage, and embedding-store authorisation.
The model feature and the classic app around it: APIs, authn/z, business logic, and places model output becomes a side effect.
Multistep adversary emulation against the model pipeline: OSINT, operator phishing, cloud pivots to artefacts, exfiltration.
Fine-tune poisoning, CI/CD gate tampering, model provenance, secrets in training and inference, isolation review.
Attack surface
Direct and indirect prompt injection, multi-turn jailbreaks, encoded payloads, system prompt leakage, and sensitive data regurgitation.
Unsafe tool invocation, privilege escalation through agents, plugin and MCP compromise, chained actions that look benign alone.
RAG corpus poisoning, connector over-permission, tenant isolation failure, training-data extraction, context-window smuggling.
Inference endpoints, vector databases, GPU estates, secrets management, resource exhaustion, supply-chain compromise of weights.
Operators, MLOps, and vendors. The shortest path to a model artefact is still often a person with a deploy key.
Why a house
Every finding is human-validated. No unconfirmed scanner output reaches a client. Reports are written for counsel, CISO, and the engineer who has to fix it.
Engagements
Tell us the model, the risk you cannot sleep on, and the date you ship.