Last updated
What is an LLM penetration test?
An LLM penetration test is a time-boxed, authorised attack on an application built on a large language model, a chatbot, copilot, RAG search product or LLM-powered workflow. It tests the whole system: the model's behaviour, the data and tools connected to it, and the application code that trusts its output. The goal is the same as any Ironbark Cyber pentest: prove real-world impact, then show you exactly how to close it.
What do you test?
- Prompt injection, direct and indirect. Instructions smuggled through user input, retrieved documents, emails, web pages and file uploads, anything that reaches the context window.
- Data exfiltration. System prompt disclosure, cross-tenant leakage through shared context or caches, and extraction of sensitive data from RAG sources the user shouldn't see.
- Tool and function abuse. What the model is allowed to call, with whose privileges, and what happens when an attacker steers it, the "excessive agency" class of the OWASP LLM Top 10.
- Guardrail and filter bypass. Jailbreaks that matter for your risk profile: not "will it say a bad word", but "will it approve the refund, leak the record, or execute the action".
- Output handling. Downstream code that trusts model output: markdown/HTML injection, SSRF via generated URLs, injection into queries and shell commands built from responses.
- The classic layer. Authentication, authorisation, rate limiting and business logic around the AI feature, because attackers don't respect the boundary between "AI bug" and "web bug".
How do you test it?
Manually, with tooling for breadth. We build a threat model of your application on the scoping call, then combine adversarial prompt campaigns with conventional application testing, chaining findings across both layers. Testing maps to the OWASP Top 10 for LLM Applications and MITRE ATLAS, and is driven by consultants who research these attacks, Ironbark Cyber's founder has been featured by HackerOne on exactly this work.
What do we receive?
- An executive summary written for non-technical stakeholders.
- A technical report: evidence, reproduction steps and remediation guidance for every finding, ranked by real-world impact.
- Mapping to the OWASP LLM Top 10, MITRE ATLAS and NIST AI RMF where your stakeholders need it.
- A debrief call with your engineering team, an attestation letter, and a free retest of fixed issues within 90 days.
How long does it take and what does it cost?
Most single-application engagements run 5–10 testing days, with the report inside a week of finishing. Pricing is fixed-fee: indicatively from AU$20,000 for a focused LLM application test, quoted exactly within one business day of scoping.
FAQ
Frequently asked questions
How much does an LLM penetration test cost?
LLM penetration tests are fixed-fee, quoted within one business day of a free scoping call. The price depends on the number of applications, the tools and data sources the model can reach, and whether agentic behaviour is in scope. Indicative pricing is published on this page, and the quote never changes mid-engagement.
What is the difference between an LLM pentest and a regular web application pentest?
A web application pentest targets deterministic code paths: authentication, access control, injection. An LLM pentest adds the probabilistic layer: prompt injection, jailbreaks, retrieval poisoning, tool abuse and output-handling flaws. It tests how the model layer and the application layer fail together. Most serious LLM findings chain both, which is why Ironbark Cyber tests them as one system.
We use OpenAI, Anthropic or a hosted model, is testing still useful?
Yes, and it is where most findings come from. The model vendor secures the model; you are responsible for what you connect to it, system prompts, RAG data, tools, tenant separation and how outputs are used downstream. Those are application-layer decisions, and they are what we attack.
Do you follow the OWASP Top 10 for LLM Applications?
Yes. Testing covers every OWASP LLM Top 10 category and maps findings to it, alongside MITRE ATLAS techniques where relevant. But the checklist is only the floor. The highest-impact findings usually come from your application's specific business logic and tool integrations.
How long does an LLM penetration test take?
Typically 5–10 testing days for a single application, with the report delivered within a week of testing finishing. Critical findings are escalated live during the engagement.
Will the results satisfy customer security reviews and auditors?
Yes. You receive an executive summary, a technical report with reproduction steps, framework mapping (OWASP LLM Top 10, MITRE ATLAS, NIST AI RMF) and an attestation letter you can share externally. Enterprise customers increasingly ask specifically whether AI features have been tested, this answers that question with evidence.
Contact
Talk to us
Tell us what you're trying to protect, secure or build. We'll come back with a plan.