How to scope an AI red team engagement (and what it should cost in Australia)

A practical guide to scoping an AI red team: choosing objectives, setting boundaries, and what a well-run engagement should cost an Australian organisation.

The fastest way to waste money on an AI red team is to buy one without deciding what you want it to prove. "Red team our chatbot" is not an objective. It lets a vendor run a jailbreak script, screenshot the model saying something embarrassing, and invoice you for it. That screenshot makes a good slide and tells you very little about your actual risk.

A red team is objective-driven by definition. Before you talk to any firm, including Ironbark Cyber, you should be able to complete this sentence: "if an adversary could ___, it would be a board-level incident." Extract another tenant's data through the retrieval layer. Steer the support agent into issuing refunds. Recover the system prompt and the API keys embedded in it. Poison the knowledge base so the model misleads your customers for a month. Objectives like these map to your risk register, they have owners, and they let you judge whether the engagement achieved something other than producing a PDF.

Scoping conversations then work backwards from those objectives to the boundaries: which systems are in play, which data is real versus synthetic, what the testers may persist and for how long, who holds the kill switch, and how a critical finding gets escalated mid-engagement rather than surfacing two weeks later in the report. A vendor who doesn't push you on these questions is planning to run a checklist rather than a campaign.

Cost follows scope. In the Australian market, AI red team pricing lands in a fairly predictable band once objectives and phases are fixed, and anything quoted before that conversation should make you suspicious in either direction. This post walks through how we scope these engagements at Ironbark Cyber, what drives the price up or down, and the questions worth asking any firm before you sign.

Objectives first: how do you turn a risk register into attack goals?

Start with the outcomes you fear rather than the technology you happen to have bought. A useful objective names an adversary, an action, and a consequence: "an external user extracts another customer's records through the assistant", "a support agent is manipulated into issuing refunds without authorisation", "the system prompt and its embedded secrets are recovered by a member of the public". Each of those is testable, since we either achieve it or we don't, and each maps to something a risk owner already loses sleep over.

Vague objectives produce vague engagements. "Make sure the AI is safe" cannot be passed or failed, so it gets quietly redefined mid-engagement into whatever the tester found easiest to demonstrate. Three to six sharp objectives, ranked, will shape a far better campaign than a twenty-item wishlist. If you can't yet write them, that is worth knowing in itself: a short architecture review or threat-modelling session is a cheaper way to find your objectives than paying a red team to discover them for you.

What is actually in scope?

An AI red team is broader than an application test because the system is broader than the application. Scope usually spans several layers, and you should decide deliberately which are in and which are out:

  • The models and their configuration. System prompts, temperature, and the guardrails and safety classifiers around them.
  • The application. The retrieval pipeline, the tools and functions the model can call, and the code that trusts model output downstream.
  • Monitoring and response. Whether your telemetry would see the attack, and whether anyone would act on it. A red team tests detection as much as prevention.
  • The supply chain, where relevant. Model registries, fine-tuning data, prompt and template stores, and the MLOps tooling around them.
  • The people. Operators who approve agent actions or review flagged output are part of the system, and with your written agreement they can be part of the scope.

You don't have to include all of it. You do have to decide, because an unstated boundary is where engagements go wrong.

Rules of engagement: production or staging?

The honest answer is "it depends, and we'll argue for whichever gives you real evidence with the least risk". A staging environment that faithfully mirrors production is ideal, but many AI-specific risks (cross-tenant retrieval, real tool integrations, live data flows) only show up against production. Where we test production, we agree rules of engagement up front: rate limits, which data is real versus synthetic, what the testers may persist and for how long, who holds the kill switch, and the stop conditions that end the campaign immediately. Destructive actions are simulated or performed against replicas. Critical objective achievements are reported the day they happen rather than saved for the report.

How is the campaign phased?

A serious campaign runs in phases so you can steer it and, if you want, commit one phase at a time. Reconnaissance and threat modelling map the attack surface and refine the objectives. Attack development builds and iterates the techniques. Execution runs against the agreed objectives. Replay walks your team back through what worked so they can reproduce and fix it. Phasing also protects your budget: if reconnaissance shows an objective is unreachable or trivially reachable, you find out before you have paid for a fortnight of execution against it.

What should an AI red team cost in Australia?

Because red teaming is scoped to objectives rather than assets, pricing varies more than a fixed-scope pentest, but it is not a mystery. Once objectives and phases are fixed, Ironbark Cyber engagements start around AU$30,000, with multi-phase campaigns against several objectives scoped to AU$100,000 and above. What moves the number is the count and difficulty of objectives, how many layers are in scope, whether the human and monitoring layers are included, and how long testers may persist.

Be suspicious in both directions. A quote that arrives before anyone has asked what you want to prove is a quote for a script rather than a campaign. And a five-figure "red team" that turns out to be a day of jailbreak prompts is cheap because it is doing very little. You are paying for a slide rather than evidence.

Do you actually need a red team, or a pentest first?

For most teams, the honest answer is a penetration test first. If you have never had the AI application tested at all, an LLM penetration test will find more exploitable issues per dollar, because it enumerates vulnerabilities across the whole application rather than pursuing a handful of objectives in depth. A red team earns its premium once you have closed the obvious findings and need to know whether a motivated adversary can still achieve a specific goal against your defences, including your detection and response. We cover the trade-off in detail in AI red teaming vs penetration testing. If you are unsure, say so on the scoping call and we will tell you honestly which one fits.

What does a good report look like?

A good AI red team report reads as a narrative rather than a vulnerability dump: what we tried, what worked, what your systems caught, and what an attacker would do with each result. Every technique is mapped to MITRE ATLAS so your defenders can look it up and build detections against a named adversary behaviour rather than a one-off anecdote. Findings are ranked by real-world impact and tied back to the objectives you set, so the report answers the question you asked, "could this happen to us, and would we notice?", rather than the question that was easiest to answer. Fixed issues get a free retest within 90 days.

Questions to ask any vendor before you sign

  • How will you translate our risk register into attack objectives, and who signs off on them?
  • What is in and out of scope across models, application, monitoring, supply chain and people?
  • How do you run production safely, and what are the stop conditions?
  • How is the engagement phased, and can we commit one phase at a time?
  • How do you report, and what framework do you map to?
  • Who does the work: the person on this call, or a junior bench?

Ironbark Cyber runs objective-driven AI red teaming for Australian organisations, with senior-only delivery, a fixed quote within one business day of scoping, findings mapped to MITRE ATLAS, and critical results escalated live. If you can finish the sentence "if an adversary could ___, it would be a board-level incident", we can help you find out whether they can.

FAQ

Frequently asked questions

What is the difference between an AI red team and an LLM penetration test?

A penetration test enumerates vulnerabilities across the whole application; a red team pursues specific attacker objectives in depth and tests your detection and response as well as your prevention. Most teams should do a pentest first and add red teaming once the obvious findings are closed.

How much does an AI red team cost in Australia?

Ironbark Cyber engagements start around AU$30,000 once objectives and phases are fixed, with multi-phase campaigns scoped to AU$100,000 and above. Price is driven by the number and difficulty of objectives, the layers in scope, and how long testers may persist.

Can an AI red team run safely against production?

Yes, with agreed rules of engagement. Rate limits, real-versus-synthetic data, persistence limits, a named kill-switch holder and clear stop conditions. Destructive actions are simulated or run against replicas, and critical results are escalated the day they happen.

How long does an AI red team take?

Typically two to six weeks, phased as reconnaissance, attack development, execution and replay, so you can steer the campaign and commit one phase at a time.

Put this into practice

A senior Ironbark Cyber consultant will scope your engagement on a free 30-minute call and give you a fixed quote within one business day.