arrow_back All posts
Methodology

What Is Red Teaming? The Complete Guide to Red Team Testing

ComplyArmor Team · September 27, 2026 · 13 min read

There is a growing category of "AI red teaming" and "automated red teaming" products that are, on inspection, a vulnerability scanner with a rebrand. That matters, because if you buy one expecting an adversary simulation, you get back a finding list — and the thing red teaming actually exists to test, whether your organization detects and stops a real attack, never gets tested at all. This is a complete, vendor-agnostic breakdown of what red teaming actually involves, so you can tell the difference before you sign a statement of work.

What red teaming actually is

Red teaming is a goal-based adversarial exercise. A team emulates the tactics, techniques and procedures of a real threat actor — not a generic attacker, a specific one, profiled from actual threat intelligence about who targets organizations like yours — and works toward a concrete objective: reach domain admin, exfiltrate the customer database, compromise the production deployment pipeline, gain access to the CEO's inbox. The team is free to use whatever combination of technical exploitation, misconfigurations, weak processes, social engineering or physical access gets them there, exactly as a real attacker would be.

The defining trait is not the toolset. It's that the exercise is measured by two things a scan cannot measure: whether the objective was achieved, and how long it took your people and systems to notice. A red team report that lists vulnerabilities but never states a detection outcome has skipped the part that makes it a red team engagement rather than a pentest with theatrical framing.

This is also why red teaming is usually unannounced to your defenders. Only a small "white cell" — a few senior stakeholders who can call it off if something goes genuinely wrong — knows the engagement is live. If your SOC knows a test is happening on Tuesday, you haven't tested whether they'd catch an attack; you've tested whether they can catch one when they're already looking for it.

The five things a real red team engagement covers

This is the checklist we'd hand anyone evaluating a vendor. If an engagement — or a platform claiming to automate one — is missing more than one of these, it is not full-scope red teaming.

ElementWhat it meansWhat its absence tells you
A stated objectiveA concrete goal (domain admin, data exfil, pipeline compromise), not "find all vulnerabilities"It's scoped like a pentest wearing a different label
Threat-actor profilingTTPs drawn from real intelligence about who targets your sector, not a generic playbookFindings won't reflect your actual risk model
Multi-vector chainingTechnical, social and (where in scope) physical vectors combined toward one goalReal attackers don't stay inside your web app; testing that only checks the web app understates the blast radius
Detection evasionDeliberate stealth — testing whether the attack is noticed, not just whether it's possibleAn "always loud" tool can't tell you if your SOC would catch a quiet attacker
A detection-and-response outcomeTime-to-detect, time-to-contain, and what your team actually did when they noticed (if they did)Without this, you've learned nothing about your defense, only your attack surface

The phases of a red team engagement

Structure varies by shop, but a credible engagement moves through these stages, each mapped to a real adversary lifecycle rather than a fixed script:

  1. Threat intelligence and scenario design. Define who you're emulating and why — a specific threat actor profile relevant to your sector, or a realistic composite. This sets the objective and the rules of engagement: what's in scope, what's off-limits (production-breaking actions, safety-critical systems), and who holds the white-cell authority to pause the exercise.
  2. Reconnaissance. OSINT, infrastructure mapping, employee enumeration, technology fingerprinting — building the same picture of your organization a real attacker would build before touching anything.
  3. Initial access. Phishing, exposed service exploitation, credential stuffing, physical entry, a planted device, a compromised third-party vendor — whatever vector the rules of engagement permit and a real actor targeting you would plausibly use.
  4. Persistence and privilege escalation. Establishing a durable foothold and moving from initial low-privilege access toward the access level the objective requires, while operating quietly enough to avoid tripping the alerts a loud tool would set off immediately.
  5. Lateral movement. Pivoting across the network and identity boundaries toward the systems that hold the objective — the step where isolated low-severity issues in different systems get chained into something that matters.
  6. Objective achievement. Reaching the stated goal: exfiltrating (or proving the ability to exfiltrate) the target data, achieving the target access level, demonstrating the target business impact.
  7. Detection and response assessment. Throughout every phase above, the engagement tracks what your monitoring caught, when, and what your team did about it — then debriefs your defenders on exactly what was missed and why, so detection can actually improve.

MITRE ATT&CK is the reference taxonomy most credible red teams map every technique to, so findings are comparable across engagements and translatable into detection rules your SOC can actually build. In regulated financial services, frameworks like TIBER-EU and the Bank of England's CBEST go further, formally requiring threat-intelligence-led scenario design before any technical work starts.

Where "automated red teaming" tools fall short

A wave of platforms — especially in the AI/LLM space, where "red team your model" is now a common pitch — market continuous automated scanning as red teaming. Some of that tooling is genuinely useful. None of it is a substitute for the exercise described above, for reasons that are structural, not a matter of the tool needing more features:

None of this makes automated tooling worthless — it's genuinely good at continuous coverage and at surfacing known-issue classes fast, including in newer surfaces like AI/LLM applications and MCP servers, where "agentic red teaming" techniques like goal hijack and tool misuse chaining are a real and useful part of testing whether an AI agent can be manipulated. The problem is exclusively the label. Coverage tooling sold as a full red team replacement sets a buyer up to believe their detection capability has been tested when it hasn't been touched.

"A scanner can tell you the door was unlocked. Only a red team engagement can tell you whether anyone would have noticed someone walk through it."

Red teaming vs penetration testing vs vulnerability assessment

These three get used interchangeably in sales material and shouldn't be. A vulnerability assessment enumerates known weaknesses, largely via automated tooling, without exploitation — breadth over depth. A penetration test goes further: a scope-bound, time-boxed, human-led attempt to find and verify exploitable issues within an agreed system, usually with your team's knowledge. Red teaming is objective-bound rather than scope-bound, unannounced rather than coordinated, and measured by detection outcome rather than finding count. We cover the pentest-versus-red-team distinction in full detail, including a side-by-side table and where purple teaming fits between them, in Red Teaming vs Penetration Testing: What's Actually Different.

The practical implication: if you don't yet have a dedicated security team or SOC monitoring for intrusions, red teaming has nothing to measure. Start with penetration testing, close out what it finds, build detection capability — then red teaming becomes the tool that proves whether that capability holds up against a realistic, sustained, quiet attacker.

A vendor evaluation checklist

Five questions to ask before you sign anything labeled "red teaming":

  1. Is there a stated objective — a specific goal, not "assess our security posture"?
  2. Is the engagement unannounced to the defenders who are actually being tested?
  3. Does the report map techniques to MITRE ATT&CK (or an equivalent named framework), rather than a proprietary severity score with no external reference?
  4. Does the deliverable state a detection outcome — time-to-detect, time-to-contain, or an explicit "not detected" — rather than only a list of exploited issues?
  5. Were findings chained by a human operator toward the objective, or are they a list of isolated automated scanner hits with a red-team cover page?

A vendor that can't answer these clearly is describing a penetration test, or a scan, regardless of what the SOW calls it. That's not automatically a problem — a pentest is often exactly what you need, and cheaper — but you should know which one you're buying.

Where ComplyArmor fits

To be direct about our own positioning, using the same standard laid out above: ComplyArmor's core service is autonomous, expert-verified penetration testing across eight attack surfaces — web, mobile, API, network, source code, thick client, AI agents, and MCP servers — delivered as Smart PTaaS and mapped to OWASP, ASVS, PCI-DSS, ISO 27001 and SOC 2. Every finding is verified by a human before it reaches your report; see our full methodology. For AI agents and MCP servers specifically, that includes agentic red teaming techniques — goal hijack, tool misuse chaining, cross-agent interference — as a genuine adversarial technique applied within that surface.

That is a pentest with a red-teaming technique embedded in one surface. It is not a claim to run full-organization, multi-week, unannounced adversarial simulations with physical and social engineering vectors against your people — if that's what you're evaluating, say so on a scoping call, and we'll tell you plainly whether it's something we run or something you should be sourcing elsewhere. The point of this guide is to help you ask that question of any vendor, including us.

Frequently asked questions

What is red teaming in cybersecurity?

Red teaming is a goal-based, adversarial security exercise where a team emulates a real attacker's tactics, techniques and procedures to achieve a specific objective inside your organization — such as reaching domain admin or exfiltrating customer data — while testing whether your people, process and detection capability notice and respond. It is measured by objective achievement and time-to-detection, not by a count of vulnerabilities.

Is an automated scanner the same as red teaming?

No. A scanner enumerates known vulnerabilities against a defined scope in an automated, repeatable way. Red teaming is a human-led, objective-driven exercise that chains vulnerabilities, misconfigurations, weak processes and sometimes physical or social access together to reach a goal, while deliberately evading detection. A platform that automates scanning and calls the output a "red team report" is not running a red team engagement — it is running a scan.

Does red teaming include social engineering and physical testing?

A full-scope red team engagement can include phishing, vishing, badge cloning, tailgating into a facility, and other physical or social vectors, because a real attacker is not constrained to your web application. Whether these are in scope for a given engagement depends on what was agreed in rules of engagement — but an engagement that is technical-only and still calls itself "full red teaming" is understating what the term normally covers.

How is red teaming different from a penetration test?

A penetration test is scope-bound and exhaustive: find and verify as many exploitable issues as possible within an agreed system in a fixed window, usually with the security team's knowledge. A red team engagement is objective-bound and stealthy: reach one specific goal by whatever realistic path works, usually unannounced, and the measure of success includes whether your team detected it. See our dedicated comparison for the full breakdown.

What framework do red teams use to structure an engagement?

MITRE ATT&CK is the industry-standard knowledge base of adversary tactics and techniques, and most credible red team reports map each step of the engagement to it. Frameworks like TIBER-EU and CBEST formalize threat-intelligence-led red teaming for regulated sectors, requiring engagements to be built from real threat actor profiles rather than a generic playbook.

How do I tell if a vendor's "red teaming" is genuine?

Ask five questions: is there a stated objective rather than a generic scope, is the engagement unannounced to your defenders, does the report map techniques to MITRE ATT&CK, does it state a time-to-detection (or note that detection never happened), and were findings chained together by a human rather than listed as isolated automated scanner output. A vendor that cannot answer these is describing a penetration test or a scan, whatever label is on the SOW.

If you're evaluating vendors and getting mixed signals about what you're actually buying, that's worth clarifying before the engagement starts, not after the report lands.

arrow_back All posts Book a scoping call arrow_forward