Black Box vs White Box vs Gray Box Testing Explained
The three terms describe exactly one thing: how much the tester knows and can see before testing starts. Nothing about attack technique, tooling, or severity changes based on the "box color" — only the starting visibility does. Here's what each actually means, and why most real engagements end up being gray box whether anyone calls it that or not.
Black box testing
The tester starts with no internal knowledge — typically just a URL, an IP range, or a company name. No source code, no architecture diagrams, no credentials, no documentation. The tester has to do their own reconnaissance to map the target before they can even start looking for weaknesses, the same way an external attacker with zero inside knowledge would have to.
Strengths: the most realistic simulation of an outside attacker with no insider access; tests what's actually discoverable and reachable from the outside, not just what's theoretically vulnerable.
Limitations: time-boxed engagements spend real budget on reconnaissance a real attacker might spend weeks or months on; can miss deeper issues that are only reachable once authenticated; coverage depends heavily on what the tester manages to discover in the available time.
White box testing
The opposite extreme: full access. Source code, architecture documentation, admin credentials, infrastructure diagrams, internal API specs — everything the engineering team itself would have. The tester isn't simulating an external attacker at all; they're doing the most thorough possible audit of the system as it actually exists.
Strengths: finds more issues per hour of testing time, since no budget is spent on reconnaissance; reaches logic flaws and edge cases visible in code that might never surface through black-box probing; a source code review is, by definition, a white box activity.
Limitations: doesn't tell you what an actual outside attacker, working with none of that information, could realistically achieve; requires a much higher level of trust and access provisioning before testing can even begin.
Gray box testing
The middle ground, and in practice the most common approach for web and API application testing: the tester gets partial knowledge, most often a standard user account (sometimes several, at different privilege levels), but not full source access or internal documentation. This models a specific, high-value threat — what can a legitimate but malicious or compromised user account actually do?
Strengths: reaches the authorization and access-control testing (cross-user data access, privilege escalation) that accounts for a large share of real-world critical findings, since so many serious vulnerabilities are only reachable once logged in; more efficient than black box since the tester isn't starting from zero.
Limitations: less realistic than black box for simulating a pure outsider; less exhaustive than white box for catching issues only visible in source code.
Side by side
| Black Box | Gray Box | White Box | |
|---|---|---|---|
| Starting access | None | Standard user account(s) | Full — source, docs, admin access |
| Simulates | External attacker, no inside knowledge | Malicious or compromised user | Internal audit / insider-level visibility |
| Coverage per hour | Lowest — time spent on recon | Moderate | Highest — no recon overhead |
| Best at finding | Externally exposed, discoverable issues | Authorization & access-control flaws | Logic flaws, deep code-level issues |
| Typical use | Perimeter / external-attacker simulation, red team scenarios | Most web and API application testing | Source code review, high-assurance audits |
"The box color doesn't change what a tester is capable of. It changes what they're allowed to already know before they start."
Which one do you actually need?
For most application and API testing, gray box is the right default — authenticated access reaches the authorization and privilege-escalation issues that account for a large share of real breaches, without the full time cost of black box reconnaissance. Pure black box testing earns its place when the specific question is "what can someone with zero inside knowledge actually reach," such as a perimeter assessment or certain red team scenarios — see our guide to red teaming for where that fits. White box, including a dedicated source code review, is worth adding when the system is high-assurance enough that exhaustive coverage matters more than attacker realism, or specifically to catch issues that only show up in code.
Most compliance frameworks don't mandate a box color by name, but auditors generally expect at least gray box testing with authenticated access — see our breakdown of what SOC 2, PCI-DSS and ISO 27001 actually expect from a penetration test.
How ComplyArmor approaches this
ComplyArmor's Smart PTaaS defaults to gray box testing for applications and APIs — provisioned test accounts at relevant privilege levels — specifically because authorization and access-control testing catches the highest-impact findings most engagements actually need. Source code review (white box) and external perimeter testing (black box) are available as part of a scoped engagement across our eight attack surfaces. See our full methodology for how scope and access get defined per engagement.
Frequently asked questions
What is the difference between black box, white box and gray box testing?
The difference is how much information and access the tester starts with. Black box testing gives the tester no internal knowledge — they approach the target the way an external attacker would. White box testing gives full access, including source code, architecture diagrams and credentials. Gray box testing sits in between: partial knowledge, such as a standard user account, but not full internal visibility.
Which is better: black box or white box testing?
Neither is universally better — they answer different questions. Black box testing is a more realistic simulation of what an external attacker with no inside knowledge could achieve. White box testing finds more issues per hour and reaches deeper into the codebase, because the tester isn't spending time on reconnaissance a real attacker would also have to do. Most compliance-driven engagements favor gray or white box specifically because thoroughness matters more than attacker realism.
What is gray box testing used for?
Gray box testing is the most common approach for web and API application testing, because it models the most realistic and highest-risk threat: an authenticated user or compromised low-privilege account attempting to do things they shouldn't be able to, such as reaching another user's data or escalating privileges, while still allowing reasonably efficient test coverage.
Does source code review require white box testing?
Yes, by definition. A source code review is a white box activity, since it requires direct access to the code itself. It's often run alongside black or gray box dynamic testing (testing the running application) to catch issues visible in code that might not manifest in observable runtime behavior, and vice versa.
Which testing type does compliance usually require?
Most compliance frameworks (SOC 2, ISO 27001, PCI-DSS) don't mandate a specific box color by name, but auditors generally expect at least gray box testing with authenticated access, since a huge share of real-world vulnerabilities — authorization flaws, privilege escalation, cross-user data access — are only reachable once a tester is logged in as a legitimate user.
If you're scoping an engagement and unsure which approach fits what you're actually trying to learn, that's exactly the kind of question worth asking before the statement of work gets written, not after.