Glossary
Penetration test
An authorised attempt to find weaknesses in a system before somebody unauthorised does, commissioned by the owner and usually summarised for customers once a year.
In plain terms
A vendor pays specialists to try to get into its own systems, with permission and within agreed limits, so that problems are found by somebody friendly. What reaches a customer is normally a short summary rather than the full findings, because the detail would be a map for anybody else. Reading the summary well is a small skill and worth having.
Why it matters
Because it is the one piece of vendor assurance that reports on what somebody actually tried rather than on what is described or certified. That makes it valuable and narrow at the same time: it is a snapshot of specific systems on specific dates, and treating it as a general statement about a vendor's security is the mistake that makes it useless.
How it works
The scope decides what the result means, and the scope is always narrower than the vendor. An exercise covers named systems within agreed limits, so a clean summary tells you about those systems and nothing about anything outside them. Checking that the product you are buying was included is the same two-minute step that matters for an audit report.
It is a point in time, and systems change continuously afterwards. A summary describes what was found on particular dates, and everything shipped since is unexamined. That is why the cadence matters more than the last result: a vendor that does this on a regular schedule is telling you something a single clean report cannot.
Customers usually receive a summary rather than the findings, and that is deliberate rather than evasive. Detailed results describe how to reach specific weaknesses, so they are held closely and shared narrowly if at all. A buyer should expect a summary and read it for scope, date, severity of what was found and whether it was resolved.
For AI products, ask whether the scope included the model's behaviour or only the system around it. These are different kinds of problem examined in different ways, and whether a model can be talked into disregarding its instructions has historically sat outside a conventional exercise, though that is changing. Where the summary does not say, it has not been established, and buyers needing assurance on model behaviour should be asking about evaluations and guardrails as well.
What the summary covers, and what a buyer often assumes
Seen in the wild
Checking whether a search deployment's connectors into internal systems were inside the scope of a vendor's most recent exercise.
GleanAsking an automation platform whether the systems holding customer credentials were included.
MakeRecognising that a clean summary for an assistant's platform says nothing about whether the model can be manipulated through its input.
ChatGPT
Common misconceptions
People assume
A clean report means the vendor is secure.
In fact
It means specified systems were examined on specified dates within agreed limits and nothing serious was found then. Everything outside the scope and everything shipped since is unaddressed, which is why the cadence and the scope carry more information than the verdict does.
People assume
We should ask for the full findings.
In fact
You will generally receive a summary, and the reason is sound: detailed results describe how to reach specific weaknesses, so circulating them widely creates the risk the exercise existed to reduce. A summary with scope, dates and how findings were resolved is the appropriate artefact.
Telling them apart
Penetration test vs SOC 2
Penetration test
What somebody actually tried against named systems, on given dates.
Whether described controls existed and operated over a period.
One tests the wall; the other tests whether anybody is maintaining it. Neither substitutes for the other.
Questions
- What should we read in the summary?
- The scope, the dates, the severity of anything found and whether it was resolved. Those four are what a summary is for. The absence of findings is much less informative than whether the product you are buying was inside the scope in the first place.
- How often should one happen?
- Regularly enough that the answer is a schedule rather than an event, since systems change continuously and any single result describes a moment that has passed. A vendor that can describe its cadence is telling you more than one presenting a single clean summary.
- Does it cover the AI part?
- Ask, because it varies and it is changing. These exercises have conventionally examined the surrounding system rather than whether a model can be induced to disregard its instructions, which is a different problem addressed by evaluations and guardrails. If the summary does not name model behaviour in its scope, a clean result is not assurance about it.
Key takeaways
- It reports what somebody actually tried, which no other artefact does.
- Scope and date decide what a clean result means.
- Expect a summary; the full findings are held closely for good reason.
- Ask whether model behaviour was in scope; conventionally it is not, and that is changing.
Last checked July 2026