Skip to content

Glossary

Penetration test

An authorised attempt to find weaknesses in a system before somebody unauthorised does, commissioned by the owner and usually summarised for customers once a year.

In plain terms

A vendor pays specialists to try to get into its own systems, with permission and within agreed limits, so that problems are found by somebody friendly. What reaches a customer is normally a short summary rather than the full findings, because the detail would be a map for anybody else. Reading the summary well is a small skill and worth having.

01

Why it matters

Because it is the one piece of vendor assurance that reports on what somebody actually tried rather than on what is described or certified. That makes it valuable and narrow at the same time: it is a snapshot of specific systems on specific dates, and treating it as a general statement about a vendor's security is the mistake that makes it useless.

02

How it works

The scope decides what the result means, and the scope is always narrower than the vendor. An exercise covers named systems within agreed limits, so a clean summary tells you about those systems and nothing about anything outside them. Checking that the product you are buying was included is the same two-minute step that matters for an audit report.

It is a point in time, and systems change continuously afterwards. A summary describes what was found on particular dates, and everything shipped since is unexamined. That is why the cadence matters more than the last result: a vendor that does this on a regular schedule is telling you something a single clean report cannot.

Customers usually receive a summary rather than the findings, and that is deliberate rather than evasive. Detailed results describe how to reach specific weaknesses, so they are held closely and shared narrowly if at all. A buyer should expect a summary and read it for scope, date, severity of what was found and whether it was resolved.

For AI products, ask whether the scope included the model's behaviour or only the system around it. These are different kinds of problem examined in different ways, and whether a model can be talked into disregarding its instructions has historically sat outside a conventional exercise, though that is changing. Where the summary does not say, it has not been established, and buyers needing assurance on model behaviour should be asking about evaluations and guardrails as well.

What the summary covers, and what a buyer often assumes

What the summary covers, and what a buyer often assumesEvery line of the left-hand column is a generalisation of a line on the right, and the generalising is done by the reader rather than by the vendor. That is worth noticing because it means the failure is recoverable simply by reading, and the specific thing to read for is whether the named systems include the product being purchased. A vendor with several offerings may have examined the platform and not the newer service you are adopting, which is unremarkable and is exactly the case where a confident left-hand reading produces an unfounded conclusion. The second habit worth forming is asking about the schedule rather than dwelling on the result. Any single exercise describes a moment that has already passed, and a vendor that repeats them on a known cadence has told you something about the future, which is the only part a buyer is really trying to predict.AssumedThe vendor was tested.Nothing was found.The product is safe.StatedThese named systems wereexamined.On these dates, within theselimits.These findings, at theseseverities, resolved thus.The right-hand column is whatthe document says and it is moreuseful than the left, because itcan be compared with what youare buying. The left cannot becompared with anything.
Every line of the left-hand column is a generalisation of a line on the right, and the generalising is done by the reader rather than by the vendor. That is worth noticing because it means the failure is recoverable simply by reading, and the specific thing to read for is whether the named systems include the product being purchased. A vendor with several offerings may have examined the platform and not the newer service you are adopting, which is unremarkable and is exactly the case where a confident left-hand reading produces an unfounded conclusion. The second habit worth forming is asking about the schedule rather than dwelling on the result. Any single exercise describes a moment that has already passed, and a vendor that repeats them on a known cadence has told you something about the future, which is the only part a buyer is really trying to predict.
03

Seen in the wild

  • Checking whether a search deployment's connectors into internal systems were inside the scope of a vendor's most recent exercise.

    Glean
  • Asking an automation platform whether the systems holding customer credentials were included.

    Make
  • Recognising that a clean summary for an assistant's platform says nothing about whether the model can be manipulated through its input.

    ChatGPT
04

Common misconceptions

People assume

A clean report means the vendor is secure.

In fact

It means specified systems were examined on specified dates within agreed limits and nothing serious was found then. Everything outside the scope and everything shipped since is unaddressed, which is why the cadence and the scope carry more information than the verdict does.

People assume

We should ask for the full findings.

In fact

You will generally receive a summary, and the reason is sound: detailed results describe how to reach specific weaknesses, so circulating them widely creates the risk the exercise existed to reduce. A summary with scope, dates and how findings were resolved is the appropriate artefact.

05

Telling them apart

Penetration test vs SOC 2

Penetration test

What somebody actually tried against named systems, on given dates.

SOC 2

Whether described controls existed and operated over a period.

One tests the wall; the other tests whether anybody is maintaining it. Neither substitutes for the other.

06

Questions

What should we read in the summary?
The scope, the dates, the severity of anything found and whether it was resolved. Those four are what a summary is for. The absence of findings is much less informative than whether the product you are buying was inside the scope in the first place.
How often should one happen?
Regularly enough that the answer is a schedule rather than an event, since systems change continuously and any single result describes a moment that has passed. A vendor that can describe its cadence is telling you more than one presenting a single clean summary.
Does it cover the AI part?
Ask, because it varies and it is changing. These exercises have conventionally examined the surrounding system rather than whether a model can be induced to disregard its instructions, which is a different problem addressed by evaluations and guardrails. If the summary does not name model behaviour in its scope, a clean result is not assurance about it.
07

Key takeaways

  • It reports what somebody actually tried, which no other artefact does.
  • Scope and date decide what a clean result means.
  • Expect a summary; the full findings are held closely for good reason.
  • Ask whether model behaviour was in scope; conventionally it is not, and that is changing.
09

Tools that use this

  • Glean

    Whether connectors into internal systems were inside the scope.

  • Make

    Whether the systems holding customer credentials were included.

  • ChatGPT

    Why a clean platform result says nothing about model manipulation.

Last checked July 2026

All glossary terms