Tests & certification · what we run

Every customer is different.
So are our tests.

Six evaluations and two certification tracks — for security vendors, enterprises, SMBs and consumers. Real-world tradecraft, in weeks not months.

6 evaluations 2 certification tracks 2 live AMTSO evaluations certification valid for VirusTotal
Four ways in

Pick the engagement, not a package.

Every methodology is public →
01 · EVALUATIONS

Evaluations

Beyond checklists — real-world tradecraft, with personalised next steps.

02 · CERTIFICATION

Certification

Efficient product validation — currently valid for VirusTotal participation.

03 · BESPOKE

Bespoke testing

Prove a new product's value under real-world methodologies.

04 · CONSULTATION

Consultation

A plan tailored to your needs — the right test, not the biggest one.

The programmes

Eight programmes. One open loop.

EVALUATION 01LIVE · AMTSO

Scam & Phishing

AI has lowered the bar for launching convincing scams. Our evaluation is platform agnostic — real-world scenarios, tested the way attackers actually work.

ConsumerAI luresLive corpusScam & Phishing 1.5
Running now · H1 2026 The world's largest AI scam & phishing test Live lure corpus, AMTSO standard. Results publish on completion. AMTSO standard listing
Verdicts

Blocked / missed, with evidence

Every lure stamped, screenshotted and reproducible.

Comparable

Same corpus, same conditions

Every product faces an identical set of live scams.

EVALUATION 02

AI SOC

Do AI assistants actually pay off in the SOC? We measure the return on investment they bring to real workflows — with key metrics, not vibes.

EnterpriseAutonomous SOCAppendix included
What we measure
Accuracy

Detection & false positives

Combining our APT and Cloud & Identity tradecraft for the most comprehensive view.

Efficiency

Alert efficiency

Alert volume change against the baselined period.

TIU

Time to initial understanding

How long until the incident is actually understood.

Dwell

Dwell time

Earliest possible detection to the moment the intrusion is caught.

TTC / TTR

Time to contain & remediate

Exposure to containment; containment to normal operations.

TTA / TTI

Time to advise & implement

From remediation to recommendations — and to them being applied.

EVALUATION 03COMMENTARY PHASE

Secure Browser

Quantify the defensive ROI of secure browsers against AI, identity and exfiltration risks — whatever the architecture.

Consumer secure browserRemote browser isolationEnterprise browser
Public commentary phase — the methodology is being shaped in the open. Challenge it before we score anything.
The modules
Module 01

Identity security

Validates browser-level defences against credential abuse.

Module 02

GenAI guardrail bypass

Tests the resilience of AI guardrails under adversarial pressure.

Module 03

Data loss prevention

Can the browser stop data loss where traditional endpoint DLP fails?

EVALUATION 04

Identity & Cloud Security

Credential abuse is the most common way attackers get in. The first ITDR and cloud-focused tradecraft — for MSSP, MDR, identity, cloud and email solutions.

MSSP / MDRITDRCloud estatesEmail
Identity scenarios
Rating

Ongoing attack

Detections during the attacker's activity.

Rating

Post compromise

Detections or enriched data after the malicious activity.

Rating

Alert efficiency

Alert volume change against baseline, across both operations.

Cloud scenarios
Stage 01

Intrusion

Detection efficacy of initial delivery methods.

Stage 02

Infiltration

Detection efficacy of actions on the initial target.

Stage 03

Propagation

Detection efficacy of spread through the organisation.

EVALUATION 05

Ransomware Impact

Benchmark your product's ability to minimise downtime during a live ransomware surge — full chain, pre-encryption to recovery.

Full chainCross-industryRansomware Impact 1.0
Why it's different
Coverage

Definitive cross-industry activity

Product efficacy assessed across multiple industries, not one lab profile.

Depth

15+ ransomware groups covered

The most comprehensive group coverage in a public test.

Realism

Real-world testing, real-world ransomware

Realistic environments for results that transfer to yours.

EVALUATION 06

Advanced Persistent Testing

The most comprehensive detection-and-response testing: efficacy, soft features and false positives. For EDR, NDR and XDR.

EDRNDRXDRFull killchain
Efficacy ratings
Rating

Detection

Measures detection capabilities, stage by stage.

Rating

Protection

Prevention-focused scenario capabilities.

Rating

False positives

Scenario-based legitimate-accuracy rating.

Capabilities assessment
Response

Customisability

How responders can tune response options to improve future detections.

Response

Options in the fight

The variety of options available to a responder during an attack.

Response

Detail

The depth of information on responses during an attack.

CERTIFICATION 01

Real World Protection

Show how your solution handles real-world threats — consumer and business tracks, live threats, no synthetics.

ConsumerBusiness & enterpriseLive threats
To certify
97%
Protection · attack rating
Evidence

Verdicts with proof

Certified, partial or failed — published with the evidence behind it.

Reply

Right of reply

Your response printed beside the verdict. Always.

CERTIFICATION 02

VirusTotal Certification

Malware scanning to AMTSO's Fundamental Principles, using Real-Time Threat List samples. Efficient, independent validation.

VendorsAMTSO principlesReal-Time Threat List
To certify
100%
Detection · malicious corpus
0%
False positives
Standard

AMTSO-aligned

Performed to AMTSO's Fundamental Principles of Testing.

Outcome

Valid for VirusTotal

Certification currently valid for VirusTotal participation.

Support

Questions we actually get.

Can't find yours? Ask us →
How are your tests different from a typical pentest or vulnerability scan?

Instead of running generic scans, tests are designed around each client's real systems, threats and goals — so they reveal practical attack paths and concrete fixes.

And our reports aren't just pentest reports: they're designed for all your stakeholders and team members, not only the engineers.

What types of organisations do you work with?

Three groups, typically: enterprises with complex or regulated environments; security vendors who need independent testing of their products; and high-growth startups preparing to sell into enterprises or regulated sectors.

How often should we repeat testing?

At minimum: after major changes in architecture or product, or annually if your environment is stable. High-change or high-risk organisations often move toward more frequent or continuous assessment.

How do you handle confidentiality and data protection?

Non-disclosure agreements are standard. Only the minimum data needed for testing is accessed, access is logged, and all artifacts are stored and retained under strict controls for agreed periods. Sensitive details in reports can be further restricted on request.

Can I give feedback on the methodologies or reports?

We pride ourselves on our transparency. Every methodology is open for challenge before a product is scored — and if something in a report is unclear, tell us. We publish the changelog.

Not sure which test fits?

Tell us what you're building and who you sell to — we'll point you at the right programme, or say honestly if none of them fit yet.

Transparency Innovation Partnership Weeks, not months Right of reply, always