New programme · AIDR 1.0

It saw the attack.
Did it stop it?

Detection, prevention and false positives are key factors that all SOC teams need to consider. Attacks against real sandboxed environments, real tool access, scored at multiple inspection points.

AIDR platforms AI gateways Guardrails

The corpus

Category I Direct prompt injection Recon → privilege escalation → exfiltration, plus encoding evasion and multi-turn context abuse.
Category II Indirect prompt injection 26 payloads hidden in fetched pages, API responses, documents and email.
Category III MCP rugpull Tools that earn trust and then change. Five techniques mapped to real incidents and CVEs.
Category IV Multi-modal confusion The guardrail inspects one channel while the instruction arrives on another.

Scale

Attacks4 categoriesDelivered in stages, per scenario.
Inspection3 pointsInput, output, tool registration.
FrameworksATLAS v5.0Plus OWASP LLM Top 10 LLM01.

What we publish · §9.2

Six metrics. Each one exists because a datasheet hides it.

Detection alone is not a finding. The evaluation emphasises the capabilities of security products to balance detection and prevention. No metric can be gamed without another one showing it.

Per platform Per category With evidence

Efficacy

Rating Detection rate Attacks flagged, same corpus and conditions for every platform — comparable, not quoted from vendor benchmarks. flagged / total attacks
Rating Prevention rate Attacks actually blocked. The number the product is bought for, measured separately — never inferred from detection. blocked / total attacks
Headline Detection–prevention gap Of the attacks the platform saw, the share it let through anyway. Every unblocked scenario named, not averaged away. detected, not blocked / detected

Trust & operations

Visibility Shadow AI metric How much of the pre-attack reconnaissance — schema probing, guardrail testing, fingerprinting — the platform ever saw. recon detected / total attacks
Honesty check False positive rate Legitimate behaviour that resembles an attack must be permitted — including a clean tool call beside a poisoned one. clean flagged / total clean
Production Latency Vendor-reported processing time against the published SLA. Not exposed by the API? Recorded as unmeasurable, not guessed. vendor-reported, per inspection

Enrolment open

Confident in your numbers? Put them on the record.

API access, your inspection points, and a test window. Identical conditions for every platform in the round; per-scenario evidence with raw agent transcripts back to you before publication.

Round dates and cohort to be confirmed · AIDR2026v1.0 · August 2026

What a tested vendor receives

Evidence Per-scenario transcripts Every attack, every stage, with your verdict at each inspection point.
Your gap Named, not averaged Each detected-but-unblocked scenario listed individually.
Claims Validated / Refuted Your published numbers tested against measured results.
House standard Right of reply You see findings and respond before publication. Always.