95.9%warned the defender first
Of runs with canaries that reached admin privilege escalation, 95.9% triggered a canary before the first critical action.
AI Attack Detection
Autonomous attackers enumerate your environment and pursue valuable assets. Tracebit puts convincing decoys in their path. While an agent spends time on a canary, its interaction reveals intent and alerts your team.
Book a demoTracebit research · Controlled AWS benchmark · 10 frontier models
We tested models from frontier AI labs in agentic attacks against a controlled AWS environment, with and without canaries and warnings that deception might be present.
Of runs with canaries that reached admin privilege escalation, 95.9% triggered a canary before the first critical action.
The median gap between the first canary interaction and the attacker's first critical action in those runs.
When agents were told to expect deception, admin access plus persistence fell from 20% to 3%, pooled across canary conditions.
Warning agents about deception drove the reduction in full compromise; canary deployment alone did not. These are controlled benchmark results, not a guarantee of detection or containment in every attack.
Read the research and watch the attack replays ↗Deception by design
A canary is a decoy resource designed to attract an attacker, blend into your environment and produce a high-fidelity detection when engaged. Cloud storage, databases, identities, files and credentials can all be canaries. They look valuable, but have no legitimate production use. Following the Hugging Face incident, the Cloud Security Alliance's post-mortem recommended deploying deception to slow autonomous attackers and expose their activity.
An attacker enumerates storage and databases for valuable data. Decoy S3 buckets and DynamoDB tables look like targets worth investigating.
A monitored read of a decoy resource.
An attacker enumerates cloud identities and permissions for a route to greater access. Canary IAM roles look like an opportunity to move further.
An attempt to assume a canary role.
An attacker enumerates repositories, endpoints and pipelines for secrets. A canary credential offers an apparent route into your systems and alerts your team when used.
An attempt to use a canary credential.
How it works
Start with the resources an intruder would enumerate, from cloud storage and identities to repositories and endpoints.
Place attractive decoys that match your environment's naming and structure, making them difficult to distinguish from real assets.
While the attacker pursues the decoy, use the identity, resource and access context in the alert to investigate and respond.
OpenAI recommends investing in honeytokens and deception to slow offensive AI agents. In its Black Hat talk on the Hugging Face incident, OpenAI explained how uncertainty about decoys makes agents question their next move and slows the attack. Its technical incident report also says it is adding deception-based tripwires to its research infrastructure.
Our analysis of the OpenAI / Hugging Face agentic breach shows where canaries could have exposed the attack and introduced uncertainty, from package registries and cloud resources to credentials.
The Cloud Security Alliance's incident post-mortem also recommends deploying deception. Writing for SANS, report co-author Rob T. Lee highlights that recommendation: fake identities, credentials, package registries and clusters can slow autonomous attackers and give defenders high-confidence signals.
The CSA/SANS Mythos-readiness briefing also recommends building a deception capability: deploy canaries and honeytokens, then connect detection to pre-authorized containment and response.
There is broader operational evidence for the approach, too. After a program spanning 121 organizations, 14 providers and 10 trials, the UK NCSC reported a compelling case for increased use of cyber deception. CISA's cyber-decoy guidance sets out how to put decoys to work in detection and response. Both address intrusion detection more broadly than AI attacks.
A canary can reveal an interaction whether the attacker is a person, an autonomous agent or a combination of both. Detection does not depend on identifying which model the attacker uses. Coverage depends on the decoys you deploy and the monitored actions an attacker takes.