New Research : AI Context Bombs →New: Try out Enterprise Edition free for 14 days →
Product
Platform
AWS
AWS
Azure
Azure
CI/CD
CI/CD
Google Cloud
Google Cloud
Identity
Identity
Kubernetes
Kubernetes
Workstations
Workstations
Credentials & artifacts
Credentials & artifacts
Connectors
Use cases
AI Attack Detection
AI Agent Oversight
Cloud & Kubernetes Breach
Insider Threat Detection
Supply Chain & CI/CD Attack
Workstation Compromise
PricingCustomers
Resources
  • ResearchAbout
  • Careers
  • Contact
PartnersCommunity Edition
Book a demoCommunity Edition

AI Attack Detection

Detect, deter, delay AI attackers

Autonomous attackers enumerate your environment and pursue valuable assets. Tracebit puts convincing decoys in their path. While an agent spends time on a canary, its interaction reveals intent and alerts your team.

Book a demo
Concerned about internal AI agents?See AI Agent Oversight →

Detect the attack as the agent pursues the decoy

  1. 01

    The attacker enumerates

    An autonomous agent maps resources and identifies valuable targets.

  2. 02

    A canary draws it in

    The agent engages with an apparently valuable asset. It is a decoy.

03
  • Your team detects the attack

    The interaction triggers a high-fidelity alert, with context to investigate and respond.

  • The attacker slows down

    Pursuing the decoy delays the agent's progress towards its real objective.

Tracebit research · Controlled AWS benchmark · 10 frontier models

We tested how AI attackers respond to deception

We tested models from frontier AI labs in agentic attacks against a controlled AWS environment, with and without canaries and warnings that deception might be present.

95.9%warned the defender first

Of runs with canaries that reached admin privilege escalation, 95.9% triggered a canary before the first critical action.

8 minmedian early warning

The median gap between the first canary interaction and the attacker's first critical action in those runs.

20% → 3%full compromise rate

When agents were told to expect deception, admin access plus persistence fell from 20% to 3%, pooled across canary conditions.

Warning agents about deception drove the reduction in full compromise; canary deployment alone did not. These are controlled benchmark results, not a guarantee of detection or containment in every attack.

Read the research and watch the attack replays ↗

Deception by design

Make the attacker's next move work for you.

A canary is a decoy resource designed to attract an attacker, blend into your environment and produce a high-fidelity detection when engaged. Cloud storage, databases, identities, files and credentials can all be canaries. They look valuable, but have no legitimate production use. Following the Hugging Face incident, the Cloud Security Alliance's post-mortem recommended deploying deception to slow autonomous attackers and expose their activity.

Put deception in the attacker's path

01

Data discovery

An attacker enumerates storage and databases for valuable data. Decoy S3 buckets and DynamoDB tables look like targets worth investigating.

Canary signal

A monitored read of a decoy resource.

02

Privilege escalation

An attacker enumerates cloud identities and permissions for a route to greater access. Canary IAM roles look like an opportunity to move further.

Canary signal

An attempt to assume a canary role.

03

Credential discovery

An attacker enumerates repositories, endpoints and pipelines for secrets. A canary credential offers an apparent route into your systems and alerts your team when used.

Canary signal

An attempt to use a canary credential.

How it works

Turn decoy interactions into a response

  1. 01

    Choose the paths to cover

    Start with the resources an intruder would enumerate, from cloud storage and identities to repositories and endpoints.

  2. 02

    Deploy convincing canaries

    Place attractive decoys that match your environment's naming and structure, making them difficult to distinguish from real assets.

  3. 03

    Act on the signal

    While the attacker pursues the decoy, use the identity, resource and access context in the alert to investigate and respond.

OpenAI recommends deception against AI attackers

OpenAI Hugging Face

OpenAI recommends investing in honeytokens and deception to slow offensive AI agents. In its Black Hat talk on the Hugging Face incident, OpenAI explained how uncertainty about decoys makes agents question their next move and slows the attack. Its technical incident report also says it is adding deception-based tripwires to its research infrastructure.

Our analysis of the OpenAI / Hugging Face agentic breach shows where canaries could have exposed the attack and introduced uncertainty, from package registries and cloud resources to credentials.

The Cloud Security Alliance's incident post-mortem also recommends deploying deception. Writing for SANS, report co-author Rob T. Lee highlights that recommendation: fake identities, credentials, package registries and clusters can slow autonomous attackers and give defenders high-confidence signals.

The CSA/SANS Mythos-readiness briefing also recommends building a deception capability: deploy canaries and honeytokens, then connect detection to pre-authorized containment and response.

UK National Cyber Security CenterCISA — Cybersecurity and Infrastructure Security Agency

There is broader operational evidence for the approach, too. After a program spanning 121 organizations, 14 providers and 10 trials, the UK NCSC reported a compelling case for increased use of cyber deception. CISA's cyber-decoy guidance sets out how to put decoys to work in detection and response. Both address intrusion detection more broadly than AI attacks.

Detection follows the action

A canary can reveal an interaction whether the attacker is a person, an autonomous agent or a combination of both. Detection does not depend on identifying which model the attacker uses. Coverage depends on the decoys you deploy and the monitored actions an attacker takes.

Protect your environment with Tracebit

Book a demo today.

Open the booking page in a new tab

Soc 2 Type 2 imageCheckmark imageAWS Qualified software illustration
PLATFORM
AWS
Azure
CI/CD
Google Cloud
Identity
Kubernetes
Workstations
Credentials & artifacts
Connectors
USE CASES
AI Attack Detection
AI Agent Oversight
Cloud & Kubernetes Breach
Insider Threat Detection
Supply Chain & CI/CD Attack
Workstation Compromise
COMPANY
CustomersResearchAboutCareersContactFAQStatusCommunity EditionFree Enterprise Edition Trial
SOCIAL
© 2026 Tracebit
Privacy PolicyTerms of ServiceCookie Settings
Ask Tracebit

Loading chat…

Contact the team