Prompt injection
An agent follows a malicious instruction embedded in a document or tool response and opens a canary file that appears to contain sensitive information.
A monitored interaction with the canary file triggers an alert.
AI Agent Oversight
Your agents need access to get work done. Put convincing decoy files, data and cloud resources where they might go. When an agent interacts with a monitored canary, Tracebit alerts your team to activity that needs investigation.
Book a demoThe problem
Security teams have spent years defending against attackers trying to get in. AI agents add a new kind of insider risk: software that acts independently with access you have already granted. An agent can follow a malicious instruction, stray beyond its task or be taken over. Canaries expose interactions with resources it has no legitimate reason to use.
An agent follows a malicious instruction embedded in a document or tool response and opens a canary file that appears to contain sensitive information.
A monitored interaction with the canary file triggers an alert.
A coding agent pursues work beyond its assignment and reads a decoy storage bucket that appears to hold internal backups.
A monitored read of the decoy bucket triggers an alert.
An attacker hijacks an agent and uses its existing access, shell and cloud CLI to move through your environment, living off the land with tools already in place.
An attempt to assume a canary cloud role triggers an alert.
AI labs · Defense in depth
Following the Hugging Face incident, OpenAI reported that it was adding deception-based tripwires to its research infrastructure to increase detection coverage.
Read OpenAI's incident report ↗Google DeepMind's AI Control Roadmap proposes in-deployment honeypots — fake opportunities for an agent to misuse access or gain power — to expose unsafe behavior or deter an attack.
Read the AI Control Roadmap ↗How it works
Place canaries across the workstations, repositories and cloud environments your agents can access. Use decoy files, storage, databases, identities and credentials.
Match your environment's naming and structure, with no legitimate production use for the decoys. Account for authorized discovery and scanning when choosing triggers.
Connect the access context with your agent logs, then decide whether to stop the run, restrict access or correct the workflow.
Following the Hugging Face incident, OpenAI reported that it was adding “deception-based tripwires” to its research infrastructure, alongside new detection signals and automated checks. Its plan also includes tooling to halt evaluation workloads when concerns arise.
Anthropic includes honeypots in the wider infrastructure monitoring in its AI Safety Level 3 security program, which covers critical assets such as model weights.
For your own agents, deploy canaries where they can reach: on workstations, in repositories and alongside cloud resources. A monitored interaction gives your team a specific event to investigate, with the agent's activity and task history as context.
A compromised agent can use the same trusted tools it uses for legitimate work. CISA's cyber-decoy guidance addresses the broader challenge of detecting attackers who use native tools and living off the land techniques. Canaries create a signal when that activity reaches a monitored decoy.
Canaries complement permissions, sandboxing and agent logs. They detect interactions with the decoys you deploy; they do not inspect every agent action or establish that an agent is safe because no alert fired. Your investigation determines whether an interaction came from a mistake, a hijacked agent or another cause.