The ultimate objective of an adversarial prompt injection attack is almost always data exfiltration or unauthorized state modification. An attacker tricks an agent into reading sensitive file contents (.env, SSH keys, or customer PII) and transmitting that data outbound through an allowed tool channel such as a web scraper, markdown image renderer, or external API call. Vark implements a multi-layered defense-in-depth model specifically designed to break this exfiltration pipeline.
┌───────────────────────┐
│ Untrusted Input Data │
└───────────┬───────────┘
│
▼
┌───────────────────────┐ ┌────────────────────────┐ ┌────────────────────────┐
│ PII & Token Masking ├─────►│ Honeytoken Seed Inject ├───►│ Output DLP Scanner │
│ (Session Pseudonyms) │ │ (Canary Detection) │ │ (Stream/Buffer Inspect)│
└───────────────────────┘ └────────────────────────┘ └───────────┬────────────┘
│
▼
┌────────────────────────┐
│ HMAC / Ed25519 Audit │
│ (Cryptographic Chain) │
└────────────────────────┘Inspecting JSON responses for leaked secrets is easy with plain strings — but high-performance agent tools stream output asynchronously via Node ReadableStream or Buffer instances. Vark's DLP engine scans binary buffer chunks using a sliding-window algorithm without consuming or mutating the stream: Shannon entropy scoring catches high-entropy strings like private keys and API credentials, targeted detectors cover Google, Azure, AWS, GitHub, npm, PyPI, Twilio, and SendGrid tokens, and credit card numbers (Luhn-validated) plus US Social Security Numbers are masked with session-scoped pseudonyms like [USER_REF_1].
To catch injections that exfiltrate silently, Vark seeds agent context and temporary filesystems with synthetic honeytokens (e.g., stripe_live_honeytoken_99f82a...). If a hijacked model echoes a canary key into an outbound call, prompt, or URL parameter, the Canary Trap halts execution instantly and locks the session.
For compliance, every tool call joins an append-only cryptographic hash chain:
H_n = HMAC-SHA256(K, H_{n-1} || Timestamp || Tool || Payload)Where H_n is the current record hash, H_{n-1} the previous record hash, and K the secret HMAC key (or asymmetric Ed25519 private key). If an attacker modifies past logs to cover their tracks, the chain breaks — and vark audit verify flags the tampering immediately.
Keep reading the source
