# Kavara — Shadow AI Governance & AI Data-Loss Prevention (Full Documentation) > Kavara solves Shadow AI — the ungoverned use of AI tools (ChatGPT, Claude, Gemini, Copilot, Perplexity, Grok, Mistral) by employees who paste sensitive data into prompts outside IT oversight. A browser extension detects sensitive data the moment it's typed or pasted into any AI tool and tokenizes it before it leaves the browser. CISOs get full visibility into Shadow AI usage without any raw sensitive data ever reaching Kavara's servers. ## What Is Shadow AI? Shadow AI is when employees use AI tools — ChatGPT, Claude, Gemini, Microsoft Copilot, Perplexity, and others — outside of IT governance and security controls. It is the AI-specific form of Shadow IT, and it's uniquely dangerous because AI tools require rich, contextual, sensitive input to produce useful output. Employees paste API keys, customer PII, source code, financial records, and internal documents into AI prompts every day, on personal accounts IT doesn't manage. Traditional DLP can't see Shadow AI: prompts go from the browser directly to AI providers over TLS, invisible to proxies and network monitoring. Blocking AI tools doesn't work — employees route around blocks on personal devices. For a comprehensive guide, see: https://www.kavara.io/shadow-ai ## The Problem Kavara Solves Shadow AI creates three simultaneous risks: (1) data exfiltration — sensitive data sent to third-party AI providers that may use it for training, (2) zero visibility — CISOs can't see what data is flowing into which AI tools, and (3) compliance exposure — regulated data flowing into uncontrolled third parties violates GDPR, HIPAA, SOC 2, and the EU AI Act. Traditional DLP blocks AI entirely, which kills productivity and drives employees to personal devices (making Shadow AI invisible). Kavara takes a different approach: it lets employees keep using AI at full speed while protecting the data automatically through on-device tokenization. ## How It Works (Technical Architecture) ### On-Device Detection Kavara runs 16 built-in detectors directly in the browser using local ML models. Categories include: - **PII**: emails, phone numbers, physical addresses, names - **Credentials**: API keys (AWS, GCP, Azure, Stripe, etc.), tokens, passwords - **Financial**: credit card numbers, bank account numbers - **Code**: source code, internal identifiers, configuration secrets - **Custom**: enterprises can add keyword, regex, and exact-match rules Detection happens entirely on the user's device. No text is ever sent to Kavara's servers for classification. ### Reversible Tokenization When sensitive data is detected in an AI prompt: 1. The sensitive span is replaced with a typed token like `[API_KEY-1]` or `[EMAIL-3]` 2. The AI tool receives the tokenized prompt — it can still reason about the structure 3. When the AI responds, tokens are rehydrated back to real values locally 4. The employee sees a normal, useful response with their actual data This is fundamentally different from blocking or masking. The AI can still reason about the data's role in the prompt without ever seeing the raw value. The employee gets a useful answer. Nothing is blocked, nothing leaks. ### What Kavara's Servers Receive Only event metadata: - Data category (e.g., "credentials", "PII") - Event count - AI tool and action type - Timestamp - Anonymous install identifier - Hashed key prefix Never: prompt text, AI responses, raw secrets, employee names, IP addresses. The database has no column for raw values. A breach of Kavara reveals nothing about customer data because we never hold it. ## Protection Modes ### Monitor Kavara observes and reports without interrupting employees. You get a true picture of Shadow AI before changing any workflow. ### Warn When sensitive data is about to be sent, the employee sees a quiet notification and decides. Awareness at the moment of risk. ### Block For highest-risk categories: tokenize automatically or stop the send outright. Precise enforcement, not blanket blocking. Organizations can use different modes for different data categories and AI tools, and change them at any time from the dashboard. ## Dashboard & Visibility The admin dashboard provides: - **Event log**: which AI tools are being used, what categories of data are detected, when, and what action was taken (monitored, warned, blocked, tokenized) - **Usage insights**: aggregate by tool, department, and data category — never per-employee surveillance - **Policy management**: configure detection rules, protection modes, and per-tool settings - **Audit log**: append-only record of all policy changes (who changed what, when) - **Enrollment**: generate activation codes for teams, manage seats ## Security Architecture - **No raw data stored**: database has no field for prompts, responses, or secrets - **Hashed credentials**: API keys stored as peppered SHA-256 hashes with display prefix - **Strict tenant isolation**: every query scoped to tenant, enforced at the data layer - **Append-only audit trail**: no update/delete path for audit records - **Encryption in transit**: TLS everywhere - **Fail-closed design**: ambiguous states deny access, never permit ## Deployment Options - **Self-serve**: generate enrollment codes, share with a team, live in minutes - **MDM/Chrome Enterprise**: managed deployment via Intune, Workspace ONE, Chrome policy - **Browsers**: Chrome, Edge, Brave (any Chromium-based browser) ## Pricing | Plan | Price | Includes | |------|-------|----------| | Pilot | Free | Up to 25 seats, Monitor mode, 16 built-in detectors, Shadow-AI dashboard, 90-day event history | | Team | $25/seat/month | Everything in Pilot + Monitor, Warn & Block modes, usage insights, custom rules, MDM deployment, audit log, email support | | Enterprise | Custom | Everything in Team + custom detectors, department rollups, configurable retention, DPA & security-review support, priority support & SLAs | ## Competitive Differentiation Kavara is the only AI DLP that uses **reversible tokenization**. Every competitor (Nightfall, Cyberhaven, LayerX, Google Chrome Enterprise Premium) either blocks or masks — destroying the employee's ability to get a useful answer from AI. Kavara tokenizes sensitive data so the AI can still reason about it, then rehydrates the real values in the response. Protection without productivity loss. Key differentiators: - **Productivity-preserving**: tokenize, don't block; rehydrate, don't mask - **On-device**: no round-trip to cloud for detection - **Browser-agnostic**: works inside Chrome, Edge, Brave — and inside competitors' secure browsers - **AI-native**: built for AI tools specifically, not adapted from email/endpoint DLP - **Covers Google AI Mode**: protects AI queries inside Google Search, a surface no competitor covers ## Coverage ### AI Tools Protected (11+ out of the box) ChatGPT, Claude, Gemini, Microsoft Copilot, Perplexity, Mistral (Le Chat), Grok, Poe, HuggingFace, You.com, Google AI Mode & AI Overviews ### Detection Categories (16 built-in) Emails, names, phone numbers, physical addresses, credit cards, bank accounts, SSNs, passport numbers, API keys (AWS, GCP, Azure, Stripe, GitHub, and more), OAuth tokens, passwords, source code, internal identifiers, custom patterns ## Company Kavara is headquartered in Sydney, Australia. Infrastructure runs in the Sydney region. ## Contact - Website: https://www.kavara.io - Dashboard: https://app.kavara.io - Email: hello@kavara.io ## Pages - Home: https://www.kavara.io/ - What Is Shadow AI? (Guide): https://www.kavara.io/shadow-ai - Features: https://www.kavara.io/features - Security: https://www.kavara.io/security - Pricing: https://www.kavara.io/pricing - About: https://www.kavara.io/about - Privacy Policy: https://www.kavara.io/privacy - Terms of Service: https://www.kavara.io/terms