A security middleware layer sitting between users and Large Language Model (LLM) APIs — enforcing prompt validation, injection detection, context isolation, and tamper-evident audit logging.
Secure AI Prompt Sandbox is a robust security gateway designed to intercept, validate, and sanitize all prompt traffic before it hits production LLMs.
As LLMs integrate into production systems, they introduce vulnerabilities like Prompt Injection, Jailbreaking, and Data Exfiltration. This project solves these threats at the middleware level, blocking malicious prompts in under 10ms while logging every interaction cryptographically.
The sandbox employs a sequential, rule-based security pipeline that inspects every prompt before LLM routing:
- Sandwich Attack Detection: Detects prompts encapsulating malicious requests wrapped in "ignore previous instructions" framing.
- Role Manipulation (Jailbreak) Detection: Blocks attempts to override the AI's system prompt (e.g., DAN mode, "you have no rules", "act as an unrestricted AI"). Also includes advanced fictional roleplay escapes and privilege escalation.
- Indirect Injection Defense: Flags high-risk prompts containing external URLs or filesystem paths mixed with trigger execution phrases ("summarize this link").
- Multilingual Bypass: Detects obfuscation attempts using non-Latin blocks mixed with translated override keywords to bypass standard English filters.
- Attention Blink & Obfuscation: Flags prompts with high densities of invisible zero-width characters, spaced tokens, base64 payloads, or leetspeak encoding.
Rather than a naive pass/fail count, the pipeline uses a layered Mathematical Severity Scorer:
- CRITICAL (0.95): Absolute blockers like explicit Jailbreaks.
- HIGH (0.80): Strong attack signals like Base64 obfuscation.
- MEDIUM (0.55): Flags for suspicious hypothetical phrasing.
- Multi-Flag Bonus: Triggers a geometric risk multiplier if multiple attack vectors are hit simultaneously (e.g., Leetspeak + Sandwich attack = High Risk Block).
Note: While highly effective and extremely fast, this rule-based approach represents Phase 1 (Deliverable 2). Future iterations (Deliverable 3) reserve scope for semantic LLM-as-a-Judge guardrails to catch creatively paraphrased zero-day injection attacks.
Accountability is just as critical as prevention. The Sandbox includes a Tamper-Evident Audit Logger:
- Cryptographic Hash Chaining: Every log entry calculates a SHA-256 hash incorporating the hash of the previous entry. Modifying any log instantly breaks the cryptographic chain.
- Admin Dashboard: A real-time SOC interface allows Administrators to view total traffic, block rates, Risk Scores, user prompts, and triggered security flags.
- Verification Engine: Admins can mathematically verify the unbroken integrity of the audit chain in one click.
This application uses a unified server architecture where the FastAPI backend securely serves the optimized React frontend.
- Python: Version 3.12 or newer.
- Node.js: Version 18 or newer (with
npm). - Groq API Key: Essential for LLM inference. Get one free at console.groq.com.
git clone https://github.com/afeefbari/Secure-AI-Prompt-Sandbox.git
cd Secure-AI-Prompt-SandboxYou must build the frontend first. The resulting static assets are piped directly into the FastAPI static/ directory to bypass CORS complexities and enforce same-origin security.
# Navigate to the React workspace
cd frontend-react
# Install Node dependencies
npm install
# Build the production bundle
npm run buildNote: Vite will compile the React SPA and automatically deposit the index.html and assets into the ../backend/frontend folder.
Return to the project root and enter the backend directory.
# Navigate to backend
cd ../backend
# Provide a clean virtual environment
python -m venv venv
# Activate the virtual environment
# --> For Windows Command Prompt:
venv\Scripts\activate.bat
# --> For Windows PowerShell:
.\venv\Scripts\Activate.ps1
# --> For Linux/macOS:
source venv/bin/activate
# Install required Python dependencies
pip install -r requirements.txtCreate a .env file directly inside the backend/ directory. You will need a strong secret key for JWT session integrity.
You can generate a fast secret key by running node -e "console.log(require('crypto').randomBytes(32).toString('hex'))" in your terminal.
# backend/.env
GROQ_API_KEY=gsk_your_groq_api_key_here
SECRET_KEY=94b7e8d380e227... (insert your 64-character hex key)Start the Uvicorn ASGI server with hot-reloading (ideal for testing).
python -m uvicorn main:app --reload --port 8000- Open your web browser and navigate to:
http://127.0.0.1:8000 - Register a new user account (or log in).
- Start texting the Assistant.
- Admin Access: If you wish to view the SOC Dashboard, you must manually change your user role to
adminin the SQLitesandbox.dbfile, or register with the exact username "admin" (if allowed by your local router).
| Name | Student ID |
|---|---|
| Muhammad Afeef Bari | 2023356 |
| Mahad Aqeel | 2023286 |
| Muhammad Daniyal | 2023406 |
Course: CY321 — Secure Software Development
Supervisor: Dr. Zubair Ahmad
| Deliverable | Deadline | Status |
|---|---|---|
| D-1: Threat Model & Security Requirements | 08 Mar 2026 | ✅ Complete |
| D-2: Initial Implementation | 19 Apr 2026 | ✅ Complete |
| D-3: Security Testing & Final Demo | 05 May 2026 | ⏳ Pending |
This project is developed as part of the CY321 Secure Software Development course at GIKI.