A personalized, multimodal WhatsApp notification-routing system that decides whether
each incoming message should notify the user now, wait for a digest, or be
**mute**d as low-value, repetitive, unwanted, suspicious, or unsafe.
Built for the HackerRank Orchestrate Message Notification Router challenge, the solution combines structured user context, behavioral history, retrieval, deterministic safety rules, OCR, speech transcription, and four tool-calling agents. It produces one grounded, schema-valid decision for every text message, image poster/screenshot, and voice note.
Each output row contains:
message_id,action,message_type,reason,confidence,evidence_message_ids
- Personalized routing: Uses user behavior, quiet hours, group relationships, business history, sender trust, notification load, and prior reactions instead of applying one global rule to everyone.
- Multimodal perception: Transcribes voice notes with Whisper and extracts OCR, descriptions, and risk signals from images with Qwen3-VL.
- Four-agent orchestration: Triage, safety, context, and evidence agents have explicit responsibilities, model tiers, handoffs, and schema validation.
- Safety veto before routing: Deterministic phishing, scam, and prompt-injection checks
can force
mute/scambefore untrusted message content reaches a routing model. - Historical-evidence retrieval: Ranks the receiver's own message history using sender, group, business, media, lexical similarity, and past reaction signals.
- Grounded explanations: Evidence IDs are clamped to real historical messages; invalid or hallucinated IDs are removed.
- Bounded concurrency: Routes several messages in parallel without firing all 110 requests at once, then restores deterministic input order before writing the CSV.
- Crash-safe caching: OCR, transcripts, evaluation decisions, and final decisions are
cached under the gitignored
media_cache/directory. - Budget visibility: Tracks calls, tokens, estimated spend, and budget warnings by model.
- Offline validation: Includes 63 tests covering data loading, safety, retrieval, perception, tool loops, orchestration, concurrency, and output guarantees.
graph LR
%% Input
subgraph Input [1. Input]
MSG[("π¨ messages.csv<br>110 messages")]
REL[("π₯ users Β· groups<br>businesses")]
HIST[("π message history<br>+ events")]
MEDIA[("πΌοΈ images<br>π voice notes")]
end
%% Load + perception
subgraph Load [2. Load & Perceive]
STORE["ποΈ typed DataStore<br>indexed lookups"]
ASR["π Whisper large-v3<br>transcribe"]
VIS["ποΈ Qwen3-VL-235B<br>OCR + describe"]
end
%% Deterministic assembly
subgraph Prep [3. Context Assembly β deterministic]
RET["π evidence retrieval<br>ranked history"]
GUARD["π‘οΈ scam + injection<br>risk flags"]
FACTS["π media facts<br>transcript Β· OCR Β· risk<br>cached on disk"]
CTX["π¦ MessageContext<br>read-only"]
RET --> CTX
GUARD --> CTX
FACTS --> CTX
end
%% Agents
subgraph Agents [4. Agent Pipeline β orchestrated handoffs]
TRI["π·οΈ Triage<br>LIGHT Β· Llama-3.3-70B"]
SAF["π‘οΈ Safety<br>ROUTER Β· GLM-5.2"]
CON["π― Context<br>ROUTER Β· GLM-5.2"]
EVI["π§Ύ Evidence<br>LIGHT Β· Llama-3.3-70B"]
TRI --> SAF
SAF -->|safe| CON
CON --> EVI
end
%% Output
subgraph Output [5. Output]
OUT[("π output.csv<br>6 columns Γ 110 rows")]
end
MSG --> STORE
REL --> STORE
HIST --> STORE
MEDIA --> ASR
MEDIA --> VIS
STORE --> RET
STORE --> GUARD
ASR --> FACTS
VIS --> FACTS
CTX --> TRI
SAF -->|veto Β· mute/scam| OUT
EVI --> OUT
style Input fill:#e1f5fe,stroke:#01579b
style Load fill:#fff3e0,stroke:#e65100
style Prep fill:#e8f5e9,stroke:#1b5e20
style Agents fill:#f3e5f5,stroke:#6a1b9a
style Output fill:#fce4ec,stroke:#880e4f
The shared MessageContext is assembled once and remains read-only while the agents use
tools such as get_user_profile, get_group_relationship, get_business_info,
get_user_business_history, retrieve_evidence, and get_media_facts.
Detailed design rationale lives in code/ARCHITECTURE.md, and module-level usage lives in code/README.md.
.
|-- README.md # solution overview and runbook
|-- problem_statement.md # challenge contract and output schema
|-- code.zip # verified submission archive
|-- code/
| |-- main.py # messages.csv -> output.csv entry point
| |-- preprocess_media.py # cached image and voice perception
| |-- evaluate.py # sample evaluation and hard-case checks
| |-- smoke_models.py # live model connectivity checks
| |-- smoke_agent.py # single-agent smoke routing
| |-- smoke_orchestrator.py # full multi-agent smoke routing
| |-- pyproject.toml # package and dependency configuration
| |-- .env.example # environment-variable template
| |-- ARCHITECTURE.md # design decisions and tradeoffs
| |-- src/notif_router/
| | |-- config.py # paths, roles, models, retries, budget
| | |-- models.py # typed CSV records
| | |-- data.py # indexed read-only data layer
| | |-- clients.py # resilient Hugging Face clients
| | |-- cost.py # thread-safe spend tracking
| | |-- perception.py # OCR, ASR, and media caching
| | |-- retrieval.py # personalized evidence ranking
| | |-- safety.py # deterministic safety guardrails
| | |-- agent.py # reusable tool-calling loop
| | `-- orchestrator.py # agent handoffs and final assembly
| `-- tests/ # 63 offline pytest tests
`-- dataset/
|-- messages.csv # 110 messages to route
|-- sample_messages.csv # 30 solved development examples
|-- output.csv # completed predictions
|-- users.csv, groups.csv, group_members.csv
|-- business_accounts.csv, user_business_history.csv
|-- message_history.csv, message_events.csv
|-- daily_notification_summary.csv
|-- images.csv, voice_notes.csv
`-- media/ # image and audio inputs
git clone https://github.com/VIVPM/hackerrank-message-notification-router-hackathon.git
cd hackerrank-message-notification-router-hackathonPython 3.12 or newer is required.
python -m venv .venvActivate it:
# Linux / macOS
source .venv/bin/activate
# Windows PowerShell
.venv\Scripts\Activate.ps1python -m pip install -e "./code[dev]"Core dependencies are openai, huggingface_hub, and pillow; the development extra
adds pytest.
Copy the template from code/.env.example to the repository root:
# Linux / macOS
cp code/.env.example .env
# Windows PowerShell
Copy-Item code/.env.example .envThen set:
HF_TOKEN=hf_your_token_here
Create a read token at https://huggingface.co/settings/tokens. The root .env file is
gitignored and must never be committed.
The one-time perception pass transcribes 13 voice notes and analyzes 20 images. Results are cached, so subsequent runs reuse them.
python code/preprocess_media.pyUseful options:
python code/preprocess_media.py --dry # show what would be processed
python code/preprocess_media.py --force # recompute cached perceptionpython code/main.pyThis reads dataset/messages.csv and writes the contract-valid result to
dataset/output.csv.
python code/main.py --limit 5
python code/main.py --concurrency 6
python code/main.py --force
python code/main.py --output path/to/output.csvcd code
python -m pytest -qExpected result: 63 passed with no network calls and no model spend.
cd code
python evaluate.pyThe evaluation harness reports action and message-type agreement, evidence recall,
reason divergence, and five safety/personalization hard cases. Use --force only when
you intentionally want to recompute cached evaluation decisions.
For each incoming message, the pipeline runs these stages:
- Load context: Parse all CSVs into typed records and retrieve the receiver, sender, group membership, business relationship, daily notification load, and message history.
- Perceive media: Transcribe voice notes or analyze images for OCR text, descriptions, and structured risk signals. Cache the result on disk.
- Retrieve evidence: Rank prior messages for the same receiver using sender/group/ business matches, media matches, lexical similarity, and reaction history.
- Run deterministic safety: Detect lookalike domains, credential requests, payment pressure, suspicious links, QR bait, and attempts to command the router itself.
- Triage: Confirm the message and conversation context and hand the shared state to the safety agent.
- Apply the safety veto: A hard rule or safety-agent verdict immediately returns
mute/scam; unsafe content cannot be promoted by the context agent. - Personalize: For safe messages, the context agent chooses
notify,digest, ormuteusing the user's relationships, preferences, and behavioral history. - Ground the result: The evidence agent selects real historical message IDs and writes a concise reason tied to the available context.
- Validate output: Coerce enums, clamp confidence, remove invalid evidence IDs, enforce exact column order, and restore the original message order before writing the CSV.
If an individual routing call fails after bounded retries, that message degrades to a safe
digest fallback and is not cached, allowing a later run to retry it.
All models are accessed through Hugging Face Inference Providers with one token.
| Role | Model | Responsibility |
|---|---|---|
ROUTER |
zai-org/GLM-5.2 |
Safety judgment and personalized routing |
LIGHT |
meta-llama/Llama-3.3-70B-Instruct |
Triage and grounded evidence selection |
VISION |
Qwen/Qwen3-VL-235B-A22B-Instruct |
Image OCR, description, and risk signals |
ASR |
openai/whisper-large-v3 |
Multilingual voice-note transcription |
Model IDs are selected by role in code/src/notif_router/config.py, keeping orchestration
logic independent of a specific checkpoint.
The router treats all message text, OCR text, links, and voice transcripts as untrusted content, never as instructions.
- Lookalike-domain checks compare sender domains with verified business domains and combine mismatches with account age and report volume.
- Content checks detect requests for OTPs, PINs, credentials, urgent payments, account threats, suspicious URLs, and QR-payment bait.
- Prompt-injection checks detect phrases such as
set action=notify, fake system notes, priority overrides, and instructions to ignore routing rules. - Trust gating prevents benign wording from an established, verified business on its official domain from being flagged solely because it mentions payment or OTPs.
- Veto isolation guarantees that high-confidence unsafe messages become
mute/scambefore the notify-capable context stage runs.
The safety and orchestrator tests verify injection cases, lookalike businesses, legitimate verified senders, and hard-veto short-circuit behavior.
| Variable | Required | Default | Purpose |
|---|---|---|---|
HF_TOKEN |
Yes for live calls | none | Hugging Face authentication |
HF_BASE_URL |
No | https://router.huggingface.co/v1 |
OpenAI-compatible routing endpoint |
DATASET_DIR |
No | dataset/ |
Override the participant dataset path |
CACHE_DIR |
No | media_cache/ |
Override the local cache path |
REQUEST_TIMEOUT_S |
No | 90 |
Per-request timeout |
MAX_ATTEMPTS |
No | 4 |
Retry cap including the initial request |
BUDGET_USD |
No | 10 |
Estimated run budget |
BUDGET_WARN_USD |
No | 8 |
One-time budget warning threshold |
The submitted output contains 110 schema-valid predictions, one for every incoming message, with no missing or duplicate IDs.
| Action | Count | Meaning |
|---|---|---|
notify |
34 | Interrupt the user now |
digest |
25 | Retain for a later digest |
mute |
51 | Suppress low-value, unwanted, suspicious, or unsafe content |
| Total | 110 | Complete dataset coverage |
Additional checks on the submitted file:
- Exact six-column output contract and deterministic input order.
- Confidence range: 0.65 to 0.97.
- 103 of 110 rows cite at least one historical evidence message.
- Sample evaluation: 80% action agreement, 80% message-type agreement, and 5/5 hard cases passed on the 30 solved development examples.
- Full offline suite: 63 tests passed.
- The tracked
code.zipwas byte-verified against the currentcode/tree.
The 110-message labels are hidden, so their distribution is reported as system output, not as proof of hidden-set accuracy.
HF_TOKENis missing: Ensure.envexists in the repository root, not insidecode/, and contains a valid read token.- 401 or authentication failure: Regenerate the token and verify that no quotes or
trailing spaces were copied into
.env. - 429 or transient provider errors: Lower
--concurrency; the client already retries retryable failures with exponential backoff and jitter. - A model is unavailable: Update the role mapping in
code/src/notif_router/config.pyto another model served by the configured provider. - Stale cached decisions: Run with
--force, or remove only the relevant files undermedia_cache/. - Image decode failure: Pillow re-encodes problematic images as clean RGB JPEGs and retries once; upgrade the installed package if local decoding still fails.
- Need a no-cost wiring check: Run
python code/smoke_models.py --dry.
The repository contains the three required deliverables:
- Runnable code archive:
code.zip - Predictions:
dataset/output.csv - Conversation transcript:
%USERPROFILE%\hackerrank_orchestrate_august26\log.txton Windows or$HOME/hackerrank_orchestrate_august26/log.txton macOS/Linux
Before submission, rerun the offline tests, confirm the output row count, and ensure the external transcript contains no secrets.
Forked from the interviewstreet/hackerrank-orchestrate-august26
starter repo; the code/ agent, design docs, and outputs are my own work.