Skip to content

Latest commit

Β 

History

10 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

Message Notification Router

A personalized, multimodal WhatsApp notification-routing system that decides whether each incoming message should notify the user now, wait for a digest, or be **mute**d as low-value, repetitive, unwanted, suspicious, or unsafe.

Built for the HackerRank Orchestrate Message Notification Router challenge, the solution combines structured user context, behavioral history, retrieval, deterministic safety rules, OCR, speech transcription, and four tool-calling agents. It produces one grounded, schema-valid decision for every text message, image poster/screenshot, and voice note.

Each output row contains:

message_id,action,message_type,reason,confidence,evidence_message_ids

Features

  • Personalized routing: Uses user behavior, quiet hours, group relationships, business history, sender trust, notification load, and prior reactions instead of applying one global rule to everyone.
  • Multimodal perception: Transcribes voice notes with Whisper and extracts OCR, descriptions, and risk signals from images with Qwen3-VL.
  • Four-agent orchestration: Triage, safety, context, and evidence agents have explicit responsibilities, model tiers, handoffs, and schema validation.
  • Safety veto before routing: Deterministic phishing, scam, and prompt-injection checks can force mute/scam before untrusted message content reaches a routing model.
  • Historical-evidence retrieval: Ranks the receiver's own message history using sender, group, business, media, lexical similarity, and past reaction signals.
  • Grounded explanations: Evidence IDs are clamped to real historical messages; invalid or hallucinated IDs are removed.
  • Bounded concurrency: Routes several messages in parallel without firing all 110 requests at once, then restores deterministic input order before writing the CSV.
  • Crash-safe caching: OCR, transcripts, evaluation decisions, and final decisions are cached under the gitignored media_cache/ directory.
  • Budget visibility: Tracks calls, tokens, estimated spend, and budget warnings by model.
  • Offline validation: Includes 63 tests covering data loading, safety, retrieval, perception, tool loops, orchestration, concurrency, and output guarantees.

Architecture

graph LR
    %% Input
    subgraph Input [1. Input]
        MSG[("πŸ“¨ messages.csv<br>110 messages")]
        REL[("πŸ‘₯ users Β· groups<br>businesses")]
        HIST[("πŸ•’ message history<br>+ events")]
        MEDIA[("πŸ–ΌοΈ images<br>πŸ”Š voice notes")]
    end

    %% Load + perception
    subgraph Load [2. Load & Perceive]
        STORE["πŸ—‚οΈ typed DataStore<br>indexed lookups"]
        ASR["πŸ”Š Whisper large-v3<br>transcribe"]
        VIS["πŸ‘οΈ Qwen3-VL-235B<br>OCR + describe"]
    end

    %% Deterministic assembly
    subgraph Prep [3. Context Assembly β€” deterministic]
        RET["πŸ” evidence retrieval<br>ranked history"]
        GUARD["πŸ›‘οΈ scam + injection<br>risk flags"]
        FACTS["πŸ“‘ media facts<br>transcript Β· OCR Β· risk<br>cached on disk"]
        CTX["πŸ“¦ MessageContext<br>read-only"]
        RET --> CTX
        GUARD --> CTX
        FACTS --> CTX
    end

    %% Agents
    subgraph Agents [4. Agent Pipeline β€” orchestrated handoffs]
        TRI["🏷️ Triage<br>LIGHT · Llama-3.3-70B"]
        SAF["πŸ›‘οΈ Safety<br>ROUTER Β· GLM-5.2"]
        CON["🎯 Context<br>ROUTER · GLM-5.2"]
        EVI["🧾 Evidence<br>LIGHT · Llama-3.3-70B"]
        TRI --> SAF
        SAF -->|safe| CON
        CON --> EVI
    end

    %% Output
    subgraph Output [5. Output]
        OUT[("πŸ“„ output.csv<br>6 columns Γ— 110 rows")]
    end

    MSG --> STORE
    REL --> STORE
    HIST --> STORE
    MEDIA --> ASR
    MEDIA --> VIS
    STORE --> RET
    STORE --> GUARD
    ASR --> FACTS
    VIS --> FACTS
    CTX --> TRI
    SAF -->|veto Β· mute/scam| OUT
    EVI --> OUT

    style Input fill:#e1f5fe,stroke:#01579b
    style Load fill:#fff3e0,stroke:#e65100
    style Prep fill:#e8f5e9,stroke:#1b5e20
    style Agents fill:#f3e5f5,stroke:#6a1b9a
    style Output fill:#fce4ec,stroke:#880e4f
Loading

The shared MessageContext is assembled once and remains read-only while the agents use tools such as get_user_profile, get_group_relationship, get_business_info, get_user_business_history, retrieve_evidence, and get_media_facts.

Detailed design rationale lives in code/ARCHITECTURE.md, and module-level usage lives in code/README.md.


Project Structure

.
|-- README.md                         # solution overview and runbook
|-- problem_statement.md              # challenge contract and output schema
|-- code.zip                          # verified submission archive
|-- code/
|   |-- main.py                       # messages.csv -> output.csv entry point
|   |-- preprocess_media.py           # cached image and voice perception
|   |-- evaluate.py                   # sample evaluation and hard-case checks
|   |-- smoke_models.py               # live model connectivity checks
|   |-- smoke_agent.py                # single-agent smoke routing
|   |-- smoke_orchestrator.py         # full multi-agent smoke routing
|   |-- pyproject.toml                # package and dependency configuration
|   |-- .env.example                  # environment-variable template
|   |-- ARCHITECTURE.md               # design decisions and tradeoffs
|   |-- src/notif_router/
|   |   |-- config.py                 # paths, roles, models, retries, budget
|   |   |-- models.py                 # typed CSV records
|   |   |-- data.py                   # indexed read-only data layer
|   |   |-- clients.py                # resilient Hugging Face clients
|   |   |-- cost.py                   # thread-safe spend tracking
|   |   |-- perception.py             # OCR, ASR, and media caching
|   |   |-- retrieval.py              # personalized evidence ranking
|   |   |-- safety.py                 # deterministic safety guardrails
|   |   |-- agent.py                  # reusable tool-calling loop
|   |   `-- orchestrator.py           # agent handoffs and final assembly
|   `-- tests/                         # 63 offline pytest tests
`-- dataset/
    |-- messages.csv                   # 110 messages to route
    |-- sample_messages.csv            # 30 solved development examples
    |-- output.csv                     # completed predictions
    |-- users.csv, groups.csv, group_members.csv
    |-- business_accounts.csv, user_business_history.csv
    |-- message_history.csv, message_events.csv
    |-- daily_notification_summary.csv
    |-- images.csv, voice_notes.csv
    `-- media/                          # image and audio inputs

Setup Instructions

1. Clone the repository

git clone https://github.com/VIVPM/hackerrank-message-notification-router-hackathon.git
cd hackerrank-message-notification-router-hackathon

2. Create a Python environment

Python 3.12 or newer is required.

python -m venv .venv

Activate it:

# Linux / macOS
source .venv/bin/activate

# Windows PowerShell
.venv\Scripts\Activate.ps1

3. Install the package and test dependencies

python -m pip install -e "./code[dev]"

Core dependencies are openai, huggingface_hub, and pillow; the development extra adds pytest.

4. Configure the Hugging Face token

Copy the template from code/.env.example to the repository root:

# Linux / macOS
cp code/.env.example .env

# Windows PowerShell
Copy-Item code/.env.example .env

Then set:

HF_TOKEN=hf_your_token_here

Create a read token at https://huggingface.co/settings/tokens. The root .env file is gitignored and must never be committed.


Running the Application

Preprocess multimodal inputs

The one-time perception pass transcribes 13 voice notes and analyzes 20 images. Results are cached, so subsequent runs reuse them.

python code/preprocess_media.py

Useful options:

python code/preprocess_media.py --dry      # show what would be processed
python code/preprocess_media.py --force    # recompute cached perception

Route all messages

python code/main.py

This reads dataset/messages.csv and writes the contract-valid result to dataset/output.csv.

python code/main.py --limit 5
python code/main.py --concurrency 6
python code/main.py --force
python code/main.py --output path/to/output.csv

Run offline tests

cd code
python -m pytest -q

Expected result: 63 passed with no network calls and no model spend.

Evaluate against solved samples

cd code
python evaluate.py

The evaluation harness reports action and message-type agreement, evidence recall, reason divergence, and five safety/personalization hard cases. Use --force only when you intentionally want to recompute cached evaluation decisions.


How It Works

For each incoming message, the pipeline runs these stages:

  1. Load context: Parse all CSVs into typed records and retrieve the receiver, sender, group membership, business relationship, daily notification load, and message history.
  2. Perceive media: Transcribe voice notes or analyze images for OCR text, descriptions, and structured risk signals. Cache the result on disk.
  3. Retrieve evidence: Rank prior messages for the same receiver using sender/group/ business matches, media matches, lexical similarity, and reaction history.
  4. Run deterministic safety: Detect lookalike domains, credential requests, payment pressure, suspicious links, QR bait, and attempts to command the router itself.
  5. Triage: Confirm the message and conversation context and hand the shared state to the safety agent.
  6. Apply the safety veto: A hard rule or safety-agent verdict immediately returns mute/scam; unsafe content cannot be promoted by the context agent.
  7. Personalize: For safe messages, the context agent chooses notify, digest, or mute using the user's relationships, preferences, and behavioral history.
  8. Ground the result: The evidence agent selects real historical message IDs and writes a concise reason tied to the available context.
  9. Validate output: Coerce enums, clamp confidence, remove invalid evidence IDs, enforce exact column order, and restore the original message order before writing the CSV.

If an individual routing call fails after bounded retries, that message degrades to a safe digest fallback and is not cached, allowing a later run to retry it.


Models

All models are accessed through Hugging Face Inference Providers with one token.

Role Model Responsibility
ROUTER zai-org/GLM-5.2 Safety judgment and personalized routing
LIGHT meta-llama/Llama-3.3-70B-Instruct Triage and grounded evidence selection
VISION Qwen/Qwen3-VL-235B-A22B-Instruct Image OCR, description, and risk signals
ASR openai/whisper-large-v3 Multilingual voice-note transcription

Model IDs are selected by role in code/src/notif_router/config.py, keeping orchestration logic independent of a specific checkpoint.


Safety Design

The router treats all message text, OCR text, links, and voice transcripts as untrusted content, never as instructions.

  • Lookalike-domain checks compare sender domains with verified business domains and combine mismatches with account age and report volume.
  • Content checks detect requests for OTPs, PINs, credentials, urgent payments, account threats, suspicious URLs, and QR-payment bait.
  • Prompt-injection checks detect phrases such as set action=notify, fake system notes, priority overrides, and instructions to ignore routing rules.
  • Trust gating prevents benign wording from an established, verified business on its official domain from being flagged solely because it mentions payment or OTPs.
  • Veto isolation guarantees that high-confidence unsafe messages become mute/scam before the notify-capable context stage runs.

The safety and orchestrator tests verify injection cases, lookalike businesses, legitimate verified senders, and hard-veto short-circuit behavior.


Environment Variables

Variable Required Default Purpose
HF_TOKEN Yes for live calls none Hugging Face authentication
HF_BASE_URL No https://router.huggingface.co/v1 OpenAI-compatible routing endpoint
DATASET_DIR No dataset/ Override the participant dataset path
CACHE_DIR No media_cache/ Override the local cache path
REQUEST_TIMEOUT_S No 90 Per-request timeout
MAX_ATTEMPTS No 4 Retry cap including the initial request
BUDGET_USD No 10 Estimated run budget
BUDGET_WARN_USD No 8 One-time budget warning threshold

Results

The submitted output contains 110 schema-valid predictions, one for every incoming message, with no missing or duplicate IDs.

Routing distribution

Action Count Meaning
notify 34 Interrupt the user now
digest 25 Retain for a later digest
mute 51 Suppress low-value, unwanted, suspicious, or unsafe content
Total 110 Complete dataset coverage

Additional checks on the submitted file:

  • Exact six-column output contract and deterministic input order.
  • Confidence range: 0.65 to 0.97.
  • 103 of 110 rows cite at least one historical evidence message.
  • Sample evaluation: 80% action agreement, 80% message-type agreement, and 5/5 hard cases passed on the 30 solved development examples.
  • Full offline suite: 63 tests passed.
  • The tracked code.zip was byte-verified against the current code/ tree.

The 110-message labels are hidden, so their distribution is reported as system output, not as proof of hidden-set accuracy.


Troubleshooting

  • HF_TOKEN is missing: Ensure .env exists in the repository root, not inside code/, and contains a valid read token.
  • 401 or authentication failure: Regenerate the token and verify that no quotes or trailing spaces were copied into .env.
  • 429 or transient provider errors: Lower --concurrency; the client already retries retryable failures with exponential backoff and jitter.
  • A model is unavailable: Update the role mapping in code/src/notif_router/config.py to another model served by the configured provider.
  • Stale cached decisions: Run with --force, or remove only the relevant files under media_cache/.
  • Image decode failure: Pillow re-encodes problematic images as clean RGB JPEGs and retries once; upgrade the installed package if local decoding still fails.
  • Need a no-cost wiring check: Run python code/smoke_models.py --dry.

Submission Artifacts

The repository contains the three required deliverables:

  1. Runnable code archive: code.zip
  2. Predictions: dataset/output.csv
  3. Conversation transcript: %USERPROFILE%\hackerrank_orchestrate_august26\log.txt on Windows or $HOME/hackerrank_orchestrate_august26/log.txt on macOS/Linux

Before submission, rerun the offline tests, confirm the output row count, and ensure the external transcript contains no secrets.


Forked from the interviewstreet/hackerrank-orchestrate-august26 starter repo; the code/ agent, design docs, and outputs are my own work.

About

Personalized multimodal WhatsApp notification router using multi-agent AI, safety checks, OCR, voice transcription, retrieval, and user context to intelligently notify, digest, or mute messages.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages