Skip to content

Latest commit

 

History

28 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

murmr

Platform License Built with Rust

img-ezgif com-instagif

Speak sloppy, prompt sharp.

murmr turns a rambled, half-formed thought into a sharp, structured prompt you can paste straight into a coding agent like Claude. Hold a hotkey, talk, release — murmr transcribes locally with whisper and compiles your speech into a well-formed prompt.

  • Not an answer, a prompt — murmr never does the task itself; it builds the ask that gets the task done well. You stay the driver.
  • Local STT — audio is transcribed on-device with whisper. Nothing leaves your machine until you choose a cloud LLM for the compile step.
  • Bring your own LLM — cloud (Bedrock / Anthropic / OpenAI-compatible) or fully local (Ollama).
  • Push-to-talk, not always-on — hold to record, release to process.
  • Prompt templates as a voice primitive — trigger phrases ("loop this…", "review this…") compile to purpose-built prompt formats, not just cleaned-up text.
Platform Desktop app Status
macOS menu-bar app primary target, .app/.dmg via cargo tauri build
Windows tray app new — see INSTALL_WINDOWS.md
Linux headless CLI only scripts/install-linux.sh; no bundled desktop app yet

No notarized/signed release exists for any platform yet — you build locally.

Contents

The core idea: talk → get a prompt, not an answer

Hold the command hotkey and mumble a request:

"can you help me figure out why the login page is really slow, i think it's the api calls but not sure, dig into it and fix it"

murmr doesn't try to debug anything. It hands you a prompt, ready to paste into Claude:

TASK: Investigate and fix the performance issues on the login page, with focus on
API call optimization.
CONTEXT: The login page is slow; API calls are suspected as the primary bottleneck
but the root cause needs confirmation.
CONSTRAINTS: Preserve existing login functionality and security. Don't break the auth
flow beyond the performance improvements.
DELIVERABLE: A faster login page with the bottleneck identified and resolved, plus
before/after measurements.

Or start with a trigger phrase to compile a different, purpose-built template. Say "loop this: get the integration tests passing" and murmr produces a persistence-gated long-horizon brief instead:

OBJECTIVE: Get the integration tests passing.
SUCCESS PREDICATE: Every test in the integration suite passes on a clean run. This
is a property of the finished artifact, not of your confidence in it.
DOES NOT COUNT:
- Deleting, skipping, or weakening tests to make the suite green.
- A pass you cannot reproduce on a fresh run.
- Fixing some tests while leaving others broken.
VERIFICATION: Re-run the full suite after each fix. A flaky pass does not count.
PERSISTENCE: Assume a solution exists. Do not stop because it's hard or slow; stop
only when the predicate holds under verification.
RETURN: Return only the passing suite — no partial progress, plans, or excuses.

How it works

        hold hotkey, speak
                │
                ▼
   ┌─────────────────────────────────────────────┐
   │  local whisper STT → LLM cleanup / compile   │
   └─────────────────────────────────────────────┘
                │  release
                ▼
       clipboard + paste at cursor
  • Command (Super+Shift+L): speak any task; murmr compiles it into a structured TASK / CONTEXT / CONSTRAINTS / DELIVERABLE prompt. It never does the task itself.
  • Dictate (Super+Shift+K): plain dictation — strips fillers, fixes punctuation/capitalization, honors self-corrections. If your speech starts with a mode trigger ("loop this", "review this"…), it compiles that template instead.

While recording, a pill drops from the top of the screen with a live waveform and timer; it switches to "Transcribing…" while the LLM works, then copies the result to your clipboard (and pastes it at your cursor, if the platform allows it).

Quick start (CLI)

The headless CLI is the fastest way to try murmr — no bundling, no permissions, works on macOS/Windows/Linux, and needs no API key if you're running Ollama:

$ git clone https://github.com/arvmaan/murmr.git && cd murmr

$ cargo run -p murmer-core --bin murmer -- --download-model base.en
  Downloading base.en to ~/.local/share/murmer/models/ggml-base.en.bin

$ ollama pull qwen3:1.7b && ollama pull phi4-mini   # cleanup + command models

$ cargo run -p murmer-core --bin murmer -- --check
  [OK] Ollama reachable at http://localhost:11434
  [OK] Whisper model found at ~/.local/share/murmer/models/ggml-base.en.bin
  ...

$ cargo run -p murmer-core --bin murmer
  murmer ready — press your hotkey to dictate

For the full desktop experience (tray icon, recording pill, Settings UI, prompt-template editor), see Download & install.

Voice template modes

Built-in modes match a trigger phrase at the start of your speech, then compile the rest into a rigorous prompt:

Mode Triggers (start of speech) Turns speech into…
loop "loop this", "ralph this", "iterate on" a persistence-gated brief with a success predicate + verification gate
review "review this", "audit" an adversarial review brief with a failure-mode checklist
spec "spec this", "specify" a pseudo-formal specification (definitions, predicate, non-counting outcomes)
fan "fan out", "parallel" a diverse parallel-search orchestration brief

The command hotkey (Super+Shift+L) is the general case — it compiles any spoken task into a TASK / CONTEXT / CONSTRAINTS / DELIVERABLE prompt without needing a trigger word, and never executes the task.

Modes are plain config — override a built-in or add your own in Settings (or config.toml).

Download & install

Prerequisites common to every platform:

# 1. Rust:        https://rustup.rs
# 2. Tauri CLI:   cargo install tauri-cli --version "^2"

# 3. Clone and download a whisper model (~150 MB for base.en)
git clone https://github.com/arvmaan/murmr.git && cd murmr
cargo run -p murmer-core --bin murmer --features bedrock -- --download-model base.en

# 4. Create your config (see Configuration below)

Then build for your platform:

Platform Build & install Notes
macOS ./scripts/bundle-macos.sh --install grants two Privacy & Security permissions on first launch — see INSTALL.md
Windows ./scripts/bundle-windows.ps1 -Install no permission model to grant, but installer is unsigned (SmartScreen warning) — see INSTALL_WINDOWS.md
Linux ./scripts/install-linux.sh headless CLI only for now; needs wtype/wl-clipboard (Wayland) or xdotool/xclip (X11)

Or for local development:

cargo tauri dev --features bedrock     # desktop app, dev mode
cargo tauri build --features bedrock   # .app/.dmg on macOS, NSIS/MSI on Windows

Once installed, hold the dictate hotkey (Super+Shift+K by default), speak, and release.

Codebase awareness

Point murmr at your repo (Settings → Codebase awareness → set the path → Re-index) and it scans your source for identifiers — IngestedBytes, parseConfig, DictionaryStore — ranked by frequency. The top terms are injected into the cleanup prompt so speech-to-text output is corrected to your project's real symbols:

you say "the ingested bytes counter" → murmr writes "the IngestedBytes counter"

It also learns over time: recurring terms from your dictations are picked up and remembered automatically. Both paths stay off the hot path — indexing happens on demand, and only a bounded slice of the vocabulary is injected, so dictation stays fast.

LLM backends

murmr auto-detects the protocol from your config. Supported:

  • AWS Bedrock — uses your AWS credentials (no API key), just set the region
  • Anthropic — API key
  • OpenAI-compatible — endpoint + API key
  • Ollama — local, no key (fully offline with local whisper)

Repo layout (Cargo workspace)

crates/
  murmer-core/            # library: all the logic; also ships a headless CLI (bin: murmer)
    src/
      audio/              # cpal capture, Silero VAD
      stt/                # whisper-rs transcription
      llm/                # LlmClient (Ollama/OpenAI/Anthropic/Bedrock) + prompts
      modes/              # voice-template engine: registry, extractor, context, engine
      dictionary/         # adaptive vocabulary learning
      input/              # hotkeys (rdev), paste (wtype/xdotool/pbcopy/powershell)
      config.rs           # TOML config
  murmer-app/             # Tauri v2 desktop app (macOS + Windows)
    src/
      main.rs             # entry, tray, pill window, reopen handling
      recording.rs        # hotkey → capture → transcribe → LLM → paste pipeline
      commands.rs         # IPC commands for the UI
      state.rs            # shared app state
      transcripts.rs      # transcript history persistence
ui/                       # vanilla HTML/CSS/JS frontend (no build step)
  index.html style.css app.js   # main window (Transcripts / Settings)
  pill.html  pill.css  pill.js   # recording pill overlay

See INSTALL.md (macOS) or INSTALL_WINDOWS.md (Windows) for the full install + troubleshooting guide.

Configuration

Config lives at ~/.config/murmer/config.toml on macOS/Linux, or %APPDATA%\murmer\config.toml on Windows. Transcript history persists to transcripts.json alongside it.

[llm]
protocol = "bedrock"          # bedrock | anthropic | openai | ollama
region = "us-west-2"          # bedrock
cleanup_model = "us.anthropic.claude-haiku-4-5-20251001-v1:0"
command_model = "us.anthropic.claude-sonnet-4-20250514-v1:0"

[hotkeys]
dictate = "Super+Shift+K"
command = "Super+Shift+L"

[stt]
model_path = ""               # platform default under the OS's app-data dir
language = "en"

[dictionary]
entries = { "k8s" = "Kubernetes", "pg" = "Postgres" }

License

MIT

About

No description, website, or topics provided.

Resources

Stars

4 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages