Skip to content

Add hermes-jailbench to Benchmarks & Datasets - #129

Open
roli-lpci wants to merge 1 commit into
ProjectRecon:mainfrom
roli-lpci:add-hermes-jailbench-20260920
Open

roli-lpci wants to merge 1 commit into
ProjectRecon:mainfrom
roli-lpci:add-hermes-jailbench-20260920

Conversation

@roli-lpci

Copy link
Copy Markdown

What

Adds hermes-jailbench to the Benchmarks & Datasets section.

What it does

hermes-jailbench is a deterministic, single-turn regression benchmark for known-pattern jailbreaks. It runs a repeatable attack battery against an authorized LLM endpoint and classifies responses as refusal, partial, or compliance with an auditable keyword scorer.

Note: this is a known-pattern regression benchmark, not proof that a model or agent is safe and not coverage of novel or multi-turn attacks.

Checklist (per CONTRIBUTING.md)

  • Scope: directly related to evaluating autonomous-agent security performance.
  • Open source: MIT, public repository.
  • Maintenance: current release v0.2.2, with the repository updated this week.
  • Format: - **[Tool](URL)** - description. matching existing entries.
  • Searched the live README at main commit f0252697b78f768233ef01a1f5f58ccd117223d1 and the live open-PR search; no existing hermes-jailbench entry or matching open duplicate PR (total_count=0).

Maintained by Hermes Labs; submitted by its maintainer.

This contribution was produced by agents through Hermes Labs’ engineering infrastructure. Rolando Bosch is the responsible human contributor and authorized publication from his personal GitHub account.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant