A safe, local adversarial-learning sandbox where a defender creates increasingly difficult synthetic passwords and an attacker adapts its bounded guessing strategy across rounds.
Password Arena is an educational simulation. Never enter real credentials or connect it to a login system.
The project explores a practical question:
Can an attacker agent improve its strategy while a defender agent learns that predictable human-style passwords are weaker than cryptographically secure randomness?
The project includes transparent, rule-based agents to establish a reproducible and trustworthy baseline (when using deterministic-test generator mode), alongside support for local and hosted LLM agents.
- Estimated entropy and structural penalties
- Attacker success rate
- Guesses used per round
- Runtime per attack
- Attacker strategy selection
- Defender password-family progression
- Defender entropy change per recorded model token for complete comparable tournament trials
- Agent observations across rounds
- Two-sided audit reports showing defender decisions, attacker budget allocation, outcomes, and learning updates
- Adaptive defender: escalates from dictionary words to passphrases and finally CSPRNG-generated passwords.
- Adaptive attacker: ranks common-password, mutation, passphrase, and bounded-random strategies using prior synthetic results.
- Multi-Model Support: Configure AI-vs-AI matchups using Gemini, OpenAI, Anthropic, Ollama, and Rule-Based agents. Thinking-level choices are restricted to what the selected model's own capability registry accepts, not shown unconditionally.
- Optional Hugging Face discovery: Search public model metadata only after an explicit button click. A discovered model ID is copied into manual input; it is not registered as an execution provider and is never downloaded or run.
- Tournament Engine: Build robust multi-model evaluation matrices, run repeated trials, and enforce budget constraints (cost, time, tokens). Provider availability is checked only on an explicit "Test connections" click, cached until the configuration changes -- never on every widget interaction.
- Evaluator: calculates strength indicators and captures experiment metrics.
- Dashboard: visualizes learning curves, generates weighted leaderboards, plots per-role efficiency (including defender entropy gain per 1K tokens when measured), filters results by role/provider/model/thinking-level/comparability, compares saved execution versions, exports experiment and tournament JSON/Markdown/CSV, and creates fail-closed public benchmark JSONL/CSV files plus a Dataset Card.
The MVP performs persistent-state adaptation; it does not claim to retrain model weights.
Requires Python 3.11 or newer.
python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -e ".[all]"The core package has no required third-party dependencies. Install only the features you need:
pip install -e "." # Offline rule-based CLI
pip install -e ".[dashboard]" # Dashboard without Hub discovery
pip install -e ".[dashboard,hf]" # Dashboard with Hub catalog searchHF_TOKEN is optional. It is read only when Search Hugging Face is clicked.
When absent, discovery explicitly disables use of any machine-persisted Hub login.
Discovery calls only the model-listing API. Password Arena never downloads model
weights, performs Hub inference, uploads a dataset, or adds a Hugging Face execution
provider.
Run the CLI:
password-arena --rounds 8 --max-guesses 5000 \
--output results/run.json \
--report results/run.mdLaunch the dashboard (Arena and Tournament modes):
streamlit run src/password_arena/dashboard.pyOr run the dashboard with Docker:
docker build -t password-arena .
docker run --rm -p 8501:8501 password-arenaRun quality checks:
ruff check .
mypy src/password_arena
pytest- The defender generates a synthetic password for the current difficulty.
- The evaluator records strength characteristics.
- The attacker spends a fixed guess budget across ranked strategies.
- Both agents observe the synthetic result and update local state.
- The evaluator creates a factual two-sided arena report from the recorded actions.
- Difficulty increases and the next round begins.
Tournament results also offer an explicit public benchmark boundary. JSONL and CSV contain one scalar, allowlisted row per recorded round, including recorded excluded rounds. Dataset generation never exports passwords, candidates, prompts, notes, events, model prose, or reasoning content. Saved tournaments require every linked experiment to be present before public downloads are enabled.
Every round documents the experiment from three views:
- Defender: selected password family, concrete actions, observed outcome, and state update.
- Attacker: ranked strategy, exact bounded guess allocation, attempted strategies, and adaptation.
- Evaluator: measured result and a security lesson grounded in the round metrics.
The journal is an audit log, not a request for private model chain-of-thought. Passwords and matched candidates remain redacted unless reveal_passwords is explicitly enabled. Export it from the dashboard or with --report results/run.md.
Methodology Warning: Results generated under different protocol versions should not be interpreted as directly interchangeable. Protocol and strategy changes are versioned and documented.
Protocol generation Benchmarks Key difference 1.0 001-002 initial bounded benchmark 1.1 003+ calibration-aware attacker coverage later extensions 004+ information policy / privilege controls
- rule vs rule
- 15 comparable rounds
- establishes deterministic baseline
- first local LLM benchmark
- 3 role configurations
- 45 comparable rounds
- no exclusions
- 0 actual attack solves
- exposed an important calibration finding where weak targets can survive if the bounded attacker lacks the right strategy.
Read Benchmark 002 Report | Dataset CSV | Dataset JSONL
- added benchmark protocol version 1.1
- integrated CalibrationPolicy to explicitly flag weak targets
- bounded attacker configuration added exhaustive strategy for lengths 1-3
- 45 comparable rounds
- 0 actual attack solves
- surfaced weak target survival as explicitly flagged calibration warnings rather than presenting misleading LLM success rates.
Read Benchmark 003 Report | Dataset CSV | Dataset JSONL
- benchmarked 7 information-sharing policies using
qwen3:4blocally - matrix of primary (004), zero-knowledge replication (005), and cross-run accumulated knowledge (006)
- 420 comparable rounds across multiple seeds
- explicitly measured the impact of model context and transparency on adversarial success
-
Does information sharing improve adversarial adaptation? For this local 4B parameter model, information sharing did not lead to any attacker success (0 solves). For the defender, providing more information (mutual sharing) paradoxically reduced overall password entropy gain (from ~106 bits in
frozento ~34-36 bits inmutualsharing), suggesting the model was overwhelmed or distracted by the additional structured context. -
Does self-learning help when neither agent sees the other's detailed behavior? No, the
self_onlypolicy performed significantly worse for the defender (24.77 bits entropy gain) compared to thefrozenbaseline (106.49 bits), showing that self-observation alone actually degraded generator performance for this specific 4B model. -
Does one-way learning favor attacker or defender? One-way learning favored the defender only when the attacker received the information (
attacker_observes_defender), which surprisingly resulted in a higher defender entropy gain (65.0 bits) than when the defender received the attacker's information (defender_observes_attackerat 34.29 bits). -
Does mutual information sharing create useful co-adaptation? No evidence of useful co-adaptation was found. Mutual sharing policies consumed drastically more tokens (~10k-14k vs ~2k) but produced lower quality passwords and identical (zero) solve rates.
-
Is full transparency actually better than bounded structured information? Full transparency (
mutual_full) performed similarly to bounded sharing (mutual_bounded) in entropy gain (36.35 vs 34.29) but heavily increased token cost (14.5k vs 10k tokens), showing no meaningful performance benefit over bounded structured information. -
If we repeat the exact same experiment from zero knowledge, how reproducible are the results? Results are highly reproducible at a macro level (0.00 solve rates across all policies). However, token usage, latency, and exact generated entropy fluctuate across replications due to token generation variance.
-
If agents are allowed to carry SAFE knowledge from a completed campaign into a new campaign, how much do they improve? Continuous multi-campaign learning (
benchmark-006) yielded an average entropy gain of 55.18, which is an improvement over the single-campaignmutual_boundedrun (34.29). Accumulating safe knowledge over longer time horizons helped the defender model adjust gradually better than in single runs. -
How much additional efficiency/entropy is gained by carrying over safe learning? The extended multi-run campaigns consumed roughly 70k tokens to gain ~55 bits of entropy. This is drastically less efficient per-token than the single-run
frozenbaseline (106 bits for ~2.3k tokens), proving that while learning is possible, this small model struggles to leverage historical observations efficiently.
Read Benchmark 004 Report | Dataset CSV | Dataset JSONL Read Benchmark 005 Report | Dataset CSV | Dataset JSONL Read Benchmark 006 Report | Dataset CSV | Dataset JSONL
- normal current-generation control produced measurable solves (13.3%)
- attacker privilege changed solve rate modestly (16.7%)
- defender privilege altered attacker success (6.7%)
- mutual privilege did not improve the attacker (0.0%)
- oracle achieved 100% one-guess solves (oracle != leaderboard result)
Read Benchmark 007 Report | Dataset CSV | Dataset JSONL
- first cross-model-family comparison; introduces
gemma3:4b(Google) alongsideqwen3:4b(Alibaba) under an identical matched protocol - Gemma qualified cleanly: 100% schema-valid (20/20) structured-output calls before benchmarking
- re-run in full on 2026-08-13 (identical code, protocol, seeds, budgets) as a reproducibility check; numbers below are from that re-run — 179 comparable rounds across the matched 6-scenario matrix (89 Qwen, 90 Gemma), plus a 30-round exploratory cross-model pilot
- context-overload-style defender degradation generalized, but asymmetrically, and reproduced across both runs: Qwen's entropy collapsed from ~36-92 bits to ~36 bits under mutual-information sharing; Gemma stayed flat at ~38-57 bits throughout
- Gemma was the more reliable participant in both runs (0/18 trials interrupted vs Qwen's 1/18 this run, 3/18 originally)
- each run's single non-zero solve rate landed in a different (model, scenario) cell — informative in itself: those solves read as noise, not a stable per-scenario effect (see the Reproducibility check section in the report)
- discovered and documented that the local Ollama container runs GPU-accelerated (Vulkan/AMD), not CPU-only as intended — likely true of Benchmarks 002-007 as well
Read Benchmark 008 Report | Model Comparison | Dataset CSV | Dataset JSONL
These experiments started as Qwen3 4B-only local runs; Benchmark 008 added Gemma 3 4B to test generalization. The following patterns have emerged:
Benchmark 002 showed very weak targets can survive if attacker strategy coverage is incomplete.
Benchmark 003 added protocol/calibration controls so weak-target survival is explicitly surfaced.
Benchmarks 004–006 showed that qwen3:4b often produced lower defender entropy and much higher token usage under self-learning/mutual-context modes than under frozen.
Cross-run memory improved some defender behavior but at a large token cost.
Fresh replication preserved some macro outcomes while latency/tokens/entropy varied.
Benchmark 007 shows privileged information changes outcomes.
The oracle path proves the solve pipeline itself can succeed deterministically when the target is known.
Benchmark 008 reran the compact 6-scenario matrix on gemma3:4b under the same code and protocol, then repeated the full matrix again on 2026-08-13 to check reproducibility. Qwen's defender entropy collapse under mutual-information sharing replicated in both runs (~36-92 bits down to ~36 bits). Gemma's defender stayed flat and low (~38-57 bits) across every scenario in both runs, so the same mechanism — added cross-agent context correlating with weaker defender behavior, not better attacker or defender outcomes — appears in both model families, but each model's starting point and sensitivity differ. Gemma also completed the matrix with zero interrupted trials in both runs, versus one to three for Qwen. Individual entropy values drifted by single-digit-to-~18-bit amounts run to run, and the one non-zero solve rate landed in a different scenario each time — reinforcing that the mechanism is reproducible even where exact numbers are not.
| Setting | Purpose | Default |
|---|---|---|
rounds |
Number of attacker-versus-defender rounds | 8 |
start_difficulty |
Initial defender level from 1–10 | 1 |
difficulty_step |
Increase after each round | 1 |
max_guesses |
Hard attack budget per round | 5,000 |
seed |
Reproducible baseline behavior | 42 |
reveal_passwords |
Display synthetic passwords | false |
generator_mode |
secure (CSPRNG) or deterministic-test (PRNG) |
secure |
- blueprint.md defines the product vision, architecture, safety model, and delivery phases.
- improvements.md is the prioritized feature and engineering backlog.
- bugs.md tracks confirmed defects, reproduction steps, and acceptance criteria.
- AGENTS.md defines the workflow coding agents must follow when changing the repository.
- docs/MODEL_PROVIDERS.md specifies planned OpenAI, Anthropic, Gemini, Ollama/local, thinking-level, and availability behavior.
- docs/DATASET_EXPORT.md defines the public benchmark schema, null semantics, fail-closed validation, Dataset Card, and no-upload policy.
- Synthetic passwords only
- Local comparisons only
- No login endpoints or credential datasets
- Bounded guess budgets
- Passwords hidden by default
- Cryptographically secure generation at high defender levels
See SECURITY.md for the full policy, docs/REPORTING.md for journal semantics, docs/MODEL_PROVIDERS.md for multi-model design, docs/DATASET_EXPORT.md for public-export guarantees, and docs/ROADMAP.md for planned delivery phases.