Skip to content

Latest commit

 

History

466 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

hopper

hopper is the sample registry, job queue, and result store for the Atomdrift malware-analysis pipeline. It catalogs files and provenance, serves bytes to authorized workers, accepts Atomdrift Scan results, and exposes the labeled data used to train Azoth.

This is infrastructure for running a scan fleet or research corpus. To scan a file on one machine, install Atomdrift Scan instead.

collectors ──► hopper ──► atomscan workers
                 │              │
                 └◄── results ──┘
                 │
                 └──► collimator / cyclotron / prism

What it provides

  • PostgreSQL and SQLite storage for samples, labels, provenance, and reports
  • Pull-based jobs for horizontally scaled atomscan worker processes
  • A file and result API with worker liveness and retry handling
  • Local filesystem ingestion from bad/, good/, sighted/, purgatory/, pending/, review/, and hot incoming/ pools
  • A dashboard for queue depth, workers, and analysis rates
  • Review, rescan, reconciliation, import, and backfill commands

Requirements

  • Go 1.25.4 or newer and CGO to build hopper
  • cleave to enumerate recognized files during ingestion
  • atomscan for the default local analysis worker
  • PostgreSQL 17 or newer for production; several corpus operations use JSON_TABLE

SQLite is useful for development and small datasets. Production deployments should use PostgreSQL and normal database authentication, backup, and network controls.

Build and test

make build
make test
./hopper

Local quick start

Create a small labeled pool:

samples/
├── bad/
├── good/
├── incoming/
├── pending/
├── purgatory/
├── review/
└── sighted/

bad/, good/, sighted/ and purgatory/ carry classification labels; pending/, review/ and incoming/ are workflow roots holding label unknown. purgatory/ is greyware — dual-use tooling and artifacts that are neither certified benign nor malicious. Every training and triage selector names the labels it wants, so a purgatory sample is outside all of them and the corpus trains on it in neither direction.

incoming/, pending/, and review/ are physical workflow roots. Samples in all three retain the catalog label unknown, and moves between them preserve the complete suffix below the root. New API uploads land in incoming/.

Then initialize SQLite and ingest it:

./hopper init --db samples.db
./hopper load \
  --db samples.db \
  --data ./samples \
  --local \
  --workers 1

load remains running: it watches the corpus, runs the local worker, and serves two listeners — the work API (--api-addr) and the HTML dashboard (--dashboard-addr, every interface by default). They are separate so each carries one access policy: the API requires a bearer token, the dashboard has none — it is read-only, so keep it on a trusted network and never route the tunnel to it. --local, or --dashboard-addr 127.0.0.1:8082, binds it back to loopback for an SSH forward.

To run a disposable local PostgreSQL instance instead, install PostgreSQL's initdb, postgres, createdb, and pg_isready tools, then run:

./hopper serve --dir ~/.hopper --port 5433
./hopper init --db postgres://localhost:5433/hopper

hopper serve uses trust authentication and is intended only for local development. Do not expose that database to a network.

Production outline

./hopper init --db "$DATABASE_URL"
./hopper load \
  --db "$DATABASE_URL" \
  --data /srv/samples \
  --api-addr 0.0.0.0:8081 \
  --token-file ~/.tok/hopper \
  --dashboard-addr 0.0.0.0:8082

--token-file requires Authorization: Bearer <token> on every API route except the liveness, readiness, and metrics probes — loopback callers included, because a Cloudflare tunnel terminates on loopback and a loopback exemption would be an internet exemption. Clients (scan's workers and uploader, hopper's own triage and post-triage) read the same token from ~/.tok/hopper. The deploy scripts generate one on first run and install it for the service user; rotation is an edit plus a restart, since it is read once at startup.

Run ./hopper <command> -h for command-specific flags. Review the deployment scripts before using them: they encode Atomdrift's own topology, database roles, replication, and service assumptions.

Protect the worker/file API, database, and dashboard as sensitive infrastructure. Hopper can accept and serve malware bytes and store authoritative labels. Only the API listener is designed to be published — and only with --token-file set. The dashboard has no authentication of its own, so it must stay on loopback and be reached over an SSH forward; never give it a tunnel hostname.

Native FreeBSD deployment

On FreeBSD, make deploy installs Hopper and the Scan worker as separate rc.d services. Hopper runs the ingestion/API process with its local worker disabled; scan-worker pulls jobs from Hopper and can read the same sample tree directly.

make deploy \
  DATA_DIR=/data/samples \
  DB='postgres://hopper@hopper-db/hopper?sslmode=disable' \
  SCAN_DIR=../scan \
  FREEBSD_WORKERS=96 \
  FREEBSD_MAX_MEMORY_GB=0

FREEBSD_MAX_MEMORY_GB=0 leaves Scan's RSS admission threshold automatic.

The deploy also builds and installs cleave, runs Hopper migrations, refreshes the tool rules, and configures hopper to listen on port 8081. The worker is supervised independently, so a worker crash or memory failure does not stop the Hopper API. The scan account must be able to read DATA_DIR; the installer adds both service accounts to the samples group by default.

Operations

Metrics are exported for Prometheus at GET /_/metrik on the API address (--api-addr, default 0.0.0.0:8081), not the dashboard one. Like the liveness and readiness probes it is exempt from the bearer token so a scraper needs no credential — keep the scrape path reachable only from your monitoring network, and scope the tunnel's ingress to the routes you mean to publish. Alert rules live in scripts/prometheus-hopper-alerts.yml and a Grafana dashboard in scripts/grafana-hopper-dashboard.json.

Useful commands

./hopper stats --db "$DATABASE_URL"
./hopper false-positives --db "$DATABASE_URL"
./hopper false-negatives --db "$DATABASE_URL"
./hopper rescan --db "$DATABASE_URL" <sha256>
./hopper import --from old.db --db "$DATABASE_URL"

# Relabel by SHA-256. The master moves the bytes into the corrected pool
# bucket and flips the label in one operation; nothing is uploaded.
./hopper mv -url "$HOPPER_URL" -target=bad <sha256> <sha256>
./hopper mv -url "$HOPPER_URL" -target=good -dry-run < shas.txt

# Park greyware where training sees it as neither class.
./hopper mv -url "$HOPPER_URL" -target=purgatory < shas.txt

# Audit hot samples queries for seq-scan regressions (needs production-like stats):
HOPPER_PLAN_DSN="$DATABASE_URL" go test ./... -run TestPlanAudit -count=1

The bare ./hopper command prints the maintained command list, including triage, cleanup, backfill, and corpus-reconciliation operations.

About

scalable malware sample store, designed to work with #atomscan workers

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages