OpenMultimodalLab is the open-source benchmark engine behind AlvenX.
Version status: current software is the
v1.1.2maintenance patch. The research and evidence baseline remains the immutablev1.0.0public release. This patch aligns the packaged Studio header with the canonical cross-product geometry and clarifies measurement and future-runtime boundaries; it does not add benchmark evidence or change published results. See the maintenance policy and portfolio evidence.
A local-first, reproducible benchmark toolkit for answering a practical question: which vision-language model works best for this task and this hardware, and what evidence supports that choice?
OpenMultimodalLab runs versioned multimodal tasks through interchangeable model adapters, preserves every output and failure as JSONL, applies deterministic task-selected scoring, and records enough configuration and environment data to rebuild a report without rerunning the model.
Status: OpenMultimodalLab v1.0.0 is the first public release. Both pinned models completed the same formal 102-task image, document, short-video, and robustness grid. Raw results, deterministic reports, video demo, license and security audits, Python 3.11/3.13 checks, fresh Windows wheel verification, and GitHub Linux CI evidence are preserved. The repository and formal v1.0.0 GitHub Release were published with explicit owner approval on 2026-08-10.
Project boundary: OpenMultimodalLab evaluates model capability under fixed tasks and inference conditions. Agent harnesses, tool-use loops, memory, and system-level agent reliability belong outside this repository. The shared AlvenX project standard and this project's conformance record make that boundary explicit.
One NVIDIA RTX 4060 Laptop GPU (8,188 MiB), the same 102 human-checked tasks, one warm-up, three complete measured repetitions, greedy decoding, batch size 1, and one clean Git commit:
| Model | Mean task score | Median TTFT | Median task latency | Peak allocated GPU memory | Runtime failures |
|---|---|---|---|---|---|
| Qwen3-VL-2B | 0.784 | 120.5 ms | 212.9 ms | 4,180.5 MiB | 0/306 |
| SmolVLM2-500M | 0.690 | 260.0 ms | 471.5 ms | 1,265.3 MiB | 0/306 |
Qwen achieved the higher aggregate score and lower median latency. SmolVLM2 used about 70% less peak allocated GPU memory and led some categories, including OCR and event order. These results describe the pinned models, 102 controlled synthetic tasks, recorded hardware, and decoding protocol; cross-family token throughput uses different token definitions.
Read the byte-rebuildable report or inspect the preserved Qwen JSONL, SmolVLM2 JSONL, and their SHA-bound manifests. Historical ten-task and document-only reports remain available in the evidence index below.
Post-release evidence note (2026-08-14): the owner re-reviewed the 10 base
image and 32 document tasks against their original media. All 102 canonical
tasks now have task-by-task reviewer/date/checklist/SHA records in
docs/reviews/. The recorded dataset hashes match the
inputs used by the formal runs; the published v1.0.0 tag remains immutable.
This is a deliberately non-cherry-picked temporal-position example from the formal grid: Qwen answered the final position correctly in all three measured repetitions, while SmolVLM2 failed all three. The copyable video tutorial connects this GIF to the exact task, commands, raw records, and deterministic rebuild script.
- A dependency-free core and deterministic
mockbackend for offline CI. - Real, lazy-loaded Qwen3-VL-2B and SmolVLM2-500M Transformers backends.
- Immutable model revisions and native processor/chat-template metadata.
- Versioned UTF-8 JSONL tasks with validation and licensed generated media.
- Exact-match, numeric-tolerance, keyword-coverage, and ordered/unordered attribute-group scoring through backward-compatible task schemas 1.0–1.2.
synthetic-v1.1: 10 owner-reviewed image-description, counting, spatial, and visual-comparison tasks over ten reproducible PNGs.synthetic-docs-v1: 32 owner-reviewed tasks over eight reproducible OCR, key-value, table, bar-chart, and line-chart images.synthetic-video-v1: 24 owner-reviewed tasks over eight deterministic, project-generated short videos.synthetic-robustness-v1: 36 owner-reviewed tasks covering small objects, low contrast, visual clutter, and partial occlusion.- A shared local-video path for both real backends: PyAV decoding, eight uniformly sampled frames, preserved sampling metadata, and no hidden processor resampling.
- Warm-up plus repeated measurement with CUDA-synchronized TTFT, generation time, throughput, preprocessing time, and peak allocated memory.
- A deterministic multi-model report-bundle builder that rejects incomplete formal grids and emits Markdown, CSV, failure data, SVG, and a self-hashed build manifest from preserved JSONL without rerunning a model.
- Typed model-load, timeout, out-of-memory, generation, and evaluation failures.
- Run record schema 0.4 with durable invocation indexes, terminal/retryable state, cumulative latency, retry policy, and cooperative deadlines.
- Strict
--resume, explicit--overwrite, output SHA-256, and atomic per-record checkpoints. - Backend-aware
doctorchecks for Python, CUDA, BF16, optional packages, and available working/model-cache disk without printing the cache path. - Bounded local input parsing for dataset/result JSONL, images, and short video, plus portable media references and path-redacted durable errors.
- Python 3.11/3.12 Linux CI, wheel builds, link/JSON/privacy checks, and an offline test suite.
flowchart LR
A["Versioned task JSONL"] --> B["Loader + validator"]
B --> C["Benchmark runner"]
C --> D["Model adapter"]
D --> E["Local VLM"]
D --> C
C --> F["Task-selected scorer"]
C --> G["Durable JSONL + manifest"]
G --> H["Reporter"]
H --> I["Rebuilt summary / comparison"]
Inference and reporting are deliberately separate. Charts and summaries are derived artifacts; raw task-level evidence remains the source of truth.
For release-grade reconstruction, the deterministic report-bundle workflow validates exact one-warm-up/three-repeat grids, source manifests, model and dataset identities, and dataset/media hashes before generating a complete comparison bundle. The committed v1.0.0 release report applies that path to the complete 102-task, two-model formal grid. The older rebuilt baseline remains available for historical auditability.
The core path does not download a model and works on Python 3.11, 3.12, or 3.13. The primary examples use Windows PowerShell:
git clone --branch v1.1.2 --depth 1 https://github.com/AlbertXXuu/OpenMultimodalLab.git
cd OpenMultimodalLab
py -3.11 -m venv .venv
.\.venv\Scripts\python.exe -m pip install --upgrade "pip>=26.2"
.\.venv\Scripts\python.exe -m pip install -e .
.\.venv\Scripts\oml.exe doctor
.\.venv\Scripts\oml.exe run `
--dataset examples/tasks/smoke.jsonl `
--output runs/smoke-001.jsonl
.\.venv\Scripts\oml.exe report `
--input runs/smoke-001.jsonl
.\.venv\Scripts\python.exe -m unittest discover -s tests -vLinux uses the same CLI arguments with POSIX environment paths:
python3.11 -m venv .venv
.venv/bin/python -m pip install --upgrade "pip>=26.2"
.venv/bin/python -m pip install -e .
.venv/bin/oml doctor
.venv/bin/oml run \
--dataset examples/tasks/smoke.jsonl \
--output runs/smoke-001.jsonl
.venv/bin/oml report --input runs/smoke-001.jsonl
.venv/bin/python -m unittest discover -s tests -vThe mock backend verifies infrastructure only. Do not use its score as a
model-quality result.
OpenMultimodalLab includes an optional local evidence workbench without changing the reproducible CLI workflow. Its Run view guides one source through model and prompt configuration into a response with live runtime evidence; Reports opens preserved JSONL records read-only, and Method explains the v1 measurement boundary. Formal benchmark evidence remains separate.
Run it inside an existing Python 3.11/3.12 model environment; the studio extra
installs the interface, not the selected model backend or a CUDA-enabled
PyTorch build.
.\.venv-ml\Scripts\python.exe -m pip install --upgrade "pip>=26.2"
.\.venv-ml\Scripts\python.exe -m pip install -e ".[studio]"
.\.venv-ml\Scripts\oml.exe doctor --backend qwen3-vl
.\.venv-ml\Scripts\oml.exe studioThe interface listens on 127.0.0.1, disables public sharing and analytics, and
serializes model calls for an 8 GB GPU. Interactive Run responses are
explicitly unscored; the interface distinguishes the first cold model load
from later warm reuse, and Clear resets both media tabs, the prompt, response,
and metrics without unloading the model. Running with that cleared prompt
restores the default visual-description prompt automatically. The AlvenX
wordmark button returns to the document top while preserving the URL and current
workspace state. It stays still at the top and scrolls immediately when reduced
motion is enabled. Use oml run for durable, comparable evidence.
See the local interface guide and security boundary.
Use a separate Python 3.11 or 3.12 environment. Install a CUDA-enabled PyTorch
build appropriate for your platform before the project extra; doctor
detects a CPU-only mismatch when an NVIDIA GPU is visible.
Example after the Qwen environment is ready:
.\.venv-ml\Scripts\oml.exe doctor --backend qwen3-vl
.\.venv-ml\Scripts\oml.exe run `
--backend qwen3-vl `
--dataset examples/tasks/synthetic-v1.1.jsonl `
--warmup 1 `
--repetitions 3 `
--max-new-tokens 64 `
--output runs/qwen3-vl-formal-001.jsonlThe first real run downloads the pinned model into the Hugging Face user
cache. Model weights and raw runs/ outputs are excluded from Git.
oml run refuses to replace an existing output or manifest. After a compatible
interruption, repeat the exact command with --resume:
.\.venv-ml\Scripts\oml.exe run `
--backend qwen3-vl `
--dataset examples/tasks/synthetic-v1.1.jsonl `
--warmup 1 `
--repetitions 3 `
--max-new-tokens 64 `
--output runs/qwen3-vl-formal-001.jsonl `
--resumeResume verifies dataset/media hashes, tasks and order, model revision,
generation settings, environment, Git state, output hash and size, record
count, and the exact attempt prefix before appending anything. Use
--overwrite only when replacing the old evidence is intentional.
See the strict resume report for the crash-consistency model and failure-injection evidence.
For long local runs, bounded retry and timeout policy is explicit:
.\.venv-ml\Scripts\oml.exe run `
--backend qwen3-vl `
--dataset examples/tasks/synthetic-docs-v1.jsonl `
--attempt-timeout-seconds 120 `
--max-retries 1 `
--output runs/qwen3-vl-docs-001.jsonlRetries apply only to timeout and generation_error; invalid input, model
load failure, and out-of-memory are terminal. Built-in model deadlines are
cooperative and begin after one-time model loading. They can bound preprocessing
and Transformers generation, but cannot safely preempt a running CUDA kernel.
A publishable run should:
- use immutable task and model revisions;
- preserve the stored prompts and semantic media;
- use deterministic decoding and batch size 1;
- perform exactly one warm-up and three complete measured repetitions;
- retain slow attempts and failures;
- publish raw JSONL, the manifest, metric definitions, and limitations.
The evaluation protocol defines timing and comparison boundaries. Cross-model token throughput is not treated as directly equivalent because tokenizers differ.
| Area | Current state |
|---|---|
| Real image backends | Qwen3-VL-2B and SmolVLM2-500M verified locally |
| Current versioned task corpus | 102 licensed, human-checked image, document, short-video, and robustness tasks |
| Preserved real-model comparison | Both pinned models, 102 tasks, 1 warm-up + 3 repetitions, 612 measured attempts |
| Performance protocol | Warm-up, three repetitions, TTFT, throughput, latency, peak memory |
| Reliability | Durable records, strict resume, integrity hashes, typed failures |
| Automated quality | Python 3.11/3.12 Linux CI, local 3.11/3.13 tests, repository audit, fresh Windows wheel smoke |
| Release status | Public repository and formal v1.0.0 GitHub Release |
The target of at least 100 human-checked tasks is complete, and the repository
and formal v1.0.0 Release are public. Future versions must repeat the same
evidence and approval workflow rather than modifying this release in place.
Post-v1 work focuses on questions the released evidence does not yet answer: first-user reproducibility, stricter artifact validation, and one justified low-VRAM comparison at a time. New adapters must bring immutable revisions, license evidence, contract tests, and a distinct evaluation question. The maintenance roadmap separates measured results from targets and defines when proposed work should be rejected.
Small, testable contributions that improve reproducibility are welcome. Before proposing another model, include its exact revision, license, installation path, verified hardware, adapter contract tests, and known limitations.
After installing the core package, run the bounded contributor smoke first. It uses the mock backend, forces the supported offline environment flags, validates the generated JSONL and manifest, and removes its temporary artifacts:
.\.venv\Scripts\python.exe scripts\contributor_smoke.py.\.venv\Scripts\python.exe -m unittest discover -s tests -v
.\.venv\Scripts\python.exe scripts/check_repository.py
.\.venv\Scripts\python.exe scripts/check_release_readiness.pyRead CONTRIBUTING.md for the full checklist. Sensitive vulnerabilities follow SECURITY.md, not a public bug report.
Trying the project for the first time is also a contribution. Share a structured first-run report with your environment, the path you tried, what worked, and the single most valuable improvement. Remove secrets, private media, usernames, and absolute local paths before posting. Independently verified outcomes are counted under the explicit rules in the adoption ledger; Stars and views are not treated as usage.
Project code and project-generated synthetic media use Apache-2.0. Model weights and optional runtime dependencies retain their own licenses and are not distributed in this repository. See THIRD_PARTY_NOTICES.md.
Apache-2.0 does not grant project-name or trademark rights beyond the limits in its Section 6. The current brand posture and owner decisions that must precede stronger public claims are recorded in IP and brand readiness.
Copyright and project attribution: AlbertXXuu. AlvenX and the
AlvenX wordmark are unregistered project trademarks claimed by
AlbertXXuu; ALONICA is the developer ID used in selected project surfaces,
not a separate company or legal owner. See TRADEMARKS.md and
NOTICE.
