Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
9 changes: 5 additions & 4 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,7 +12,7 @@
![Python](https://img.shields.io/badge/Python-3.11%20%7C%203.12-3776AB)
[![License](https://img.shields.io/badge/License-Apache--2.0-blue.svg)](LICENSE)

[简体中文](README.zh-CN.md)
[English](README.md) · [简体中文](README.zh-CN.md)

> Version status: current software is the `v1.1.2` maintenance patch. The research and evidence
> baseline remains the immutable `v1.0.0` public release. This patch aligns the packaged Studio
Expand Down Expand Up @@ -63,8 +63,9 @@ cross-family token throughput uses different token definitions.

![Formal 102-task Qwen3-VL-2B and SmolVLM2-500M comparison](docs/reports/v1.0.0-candidate/overview.svg)

Read the [byte-rebuildable report](docs/reports/v1.0.0-candidate/report.md) or
inspect the preserved
Read the benchmark report in [English](docs/reports/v1.0.0-candidate/report.md) or
[简体中文](docs/reports/v1.0.0-candidate.zh-CN.md). The English original remains byte-rebuildable.
Inspect the preserved
[Qwen JSONL](docs/reports/results/2026-08-10-qwen3-vl-v1.0.0-formal.jsonl),
[SmolVLM2 JSONL](docs/reports/results/2026-08-10-smolvlm2-v1.0.0-formal.jsonl),
and their SHA-bound manifests. Historical ten-task and document-only reports
Expand Down Expand Up @@ -311,7 +312,7 @@ equivalent because tokenizers differ.
| Run records, manifests, and resume | [Artifact contract](docs/run-records-and-manifests.md) |
| Two-model formal result | [Qwen3-VL vs SmolVLM2](docs/reports/2026-07-31-qwen3-vl-vs-smolvlm2.md) |
| 32-task document comparison | [Qwen3-VL vs SmolVLM2 on documents](docs/reports/2026-08-02-document-model-comparison.md) |
| 102-task v1.0.0 benchmark comparison | [Byte-rebuildable final-corpus bundle](docs/reports/v1.0.0-candidate/report.md) |
| 102-task v1.0.0 benchmark comparison | [English (byte-rebuildable)](docs/reports/v1.0.0-candidate/report.md) · [简体中文](docs/reports/v1.0.0-candidate.zh-CN.md) |
| Performance methodology | [Qwen formal performance baseline](docs/reports/2026-07-31-qwen3-vl-formal-performance.md) |
| Quality and public-release gates | [Quality standard](docs/06-quality-and-open-source.md) |
| Live public-release status | [Evidence matrix and strict readiness check](docs/public-release-readiness.md) |
Expand Down
9 changes: 5 additions & 4 deletions README.zh-CN.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,7 +12,7 @@
![Python](https://img.shields.io/badge/Python-3.11%20%7C%203.12-3776AB)
[![License](https://img.shields.io/badge/License-Apache--2.0-blue.svg)](LICENSE)

[English](README.md)
[English](README.md) · [简体中文](README.zh-CN.md)

> 版本状态:当前软件是 `v1.1.2` 维护修补版。研究与证据基线仍是不可变的 `v1.0.0`
> 公开正式版。本次修补让打包后的 Studio Header 回到跨产品统一几何,并澄清测量与
Expand Down Expand Up @@ -56,7 +56,8 @@ OCR、事件顺序等部分分类领先。结果适用于固定模型 revision

![102 条任务正式对比](docs/reports/v1.0.0-candidate/overview.svg)

完整证据见[可逐字节重建报告](docs/reports/v1.0.0-candidate/report.md),原始
完整报告:[English(可逐字节重建原件)](docs/reports/v1.0.0-candidate/report.md) ·
[简体中文](docs/reports/v1.0.0-candidate.zh-CN.md)。原始
[Qwen JSONL](docs/reports/results/2026-08-10-qwen3-vl-v1.0.0-formal.jsonl)、
[SmolVLM2 JSONL](docs/reports/results/2026-08-10-smolvlm2-v1.0.0-formal.jsonl)
及其 SHA 绑定 manifest 均已保存。早期 10 条任务和文档专项报告仍保留在下方
Expand Down Expand Up @@ -125,7 +126,7 @@ flowchart LR
面向发布的重建流程见[确定性报告包说明](docs/report-bundles.md)。它会先验证
“恰好一次 warm-up + 三次完整重复”、来源 manifest、模型/数据集身份和媒体哈希,
再生成完整对比包。仓库中的
[v1.0.0 正式报告](docs/reports/v1.0.0-candidate/report.md)已经覆盖完整的
[v1.0.0 正式报告(简体中文)](docs/reports/v1.0.0-candidate.zh-CN.md)已经覆盖完整的
102 条任务双模型正式网格;旧的
[重建基线](docs/reports/rebuilt-baseline/report.md)仅作为历史审计证据保留。

Expand Down Expand Up @@ -270,7 +271,7 @@ Git 状态、输出哈希与大小、记录数和严格尝试前缀。只有明
| 运行记录、manifest 和恢复 | [产物契约](docs/run-records-and-manifests.md) |
| 双模型正式结果 | [Qwen3-VL vs SmolVLM2](docs/reports/2026-07-31-qwen3-vl-vs-smolvlm2.md) |
| 32 条文档任务对比 | [Qwen3-VL 与 SmolVLM2 文档评测](docs/reports/2026-08-02-document-model-comparison.md) |
| 102 条 v1.0.0 正式对比 | [可逐字节重建的完整语料报告](docs/reports/v1.0.0-candidate/report.md) |
| 102 条 v1.0.0 正式对比 | [English(可逐字节重建)](docs/reports/v1.0.0-candidate/report.md) · [简体中文](docs/reports/v1.0.0-candidate.zh-CN.md) |
| 性能方法 | [Qwen 正式性能基线](docs/reports/2026-07-31-qwen3-vl-formal-performance.md) |
| 质量与公开门槛 | [质量标准](docs/06-quality-and-open-source.md) |
| 实时公开准备状态 | [证据矩阵与严格验收命令](docs/public-release-readiness.md) |
Expand Down
96 changes: 96 additions & 0 deletions docs/reports/v1.0.0-candidate.zh-CN.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,96 @@
# 重建的正式多模态基准报告

[English](v1.0.0-candidate/report.md) · 简体中文

> 本文完整翻译冻结的 [`docs/reports/v1.0.0-candidate/report.md`](v1.0.0-candidate/report.md),来源版本为 `5a31514dda9a43235e6073ed2ba6b0495c90657d`。译文置于证据包之外;可逐字节重建的英文原件、原始结果与清单为核验依据。模型、数据集和类别标识保持原样,便于对照机器可读结果。

本报告根据保存的结果 JSONL 及其配套清单,以确定性方式生成,没有重新运行模型。汇总前,每份来源均通过哈希、数据集与媒体、干净提交状态、模型身份和正式协议验证。

## 范围

- 2 个模型后端
- 1 份由 SHA 绑定的评测输入
- 4 个保留的数据集版本
- 102 个不同的数据集任务
- 每份来源均先完成且仅完成 1 次成功预热,再进行 3 轮完整的计量重复
- 批量大小为 1,采用确定性解码;计量期间不重试、不重新加载模型

![正式基准概览](v1.0.0-candidate/overview.svg)

## 运行汇总

| 数据集 | 后端 | 任务数 | 成功次数 | 平均得分 | 首 token 延迟(TTFT)中位数 | 吞吐量中位数 | GPU 显存峰值 |
|---|---|---:|---:|---:|---:|---:|---:|
| synthetic-docs-v1, synthetic-robustness-v1, synthetic-v1.1, synthetic-video-v1 | `qwen3-vl` | 102 | 306/306 | 0.784 | 120.5 ms | 12.7 tok/s | 4180.5 MiB |
| synthetic-docs-v1, synthetic-robustness-v1, synthetic-v1.1, synthetic-video-v1 | `smolvlm2` | 102 | 306/306 | 0.690 | 260.0 ms | 10.1 tok/s | 1265.3 MiB |

完整的机器可读数值见 [`run-summary.csv`](v1.0.0-candidate/run-summary.csv)。吞吐量使用各模型原生 tokenizer 计算,最适合比较同一固定版本模型的重复运行。

## 类别汇总

| 数据集 | 后端 | 类别标识 | 任务数 | 成功次数 | 平均得分 | TTFT 中位数 | GPU 显存峰值 |
|---|---|---|---:|---:|---:|---:|---:|
| synthetic-docs-v1, synthetic-robustness-v1, synthetic-v1.1, synthetic-video-v1 | `qwen3-vl` | `attribute-recognition` | 24 | 72/72 | 0.958 | 109.2 ms | 4089.5 MiB |
| synthetic-docs-v1, synthetic-robustness-v1, synthetic-v1.1, synthetic-video-v1 | `qwen3-vl` | `chart-qa` | 8 | 24/24 | 0.625 | 164.4 ms | 4179.4 MiB |
| synthetic-docs-v1, synthetic-robustness-v1, synthetic-v1.1, synthetic-video-v1 | `qwen3-vl` | `counting` | 6 | 18/18 | 1.000 | 105.5 ms | 4089.5 MiB |
| synthetic-docs-v1, synthetic-robustness-v1, synthetic-v1.1, synthetic-video-v1 | `qwen3-vl` | `document-key-value` | 10 | 30/30 | 0.900 | 165.9 ms | 4180.5 MiB |
| synthetic-docs-v1, synthetic-robustness-v1, synthetic-v1.1, synthetic-video-v1 | `qwen3-vl` | `document-ocr` | 6 | 18/18 | 0.833 | 172.7 ms | 4179.1 MiB |
| synthetic-docs-v1, synthetic-robustness-v1, synthetic-v1.1, synthetic-video-v1 | `qwen3-vl` | `event-order` | 4 | 12/12 | 0.000 | 100.5 ms | 4096.0 MiB |
| synthetic-docs-v1, synthetic-robustness-v1, synthetic-v1.1, synthetic-video-v1 | `qwen3-vl` | `image-description` | 2 | 6/6 | 1.000 | 110.1 ms | 4092.1 MiB |
| synthetic-docs-v1, synthetic-robustness-v1, synthetic-v1.1, synthetic-video-v1 | `qwen3-vl` | `motion-direction` | 4 | 12/12 | 0.500 | 95.6 ms | 4096.2 MiB |
| synthetic-docs-v1, synthetic-robustness-v1, synthetic-v1.1, synthetic-video-v1 | `qwen3-vl` | `occlusion-reasoning` | 3 | 9/9 | 0.667 | 108.6 ms | 4089.8 MiB |
| synthetic-docs-v1, synthetic-robustness-v1, synthetic-v1.1, synthetic-video-v1 | `qwen3-vl` | `spatial-reasoning` | 10 | 30/30 | 0.800 | 109.5 ms | 4093.4 MiB |
| synthetic-docs-v1, synthetic-robustness-v1, synthetic-v1.1, synthetic-video-v1 | `qwen3-vl` | `state-change` | 3 | 9/9 | 1.000 | 108.9 ms | 4097.3 MiB |
| synthetic-docs-v1, synthetic-robustness-v1, synthetic-v1.1, synthetic-video-v1 | `qwen3-vl` | `table-qa` | 8 | 24/24 | 0.500 | 167.5 ms | 4179.9 MiB |
| synthetic-docs-v1, synthetic-robustness-v1, synthetic-v1.1, synthetic-video-v1 | `qwen3-vl` | `temporal-counting` | 5 | 15/15 | 0.800 | 99.6 ms | 4097.5 MiB |
| synthetic-docs-v1, synthetic-robustness-v1, synthetic-v1.1, synthetic-video-v1 | `qwen3-vl` | `temporal-position` | 8 | 24/24 | 0.750 | 110.0 ms | 4097.3 MiB |
| synthetic-docs-v1, synthetic-robustness-v1, synthetic-v1.1, synthetic-video-v1 | `qwen3-vl` | `visual-comparison` | 1 | 3/3 | 1.000 | 117.8 ms | 4090.2 MiB |
| synthetic-docs-v1, synthetic-robustness-v1, synthetic-v1.1, synthetic-video-v1 | `smolvlm2` | `attribute-recognition` | 24 | 72/72 | 1.000 | 260.7 ms | 1265.2 MiB |
| synthetic-docs-v1, synthetic-robustness-v1, synthetic-v1.1, synthetic-video-v1 | `smolvlm2` | `chart-qa` | 8 | 24/24 | 0.375 | 264.2 ms | 1265.2 MiB |
| synthetic-docs-v1, synthetic-robustness-v1, synthetic-v1.1, synthetic-video-v1 | `smolvlm2` | `counting` | 6 | 18/18 | 1.000 | 261.0 ms | 1265.2 MiB |
| synthetic-docs-v1, synthetic-robustness-v1, synthetic-v1.1, synthetic-video-v1 | `smolvlm2` | `document-key-value` | 10 | 30/30 | 0.900 | 262.1 ms | 1265.2 MiB |
| synthetic-docs-v1, synthetic-robustness-v1, synthetic-v1.1, synthetic-video-v1 | `smolvlm2` | `document-ocr` | 6 | 18/18 | 1.000 | 264.6 ms | 1265.2 MiB |
| synthetic-docs-v1, synthetic-robustness-v1, synthetic-v1.1, synthetic-video-v1 | `smolvlm2` | `event-order` | 4 | 12/12 | 0.750 | 174.5 ms | 1177.3 MiB |
| synthetic-docs-v1, synthetic-robustness-v1, synthetic-v1.1, synthetic-video-v1 | `smolvlm2` | `image-description` | 2 | 6/6 | 0.500 | 262.3 ms | 1265.2 MiB |
| synthetic-docs-v1, synthetic-robustness-v1, synthetic-v1.1, synthetic-video-v1 | `smolvlm2` | `motion-direction` | 4 | 12/12 | 0.250 | 180.5 ms | 1177.3 MiB |
| synthetic-docs-v1, synthetic-robustness-v1, synthetic-v1.1, synthetic-video-v1 | `smolvlm2` | `occlusion-reasoning` | 3 | 9/9 | 0.333 | 261.3 ms | 1265.2 MiB |
| synthetic-docs-v1, synthetic-robustness-v1, synthetic-v1.1, synthetic-video-v1 | `smolvlm2` | `spatial-reasoning` | 10 | 30/30 | 0.533 | 264.0 ms | 1265.3 MiB |
| synthetic-docs-v1, synthetic-robustness-v1, synthetic-v1.1, synthetic-video-v1 | `smolvlm2` | `state-change` | 3 | 9/9 | 0.667 | 167.8 ms | 1177.3 MiB |
| synthetic-docs-v1, synthetic-robustness-v1, synthetic-v1.1, synthetic-video-v1 | `smolvlm2` | `table-qa` | 8 | 24/24 | 0.250 | 262.1 ms | 1265.2 MiB |
| synthetic-docs-v1, synthetic-robustness-v1, synthetic-v1.1, synthetic-video-v1 | `smolvlm2` | `temporal-counting` | 5 | 15/15 | 0.400 | 180.9 ms | 1177.3 MiB |
| synthetic-docs-v1, synthetic-robustness-v1, synthetic-v1.1, synthetic-video-v1 | `smolvlm2` | `temporal-position` | 8 | 24/24 | 0.500 | 173.1 ms | 1177.3 MiB |
| synthetic-docs-v1, synthetic-robustness-v1, synthetic-v1.1, synthetic-video-v1 | `smolvlm2` | `visual-comparison` | 1 | 3/3 | 1.000 | 267.8 ms | 1265.2 MiB |

## 失败记录

未记录到失败的计量尝试。仍按固定结构输出 [`failures.csv`](v1.0.0-candidate/failures.csv),以便下游自动化按统一方式处理。

## 保存的证据

| 数据集 | 后端 | 结果文件 | 结果 SHA-256 | 清单 | 清单 SHA-256 |
|---|---|---|---|---|---|
| synthetic-docs-v1, synthetic-robustness-v1, synthetic-v1.1, synthetic-video-v1 | `qwen3-vl` | `docs/reports/results/2026-08-10-qwen3-vl-v1.0.0-formal.jsonl` | `a6574423770718fe20f7bd308d09fb8246ed9d0f71a70bd5b15138ff9c90908c` | `docs/reports/results/2026-08-10-qwen3-vl-v1.0.0-formal.manifest.json` | `7b73efc0dececc8d03f998040844ffd122284dd24745907035eb33c8c825cc78` |
| synthetic-docs-v1, synthetic-robustness-v1, synthetic-v1.1, synthetic-video-v1 | `smolvlm2` | `docs/reports/results/2026-08-10-smolvlm2-v1.0.0-formal.jsonl` | `b195dc43e7d0f719c02f819c4ffab8017dd905b74cd7981d4e39bf8527fdb027` | `docs/reports/results/2026-08-10-smolvlm2-v1.0.0-formal.manifest.json` | `cc2c3b4e66b2814475d39c292ce3aec874f0940dac0f5df978777bf6aae9bca9` |

## 重建

在仓库根目录执行:

```powershell
python scripts/build_benchmark_report.py `
--input docs/reports/results/2026-08-10-qwen3-vl-v1.0.0-formal.jsonl `
--input docs/reports/results/2026-08-10-smolvlm2-v1.0.0-formal.jsonl `
--output-dir <output-directory>
```

然后验证全部来源、输出哈希、生成器哈希,以及包含自身哈希的构建清单:

```powershell
python scripts/build_benchmark_report.py `
--verify `
--output-dir <output-directory>
```

## 解释边界

这些结果仅适用于来源清单记录的固定模型、任务文件、媒体哈希、软件环境和硬件,不能视为普适模型排名,也不代表用户偏好或生产环境质量。
Loading