From 9e4e728e73fd4e26b1f49d7ea063720d128bf31f Mon Sep 17 00:00:00 2001 From: armstrongttwalker-alt Date: Sat, 5 Sep 2026 07:52:42 +0000 Subject: [PATCH] Auto-update ModelScope documentation [$(TZ='Asia/Shanghai' date +'%Y-%m-%d %H:%M')] --- docs/flagrelease_en/model_list.txt | 28 +- ...elease_GLM-5.3-Flash-BF16-ascend-FlagOS.md | 142 ++++++++++ ...Release_GLM-5.3-Flash-BF16-hygon-FlagOS.md | 159 +++++++++++ ...ase_GLM-5.3-Flash-BF16-kunlunxin-FlagOS.md | 167 +++++++++++ ...Release_GLM-5.3-Flash-BF16-metax-FlagOS.md | 232 +++++++++++++++ ...elease_GLM-5.3-Flash-BF16-nvidia-FlagOS.md | 263 ++++++++++++++++++ ...se_GLM-5.3-Flash-BF16-tsingmicro-FlagOS.md | 125 +++++++++ ...elease_GLM-5.3-Flash-BF16-zhenwu-FlagOS.md | 149 ++++++++++ ...lease_GLM-5.3-Flash-FP8-mthreads-FlagOS.md | 134 +++++++++ ...Release_Hy4-preview-FP8-mthreads-FlagOS.md | 140 ++++++++++ ...agRelease_Hy4-preview-FP8-nvidia-FlagOS.md | 233 ++++++++++++++++ ...gRelease_Hy4-preview-INT8-ascend-FlagOS.md | 140 ++++++++++ ...agRelease_Hy4-preview-INT8-hygon-FlagOS.md | 155 +++++++++++ ...lease_Hy4-preview-INT8-kunlunxin-FlagOS.md | 176 ++++++++++++ ...agRelease_Hy4-preview-INT8-metax-FlagOS.md | 216 ++++++++++++++ ...gRelease_Hy4-preview-INT8-zhenwu-FlagOS.md | 149 ++++++++++ ...Qwen2.5-Coder-7B-Instruct-ascend-FlagOS.md | 113 ++++++++ ...e_Qwen3.8-Flash-Next-BF16-ascend-FlagOS.md | 226 +++++++++++++++ ...se_Qwen3.8-Flash-Next-BF16-hygon-FlagOS.md | 180 ++++++++++++ ...Qwen3.8-Flash-Next-BF16-iluvatar-FlagOS.md | 157 +++++++++++ ...wen3.8-Flash-Next-BF16-kunlunxin-FlagOS.md | 166 +++++++++++ ...se_Qwen3.8-Flash-Next-BF16-metax-FlagOS.md | 146 ++++++++++ ...Qwen3.8-Flash-Next-BF16-mthreads-FlagOS.md | 142 ++++++++++ ...e_Qwen3.8-Flash-Next-BF16-nvidia-FlagOS.md | 134 +++++++++ ..._Qwen3.8-Flash-Next-BF16-sunrise-FlagOS.md | 125 +++++++++ ...en3.8-Flash-Next-BF16-tsingmicro-FlagOS.md | 136 +++++++++ ...e_Qwen3.8-Flash-Next-BF16-zhenwu-FlagOS.md | 138 +++++++++ 27 files changed, 4269 insertions(+), 2 deletions(-) create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_GLM-5.3-Flash-BF16-ascend-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_GLM-5.3-Flash-BF16-hygon-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_GLM-5.3-Flash-BF16-kunlunxin-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_GLM-5.3-Flash-BF16-metax-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_GLM-5.3-Flash-BF16-nvidia-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_GLM-5.3-Flash-BF16-tsingmicro-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_GLM-5.3-Flash-BF16-zhenwu-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_GLM-5.3-Flash-FP8-mthreads-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_Hy4-preview-FP8-mthreads-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_Hy4-preview-FP8-nvidia-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_Hy4-preview-INT8-ascend-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_Hy4-preview-INT8-hygon-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_Hy4-preview-INT8-kunlunxin-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_Hy4-preview-INT8-metax-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_Hy4-preview-INT8-zhenwu-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_Qwen2.5-Coder-7B-Instruct-ascend-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_Qwen3.8-Flash-Next-BF16-ascend-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_Qwen3.8-Flash-Next-BF16-hygon-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_Qwen3.8-Flash-Next-BF16-iluvatar-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_Qwen3.8-Flash-Next-BF16-kunlunxin-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_Qwen3.8-Flash-Next-BF16-metax-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_Qwen3.8-Flash-Next-BF16-mthreads-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_Qwen3.8-Flash-Next-BF16-nvidia-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_Qwen3.8-Flash-Next-BF16-sunrise-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_Qwen3.8-Flash-Next-BF16-tsingmicro-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_Qwen3.8-Flash-Next-BF16-zhenwu-FlagOS.md diff --git a/docs/flagrelease_en/model_list.txt b/docs/flagrelease_en/model_list.txt index 24e5a82ad..c313ce131 100644 --- a/docs/flagrelease_en/model_list.txt +++ b/docs/flagrelease_en/model_list.txt @@ -70,6 +70,14 @@ FlagRelease/GLM-5.2-hygon-FlagOS FlagRelease/GLM-5.2-metax-FlagOS FlagRelease/GLM-5.2-mthreads-FlagOS FlagRelease/GLM-5.2-zhenwu-FlagOS +FlagRelease/GLM-5.3-Flash-BF16-ascend-FlagOS +FlagRelease/GLM-5.3-Flash-BF16-hygon-FlagOS +FlagRelease/GLM-5.3-Flash-BF16-kunlunxin-FlagOS +FlagRelease/GLM-5.3-Flash-BF16-metax-FlagOS +FlagRelease/GLM-5.3-Flash-BF16-nvidia-FlagOS +FlagRelease/GLM-5.3-Flash-BF16-tsingmicro-FlagOS +FlagRelease/GLM-5.3-Flash-BF16-zhenwu-FlagOS +FlagRelease/GLM-5.3-Flash-FP8-mthreads-FlagOS FlagRelease/HY-MT2-1.8B-ascend-FlagOS FlagRelease/HY-MT2-1.8B-hygon-FlagOS FlagRelease/HY-MT2-1.8B-metax-FlagOS @@ -96,6 +104,13 @@ FlagRelease/Hy3-mthreads-FlagOS FlagRelease/Hy3-nvidia-FlagOS-Express FlagRelease/Hy3-tsingmicro-FlagOS FlagRelease/Hy3-zhenwu-FlagOS +FlagRelease/Hy4-preview-FP8-mthreads-FlagOS +FlagRelease/Hy4-preview-FP8-nvidia-FlagOS +FlagRelease/Hy4-preview-INT8-ascend-FlagOS +FlagRelease/Hy4-preview-INT8-hygon-FlagOS +FlagRelease/Hy4-preview-INT8-kunlunxin-FlagOS +FlagRelease/Hy4-preview-INT8-metax-FlagOS +FlagRelease/Hy4-preview-INT8-zhenwu-FlagOS FlagRelease/Jan-v1-4B-hygon-FlagOS FlagRelease/Jan-v1-4B-iluvatar-FlagOS FlagRelease/Jan-v1-4B-metax-FlagOS @@ -173,6 +188,7 @@ FlagRelease/QwQ-32B-FlagOS-Nvidia FlagRelease/Qwen2-7B-FlagOS-Arm FlagRelease/Qwen2-7B-Instruct-FlagOS FlagRelease/Qwen2.5-32B-Instruct-FlagOS-Nvidia +FlagRelease/Qwen2.5-Coder-7B-Instruct-ascend-FlagOS FlagRelease/Qwen2.5-VL-32B-Instruct-FlagOS-Metax-BF16 FlagRelease/Qwen2.5-VL-32B-Instruct-FlagOS-Nvidia FlagRelease/Qwen3-235B-A22B-FlagOS-nvidia @@ -235,6 +251,16 @@ FlagRelease/Qwen3.8-27B-BF16-tsingmicro-FlagOS FlagRelease/Qwen3.8-27B-BF16-zhenwu-FlagOS-Express FlagRelease/Qwen3.8-27B-FP8-mthreads-FlagOS FlagRelease/Qwen3.8-27B-W4A8-arm-FlagOS-Express +FlagRelease/Qwen3.8-Flash-Next-BF16-ascend-FlagOS +FlagRelease/Qwen3.8-Flash-Next-BF16-hygon-FlagOS +FlagRelease/Qwen3.8-Flash-Next-BF16-iluvatar-FlagOS +FlagRelease/Qwen3.8-Flash-Next-BF16-kunlunxin-FlagOS +FlagRelease/Qwen3.8-Flash-Next-BF16-metax-FlagOS +FlagRelease/Qwen3.8-Flash-Next-BF16-mthreads-FlagOS +FlagRelease/Qwen3.8-Flash-Next-BF16-nvidia-FlagOS +FlagRelease/Qwen3.8-Flash-Next-BF16-sunrise-FlagOS +FlagRelease/Qwen3.8-Flash-Next-BF16-tsingmicro-FlagOS +FlagRelease/Qwen3.8-Flash-Next-BF16-zhenwu-FlagOS FlagRelease/RoboBrain-X0 FlagRelease/RoboBrain-X0-Preview-FlagOS FlagRelease/RoboBrain-X0-Preview-ascend-FlagOS @@ -255,8 +281,6 @@ FlagRelease/Seed-OSS-36B-Instruct-metax-FlagOS FlagRelease/Seed-OSS-36B-Instruct-mthreads-FlagOS FlagRelease/Seed-OSS-36B-Instruct-nvidia-FlagOS FlagRelease/TeleChat3-36B-Thinking-mthreads-FlagOS -FlagRelease/ZCK-Qwen3-8B-metax-FlagOS -FlagRelease/ZCK-Qwen3-8B-nvidia-FlagOS FlagRelease/deepseek-r1-1.5b-nvidia-FlagOS FlagRelease/farm_molecular_representation-hygon-FlagOS FlagRelease/farm_molecular_representation-nvidia-FlagOS diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_GLM-5.3-Flash-BF16-ascend-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_GLM-5.3-Flash-BF16-ascend-FlagOS.md new file mode 100644 index 000000000..0e02ed4c7 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_GLM-5.3-Flash-BF16-ascend-FlagOS.md @@ -0,0 +1,142 @@ +--- +license: apache-2.0 +language: +- zh +- en +--- + +# Introduction +GLM‑5.3‑Flash is the first natively multimodal model in the GLM‑5 series. It has 320B total parameters with 18B activated per token. It outperforms GLM‑5.2 and scores 57 points on the globally‑recognized Artificial Analysis Composite Intelligence Index (AA Composite Intelligence Index), ranking among the world’s frontier models and matching the score of Claude Opus 4.8. In Z.ai’s internal Code‑Bench experiential evaluation, its coding performance is on par with Claude Opus 4.8. + +The FlagOS community has completed day‑0 adaptation, precision alignment and deployment validation for AI chips from nine vendors: T‑Head (平头哥), NVIDIA(英伟达), Moore Threads(摩尔), Ascend(华为), Hygon(海光), MetaX(沐曦), Tsingmicro (清微智能), Kunlunxin(昆仑芯) and Sunrise(曦望). Model images have been published to ModelScope and HuggingFace, enabling developers to get out‑of‑the‑box solutions for respective chips directly. + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Ascend** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + +# Evaluation Results +## Benchmark Result +| Metrics | GLM-5.3-Flash-Nvidia-Origin | GLM-5.3-Flash-Ascend-FlagOS | +|--------------|--------------------------------|--------------------------------------| +| GPQA_Diamond | 89.29 | 88.38 | +| Musr | 74.6 | Evaluating | + + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker version 20.10.8 | +| Operating System | openEuler 22.03 (LTS-SP4) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/glm-5.3-flash-bf16-ascend001-gems5.3.4-tree0.6.1-cxnone-pluginnone-vllmnone-sglang0.5.17-sglangfl0.1.0-cp311-ptnpu210-cann90-a64-25.5.0:202608281251 + +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/GLM-5.3-Flash-BF16-ascend-FlagOS --local_dir /data/GLM-5.3-Flash +``` + +### Start the Container +```bash +docker run -dit \ + --name glm-5.3-sglang-plugin \ + --privileged \ + --init \ + --network=host --ipc=host --shm-size=512g \ + --device=/dev/davinci0 --device=/dev/davinci1 --device=/dev/davinci2 --device=/dev/davinci3 \ + --device=/dev/davinci4 --device=/dev/davinci5 --device=/dev/davinci6 --device=/dev/davinci7 \ + --device=/dev/davinci8 --device=/dev/davinci9 --device=/dev/davinci10 --device=/dev/davinci11 \ + --device=/dev/davinci12 --device=/dev/davinci13 --device=/dev/davinci14 --device=/dev/davinci15 \ + --device=/dev/davinci_manager \ + --device=/dev/hisi_hdc \ + --volume /usr/local/sbin:/usr/local/sbin \ + --volume /usr/local/Ascend/driver:/usr/local/Ascend/driver \ + --volume /usr/local/Ascend/firmware:/usr/local/Ascend/firmware \ + --volume /etc/ascend_install.info:/etc/ascend_install.info \ + --volume /var/queue_schedule:/var/queue_schedule \ + --volume /data:/model \ + -v /etc/localtime:/etc/localtime:ro \ + -v /etc/timezone:/etc/timezone:ro \ + -e TZ=$(cat /etc/timezone 2>/dev/null || echo Asia/Shanghai) \ + --entrypoint=bash \ + harbor.baai.ac.cn/flagrelease-public/glm-5.3-flash-bf16-ascend001-gems5.3.4-tree0.6.1-cxnone-pluginnone-vllmnone-sglang0.5.17-sglangfl0.1.0-cp311-ptnpu210-cann90-a64-25.5.0:202608281251 + + +``` +### Start the Server +```bash +docker exec -it glm-5.3-sglang-plugin /bin/bash +cd /sgl-workspace +bash launch.sh +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "glm-5.3-flash", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of . This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from ZhipuAI/GLM-5.3-Flash-BF16 and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_GLM-5.3-Flash-BF16-hygon-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_GLM-5.3-Flash-BF16-hygon-FlagOS.md new file mode 100644 index 000000000..3279244fc --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_GLM-5.3-Flash-BF16-hygon-FlagOS.md @@ -0,0 +1,159 @@ +--- +license: apache-2.0 +language: +- zh +- en +--- + +# Introduction +GLM‑5.3‑Flash is the first natively multimodal model in the GLM‑5 series. It has 320B total parameters with 18B activated per token. It outperforms GLM‑5.2 and scores 57 points on the globally‑recognized Artificial Analysis Composite Intelligence Index (AA Composite Intelligence Index), ranking among the world’s frontier models and matching the score of Claude Opus 4.8. In Z.ai’s internal Code‑Bench experiential evaluation, its coding performance is on par with Claude Opus 4.8. + +The FlagOS community has completed day‑0 adaptation, precision alignment and deployment validation for AI chips from nine vendors: T‑Head (平头哥), NVIDIA(英伟达), Moore Threads(摩尔), Ascend(华为), Hygon(海光), MetaX(沐曦), Tsingmicro (清微智能), Kunlunxin(昆仑芯) and Sunrise(曦望). Model images have been published to ModelScope and HuggingFace, enabling developers to get out‑of‑the‑box solutions for respective chips directly. + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Hygon** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + +# Evaluation Results +## Benchmark Result +| Metrics | GLM-5.3-Flash-Nvidia-Origin | GLM-5.3-Flash-Hygon-FlagOS | +|--------------|--------------------------------|--------------------------------------| +| GPQA_Diamond | 89.29 | 87.46 | +| Musr | 74.6 | Evaluating | + + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker version 27.3.1 | +| Operating System | 22.04 LTS | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/glm5.3-flash-hygon001-gems5.4.0-tree0.6.1-cxnone-plugin0.2.0-vllm0.26.1-cp310-pt210-dtk2604-x64-6.3.30-v1.4.1a:202608271850 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/GLM-5.3-Flash-BF16-hygon-FlagOS --local_dir /data/GLM-5.3-Flash +``` + +### Start the Container +```bash +docker run -d --name flagos \ + --network host --ipc host \ + --device /dev/kfd --device /dev/mkfd --device /dev/dri \ + --mount type=bind,src=/dev/infiniband,dst=/dev/infiniband \ + --mount type=bind,src=/sys/class/infiniband,dst=/sys/class/infiniband,readonly \ + --mount type=bind,src=/sys/class/infiniband_verbs,dst=/sys/class/infiniband_verbs,readonly \ + --mount type=bind,src=/sys/class/net,dst=/sys/class/net,readonly \ + --mount type=bind,src=/usr/etc/libibverbs.d,dst=/usr/etc/libibverbs.d,readonly \ + --mount type=bind,src=/lib/x86_64-linux-gnu/libibverbs.so.1.14.44.0,dst=/lib/x86_64-linux-gnu/libibverbs.so.1.14.47.0,readonly \ + --mount type=bind,src=/lib/x86_64-linux-gnu/libshca-rdmav34.so,dst=/lib/x86_64-linux-gnu/libshca-rdmav34.so,readonly \ + --mount type=bind,src=/lib/x86_64-linux-gnu/libnl-3.so.200,dst=/lib/x86_64-linux-gnu/libnl-3.so.200,readonly \ + --mount type=bind,src=/lib/x86_64-linux-gnu/libnl-route-3.so.200,dst=/lib/x86_64-linux-gnu/libnl-route-3.so.200,readonly \ + -v /opt/hyhal:/opt/hyhal \ + -v /data:/data:rslave \ + -v /public-flash:/public-flash \ + --group-add video \ + --cap-add SYS_PTRACE \ + --security-opt seccomp=unconfined \ + --security-opt label=disable \ + -e HSA_FORCE_FINE_GRAIN_PCIE=1 \ + -e NCCL_IB_DISABLE=0 \ + harbor.baai.ac.cn/flagrelease-public/glm5.3-flash-hygon001-gems5.4.0-tree0.6.1-cxnone-plugin0.2.0-vllm0.26.1-cp310-pt210-dtk2604-x64-6.3.30-v1.4.1a:202608271850 \ + bash -lc 'sleep infinity' +docker exec -it flagos /bin/bash +``` +### Start the Server +```bash +set +o nounset +source /opt/dtk-26.04-DCC2602-0317/env.sh +set -o nounset +export HIP_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 +export HSA_FORCE_FINE_GRAIN_PCIE=1 +export TRITON_HIP_CLANG_PATH=/opt/dtk-26.04-DCC2602-0317/aillvm/bin/clang-18 +export LD_LIBRARY_PATH="/opt/ucx-shca/lib:/opt/rccl-shca-net/lib:${LD_LIBRARY_PATH:-}" +export UCX_MODULE_DIR=/opt/ucx-shca/lib/ucx +export MASTER_ADDR=10.232.2.19 +export GLOO_SOCKET_IFNAME=ib0 +export NCCL_SOCKET_IFNAME=ib0 +export NCCL_NET_PLUGIN=shca +export NCCL_IB_DISABLE=0 +export NCCL_IB_HCA=shca_0,shca_1,shca_2,shca_3 +export HF_HUB_OFFLINE=1 +export TRANSFORMERS_OFFLINE=1 +export HF_DATASETS_OFFLINE=1 +# vLLM 0.24+ RPC timeout for execute_model calls (seconds). +# W8A8 quantized models need extra time on first inference for kernel compilation. +export VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=1800 +vllm serve /data/GLM-5.3-Flash --served-model-name GLM-5.3-Flash-BF16 --trust-remote-code --distributed-executor-backend mp --nnodes 2 --node-rank 0 --master-addr 10.232.2.30 --master-port 29850 --distributed-timeout-seconds 1800 --tensor-parallel-size 16 --pipeline-parallel-size 1 --data-parallel-size 1 --load-format safetensors --disable-custom-all-reduce --moe-backend triton --max-model-len 131072 --max-num-seqs 64 --gpu-memory-utilization 0.95 --no-enable-log-requests --reasoning-parser glm45 --block-size 128 --host 0.0.0.0 --port 8000 +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "GLM-5.3-Flash-BF16", + "messages": [{"role": "user", "content": "Hi"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of . This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from ZhipuAI/GLM-5.3-Flash-BF16 and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_GLM-5.3-Flash-BF16-kunlunxin-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_GLM-5.3-Flash-BF16-kunlunxin-FlagOS.md new file mode 100644 index 000000000..29d57b187 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_GLM-5.3-Flash-BF16-kunlunxin-FlagOS.md @@ -0,0 +1,167 @@ +--- +license: apache-2.0 +language: +- zh +- en +--- + +# Introduction +GLM‑5.3‑Flash is the first natively multimodal model in the GLM‑5 series. It has 320B total parameters with 18B activated per token. It outperforms GLM‑5.2 and scores 57 points on the globally‑recognized Artificial Analysis Composite Intelligence Index (AA Composite Intelligence Index), ranking among the world’s frontier models and matching the score of Claude Opus 4.8. In Z.ai’s internal Code‑Bench experiential evaluation, its coding performance is on par with Claude Opus 4.8. + +The FlagOS community has completed day‑0 adaptation, precision alignment and deployment validation for AI chips from nine vendors: T‑Head (平头哥), NVIDIA(英伟达), Moore Threads(摩尔), Ascend(华为), Hygon(海光), MetaX(沐曦), Tsingmicro (清微智能), Kunlunxin(昆仑芯) and Sunrise(曦望). Model images have been published to ModelScope and HuggingFace, enabling developers to get out‑of‑the‑box solutions for respective chips directly. + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Kunlunxin** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + +# Evaluation Results +## Benchmark Result +| Metrics | GLM-5.3-Flash-Nvidia-Origin | GLM-5.3-Flash-Kunlunxin-FlagOS | +|--------------|--------------------------------|--------------------------------------| +| GPQA_Diamond | 89.29 | Evaluating | +| Musr | Evaluating | Evaluating | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker version 20.10.12 | +| Operating System | 22.04 LTS | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/glm-5.3-flash-bf16-kunlunxin001-gems5.0.0-treenone-cx0.13.0-plugin0.2.0-vllm0.20.2-cp310-pt29-xrt513-x64-5.0.21.47:202608281449 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/GLM-5.3-Flash-BF16-kunlunxin-FlagOS --local_dir /data/GLM-5.3-Flash +``` + +### Start the Container +```bash +set -euo pipefail + +docker run -d \ + --name glm53-flash-flagos-p800 \ + --privileged \ + --network host \ + --shm-size 128g \ + --cap-add SYS_PTRACE \ + --ulimit memlock=-1:-1 \ + --ulimit nofile=120000:120000 \ + --ulimit stack=67108864:67108864 \ + -v /data:/data:ro \ + --entrypoint /bin/bash \ + harbor.baai.ac.cn/flagrelease-public/glm-5.3-flash-bf16-kunlunxin001-gems5.0.0-treenone-cx0.13.0-plugin0.2.0-vllm0.20.2-cp310-pt29-xrt513-x64-5.0.21.47:202608281449 \ + -lc 'sleep infinity' +``` +### Start the Server +```bash +set -euo pipefail + +docker exec -d glm53-flash-flagos-p800 /bin/bash -lc ' + export USE_FLAGGEMS=1 + export VLLM_DISABLE_PYNCCL=1 + export VLLM_FL_FLAGOS_WHITELIST=rms_norm + export VLLM_FL_FLAGGEMS_ATEN_PLAN_CACHE=1 + export VLLM_FL_FLAGGEMS_ATEN_PLAN_CACHE_REPORT=1 + export FLAGGEMS_ATEN_PLAN_CACHE=1 + export FLAGGEMS_ATEN_PLAN_CACHE_SIZE=128 + export VLLM_FL_COMMON_SLOT_MAPPING_GRAPH=0 + export VLLM_FL_GLM5_PROVIDER=auto + export VLLM_PLUGINS=fl + export VLLM_WORKER_MULTIPROC_METHOD=spawn + export VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=1800 + export USE_RESHAPE_AND_CACHE_FLASH=1 + export PYTHONUNBUFFERED=1 + export CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 + exec /root/miniconda/envs/python310_torch29_cuda/bin/vllm serve /data/GLM-5.3-Flash \ + --served-model-name glm5.3-flash-bf16 \ + --host 0.0.0.0 \ + --port 8000 \ + --tensor-parallel-size 8 \ + --enable-expert-parallel \ + --all2all-backend allgather_reducescatter \ + --distributed-executor-backend mp \ + --load-format safetensors \ + --dtype bfloat16 \ + --disable-custom-all-reduce \ + --max-model-len 4096 \ + --max-num-seqs 1 \ + --max-num-batched-tokens 2048 \ + --block-size 128 \ + --gpu-memory-utilization 0.90 \ + --enable-chunked-prefill \ + --no-enable-prefix-caching \ + --language-model-only \ + --enforce-eager >>/var/log/glm53-flash-vllm.log 2>&1 +' +``` + +## Service Invocation +### Invocation Script +```bash +set -euo pipefail + +curl --noproxy '*' -fsS http://127.0.0.1:8000/v1/chat/completions \ + -H 'Content-Type: application/json' \ + -d '{"model":"glm5.3-flash-bf16","messages":[{"role":"user","content":"Compute 17 * 19. Return only the integer."}],"temperature":0,"max_tokens":32}' + +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of . This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from ZhipuAI/GLM-5.3-Flash-BF16 and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_GLM-5.3-Flash-BF16-metax-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_GLM-5.3-Flash-BF16-metax-FlagOS.md new file mode 100644 index 000000000..412695fd0 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_GLM-5.3-Flash-BF16-metax-FlagOS.md @@ -0,0 +1,232 @@ +--- +license: apache-2.0 +language: +- zh +- en +--- + +# Introduction +GLM‑5.3‑Flash is the first natively multimodal model in the GLM‑5 series. It has 320B total parameters with 18B activated per token. It outperforms GLM‑5.2 and scores 57 points on the globally‑recognized Artificial Analysis Composite Intelligence Index (AA Composite Intelligence Index), ranking among the world’s frontier models and matching the score of Claude Opus 4.8. In Z.ai’s internal Code‑Bench experiential evaluation, its coding performance is on par with Claude Opus 4.8. + +The FlagOS community has completed day‑0 adaptation, precision alignment and deployment validation for AI chips from nine vendors: T‑Head (平头哥), NVIDIA(英伟达), Moore Threads(摩尔), Ascend(华为), Hygon(海光), MetaX(沐曦), Tsingmicro (清微智能), Kunlunxin(昆仑芯) and Sunrise(曦望). Model images have been published to ModelScope and HuggingFace, enabling developers to get out‑of‑the‑box solutions for respective chips directly. + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Metax** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + +# Evaluation Results +## Benchmark Result +| Metrics | GLM-5.3-Flash-Nvidia-Origin | GLM-5.3-Flash-Metax-FlagOS | +|--------------|--------------------------------|--------------------------------------| +| GPQA_Diamond | 89.29 | Evaluating | +|Musr| 74.6 | Evaluating | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker version 29.3.1 | +| Operating System | 22.04 LTS | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/glm5.3-flash-metax001-gems5.4.0-treenone-cxnone-plugin3.0.0-vllm0.24.0-cp312-pt28-maca37-x64-3.8.1:202608301009 + +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/GLM-5.3-Flash-BF16-metax-FlagOS --local_dir /data/GLM-5.3-Flash +``` + +### Start the Container +```bash +IMAGE="harbor.baai.ac.cn/flagrelease-public/glm5.3-flash-metax001-gems5.4.0-treenone-cxnone-plugin3.0.0-vllm0.24.0-cp312-pt28-maca37-x64-3.8.1:202608301009" +docker run -d \ + --name glm5.3-flash \ + --network host \ + --shm-size 64g \ + --device /dev/dri:/dev/dri:rwm \ + --device /dev/mxcd:/dev/mxcd:rwm \ + -v /public-flash/models:/models \ + ${IMAGE} \ + sleep infinity +docker exec -it glm5.3-flash /bin/bash +``` +### Start the Server +```bash +# NOTICE! Environment variables must be set first, then start Ray on two Metax 8*64G machines: one as head node and the other as worker node +# GLM‑5.3‑Flash Multi‑node Ray Common Environment Variables +# Execute `source ray_env.sh` on all nodes (head + worker) before starting Ray +# ============ Network Configuration ============ +# View network interface: ip addr / ifconfig; --network=host is used, container network interface aligns with host +# MetaX implements MCCL as its NCCL counterpart. Set NCCL_/MCCL_/GLOO_ variables all together +export NIC="${NIC:-inbond1}" +export GLOO_SOCKET_IFNAME=${NIC} +export MCCL_SOCKET_IFNAME=${NIC} +export NCCL_SOCKET_IFNAME=${NIC} + +# ============ Ray Cluster Configuration ============ +# Fill in HEAD_IP with actual value (inbond1 IP of head host, query via `ip -br addr show inbond1`) +export HEAD_IP="${HEAD_IP:-192.168.2.109}" +export HEAD_PORT="${HEAD_PORT:-6379}" +export RAY_ADDRESS="${HEAD_IP}:${HEAD_PORT}" + +# Number of GPUs per machine +export NUM_GPUS="${NUM_GPUS:-8}" + +# ============ vLLM / MetaX Runtime ============ +export GEMS_VENDOR=metax +export VLLM_PLUGINS=fl +export VLLM_WORKER_MULTIPROC_METHOD=spawn + +# Timeout settings (for multi‑node communication, initial Triton compilation and loading of large 288‑expert model) +export VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=7200 +export VLLM_ENGINE_ITERATION_TIMEOUT_S=7200 +export VLLM_ENGINE_READY_TIMEOUT_S=3600 +export VLLM_RPC_TIMEOUT=36000000 + +export VLLM_FL_FLAGOS_BLACKLIST="mm,mm_out,bmm,bmm_out,sort,stable_sort,masked_fill,masked_fill_,log_softmax,log_softmax_out,log_softmax_backward,log_softmax_backward_out,pad,constant_pad_nd,copy_,topk,layer_norm" + +# ============ Ray Runtime Env ============ +# Instruct Ray to inject environment variables automatically across all worker processes +# (export only required before vllm serve on head node) +export RAY_RUNTIME_ENV='{ + "env_vars": { + "VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS": "7200", + "VLLM_ENGINE_ITERATION_TIMEOUT_S": "7200", + "VLLM_RPC_TIMEOUT": "36000000", + "GLOO_SOCKET_IFNAME": "'"${NIC}"'", + "MCCL_SOCKET_IFNAME": "'"${NIC}"'", + "NCCL_SOCKET_IFNAME": "'"${NIC}"'", + "GEMS_VENDOR": "metax", + "VLLM_PLUGINS": "fl", + "VLLM_FL_FLAGOS_BLACKLIST": "'"${VLLM_FL_FLAGOS_BLACKLIST}"'" + } +}' + + +# Commands below are executed ONLY on the head node +# ============ Model Configuration ============ +export MODEL_PATH="${MODEL_PATH:-/data/GLM-5.3-Flash}" +export MODEL_NAME="${MODEL_NAME:-glm53-flash}" +export PORT="${PORT:-8000}" +export TP_SIZE="${TP_SIZE:-16}" # 2 nodes × 8 GPUs +export PP_SIZE="${PP_SIZE:-1}" # Pipeline parallel disabled for now; adjust when out‑of‑memory occurs +export MAX_MODEL_LEN="${MAX_MODEL_LEN:-8192}" + +# GLM‑5.3‑Flash Multi‑node — launch vllm serve only (do NOT rebuild Ray cluster) +# Used for restarting vllm process after failure while Ray cluster remains alive +# +# Usage (run on HEAD node): + +set -e + +SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)" +source "${SCRIPT_DIR}/ray_env.sh" + +# Verify Ray cluster health +echo "[SERVE] Checking Ray cluster status..." +ray status || { echo "[SERVE] ERROR: Ray cluster unavailable. Please run ray_start_head.sh + ray_start_worker.sh first"; exit 1; } +echo "" + +# Launch vllm serve +LOG_DIR="/models/glm53_flash_running/logs" +LOG_FILE="${LOG_DIR}/glm53-ray-serve-$(date +%Y%m%d-%H%M%S).log" +mkdir -p ${LOG_DIR} + +echo "[SERVE] Starting vllm serve..." | tee -a ${LOG_FILE} +echo "[SERVE] Log file: ${LOG_FILE}" | tee -a ${LOG_FILE} +echo "[SERVE] Start time: $(date)" | tee -a ${LOG_FILE} +echo "[SERVE] TP=${TP_SIZE} PP=${PP_SIZE} max_model_len=${MAX_MODEL_LEN}" | tee -a ${LOG_FILE} +echo "[SERVE] Model path: ${MODEL_PATH}" | tee -a ${LOG_FILE} +echo "" + +/opt/conda/bin/vllm serve ${MODEL_PATH} \ + --port ${PORT} \ + --served-model-name ${MODEL_NAME} \ + --tensor-parallel-size ${TP_SIZE} \ + --pipeline-parallel-size ${PP_SIZE} \ + --distributed-executor-backend ray \ + --dtype bfloat16 \ + --gpu-memory-utilization 0.9 \ + --no-async-scheduling \ + --enforce-eager \ + --trust-remote-code \ + --compilation-config '{"pass_config":{"fuse_allreduce_rms":false}}' \ + --max-model-len 32768 \ + 2>&1 | tee -a ${LOG_FILE} + +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ +-H "Content-Type: application/json" \ +-d '{ + "model": "glm53-flash", + "messages": [{"role": "user", "content": "中国的首都是哪里?"}], + "temperature": 0.7, + "max_tokens": 128, + "chat_template_kwargs": { + "enable_thinking": false + } +}' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of . This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from ZhipuAI/GLM-5.3-Flash-BF16 and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_GLM-5.3-Flash-BF16-nvidia-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_GLM-5.3-Flash-BF16-nvidia-FlagOS.md new file mode 100644 index 000000000..6866b5e4c --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_GLM-5.3-Flash-BF16-nvidia-FlagOS.md @@ -0,0 +1,263 @@ +--- +license: apache-2.0 +language: +- zh +- en +--- + +# Introduction +GLM‑5.3‑Flash is the first natively multimodal model in the GLM‑5 series. It has 320B total parameters with 18B activated per token. It outperforms GLM‑5.2 and scores 57 points on the globally‑recognized Artificial Analysis Composite Intelligence Index (AA Composite Intelligence Index), ranking among the world’s frontier models and matching the score of Claude Opus 4.8. In Z.ai’s internal Code‑Bench experiential evaluation, its coding performance is on par with Claude Opus 4.8. + +The FlagOS community has completed day‑0 adaptation, precision alignment and deployment validation for AI chips from nine vendors: T‑Head (平头哥), NVIDIA(英伟达), Moore Threads(摩尔), Ascend(华为), Hygon(海光), MetaX(沐曦), Tsingmicro (清微智能), Kunlunxin(昆仑芯) and Sunrise(曦望). Model images have been published to ModelScope and HuggingFace, enabling developers to get out‑of‑the‑box solutions for respective chips directly. + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Nvidia** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + +# Evaluation Results +## Benchmark Result +| Metrics | GLM-5.3-Flash-Nvidia-Origin | GLM-5.3-Flash-Nvidia-FlagOS | +|--------------|--------------------------------|--------------------------------------| +| GPQA_Diamond | 89.29 | 90.36 | +| Musr | 74.6 | 73.5 | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker version 24.0.0 | +| Operating System | 22.04.4 LTS | + +## Operation Steps + +### Download FlagOS Image +```bash +set -euo pipefail + +IMAGE=harbor.baai.ac.cn/flagrelease-public/glm5.3-flash-bf16-nvidia003-gems5.3.3-tree0.5.0-cxnone-plugin0.3.0-vllm0.24.0-cp312-pt211-cu129-x64-580.126.20:202608271300 +docker pull "${IMAGE}" +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/GLM-5.3-Flash-BF16-nvidia-FlagOS --local_dir /data/GLM-5.3-Flash +``` + +### Start the Container +```bash +set -euo pipefail + +: "${NODE_RANK:?set NODE_RANK to 0 on the API node or 1 on the headless node}" +[[ "${NODE_RANK}" =~ ^[01]$ ]] || { echo "NODE_RANK must be 0 or 1" >&2; exit 2; } + +IMAGE=${IMAGE:-harbor.baai.ac.cn/flagos-inner-models-release/glm5.3-flash-bf16-nvidia003-gems5.3.3-tree0.5.0-cxnone-plugin0.3.0-vllm0.24.0-cp312-pt211-cu129-x64-580.126.20@sha256:aa04dfbcf07793d441a1f55713c83f676cdf5e3d77ca39d6ce9bdbf1c9ee8084} +CONTAINER=${CONTAINER:-glm53-flash-flagos-n${NODE_RANK}} +MODEL_ROOT=${MODEL_ROOT:-/data} + +driver_libcuda=$(readlink -f "$(ldconfig -p | awk "/libcuda\\.so\\.1 \\(libc6,x86-64\\)/ {print \\$NF; exit}")") +driver_libnvml=$(readlink -f "$(ldconfig -p | awk "/libnvidia-ml\\.so\\.1 \\(libc6,x86-64\\)/ {print \\$NF; exit}")") +driver_libptxjit=$(readlink -f "$(ldconfig -p | awk "/libnvidia-ptxjitcompiler\\.so\\.1 \\(libc6,x86-64\\)/ {print \\$NF; exit}")") +test -f "${driver_libcuda}" +test -f "${driver_libnvml}" +test -f "${driver_libptxjit}" + +device_args=() +for dev in /dev/infiniband/* /dev/nvidia[0-9]* /dev/nvidiactl \ + /dev/nvidia-uvm /dev/nvidia-uvm-tools /dev/nvidia-nvlink \ + /dev/nvidia-nvswitch* /dev/nvidia-caps/*; do + [[ -e "${dev}" ]] && device_args+=(--device "${dev}:${dev}") +done + +docker run -d \ + --name "${CONTAINER}" \ + --network host \ + --ipc host \ + --pid host \ + --pids-limit=-1 \ + --shm-size=128g \ + --ulimit memlock=-1:-1 \ + --ulimit stack=67108864:67108864 \ + --cap-add SYS_PTRACE \ + --security-opt seccomp=unconfined \ + --runtime runc \ + "${device_args[@]}" \ + -v "${driver_libcuda}:/driver/libcuda.so.1:ro" \ + -v "${driver_libcuda}:/driver/libcuda.so:ro" \ + -v "${driver_libnvml}:/driver/libnvidia-ml.so.1:ro" \ + -v "${driver_libnvml}:/driver/libnvidia-ml.so:ro" \ + -v "${driver_libptxjit}:/driver/libnvidia-ptxjitcompiler.so.1:ro" \ + -v "${driver_libptxjit}:/driver/libnvidia-ptxjitcompiler.so:ro" \ + -v /data:/models:ro" \ + -e NODE_RANK="${NODE_RANK}" \ + -e CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 \ + "${IMAGE}" -lc "exec sleep infinity" +``` +### Start the Server +```bash +set -euo pipefail + +: "${NODE_RANK:?set NODE_RANK to 0 on the API node or 1 on the headless node}" +: "${MASTER_ADDR:?set MASTER_ADDR to a resolvable address of NODE_RANK=0}" +[[ "${NODE_RANK}" =~ ^[01]$ ]] || { echo "NODE_RANK must be 0 or 1" >&2; exit 2; } + +CONTAINER=${CONTAINER:-glm53-flash-flagos-n${NODE_RANK}} +MASTER_PORT=${MASTER_PORT:-29873} +PORT=${PORT:-8000} +GLOO_SOCKET_IFNAME=${GLOO_SOCKET_IFNAME:-bond0} +NCCL_SOCKET_IFNAME=${NCCL_SOCKET_IFNAME:-bond0} +NCCL_IB_HCA=${NCCL_IB_HCA:-=mlx5_100:1,mlx5_101:1,mlx5_102:1,mlx5_103:1,mlx5_104:1,mlx5_105:1,mlx5_106:1,mlx5_107:1} + +role_args=(--headless) +if [[ "${NODE_RANK}" == 0 ]]; then + role_args=(--host 0.0.0.0 --port "${PORT}") +fi + +docker exec -d \ + -e NODE_RANK="${NODE_RANK}" \ + -e MASTER_ADDR="${MASTER_ADDR}" \ + -e MASTER_PORT="${MASTER_PORT}" \ + -e PORT="${PORT}" \ + -e GLOO_SOCKET_IFNAME="${GLOO_SOCKET_IFNAME}" \ + -e NCCL_SOCKET_IFNAME="${NCCL_SOCKET_IFNAME}" \ + -e NCCL_IB_HCA="${NCCL_IB_HCA}" \ + -e NCCL_IB_DISABLE=0 \ + -e NCCL_CROSS_NIC=0 \ + -e NCCL_NVLS_ENABLE=0 \ + -e LD_LIBRARY_PATH=/driver:/usr/local/cuda/lib64 \ + -e CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 \ + -e HF_HUB_OFFLINE=1 \ + -e TRANSFORMERS_OFFLINE=1 \ + -e TOKENIZERS_PARALLELISM=false \ + -e VLLM_PLUGINS=fl \ + -e USE_FLAGGEMS=1 \ + -e VLLM_FL_PREFER=flagos \ + -e VLLM_FL_PREFER_ENABLED=true \ + -e VLLM_FL_OOT_ENABLED=1 \ + -e VLLM_FL_STRICT=0 \ + -e VLLM_FL_GLM5_PROVIDER=auto \ + -e VLLM_FL_FLAGOS_WHITELIST=grouped_topk,moe_sum \ + -e VLLM_USE_BREAKABLE_CUDAGRAPH=1 \ + "${CONTAINER}" /usr/local/bin/vllm serve /models/GLM-5.3-Flash \ + --tokenizer /models/GLM-5.3-Flash \ + --served-model-name GLM-5.3-Flash \ + --tensor-parallel-size 16 \ + --pipeline-parallel-size 1 \ + --distributed-executor-backend mp \ + --nnodes 2 \ + --node-rank "${NODE_RANK}" \ + --master-addr "${MASTER_ADDR}" \ + --master-port "${MASTER_PORT}" \ + --distributed-timeout-seconds 1800 \ + --cpu-distributed-timeout-seconds 1800 \ + --load-format auto \ + --dtype bfloat16 \ + --disable-custom-all-reduce \ + --max-model-len 102400 \ + --max-num-batched-tokens 32768 \ + --max-num-seqs 128 \ + --gpu-memory-utilization 0.80 \ + --seed 1234 \ + --compilation-config "{\"cudagraph_mode\":\"PIECEWISE\",\"pass_config\":{\"fuse_allreduce_rms\":false}}" \ + --enable-expert-parallel \ + --mm-encoder-tp-mode data \ + --limit-mm-per-prompt "{\"image\":16,\"video\":0}" \ + --reasoning-parser glm47 \ + --enable-auto-tool-choice \ + --tool-call-parser glm47 \ + "${role_args[@]}" + +if [[ "${NODE_RANK}" == 0 ]]; then + for _ in $(seq 1 360); do + if curl --noproxy "*" -fsS "http://127.0.0.1:${PORT}/health" >/dev/null; then + echo "health=200 port=${PORT}" + exit 0 + fi + sleep 10 + done + docker logs --tail 200 "${CONTAINER}" >&2 || true + exit 1 +fi + +echo "headless node started; start NODE_RANK=0 and wait for its health check" + +``` + +## Service Invocation +### Invocation Script +```bash +set -euo pipefail + +PORT=${PORT:-8000} +curl --noproxy "*" -fsS "http://127.0.0.1:${PORT}/health" +printf "\n" +curl --noproxy "*" -fsS "http://127.0.0.1:${PORT}/v1/models" +printf "\n" +curl --noproxy "*" -fsS "http://127.0.0.1:${PORT}/v1/chat/completions" \ + -H "Content-Type: application/json" \ + --data-binary @- <. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from ZhipuAI/GLM-5.3-Flash-BF16 and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_GLM-5.3-Flash-BF16-tsingmicro-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_GLM-5.3-Flash-BF16-tsingmicro-FlagOS.md new file mode 100644 index 000000000..41915f95f --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_GLM-5.3-Flash-BF16-tsingmicro-FlagOS.md @@ -0,0 +1,125 @@ +--- +license: apache-2.0 +language: +- zh +- en +--- + +# Introduction +GLM‑5.3‑Flash is the first natively multimodal model in the GLM‑5 series. It has 320B total parameters with 18B activated per token. It outperforms GLM‑5.2 and scores 57 points on the globally‑recognized Artificial Analysis Composite Intelligence Index (AA Composite Intelligence Index), ranking among the world’s frontier models and matching the score of Claude Opus 4.8. In Z.ai’s internal Code‑Bench experiential evaluation, its coding performance is on par with Claude Opus 4.8. + +The FlagOS community has completed day‑0 adaptation, precision alignment and deployment validation for AI chips from nine vendors: T‑Head (平头哥), NVIDIA(英伟达), Moore Threads(摩尔), Ascend(华为), Hygon(海光), MetaX(沐曦), Tsingmicro (清微智能), Kunlunxin(昆仑芯) and Sunrise(曦望). Model images have been published to ModelScope and HuggingFace, enabling developers to get out‑of‑the‑box solutions for respective chips directly. + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Tsingmicro** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + +# Evaluation Results +## Benchmark Result +| Metrics | GLM-5.3-Flash-Nvidia-Origin | GLM-5.3-Flash-Tsingmicro-FlagOS | +|--------------|--------------------------------|--------------------------------------| +| GPQA_Diamond | 89.29 | 90.36 | +| Musr | Evaluating | Evaluating | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker version 20.10.12 | +| Operating System | 22.04 LTS | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/glm53-flash-tsingmicro001-gems4.2.1-treenone-cx0.1.0-plugin0.0.0-vllm0.20.2-cp310-pt211-raisa0.2927-x64-v0.29277.8:202608280932 + +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/GLM-5.3-Flash-BF16-tsingmicro-FlagOS --local_dir /data/GLM-5.3-Flash +``` + +### Start the Container +```bash +docker run --privileged -dit \ + --name glm53-flash \ + --shm-size=128g \ + --network host \ + --ipc=host \ + -v /sys:/sys \ + -v /dev:/dev \ + -v /lib/modules:/lib/modules \ + -v /mnt/nvme_data:/mnt/nvme_data \ + -v /data:/data \ + harbor.baai.ac.cn/flagrelease-public/glm53-flash-tsingmicro001-gems4.2.1-treenone-cx0.1.0-plugin0.0.0-vllm0.20.2-cp310-pt211-raisa0.2927-x64-v0.29277.8:202608280932 \ + /bin/bash +docker exec -it glm53-flash /bin/bash +``` +### Start the Server +```bash +cd /home/secure/260629145101/glm-5.3-flash +source env.sh +bash run_offline.sh +``` + +## Service Invocation +### Invocation Script +```bash +bash run_offline.sh +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of . This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from ZhipuAI/GLM-5.3-Flash-BF16 and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_GLM-5.3-Flash-BF16-zhenwu-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_GLM-5.3-Flash-BF16-zhenwu-FlagOS.md new file mode 100644 index 000000000..733a6648a --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_GLM-5.3-Flash-BF16-zhenwu-FlagOS.md @@ -0,0 +1,149 @@ +--- +license: apache-2.0 +language: +- zh +- en +--- + +# Introduction +GLM‑5.3‑Flash is the first natively multimodal model in the GLM‑5 series. It has 320B total parameters with 18B activated per token. It outperforms GLM‑5.2 and scores 57 points on the globally‑recognized Artificial Analysis Composite Intelligence Index (AA Composite Intelligence Index), ranking among the world’s frontier models and matching the score of Claude Opus 4.8. In Z.ai’s internal Code‑Bench experiential evaluation, its coding performance is on par with Claude Opus 4.8. + +The FlagOS community has completed day‑0 adaptation, precision alignment and deployment validation for AI chips from nine vendors: T‑Head (平头哥), NVIDIA(英伟达), Moore Threads(摩尔), Ascend(华为), Hygon(海光), MetaX(沐曦), Tsingmicro (清微智能), Kunlunxin(昆仑芯) and Sunrise(曦望). Model images have been published to ModelScope and HuggingFace, enabling developers to get out‑of‑the‑box solutions for respective chips directly. + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Zhenwu** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + +# Evaluation Results +## Benchmark Result +| Metrics | GLM-5.3-Flash-Nvidia-Origin | GLM-5.3-Flash-Zhenwu-FlagOS | +|--------------|--------------------------------|--------------------------------------| +| GPQA_Diamond | 89.29 | 89.8 | +| Musr | 74.6 | 72.66 | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker version 29.7.2 | +| Operating System | 24.04.2 LTS | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/glm5.3-flash-pp001-gems5.4-tree0.6.1-cxnone-plugin0.3.0-vllm0.24.0-cp312-ptnone-hggcnone-x64-2.1.1-rbd225:202608271447 + +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/GLM-5.3-Flash-BF16-zhenwu-FlagOS --local_dir /data/GLM-5.3-Flash +``` + +### Start the Container +```bash +docker run --privileged -dit \ + --network=host \ + --device=/dev/infiniband \ + --ipc=host \ + --device=/dev/alixpu_ctl \ + --device=/dev/alixpu \ + --ulimit memlock=-1 \ + --ulimit stack=67108864 \ + --init \ + -v /data:/data\ + -w /data/ \ + --name flagos \ + harbor.baai.ac.cn/flagrelease-public/glm5.3-flash-pp001-gems5.4-tree0.6.1-cxnone-plugin0.3.0-vllm0.24.0-cp312-ptnone-hggcnone-x64-2.1.1-rbd225:202608271447 +docker exec -it flagos /bin/bash + +``` +### Start the Server +```bash +NCCL_IB_DISABLE=1 NCCL_SOCKET_IFNAME=eth0 NCCL_DEBUG=WARN \ + VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=3600 \ + VLLM_WORKER_MULTIPROC_METHOD=spawn \ + VLLM_USE_BREAKABLE_CUDAGRAPH=1 \ + VLLM_FL_USE_FLAGGEMS_ATTN=1 \ + VLLM_FL_GLM5_PROVIDER=flaggems \ + vllm serve /data/GLM-5.3-Flash \ + --served-model-name glm-5.3-flash \ + --trust-remote-code \ + --tensor-parallel-size 16 \ + --distributed-executor-backend mp \ + --host 0.0.0.0 \ + --port 8021 \ + --max-model-len 65536 \ + --max-num-seqs 32 \ + --gpu-memory-utilization 0.90 \ + --reasoning-parser glm45 \ + --compilation-config \ + '{"cudagraph_mode":"FULL_AND_PIECEWISE","cudagraph_capture_sizes":[1,2,4,8,16,32]}' + +``` + +## Service Invocation +### Invocation Script +```bash +curl http://127.0.0.1:8021/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "messages": [{"role": "user", "content": "Hi"}], + "max_tokens": 1024 + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of . This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from ZhipuAI/GLM-5.3-Flash-BF16 and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_GLM-5.3-Flash-FP8-mthreads-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_GLM-5.3-Flash-FP8-mthreads-FlagOS.md new file mode 100644 index 000000000..516538a6f --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_GLM-5.3-Flash-FP8-mthreads-FlagOS.md @@ -0,0 +1,134 @@ +--- +license: apache-2.0 +language: +- zh +- en +--- + +# Introduction +GLM‑5.3‑Flash is the first natively multimodal model in the GLM‑5 series. It has 320B total parameters with 18B activated per token. It outperforms GLM‑5.2 and scores 57 points on the globally‑recognized Artificial Analysis Composite Intelligence Index (AA Composite Intelligence Index), ranking among the world’s frontier models and matching the score of Claude Opus 4.8. In Z.ai’s internal Code‑Bench experiential evaluation, its coding performance is on par with Claude Opus 4.8. + +The FlagOS community has completed day‑0 adaptation, precision alignment and deployment validation for AI chips from nine vendors: T‑Head (平头哥), NVIDIA(英伟达), Moore Threads(摩尔), Ascend(华为), Hygon(海光), MetaX(沐曦), Tsingmicro (清微智能), Kunlunxin(昆仑芯) and Sunrise(曦望). Model images have been published to ModelScope and HuggingFace, enabling developers to get out‑of‑the‑box solutions for respective chips directly. + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Mthreads** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + +# Evaluation Results +## Benchmark Result +| Metrics | GLM-5.3-Flash-Nvidia-Origin | GLM-5.3-Flash-Mthreads-FlagOS | +|--------------|--------------------------------|--------------------------------------| +| GPQA_Diamond | 89.29 | 89.6 | +| Musr | 74.6 | 74.7 | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker version 20.10.12 | +| Operating System | 22.04 LTS | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/glm-5.3-flash-mthreads001-gems5.0.2-treenone-cx0.13.0-pluginnone-vllmnone-cp310-pt29-musa43-x64-3.3.5-server:202608271817 + +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/GLM-5.3-Flash-FP8-mthreads-FlagOS --local_dir /data/GLM-5.3-Flash +``` + +### Start the Container +```bash +docker run -dit \ + --name flagos \ + --privileged \ + --ipc host \ + --network host \ + --shm-size 512g \ + -w /workspace \ + -v /data/:/data/ \ + -v /etc/localtime:/etc/localtime:ro \ + -v /etc/timezone:/etc/timezone:ro \ + --env MTHREADS_VISIBLE_DEVICES=all \ +harbor.baai.ac.cn/flagrelease-public/glm-5.3-flash-mthreads001-gems5.0.2-treenone-cx0.13.0-pluginnone-vllmnone-cp310-pt29-musa43-x64-3.3.5-server:202608271817 \ + sleep infinity +docker exec -it flagos /bin/bash +``` +### Start the Server +```bash +nohup /workspace/glm5-adapt/start_glm53_tp8_breakable_sep_moe_bs16.sh \ + > /data/glm53.log 2>&1 & +``` + +## Service Invocation +### Invocation Script +```bash +curl -X POST http://127.0.0.1:31000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "GLM-5.3-Flash", + "messages": [ + {"role": "user", "content": "中国的首都是哪里?"} + ], + "temperature": 0.7, + "max_tokens": 500 + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of . This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from ZhipuAI/GLM-5.3-Flash and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Hy4-preview-FP8-mthreads-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Hy4-preview-FP8-mthreads-FlagOS.md new file mode 100644 index 000000000..d355f7ee9 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Hy4-preview-FP8-mthreads-FlagOS.md @@ -0,0 +1,140 @@ +--- +license: apache-2.0 +language: +- zh +- en +--- + +# Introduction +On August 28, Tencent Hunyuan released and open‑sourced its new‑generation flagship model Hy4 preview. The Zhongzhi(众智) FlagOS Community completed cross‑chip adaptation, precision alignment and deployment validation of Hy4 preview across AI chips. Multi‑chip‑adapted model images have been simultaneously released on ModelScope and HuggingFace. Developers can directly obtain out‑of‑the‑box solutions for their respective chips. + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Mthreads** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + +# Evaluation Results +## Benchmark Result +| Metrics | Hy4-preview-Nvidia-Origin | Hy4-preview-Mthreads-FlagOS | +|--------------|--------------------------------|--------------------------------------| +| GPQA_Diamond | 90.91 | Evaluating | +| Musr_team | 82.8 | Evaluating | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker version 20.10.12 | +| Operating System | 22.04 LTS | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/hy4-preview-mthreads001-gems5.0.2-treenone-cx0.13.0-pluginnone-vllmnone-cp310-pt29-musa43-x64-3.3.5-server:202608300624 + +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/Hy4-preview-FP8-mthreads-FlagOS --local_dir /data/Hy4-preview +``` + +### Start the Container +```bash +docker run -dit \ + --name Hy4_preview \ + --privileged \ + --ipc host \ + --network host \ + --shm-size 512g \ + -w /workspace \ + -v /data/:/data/ \ + -v /etc/localtime:/etc/localtime:ro \ + -v /etc/timezone:/etc/timezone:ro \ + --env MTHREADS_VISIBLE_DEVICES=all \ +harbor.baai.ac.cn/flagos-inner-models-release/preview-mthreads001-gems5.0.2-treenone-cx0.13.0-pluginnone-vllmnone-cp310-pt29-musa43-x64-3.3.5-server:202608300624 \ + sleep infinity +``` +### Start the Server +```bash +# first in rank 1: +docker exec -d \ + -e DIST_INIT_ADDR=:29699 \ + -e LOG_FILE=/workspace/hy4_adapter_musa/logs/hy4_gpqa_rank1.log \ + Hy4_preview \ + /opt/hy4_tp16_handoff/launch_hy4_pr36805_worker3208x.sh + +# then in rank 0: +docker exec -d \ + -e DIST_INIT_ADDR=:29699 \ + -e LOG_FILE=/workspace/hy4_adapter_musa/logs/hy4_gpqa_rank0.log \ + Hy4_preview \ + /opt/hy4_tp16_handoff/launch_hy4_pr36805_worker3208x.sh +``` + +## Service Invocation +### Invocation Script +```bash +curl http://:30100/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "Hy4-preview", + "messages": [ + {"role": "user", "content": "你好,请介绍一下自己"} + ] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of . This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from Tencent-Hunyuan/Hy4-preview-FP8 and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Hy4-preview-FP8-nvidia-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Hy4-preview-FP8-nvidia-FlagOS.md new file mode 100644 index 000000000..fbfa37fd6 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Hy4-preview-FP8-nvidia-FlagOS.md @@ -0,0 +1,233 @@ +--- +license: apache-2.0 +language: +- zh +- en +--- + +# Introduction +On August 28, Tencent Hunyuan released and open‑sourced its new‑generation flagship model Hy4 preview. The Zhongzhi(众智) FlagOS Community completed cross‑chip adaptation, precision alignment and deployment validation of Hy4 preview across AI chips. Multi‑chip‑adapted model images have been simultaneously released on ModelScope and HuggingFace. Developers can directly obtain out‑of‑the‑box solutions for their respective chips. + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Nvidia** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + +# Evaluation Results +## Benchmark Result +| Metrics | Hy4-preview-Nvidia-Origin | Hy4-preview-Nvidia-FlagOS | +|--------------|--------------------------------|--------------------------------------| +| GPQA_Diamond | 90.91 | 89.9 | +| Musr_team | 82.8 | 82.8 | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker version 24.0.0 | +| Operating System | 22.04.4 LTS | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/hy4-preview-fp8-nvidia004-gems5.3.3-tree0.5.0-cxnone-plugin0.3.0-vllm0.24.0-cp312-pt211-cu129-x64-580.126.20:202608300817 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/Hy4-preview-FP8-nvidia-FlagOS --local_dir /data/Hy4-preview +``` + +### Start the Container +```bash +set -euo pipefail + +: "${NODE_RANK:?set NODE_RANK to 0 on the API node or 1 on the headless node}" +[[ "${NODE_RANK}" =~ ^[01]$ ]] || { echo "NODE_RANK must be 0 or 1" >&2; exit 2; } +IMAGE_REF="${IMAGE_REF:-harbor.baai.ac.cn/flagos-inner-models-release/hy4-preview-fp8-testing-nvidia003-gems5.3.3-tree0.5.0-cxnone-plugin0.3.0-vllm0.24.0-cp312-pt211-cu129-x64-580.126.20@sha256:4bf8be17128e2b5cdb2ed085bc1e2c621580b08e51d3378e60f901b4fb908927}" +MODEL_SOURCE="${MODEL_SOURCE:-/data/Hy4-preview-FP8-Testing}" +CONTAINER="${CONTAINER:-hy4-flagos-n${NODE_RANK}}" +test -d "${MODEL_SOURCE}" + +driver_libcuda=$(readlink -f "$(ldconfig -p | awk "/libcuda.so.1 (libc6,x86-64)/ {print \$NF; exit}")") +driver_libnvml=$(readlink -f "$(ldconfig -p | awk "/libnvidia-ml.so.1 (libc6,x86-64)/ {print \$NF; exit}")") +driver_libptxjit=$(readlink -f "$(ldconfig -p | awk "/libnvidia-ptxjitcompiler.so.1 (libc6,x86-64)/ {print \$NF; exit}")") +test -f "${driver_libcuda}" +test -f "${driver_libnvml}" +test -f "${driver_libptxjit}" + +device_args=() +for device in /dev/infiniband/* /dev/nvidia[0-9]* /dev/nvidiactl /dev/nvidia-uvm /dev/nvidia-uvm-tools /dev/nvidia-nvlink /dev/nvidia-nvswitch* /dev/nvidia-caps/*; do + [[ -e "${device}" ]] && device_args+=(--device "${device}":${device}) +done + +docker run -d \ + --name "${CONTAINER}" \ + --network host \ + --ipc host \ + --pid host \ + --pids-limit=-1 \ + --shm-size=128g \ + --ulimit memlock=-1:-1 \ + --ulimit stack=67108864:67108864 \ + --cap-add SYS_PTRACE \ + --security-opt seccomp=unconfined \ + --runtime runc \ + --gpus all \ + ${device_args[@]} \ + -v "${driver_libcuda}":/driver/libcuda.so.1:ro \ + -v "${driver_libcuda}":/driver/libcuda.so:ro \ + -v "${driver_libnvml}":/driver/libnvidia-ml.so.1:ro \ + -v "${driver_libnvml}":/driver/libnvidia-ml.so:ro \ + -v "${driver_libptxjit}":/driver/libnvidia-ptxjitcompiler.so.1:ro \ + -v "${driver_libptxjit}":/driver/libnvidia-ptxjitcompiler.so:ro \ + -v "${MODEL_SOURCE}":/models/Hy4-preview-FP8-Testing:ro \ + -e "NODE_RANK=${NODE_RANK}" \ + -e CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 \ + -e VLLM_PLUGINS=fl \ + -e USE_FLAGGEMS=1 \ + -e HF_HUB_OFFLINE=1 \ + -e TRANSFORMERS_OFFLINE=1 \ + --entrypoint /usr/bin/bash \ + harbor.baai.ac.cn/flagrelease-public/hy4-preview-fp8-nvidia004-gems5.3.3-tree0.5.0-cxnone-plugin0.3.0-vllm0.24.0-cp312-pt211-cu129-x64-580.126.20:202608300817 -lc "exec sleep infinity" +``` +### Start the Server +```bash +set -euo pipefail + +: "${NODE_RANK:?set NODE_RANK to 0 on the API node or 1 on the headless node}" +: "${MASTER_ADDR:?set MASTER_ADDR to the address reachable by both nodes}" +[[ "${NODE_RANK}" =~ ^[01]$ ]] || { echo "NODE_RANK must be 0 or 1" >&2; exit 2; } +CONTAINER="${CONTAINER:-hy4-flagos-n${NODE_RANK}}" +MASTER_PORT="${MASTER_PORT:-29978}" +PORT="${PORT:-8000}" +MODEL_PATH=/data/Hy4-preview +reasoning_plugin=$(docker exec "${CONTAINER}" /usr/bin/python3 -c "import vllm_fl.reasoning.hy_v4_reasoning_parser as m; print(m.__file__)") +net_env=() +if [[ -n "${GLOO_SOCKET_IFNAME-}" ]]; then net_env+=(-e "GLOO_SOCKET_IFNAME=${GLOO_SOCKET_IFNAME}"); fi +if [[ -n "${NCCL_SOCKET_IFNAME-}" ]]; then net_env+=(-e "NCCL_SOCKET_IFNAME=${NCCL_SOCKET_IFNAME}"); fi +role_args=(--headless) +if [[ "${NODE_RANK}" == 0 ]]; then role_args=(--host 0.0.0.0 --port "${PORT}"); fi + +docker exec -d ${net_env[@]} \ + -v /data:/data \ + -e "NODE_RANK=${NODE_RANK}" \ + -e "MASTER_ADDR=${MASTER_ADDR}" \ + -e "MASTER_PORT=${MASTER_PORT}" \ + -e "PORT=${PORT}" \ + -e LD_LIBRARY_PATH=/driver:/usr/local/cuda/lib64 \ + -e CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 \ + -e HF_HUB_OFFLINE=1 \ + -e TRANSFORMERS_OFFLINE=1 \ + -e TOKENIZERS_PARALLELISM=false \ + -e VLLM_PLUGINS=fl \ + -e USE_FLAGGEMS=1 \ + -e VLLM_FL_PREFER=flagos \ + -e VLLM_FL_PREFER_ENABLED=true \ + -e VLLM_FL_FLAGOS_BLACKLIST_APPEND=linear \ + -e VLLM_FL_FLAGOS_MM_SHAPE_AWARE=1 \ + -e VLLM_FL_FLAGOS_MM_DECODE_MAX_M=64 \ + -e VLLM_HY4_HC_N8_PROJECTION=1 \ + -e VLLM_HY4_HC_POINTWISE_FUSION=1 \ + -e TORCHINDUCTOR_COMPILE_THREADS=1 \ + "${CONTAINER}" /usr/local/bin/vllm serve "${MODEL_PATH}" \ + --tokenizer "${MODEL_PATH}" \ + --served-model-name hy4-preview-fp8 \ + --tensor-parallel-size 16 \ + --pipeline-parallel-size 1 \ + --distributed-executor-backend mp \ + --nnodes 2 \ + --node-rank "${NODE_RANK}" \ + --master-addr "${MASTER_ADDR}" \ + --master-port "${MASTER_PORT}" \ + --distributed-timeout-seconds 1800 \ + --cpu-distributed-timeout-seconds 1800 \ + --load-format hy4_safetensors \ + --dtype bfloat16 \ + --disable-custom-all-reduce \ + --max-model-len 102400 \ + --max-num-seqs 1 \ + --max-num-batched-tokens 8192 \ + --gpu-memory-utilization 0.81 \ + --enable-chunked-prefill \ + --no-enable-prefix-caching \ + --seed 0 \ + --no-enable-log-requests \ + --compilation-config "{\"pass_config\":{\"fuse_allreduce_rms\":false}}" \ + --enable-expert-parallel \ + --enable-ep-weight-filter \ + --all2all-backend allgather_reducescatter \ + --reasoning-parser hy_v4 \ + --reasoning-parser-plugin "${reasoning_plugin}" \ + ${role_args[@]} +``` + +## Service Invocation +### Invocation Script +```bash +curl --fail-with-body --noproxy "*" 'http://127.0.0.1:8000/v1/chat/completions' \ + -H 'Content-Type: application/json' \ + --data-raw '{ + "model":"hy4-preview-fp8", + "messages":[ + {"role":"user","content":"Return only 7391"} + ], + "temperature":0, + "max_tokens":32 + }' + +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of . This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from Tencent-Hunyuan/Hy4-preview-FP8 and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Hy4-preview-INT8-ascend-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Hy4-preview-INT8-ascend-FlagOS.md new file mode 100644 index 000000000..88c4f5661 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Hy4-preview-INT8-ascend-FlagOS.md @@ -0,0 +1,140 @@ +--- +license: apache-2.0 +language: +- zh +- en +--- + +# Introduction +On August 28, Tencent Hunyuan released and open‑sourced its new‑generation flagship model Hy4 preview. The Zhongzhi(众智) FlagOS Community completed cross‑chip adaptation, precision alignment and deployment validation of Hy4 preview across AI chips. Multi‑chip‑adapted model images have been simultaneously released on ModelScope and HuggingFace. Developers can directly obtain out‑of‑the‑box solutions for their respective chips. + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Ascend** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + +# Evaluation Results +## Benchmark Result +| Metrics | Hy4-preview-Nvidia-Origin | Hy4-preview-Ascend-FlagOS | +|--------------|--------------------------------|--------------------------------------| +| GPQA_Diamond | 90.91 | Evaluating | +| Musr_team | 82.8 | Evaluating | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker version 20.10.8 | +| Operating System | openEuler 22.03 (LTS-SP4) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/hy4-preview-bf16-ascend001-gems5.3.4-tree0.6.1-cxnone-pluginnone-vllmnone-sglang0.5.17-sglangfl0.1.0-cp311-ptnpu210-cann90-a64-25.5.0:202608301205 + +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/Hy4-preview-INT8-ascend-FlagOS --local_dir /data/Hy4-preview +``` + +### Start the Container +```bash +docker run -dit \ + --name hy4-preview-sglang-plugin \ + --privileged \ + --init \ + --network=host --ipc=host --shm-size=512g \ + --device=/dev/davinci0 --device=/dev/davinci1 --device=/dev/davinci2 --device=/dev/davinci3 \ + --device=/dev/davinci4 --device=/dev/davinci5 --device=/dev/davinci6 --device=/dev/davinci7 \ + --device=/dev/davinci8 --device=/dev/davinci9 --device=/dev/davinci10 --device=/dev/davinci11 \ + --device=/dev/davinci12 --device=/dev/davinci13 --device=/dev/davinci14 --device=/dev/davinci15 \ + --device=/dev/davinci_manager \ + --device=/dev/hisi_hdc \ + --volume /usr/local/sbin:/usr/local/sbin \ + --volume /usr/local/Ascend/driver:/usr/local/Ascend/driver \ + --volume /usr/local/Ascend/firmware:/usr/local/Ascend/firmware \ + --volume /etc/ascend_install.info:/etc/ascend_install.info \ + --volume /var/queue_schedule:/var/queue_schedule \ + --volume /data:/models \ + -v /etc/localtime:/etc/localtime:ro \ + -v /etc/timezone:/etc/timezone:ro \ + -v /data:/data \ + -e TZ=$(cat /etc/timezone 2>/dev/null || echo Asia/Shanghai) \ + --entrypoint=bash \ + harbor.baai.ac.cn/flagrelease-public/hy4-preview-bf16-ascend001-gems5.3.4-tree0.6.1-cxnone-pluginnone-vllmnone-sglang0.5.17-sglangfl0.1.0-cp311-ptnpu210-cann90-a64-25.5.0:202608301205 + +docker exec -it hy4-preview-sglang-plugin /bin/bash +``` +### Start the Server +```bash +# 4 node of 16*910C servers recommended for bf16 inference +cd /sgl-workspace +bash launch.sh +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "hy4-preview", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of . This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from Tencent-Hunyuan/Hy4-preview and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Hy4-preview-INT8-hygon-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Hy4-preview-INT8-hygon-FlagOS.md new file mode 100644 index 000000000..4fbe2a3a8 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Hy4-preview-INT8-hygon-FlagOS.md @@ -0,0 +1,155 @@ +--- +license: apache-2.0 +language: +- zh +- en +--- + +# Introduction +On August 28, Tencent Hunyuan released and open‑sourced its new‑generation flagship model Hy4 preview. The Zhongzhi(众智) FlagOS Community completed cross‑chip adaptation, precision alignment and deployment validation of Hy4 preview across AI chips. Multi‑chip‑adapted model images have been simultaneously released on ModelScope and HuggingFace. Developers can directly obtain out‑of‑the‑box solutions for their respective chips. + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Hygon** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + +# Evaluation Results +## Benchmark Result +| Metrics | Hy4-preview-Nvidia-Origin | Hy4-preview-Hygon-FlagOS | +|--------------|--------------------------------|--------------------------------------| +| GPQA_Diamond | 90.91 | Evaluating | +| Musr_team | 82.8 | Evaluating | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker version 27.3.1 | +| Operating System | 22.04 LTS | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/hy4-preview-hygon001-gems5.4.0-tree0.6.1-cxnone-plugin0.2.0-vllm0.24.0-cp310-pt210-dtk2604-x64-6.3.30-v1.4.1a:202608292346 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/Hy4-preview-INT8-hygon-FlagOS --local_dir /data/Hy4-preview +``` + +### Start the Container +```bash +docker run -d --name flagos \ + --network host --ipc host \ + --device /dev/kfd --device /dev/mkfd --device /dev/dri \ + --mount type=bind,src=/dev/infiniband,dst=/dev/infiniband \ + --mount type=bind,src=/sys/class/infiniband,dst=/sys/class/infiniband,readonly \ + --mount type=bind,src=/sys/class/infiniband_verbs,dst=/sys/class/infiniband_verbs,readonly \ + --mount type=bind,src=/sys/class/net,dst=/sys/class/net,readonly \ + --mount type=bind,src=/usr/etc/libibverbs.d,dst=/usr/etc/libibverbs.d,readonly \ + --mount type=bind,src=/lib/x86_64-linux-gnu/libibverbs.so.1.14.44.0,dst=/lib/x86_64-linux-gnu/libibverbs.so.1.14.47.0,readonly \ + --mount type=bind,src=/lib/x86_64-linux-gnu/libshca-rdmav34.so,dst=/lib/x86_64-linux-gnu/libshca-rdmav34.so,readonly \ + --mount type=bind,src=/lib/x86_64-linux-gnu/libnl-3.so.200,dst=/lib/x86_64-linux-gnu/libnl-3.so.200,readonly \ + --mount type=bind,src=/lib/x86_64-linux-gnu/libnl-route-3.so.200,dst=/lib/x86_64-linux-gnu/libnl-route-3.so.200,readonly \ + -v /opt/hyhal:/opt/hyhal \ + -v /data:/data:rslave \ + -v /public-flash:/public-flash \ + --group-add video \ + --cap-add SYS_PTRACE \ + --security-opt seccomp=unconfined \ + --security-opt label=disable \ + -e HSA_FORCE_FINE_GRAIN_PCIE=1 \ + -e NCCL_IB_DISABLE=0 \ + harbor.baai.ac.cn/flagos-inner-models-release/hy4-preview-hygon001-gems5.4.0-tree0.6.1-cxnone-plugin0.2.0-vllm0.24.0-cp310-pt210-dtk2604-x64-6.3.30-v1.4.1a:202608292346 \ + bash -lc 'sleep infinity' +``` +### Start the Server +```bash +set +o nounset +source /opt/dtk-26.04-DCC2602-0317/env.sh +set -o nounset +export HIP_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 +export HSA_FORCE_FINE_GRAIN_PCIE=1 +export TRITON_HIP_CLANG_PATH=/opt/dtk-26.04-DCC2602-0317/aillvm/bin/clang-18 +export LD_LIBRARY_PATH="/opt/ucx-shca/lib:/opt/rccl-shca-net/lib:${LD_LIBRARY_PATH:-}" +export UCX_MODULE_DIR=/opt/ucx-shca/lib/ucx +export MASTER_ADDR=10.232.2.19 +export GLOO_SOCKET_IFNAME=ib0 +export NCCL_SOCKET_IFNAME=ib0 +export NCCL_NET_PLUGIN=shca +export NCCL_IB_DISABLE=0 +export NCCL_IB_HCA=shca_0,shca_1,shca_2,shca_3 +export HF_HUB_OFFLINE=1 +export TRANSFORMERS_OFFLINE=1 +export HF_DATASETS_OFFLINE=1 +# vLLM 0.24+ RPC timeout for execute_model calls (seconds). +# W8A8 quantized models need extra time on first inference for kernel compilation. +export VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=1800 +vllm serve /data/Hy4-preview --served-model-name Hy4-preview-W8A8-linear-moe --trust-remote-code --distributed-executor-backend mp --nnodes 4 --node-rank 0 --master-addr 10.232.2.30 --master-port 29864 --distributed-timeout-seconds 1800 --tensor-parallel-size 16 --pipeline-parallel-size 2 --data-parallel-size 1 --load-format hy4_safetensors --disable-custom-all-reduce --moe-backend triton --max-model-len 131072 --max-num-seqs 64 --gpu-memory-utilization 0.95 --no-enable-log-requests --enable-expert-parallel --all2all-backend allgather_reducescatter --enable-ep-weight-filter --reasoning-parser hy_v4 --host 0.0.0.0 --port 8000 +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "flagOS", + "messages": [{"role": "user", "content": "hi"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of . This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from Tencent-Hunyuan/Hy4-preview and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Hy4-preview-INT8-kunlunxin-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Hy4-preview-INT8-kunlunxin-FlagOS.md new file mode 100644 index 000000000..d763eff63 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Hy4-preview-INT8-kunlunxin-FlagOS.md @@ -0,0 +1,176 @@ +--- +license: apache-2.0 +language: +- zh +- en +--- + +# Introduction +On August 28, Tencent Hunyuan released and open‑sourced its new‑generation flagship model Hy4 preview. The Zhongzhi(众智) FlagOS Community completed cross‑chip adaptation, precision alignment and deployment validation of Hy4 preview across AI chips. Multi‑chip‑adapted model images have been simultaneously released on ModelScope and HuggingFace. Developers can directly obtain out‑of‑the‑box solutions for their respective chips. + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Kunlunxin** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + +# Evaluation Results +## Benchmark Result +| Metrics | Hy4-preview-Nvidia-Origin | Hy4-preview-Kunlunxin-FlagOS | +|--------------|--------------------------------|--------------------------------------| +| GPQA_Diamond | 90.91 | Evaluating | +| Musr_team | 82.8 | Evaluating | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker version 20.10.12 | +| Operating System | 22.04 LTS | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/hy4-preview-w8a8-linear-moe-kunlunxin001-gems5.0.0-treenone-cx0.13.0-plugin0.2.0-vllm0.20.2-cp310-pt29-xrt513-x64-5.0.21.47:202608310922 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/Hy4-preview-INT8-kunlunxin-FlagOS --local_dir /data/Hy4-preview +``` + +### Start the Container +```bash +set -euo pipefail + +# Run once on each node. Set NODE_RANK=0 on the API node and NODE_RANK=1 +# on the worker node. MASTER_ADDR must be the routable address of node 0. +: "${NODE_RANK:?set NODE_RANK to 0 or 1}" +: "${MASTER_ADDR:?set MASTER_ADDR to the address of node 0}" +MASTER_PORT="${MASTER_PORT:-29564}" +CONTAINER_NAME="${CONTAINER_NAME:-hy4-w8a8-flagos-p800}" + +docker run -d \ + --name "${CONTAINER_NAME}" \ + --privileged \ + --network host \ + --shm-size 128g \ + --cap-add SYS_PTRACE \ + --ulimit memlock=-1:-1 \ + --ulimit nofile=120000:120000 \ + --ulimit stack=67108864:67108864 \ + -e NODE_RANK="${NODE_RANK}" \ + -e MASTER_ADDR="${MASTER_ADDR}" \ + -e MASTER_PORT="${MASTER_PORT}" \ + -v /data:/data \ + --entrypoint /bin/bash \ + harbor.baai.ac.cn/flagrelease-public/hy4-preview-w8a8-linear-moe-kunlunxin001-gems5.0.0-treenone-cx0.13.0-plugin0.2.0-vllm0.20.2-cp310-pt29-xrt513-x64-5.0.21.47:202608310922 \ + -lc 'sleep infinity' +``` +### Start the Server +```bash +# Run once on each node after both containers exist. The API is exposed by +# rank 0 at port 8000 after all 16 ranks have joined. +CONTAINER_NAME="${CONTAINER_NAME:-hy4-w8a8-flagos-p800}" + +docker exec -d "${CONTAINER_NAME}" /bin/bash -lc ' +set -euo pipefail +export USE_FLAGGEMS=1 +export VLLM_DISABLE_PYNCCL=1 +export VLLM_PLUGINS=fl +export VLLM_WORKER_MULTIPROC_METHOD=spawn +export VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=1800 +export PYTHONUNBUFFERED=1 +export CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 +exec /root/miniconda/envs/python310_torch29_cuda/bin/vllm serve /data/Hy4-preview \ + --served-model-name hy4-w8a8 \ + --host 0.0.0.0 \ + --port 8000 \ + --tensor-parallel-size 16 \ + --enable-expert-parallel \ + --enable-ep-weight-filter \ + --all2all-backend allgather_reducescatter \ + --distributed-executor-backend mp \ + --nnodes 2 \ + --node-rank "${NODE_RANK}" \ + --master-addr "${MASTER_ADDR}" \ + --master-port "${MASTER_PORT}" \ + --distributed-timeout-seconds 1800 \ + --load-format hy4_safetensors \ + --dtype bfloat16 \ + --disable-custom-all-reduce \ + --max-model-len 8192 \ + --max-num-seqs 8 \ + --max-num-batched-tokens 1024 \ + --block-size 128 \ + --gpu-memory-utilization 0.90 \ + --enable-chunked-prefill \ + --no-enable-prefix-caching \ + --enforce-eager >>"/var/log/hy4-w8a8-vllm-rank${NODE_RANK}.log" 2>&1 +' +``` + +## Service Invocation +### Invocation Script +```bash +set -euo pipefail + +# Run on node 0 after /health returns HTTP 200. +curl --noproxy '*' -fsS http://127.0.0.1:8000/v1/chat/completions \ + -H 'Content-Type: application/json' \ + -d '{"model":"hy4-w8a8","messages":[{"role":"user","content":"Compute 17 * 19. Return only the integer."}],"temperature":0,"max_tokens":64}' + +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of . This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from Tencent-Hunyuan/Hy4-preview and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Hy4-preview-INT8-metax-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Hy4-preview-INT8-metax-FlagOS.md new file mode 100644 index 000000000..67111376f --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Hy4-preview-INT8-metax-FlagOS.md @@ -0,0 +1,216 @@ +--- +license: apache-2.0 +language: +- zh +- en +--- + +# Introduction +On August 28, Tencent Hunyuan released and open‑sourced its new‑generation flagship model Hy4 preview. The Zhongzhi(众智) FlagOS Community completed cross‑chip adaptation, precision alignment and deployment validation of Hy4 preview across AI chips. Multi‑chip‑adapted model images have been simultaneously released on ModelScope and HuggingFace. Developers can directly obtain out‑of‑the‑box solutions for their respective chips. + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Metax** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + +# Evaluation Results +## Benchmark Result +| Metrics | Hy4-preview-Nvidia-Origin | Hy4-preview-Metax-FlagOS | +|--------------|--------------------------------|--------------------------------------| +| GPQA_Diamond | 90.91 | Evaluating | +| Musr_team | 82.8 | Evaluating | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker version 29.3.1 | +| Operating System | 22.04 LTS | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/hy4-preview-metax001-gems5.4.0-treenone-cxnone-plugin3.0.0-vllm0.24.0-cp312-pt28-maca37-x64-3.8.1:202608301009 + +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/Hy4-preview-INT8-metax-FlagOS --local_dir /data/Hy4-preview +``` + +### Start the Container +```bash +IMAGE="harbor.baai.ac.cn/flagrelease-public/hy4-preview-metax001-gems5.4.0-treenone-cxnone-plugin3.0.0-vllm0.24.0-cp312-pt28-maca37-x64-3.8.1:202608301009" +docker run -d \ + --name flagos\ + --network host \ + --shm-size 64g \ + --device /dev/dri:/dev/dri:rwm \ + --device /dev/mxcd:/dev/mxcd:rwm \ + -v /data:/data \ + ${IMAGE} \ + sleep infinity +docker exec -it flagos/bin/bash +``` +### Start the Server +```bash +# Hy4 多机 Ray 公共环境变量 +# 所有节点(head + worker)执行 source ray_env.sh 后再启动 ray + +# ============ 网络配置 ============ +# 查看网卡: ip addr / ifconfig +# 用了 --network=host,容器里网卡和宿主机一致 +export GLOO_SOCKET_IFNAME=inbond1 +export MCCL_SOCKET_IFNAME=inbond1 + +# ============ Ray 集群配置 ============ +# HEAD_IP 需要按实际填写(主机的 inbond1 IP) +export HEAD_IP="" +export HEAD_PORT=6379 +export RAY_ADDRESS="${HEAD_IP}:${HEAD_PORT}" + +# ============ vLLM / MetaX 运行时 ============ +export GEMS_VENDOR=metax +export VLLM_PLUGINS=fl +export VLLM_WORKER_MULTIPROC_METHOD=spawn +export VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=6000 +export VLLM_ENGINE_ITERATION_TIMEOUT_S=7200 +export VLLM_ENGINE_READY_TIMEOUT_S=3600 + +export VLLM_FL_FLAGOS_BLACKLIST="mm,mm_out,bmm,bmm_out,full,broadcast_to,nonzero,nonzero_numpy,linear,diff,masked_fill,masked_fill_,sort,stable_sort,topk,layer_norm" + +export VLLM_HY4_SPARSE_TORCH_REF=0 + +# ============ Ray Runtime Env ============ +# 这个变量让 ray 在所有 worker 进程中自动注入环境变量 +# 只在 head 节点 vllm serve 之前 export 即可(worker 通过 ray 传播) +export RAY_RUNTIME_ENV='{ + "env_vars": { + "VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS": "6000", + "VLLM_ENGINE_ITERATION_TIMEOUT_S": "7200", + "GLOO_SOCKET_IFNAME": "inbond1", + "MCCL_SOCKET_IFNAME": "inbond1", + "GEMS_VENDOR": "metax", + "VLLM_PLUGINS": "fl", + "VLLM_HY4_SPARSE_TORCH_REF": "0", + "VLLM_FL_FLAGOS_BLACKLIST": "mm,mm_out,bmm,bmm_out,full,broadcast_to,nonzero,nonzero_numpy,linear,diff,masked_fill,masked_fill_,sort,stable_sort,topk,layer_norm" + } +}' + +# ============ 模型配置 ============ +export MODEL_PATH="/data/Hy4-preview" +export MODEL_NAME="hy4" +export PORT=8000 +export TP_SIZE=8 +export PP_SIZE=4 +export MAX_MODEL_LEN=8192 + +set -e + +SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)" +source "${SCRIPT_DIR}/ray_env.sh" + +# 确认 ray 集群存活 +echo "[SERVE] 检查 ray 集群状态..." +ray status || { echo "[SERVE] ERROR: ray 集群不可用,请先执行 ray_start_head.sh + ray_start_worker.sh"; exit 1; } +echo "" + +# 启动 vllm serve +LOG_DIR="/data/hy4_running/logs" +LOG_FILE="${LOG_DIR}/hy4-ray-serve-$(date +%Y%m%d-%H%M%S).log" +mkdir -p ${LOG_DIR} + +echo "[SERVE] 启动 vllm serve..." | tee -a ${LOG_FILE} +echo "[SERVE] 日志: ${LOG_FILE}" | tee -a ${LOG_FILE} +echo "[SERVE] 启动时间: $(date)" | tee -a ${LOG_FILE} +echo "[SERVE] TP=${TP_SIZE} PP=${PP_SIZE} max_model_len=${MAX_MODEL_LEN}" | tee -a ${LOG_FILE} +echo "[SERVE] 模型: ${MODEL_PATH}" | tee -a ${LOG_FILE} +echo "" + +/opt/conda/bin/vllm serve ${MODEL_PATH} \ + --port ${PORT} \ + --served-model-name ${MODEL_NAME} \ + --tensor-parallel-size ${TP_SIZE} \ + --pipeline-parallel-size ${PP_SIZE} \ + --distributed-executor-backend ray \ + --dtype bfloat16 \ + --gpu-memory-utilization 0.9 \ + --no-async-scheduling \ + --enforce-eager \ + --max-num-seqs 4 \ + --trust-remote-code \ + --compilation-config '{"pass_config":{"fuse_allreduce_rms":false}}' \ + --max-model-len 32768 \ + 2>&1 | tee -a ${LOG_FILE} +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ +-H "Content-Type: application/json" \ +-d '{ + "model": "hy4", + "messages": [{"role": "user", "content": "中国的首都是哪里?"}], + "temperature": 0.7, + "max_tokens": 128, + "chat_template_kwargs": { + "enable_thinking": false + } +}' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of . This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from Tencent-Hunyuan/Hy4-preview and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Hy4-preview-INT8-zhenwu-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Hy4-preview-INT8-zhenwu-FlagOS.md new file mode 100644 index 000000000..f308569e6 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Hy4-preview-INT8-zhenwu-FlagOS.md @@ -0,0 +1,149 @@ +--- +license: apache-2.0 +language: +- zh +- en +--- + +# Introduction +On August 28, Tencent Hunyuan released and open‑sourced its new‑generation flagship model Hy4 preview. The Zhongzhi(众智) FlagOS Community completed cross‑chip adaptation, precision alignment and deployment validation of Hy4 preview across AI chips. Multi‑chip‑adapted model images have been simultaneously released on ModelScope and HuggingFace. Developers can directly obtain out‑of‑the‑box solutions for their respective chips. + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Zhenwu** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + +# Evaluation Results +## Benchmark Result +| Metrics | Hy4-preview-Nvidia-Origin | Hy4-preview-Zhenwu-FlagOS | +|--------------|--------------------------------|--------------------------------------| +| GPQA_Diamond | 90.91 | 90.91 | +| Musr_team | 82.8 | 84.8 | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker version 29.7.2 | +| Operating System | 24.04.2 LTS | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/hy4-preview-pp001-gems0.0-tree0.6.1-cxnone-plugin0.2.0-vllm0.24.0-cp312-pt210-hggc130-x64-2.1.1-rbd225:202608281106 + +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/Hy4-preview-INT8-zhenwu-FlagOS --local_dir /data/Hy4-preview +``` + +### Start the Container +```bash +sudo docker run --privileged -dit \ + --network=host \ + --device=/dev/infiniband \ + --ipc=host \ + --device=/dev/alixpu_ctl \ + --device=/dev/alixpu \ + --ulimit memlock=-1 \ + --ulimit stack=67108864 \ + --init \ + -v /mnt:/mnt \ + -v /data:/data \ + --name hyv4 \ + harbor.baai.ac.cn/flagrelease-public/hy4-preview-pp001-gems0.0-tree0.6.1-cxnone-plugin0.2.0-vllm0.24.0-cp312-pt210-hggc130-x64-2.1.1-rbd225:202608281106 + + +``` +### Start the Server +```bash +VLLM_WORKER_MULTIPROC_METHOD=spawn \ +VLLM_FL_HYV4_INDEXER_TOPK_MODE=scoped_native \ +VLLM_FL_HYV4_UNCOMPILED_GROUPED_TOPK=1 \ +VLLM_FL_HYV4_SAFE_DETOKENIZER=1 \ +FLAGGEMS_DB_URL='sqlite:////root/.flaggems/hy4-tp16.db?timeout=600' \ +VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=1800 \ +vllm serve /data/Hy4-preview \ + --served-model-name hy4 \ + --host 0.0.0.0 \ + --port 8010 \ + --trust-remote-code \ + --reasoning-parser-plugin /workspace/vllm-plugin-FL/vllm_fl/hy_v4_reasoning_parser.py \ + --reasoning-parser hy_v4 \ + --load-format safetensors \ + --tensor-parallel-size 16 \ + --max-model-len 50000 \ + --max-num-seqs 32 \ + --max-num-batched-tokens 1024 \ + --gpu-memory-utilization 0.8 \ + --compilation-config \ + '{"mode":"none","cudagraph_mode":"FULL_DECODE_ONLY","cudagraph_capture_sizes":[1,32]}' + +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "flagOS", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of . This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from Tencent-Hunyuan/Hy4-preview and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Qwen2.5-Coder-7B-Instruct-ascend-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Qwen2.5-Coder-7B-Instruct-ascend-FlagOS.md new file mode 100644 index 000000000..b9e0aa53d --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Qwen2.5-Coder-7B-Instruct-ascend-FlagOS.md @@ -0,0 +1,113 @@ +# Introduction + + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Ascend** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | Qwen2.5-Coder-7B-Instruct-ascend-FlagOS-Origin | Qwen2.5-Coder-7B-Instruct-ascend-FlagOS-FlagOS | +|--------------|------------------------------------------------|------------------------------------------------| +| GPQA_Diamond | 38.0 | 32.0 | +| ERQA | - | - | +| Aime24 | - | - | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | 20.10.8 | +| Operating System | Ubuntu 22.04.5 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-project/qwen2.5-coder-7b-instruct-ascend001-gems5.3.4-tree0.6.0-cxnone-plugin0.2.0-vllm-ascend0.20.2-cp311-ptnpu210-cann90-a64-25.5.0:202608261757-v3-hotfix1 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/Qwen2.5-Coder-7B-Instruct-ascend-FlagOS --local_dir /data/Qwen2.5-Coder-7B-Instruct-FlagOS +``` + +### Start the Container +```bash +docker run -d --name flagos --net=host --ipc=host --privileged --shm-size=64g -v /usr/local/Ascend/driver:/usr/local/Ascend/driver -v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi -v /usr/local/dcmi:/usr/local/dcmi -v /usr/local/sbin:/usr/local/sbin -v /etc/ascend_install.info:/etc/ascend_install.info -v /data:/data -e PYTORCH_NPU_ALLOC_CONF=max_split_size_mb:256 harbor.baai.ac.cn/flagrelease-project/qwen2.5-coder-7b-instruct-ascend001-gems5.3.4-tree0.6.0-cxnone-plugin0.2.0-vllm-ascend0.20.2-cp311-ptnpu210-cann90-a64-25.5.0:202608261757-v3-hotfix1 sleep infinity +``` +### Start the Server +```bash +VLLM_PLUGINS=fl vllm serve /data/Qwen2.5-Coder-7B-Instruct-FlagOS \ +--host 0.0.0.0 --port 8000 \ +--tensor-parallel-size 1 \ +--served-model-name Qwen2.5-Coder-7B-Instruct \ +--trust-remote-code \ +--max-model-len 32768 +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "Qwen2.5-Coder-7B-Instruct", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator library implemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutral kernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multiple AI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. For upstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to support the entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ model evaluations across NLP, CV, Audio, and Multimodal fields, covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizontal evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from Qwen/Qwen2.5-Coder-7B-Instruct and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Qwen3.8-Flash-Next-BF16-ascend-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Qwen3.8-Flash-Next-BF16-ascend-FlagOS.md new file mode 100644 index 000000000..4b515f8f2 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Qwen3.8-Flash-Next-BF16-ascend-FlagOS.md @@ -0,0 +1,226 @@ +--- +language: +- zh +- en +license: apache-2.0 +--- + +# Introduction +On August 26, Alibaba open-sourced its multimodal MoE model Qwen3.8-Flash-Next. The FlagOS community completed Day-0 synchronized adaptation across multiple chips, having accomplished multi-chip adaptation, precision alignment, and deployment validation based on the FlagOS unified open-source technology stack on 8 AI chips — including T-Head (Pingtouge), NVIDIA, Moore Threads, Ascend, MetaX, Kunlunxin, Hygon, and Iluvatar CoreX. The first batch uniformly provides a BF16-precision version. The multi-chip versions have been open-sourced on the ModelScope and Hugging Face platforms, where developers can directly obtain out-of-the-box solutions for their respective chips. + + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Ascend** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + +# Evaluation Results +## Benchmark Result +| Metrics | Qwen3.8-Flash-Next-Nvidia-Origin | Qwen3.8-Flash-Next-Ascend-FlagOS | +|--------------|--------------------------------|--------------------------------------| +| GPQA_Diamond | 92.9 | 92.8 | +| MuSR| 78.57 | Evaluating | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker version 20.10.8 | +| Operating System | openEuler 22.03 (LTS-SP4) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/qwen3.8-flash-next-ascend001-gems5.3.0-tree0.6.0-cx0.13.0-pluginnone-vllmnone-sglang0.5.11-sglangfl0.1.0-cp311-ptnpu28-cann85-a64-25.5.0:202609021230 + +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/Qwen3.8-Flash-Next-BF16-ascend-FlagOS --local_dir /data/Qwen3.8-Flash-Next +``` + +### Start the Container +```bash +docker run -dit \ + --name flagos \ + --privileged \ + --init \ + --network=host --ipc=host --shm-size=512g \ + --device=/dev/davinci0 --device=/dev/davinci1 --device=/dev/davinci2 --device=/dev/davinci3 \ + --device=/dev/davinci4 --device=/dev/davinci5 --device=/dev/davinci6 --device=/dev/davinci7 \ + --device=/dev/davinci8 --device=/dev/davinci9 --device=/dev/davinci10 --device=/dev/davinci11 \ + --device=/dev/davinci12 --device=/dev/davinci13 --device=/dev/davinci14 --device=/dev/davinci15 \ + --device=/dev/davinci_manager \ + --device=/dev/hisi_hdc \ + --volume /usr/local/sbin:/usr/local/sbin \ + --volume /usr/local/Ascend/driver:/usr/local/Ascend/driver \ + --volume /usr/local/Ascend/firmware:/usr/local/Ascend/firmware \ + --volume /etc/ascend_install.info:/etc/ascend_install.info \ + --volume /var/queue_schedule:/var/queue_schedule \ + -v /etc/localtime:/etc/localtime:ro \ + -v /etc/timezone:/etc/timezone:ro \ + -v /data:/data \ + -v /public-flash/models:/models \ + --entrypoint=bash \ + harbor.baai.ac.cn/flagrelease-public/qwen3.8-flash-next-ascend001-gems5.3.0-tree0.6.0-cx0.13.0-pluginnone-vllmnone-sglang0.5.11-sglangfl0.1.0-cp311-ptnpu28-cann85-a64-25.5.0:202609021230 \ + -c "tail -f /dev/null" +docker exec -it flagos /bin/bash +``` +### Start the Server +```bash +set -euo pipefail + +# 校验并安装优化后的模型代码 +cd /opt/qwen38-best/current +sha256sum -c payload.sha256 + +install -m 0644 payload/qwen4.py \ + /sgl-workspace/sglang/python/sglang/srt/models/qwen4.py + +install -m 0644 payload/qwen4_ops.py \ + /sgl-workspace/sglang/python/sglang/srt/models/qwen4_ops.py + +install -m 0644 payload/scheduler.py \ + /sgl-workspace/sglang/python/sglang/srt/managers/scheduler.py + +# 设备与线程配置 +export ASCEND_RT_VISIBLE_DEVICES=0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15 +export OMP_NUM_THREADS=8 +export STREAMS_PER_DEVICE=32 +export SGLANG_SET_CPU_AFFINITY=1 + +# 多流与通信配置 +export SGLANG_ENABLE_OVERLAP_PLAN_STREAM=1 +export SGLANG_NPU_USE_MULTI_STREAM=1 +export HCCL_BUFFSIZE=1000 +export HCCL_OP_EXPANSION_MODE=AIV +export HCCL_SOCKET_IFNAME=business +export GLOO_SOCKET_IFNAME=business +export FLAGCX_PATH=/sgl-workspace/FlagCX + +# 内存与离线配置 +export HF_HUB_OFFLINE=1 +export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True +export PYTHONPATH="/sgl-workspace/sglang/python:${PYTHONPATH:-}" + +# FlagGems 与算子配置 +export USE_FLAGGEMS=1 +export SGLANG_FL_WATCHDOG_DIAG=1 +export SGLANG_FL_PER_OP=mrotary_embedding=reference +export SGLANG_FL_FLAGOS_BLACKLIST='conv1d,conv2d,index,index_put,index_put_,_index_put_impl_,full_like,mul,mul_,sub,sub_,remainder,remainder_,floor_divide,floor_divide_,add,add_,ge,ge_scalar,lt,lt_scalar,bitwise_and_scalar,bitwise_and_scalar_tensor,bitwise_and_tensor,bitwise_and_scalar_,bitwise_and_tensor_,bitwise_not,bitwise_not_,fill_scalar,fill_scalar_out,fill_tensor,fill_tensor_out,fill_scalar_,fill_tensor_,sum,sum_out,sum_dim,sum_dim_out,mean,mean_out,mean_dim,mean_dim_out,max,max_dim,gather,gather_backward,argmax,bincount,where_self,where_self_out,where_scalar_self,where_scalar_other,silu,silu_,silu_backward,bmm,bmm_out' + +# Qwen3.8-Flash-Next W8A16 配置 +export SGLANG_QWEN4_MOE_W8A16=1 +export SGLANG_QWEN4_FORCE_NPU_GMM=1 +export SGLANG_QWEN4_STATE_CAPACITY=40000 +export SGLANG_QWEN4_ENABLE_NEXTN=0 +export SGLANG_QWEN4_MTP_MODE=NEXTN + +# 清除可能影响当前最优配置的变量 +unset https_proxy http_proxy HTTPS_PROXY HTTP_PROXY +unset ASCEND_LAUNCH_BLOCKING +unset SGLANG_QWEN4_GRAPH_CANARY +unset SGLANG_QWEN4_GRAPH_FIXED_REQ_KEYS +unset SGLANG_QWEN4_GRAPH_INPLACE_STATE +unset SGLANG_QWEN4_GRAPH_MAX_BS +unset SGLANG_QWEN4_NOPLE + +# 检查运行环境 +python3 -c 'import tbe; print("TBE_IMPORT_OK")' + +# 启动服务 +exec python3 -m sglang.launch_server \ + --model-path /data/Qwen3.8-Flash-Next/ \ + --served-model-name Qwen3.8-Flash-Next \ + --host 0.0.0.0 \ + --port 30100 \ + --tp-size 4 \ + --dp-size 4 \ + --load-balance-method round_robin \ + --attention-backend ascend \ + --trust-remote-code \ + --reasoning-parser qwen3-thinking \ + --skip-server-warmup \ + --disable-cuda-graph \ + --disable-radix-cache \ + --chunked-prefill-size 8192 \ + --max-prefill-tokens 32768 \ + --prefill-max-requests 4 \ + --max-total-tokens 70000 \ + --max-running-requests 4 \ + --max-mamba-cache-size 4 \ + --watchdog-timeout 900 \ + --mem-fraction-static 0.88 +``` + +## Service Invocation +### Invocation Script +```bash +curl -s http://localhost:30100/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "Qwen3.8-Flash-Next", + "messages": [ + {"role": "system", "content": "You are a helpful assistant."}, + {"role": "user", "content": "请用一句话介绍北京。"} + ], + "max_tokens": 50, + "temperature": 0.7 + }' | jq . +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of . This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from Qwen/Qwen3.8-Flash-Next and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt + diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Qwen3.8-Flash-Next-BF16-hygon-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Qwen3.8-Flash-Next-BF16-hygon-FlagOS.md new file mode 100644 index 000000000..27b07281b --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Qwen3.8-Flash-Next-BF16-hygon-FlagOS.md @@ -0,0 +1,180 @@ +--- +language: +- zh +- en +license: apache-2.0 +--- + +# Introduction +On August 26, Alibaba open-sourced its multimodal MoE model Qwen3.8-Flash-Next. The FlagOS community completed Day-0 synchronized adaptation across multiple chips, having accomplished multi-chip adaptation, precision alignment, and deployment validation based on the FlagOS unified open-source technology stack on 8 AI chips — including T-Head (Pingtouge), NVIDIA, Moore Threads, Ascend, MetaX, Kunlunxin, Hygon, and Iluvatar CoreX. The first batch uniformly provides a BF16-precision version. The multi-chip versions have been open-sourced on the ModelScope and Hugging Face platforms, where developers can directly obtain out-of-the-box solutions for their respective chips. + + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Hygon** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | Qwen3.8-Flash-Next-Nvidia-Origin | Qwen3.8-Flash-Next-Hygon-FlagOS | +|--------------|----------------------------------|---------------------------------| +| GPQA_Diamond | 92.9 | 89.9 | +| MuSR | 78.57 | 77.51 | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker version 20.10.5, build 55c4c88 | +| Operating System | Ubuntu 22.04.4 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/qwen3.8-flash-next-hygon001-gems5.4.0-tree0.6.1-cxnone-plugin0.2.0-vllm0.24.0-cp310-pt210-dtk2604-x64-6.3.30-v1.4.1a:202608280655 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/Qwen3.8-Flash-Next-BF16-hygon-FlagOS --local_dir /data/Qwen3.8-Flash-Next +``` + +### Start the Container +```bash +docker run -d --name flagos \ + --network host --ipc host \ + --device /dev/kfd --device /dev/mkfd --device /dev/dri \ + --mount type=bind,src=/dev/infiniband,dst=/dev/infiniband \ + --mount type=bind,src=/sys/class/infiniband,dst=/sys/class/infiniband,readonly \ + --mount type=bind,src=/sys/class/infiniband_verbs,dst=/sys/class/infiniband_verbs,readonly \ + --mount type=bind,src=/sys/class/net,dst=/sys/class/net,readonly \ + --mount type=bind,src=/usr/etc/libibverbs.d,dst=/usr/etc/libibverbs.d,readonly \ + --mount type=bind,src=/lib/x86_64-linux-gnu/libibverbs.so.1.14.44.0,dst=/lib/x86_64-linux-gnu/libibverbs.so.1.14.47.0,readonly \ + --mount type=bind,src=/lib/x86_64-linux-gnu/libshca-rdmav34.so,dst=/lib/x86_64-linux-gnu/libshca-rdmav34.so,readonly \ + --mount type=bind,src=/lib/x86_64-linux-gnu/libnl-3.so.200,dst=/lib/x86_64-linux-gnu/libnl-3.so.200,readonly \ + --mount type=bind,src=/lib/x86_64-linux-gnu/libnl-route-3.so.200,dst=/lib/x86_64-linux-gnu/libnl-route-3.so.200,readonly \ + -v /opt/hyhal:/opt/hyhal \ + -v /data:/data:rslave \ + -v /public-flash:/public-flash \ + --group-add video \ + --cap-add SYS_PTRACE \ + --security-opt seccomp=unconfined \ + --security-opt label=disable \ + -e HSA_FORCE_FINE_GRAIN_PCIE=1 \ + -e NCCL_IB_DISABLE=0 \ + harbor.baai.ac.cn/flagrelease-public/qwen3.8-flash-next-hygon001-gems5.4.0-tree0.6.1-cxnone-plugin0.2.0-vllm0.24.0-cp310-pt210-dtk2604-x64-6.3.30-v1.4.1a:202608280655 \ + bash -lc 'sleep infinity' +docker exec -it flagos /bin/bash +``` +### Start the Server +```bash +set +o nounset +source /opt/dtk-26.04-DCC2602-0317/env.sh +set -o nounset +export HIP_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 +export HSA_FORCE_FINE_GRAIN_PCIE=1 +export TRITON_HIP_CLANG_PATH=/opt/dtk-26.04-DCC2602-0317/aillvm/bin/clang-18 +export LD_LIBRARY_PATH="/opt/ucx-shca/lib:/opt/rccl-shca-net/lib:${LD_LIBRARY_PATH:-}" +export UCX_MODULE_DIR=/opt/ucx-shca/lib/ucx +export MASTER_ADDR=10.232.2.19 +export GLOO_SOCKET_IFNAME=ib0 +export NCCL_SOCKET_IFNAME=ib0 +export NCCL_NET_PLUGIN=shca +export NCCL_IB_DISABLE=0 +export NCCL_IB_HCA=shca_0,shca_1,shca_2,shca_3 +export HF_HUB_OFFLINE=1 +export TRANSFORMERS_OFFLINE=1 +export HF_DATASETS_OFFLINE=1 +# vLLM 0.24+ RPC timeout for execute_model calls (seconds). +# W8A8 quantized models need extra time on first inference for kernel compilation. +export VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=1800 +vllm serve "$MODEL_PATH" \ + --served-model-name "$SERVED_MODEL_NAME" \ + --trust-remote-code \ + --distributed-executor-backend mp \ + --nnodes 1 \ + --node-rank "$NODE_RANK" \ + --master-addr "$MASTER_ADDR" \ + --master-port "$MASTER_PORT" \ + --distributed-timeout-seconds 1800 \ + --tensor-parallel-size "$TP" \ + --pipeline-parallel-size "$PP" \ + --data-parallel-size "$DP" \ + --load-format "$LOAD_FORMAT" \ + --mm-encoder-tp-mode data \ + --disable-custom-all-reduce \ + --moe-backend triton \ + --max-model-len "${MAX_MODEL_LEN:-32768}" \ + --reasoning-parser qwen3 \ + --max-num-seqs 16 \ + --gpu-memory-utilization 0.95 \ + --language-model-only \ + --no-enable-log-requests +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:$MASTER_PORT/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "$SERVED_MODEL_NAME", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., “Explain the basics of quantum computing”) +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a “develop once, run anywhere” workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from Qwen/Qwen3.8-Flash-Next and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt + diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Qwen3.8-Flash-Next-BF16-iluvatar-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Qwen3.8-Flash-Next-BF16-iluvatar-FlagOS.md new file mode 100644 index 000000000..6cbf44d28 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Qwen3.8-Flash-Next-BF16-iluvatar-FlagOS.md @@ -0,0 +1,157 @@ +--- +language: +- zh +- en +license: apache-2.0 +--- + +# Introduction +On August 26, Alibaba open-sourced its multimodal MoE model Qwen3.8-Flash-Next. The FlagOS community completed Day-0 synchronized adaptation across multiple chips, having accomplished multi-chip adaptation, precision alignment, and deployment validation based on the FlagOS unified open-source technology stack on 8 AI chips — including T-Head (Pingtouge), NVIDIA, Moore Threads, Ascend, MetaX, Kunlunxin, Hygon, and Iluvatar CoreX. The first batch uniformly provides a BF16-precision version. The multi-chip versions have been open-sourced on the ModelScope and Hugging Face platforms, where developers can directly obtain out-of-the-box solutions for their respective chips. + + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Iluvatar** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + +# Evaluation Results +## Benchmark Result +| Metrics | Qwen3.8-Flash-Next-Nvidia-Origin | Qwen3.8-Flash-Next-Iluvatar-FlagOS | +|--------------|--------------------------------|--------------------------------------| +| GPQA_Diamond | 92.9 | Evaluating | +| MuSR| 78.57 | Evaluating | + + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker version 20.10.12 | +| Operating System | 22.04 LTS | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/flagrelease-qwen3.8-flash-next-iluvatar-tree_none-gems_5.0.0-vllm_0.20.0-plugin_main-cx_none-python_3.12.11-torch_2.10.0_corex.4.5.0-pcp_ixml4.4.0-gpu_iluvatar001-arc_amd64-driver_4.5.0:202608261637 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/Qwen3.8-Flash-Next-BF16-iluvatar-FlagOS --local_dir /data/Qwen3.8-Flash-Next +``` + +### Start the Container +```bash +IMAGE="harbor.baai.ac.cn/flagrelease-public/flagrelease-qwen3.8-flash-next-iluvatar-tree_none-gems_5.0.0-vllm_0.20.0-plugin_main-cx_none-python_3.12.11-torch_2.10.0_corex.4.5.0-pcp_ixml4.4.0-gpu_iluvatar001-arc_amd64-driver_4.5.0:202608261637" +# 对应plugin-FL 分支 ilvita-int8-main +docker run -d --network host \ + --privileged --ipc=host \ + -v /dev:/dev -v /lib/modules:/lib/modules \ + -v /sys:/sys -v /data:/data -v /mnt:/mnt -v /mnt/share/models/:/models \ + -w /workspace \ + --name flagos -it ${IMAGE} +docker exec -it flagos bash +cd /models/running_common +``` +### Start the Server +```bash +export CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15 +export VLLM_PLUGINS=fl +export VLLM_WORKER_MULTIPROC_METHOD=spawn + +# Timeout — Iluvatar Triton compilation takes a long time +export VLLM_ENGINE_ITERATION_TIMEOUT_S=72000 +export VLLM_RPC_TIMEOUT=72000000 +export VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=7200 + +# === Modify the following parameters according to actual conditions. === +MODEL_PATH="/data/Qwen3.8-Flash-Next" +MODEL_NAME="qwen38_flash" +PORT=8030 +TP=8 +MAX_MODEL_LEN=32768 +export CUDA_VISIBLE_DEVICES=$CUDA_VISIBLE_DEVICES && \ +export VLLM_PLUGINS=$VLLM_PLUGINS && \ +export VLLM_WORKER_MULTIPROC_METHOD=$VLLM_WORKER_MULTIPROC_METHOD && \ +export VLLM_ENGINE_ITERATION_TIMEOUT_S=$VLLM_ENGINE_ITERATION_TIMEOUT_S && \ +export VLLM_RPC_TIMEOUT=$VLLM_RPC_TIMEOUT && \ +export VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=$VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS && \ +vllm serve $MODEL_PATH \ + --served-model-name $MODEL_NAME \ + --dtype bfloat16 \ + --tensor-parallel-size $TP \ + --pipeline-parallel-size 2 \ + --max-model-len $MAX_MODEL_LEN \ + --gpu-memory-utilization 0.95 \ + --port $PORT \ + --enforce-eager \ + --trust-remote-code 2>&1 | tee -a $LOG_FILE +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8030/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "qwen38_flash", + "messages": [{"role": "user", "content": "中国的首都是哪里?"}], + "temperature": 0.7, + "max_tokens": 1024 + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of . This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from Qwen/Qwen3.8-Flash-Next and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt + diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Qwen3.8-Flash-Next-BF16-kunlunxin-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Qwen3.8-Flash-Next-BF16-kunlunxin-FlagOS.md new file mode 100644 index 000000000..6da41fd94 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Qwen3.8-Flash-Next-BF16-kunlunxin-FlagOS.md @@ -0,0 +1,166 @@ +--- +language: +- zh +- en +license: apache-2.0 +--- + +# Introduction +On August 26, Alibaba open-sourced its multimodal MoE model Qwen3.8-Flash-Next. The FlagOS community completed Day-0 synchronized adaptation across multiple chips, having accomplished multi-chip adaptation, precision alignment, and deployment validation based on the FlagOS unified open-source technology stack on 8 AI chips — including T-Head (Pingtouge), NVIDIA, Moore Threads, Ascend, MetaX, Kunlunxin, Hygon, and Iluvatar CoreX. The first batch uniformly provides a BF16-precision version. The multi-chip versions have been open-sourced on the ModelScope and Hugging Face platforms, where developers can directly obtain out-of-the-box solutions for their respective chips. + + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Kunlunxin** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | Qwen3.8-Flash-Next-Nvidia-Origin | Qwen3.8-Flash-Next-Kunlunxin-FlagOS | +|--------------|--------------------------------|--------------------------------------| +| GPQA_Diamond | 92.9 | Evaluating | +| MuSR | 78.57 | Evaluating | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker version 28.2.2, build e6534b4 | +| Operating System | 22.04.4 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/qwen3.8-flash-next-kunlunxin001-gems5.0.0-treenone-cx0.13.0-plugin0.2.0-vllm0.20.2-cp310-pt29-xrt50-x64-5.0.21.47:202608262111 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/Qwen3.8-Flash-Next-BF16-kunlunxin-FlagOS --local_dir /data/Qwen3.8-Flash-Next +``` + +### Start the Container +```bash +set -euo pipefail + +docker run -d \ + --name qwen38-flash-next-flagos-p800 \ + --privileged \ + --network host \ + --shm-size 128g \ + --cap-add SYS_PTRACE \ + --ulimit memlock=-1:-1 \ + --ulimit nofile=120000:120000 \ + --ulimit stack=67108864:67108864 \ + -v /data/Qwen3.8-Flash-Next:/models/Qwen3.8-Flash-Next:ro \ + --entrypoint /bin/bash \ + harbor.baai.ac.cn/flagrelease-public/qwen3.8-flash-next-kunlunxin001-gems5.0.0-treenone-cx0.13.0-plugin0.2.0-vllm0.20.2-cp310-pt29-xrt50-x64-5.0.21.47:202608262111 \ + -lc 'sleep infinity' +``` +### Start the Server +```bash +set -euo pipefail + +docker exec -d qwen38-flash-next-flagos-p800 /bin/bash -lc ' + export USE_FLAGGEMS=1 + export VLLM_DISABLE_PYNCCL=1 + export VLLM_FL_FLAGOS_WHITELIST=rms_norm + export VLLM_FL_FLAGGEMS_ATEN_PLAN_CACHE=1 + export VLLM_FL_FLAGGEMS_ATEN_PLAN_CACHE_REPORT=1 + export FLAGGEMS_ATEN_PLAN_CACHE=1 + export FLAGGEMS_ATEN_PLAN_CACHE_SIZE=128 + export QWEN4_QSA_FUSED_COMPRESS=1 + export VLLM_FL_QWEN4_STABLE_TP_ALL_REDUCE=1 + export VLLM_FL_COMMON_SLOT_MAPPING_GRAPH=0 + export VLLM_PLUGINS=fl + export VLLM_WORKER_MULTIPROC_METHOD=spawn + export VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=1800 + export PYTHONUNBUFFERED=1 + exec /root/miniconda/envs/python310_torch29_cuda/bin/vllm serve /models/Qwen3.8-Flash-Next \ + --served-model-name Qwen3.8-Flash-Next \ + --host 0.0.0.0 \ + --port 8000 \ + --tensor-parallel-size 8 \ + --distributed-executor-backend mp \ + --load-format safetensors \ + --dtype bfloat16 \ + --disable-custom-all-reduce \ + --max-model-len 102400 \ + --max-num-seqs 64 \ + --max-num-batched-tokens 2048 \ + --limit-mm-per-prompt '\''{"image":16}'\'' \ + --block-size 128 \ + --gpu-memory-utilization 0.85 \ + --enable-chunked-prefill \ + --no-enable-prefix-caching \ + --mm-encoder-tp-mode data \ + --mm-encoder-attn-backend TORCH_SDPA \ + --enforce-eager >>/var/log/qwen38-flash-next-vllm.log 2>&1 +' +``` + +## Service Invocation +### Invocation Script +```bash +set -euo pipefail + +curl --noproxy '*' -fsS http://127.0.0.1:8000/v1/chat/completions \ + -H 'Content-Type: application/json' \ + -d '{"model":"Qwen3.8-Flash-Next","messages":[{"role":"user","content":"Compute 17 * 19. Return only the integer."}],"temperature":0,"max_tokens":32}' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., “Explain the basics of quantum computing”) +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a “develop once, run anywhere” workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from Qwen/Qwen3.8-Flash-Next and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt + diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Qwen3.8-Flash-Next-BF16-metax-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Qwen3.8-Flash-Next-BF16-metax-FlagOS.md new file mode 100644 index 000000000..e80af532f --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Qwen3.8-Flash-Next-BF16-metax-FlagOS.md @@ -0,0 +1,146 @@ +--- +language: +- zh +- en +license: apache-2.0 +--- + +# Introduction +On August 26, Alibaba open-sourced its multimodal MoE model Qwen3.8-Flash-Next. The FlagOS community completed Day-0 synchronized adaptation across multiple chips, having accomplished multi-chip adaptation, precision alignment, and deployment validation based on the FlagOS unified open-source technology stack on 8 AI chips — including T-Head (Pingtouge), NVIDIA, Moore Threads, Ascend, MetaX, Kunlunxin, Hygon, and Iluvatar CoreX. The first batch uniformly provides a BF16-precision version. The multi-chip versions have been open-sourced on the ModelScope and Hugging Face platforms, where developers can directly obtain out-of-the-box solutions for their respective chips. + + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Metax** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + +# Evaluation Results +## Benchmark Result +| Metrics | Qwen3.8-Flash-Next-Nvidia-Origin | Qwen3.8-Flash-Next-Metax-FlagOS | +|--------------|--------------------------------|--------------------------------------| +| GPQA_Diamond | 92.9 | 91.16 | +| MuSR| 78.57 | 76.58 | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker version 29.3.1 | +| Operating System | 22.04 LTS | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/qwen3.8-flash-next-metax001-gems5.4.0-treenone-cxnone-plugin3.0.0-vllm0.24.0-cp312-pt28-maca37-x64-3.8.1:202608261009 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/Qwen3.8-Flash-Next-BF16-metax-FlagOS --local_dir /data/Qwen3.8-Flash-Next +``` + +### Start the Container +```bash +IMAGE="harbor.baai.ac.cn/flagrelease-public/qwen3.8-flash-next-metax001-gems5.4.0-treenone-cxnone-plugin3.0.0-vllm0.24.0-cp312-pt28-maca37-x64-3.8.1:202608261009" +docker run -d \ + --name flagos \ + --network host \ + --shm-size 64g \ + --device /dev/dri:/dev/dri:rwm \ + --device /dev/mxcd:/dev/mxcd:rwm \ + -v /data:/data \ + ${IMAGE} \ + sleep infinity + +docker exec -it flagos /bin/bash +``` +### Start the Server +```bash +export CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 +export VLLM_WORKER_MULTIPROC_METHOD=spawn +export GEMS_VENDOR=metax +export VLLM_PLUGINS=fl +export VLLM_FL_FLAGOS_BLACKLIST="mm,mm_out,sort,stable_sort,masked_fill,masked_fill_,log_softmax,log_softmax_out,log_softmax_backward,log_softmax_backward_out,pad,constant_pad_nd,copy_,topk" +export VLLM_ENGINE_ITERATION_TIMEOUT_S=7200 +export VLLM_ENGINE_READY_TIMEOUT_S=1800 +export VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=7200 +export VLLM_RINGBUFFER_WARNING_INTERVAL=300 + +vllm serve /data/Qwen3.8-Flash-Next/ \ + --dtype bfloat16 \ + --tensor-parallel-size 8 \ + --max-model-len 32768 \ + --gpu-memory-utilization 0.9 \ + --port 8000 \ + --enforce-eager \ + --trust-remote-code \ + --served-model-name qwen38_flash +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8030/v1/chat/completions \ +-H "Content-Type: application/json" \ +-d '{ + "model": "qwen38_flash", + "messages": [{"role": "user", "content": "中国的首都是哪里?"}], + "temperature": 0.7, + "max_tokens": 1024 +}' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of . This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from Qwen/Qwen3.8-Flash-Next and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt + diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Qwen3.8-Flash-Next-BF16-mthreads-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Qwen3.8-Flash-Next-BF16-mthreads-FlagOS.md new file mode 100644 index 000000000..6e9342bbc --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Qwen3.8-Flash-Next-BF16-mthreads-FlagOS.md @@ -0,0 +1,142 @@ +--- +language: +- zh +- en +license: apache-2.0 +--- + +# Introduction +On August 26, Alibaba open-sourced its multimodal MoE model Qwen3.8-Flash-Next. The FlagOS community completed Day-0 synchronized adaptation across multiple chips, having accomplished multi-chip adaptation, precision alignment, and deployment validation based on the FlagOS unified open-source technology stack on 8 AI chips — including T-Head (Pingtouge), NVIDIA, Moore Threads, Ascend, MetaX, Kunlunxin, Hygon, and Iluvatar CoreX. The first batch uniformly provides a BF16-precision version. The multi-chip versions have been open-sourced on the ModelScope and Hugging Face platforms, where developers can directly obtain out-of-the-box solutions for their respective chips. + + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Mthreads** container image supporting deployment within minutes + +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + +# Evaluation Results +## Benchmark Result + +| Metrics | Qwen3.8-Flash-Next-Nvidia-Origin | Qwen3.8-Flash-Next-Mthreads-FlagOS | +|--------------|--------------------------------|--------------------------------------| +| GPQA_Diamond | 92.9 | 91.66 | +| MuSR | 78.57 | 75.93 | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|-------------------------| +| Docker Version | Docker version 20.10.12 | +| Operating System | 22.04 LTS | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/sglang-qwen3.8-flash-next-mthreads001-gems5.0.2-treenone-cx0.13.0-pluginnone-vllmnone-cp310-pt29-musa43-x64-3.3.5-server:202608262100 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/Qwen3.8-Flash-Next-BF16-mthreads-FlagOS --local_dir /data/Qwen3.8-Flash-Next +``` + +### Start the Container +```bash +docker run -dit \ + --name flagos \ + --privileged \ + --ipc host \ + --network host \ + --shm-size 512g \ + -w /workspace \ + -v /data/:/data/ \ + -v /etc/localtime:/etc/localtime:ro \ + -v /etc/timezone:/etc/timezone:ro \ + --env MTHREADS_VISIBLE_DEVICES=all \ + harbor.baai.ac.cn/flagrelease-public/sglang-qwen3.8-flash-next-mthreads001-gems5.0.2-treenone-cx0.13.0-pluginnone-vllmnone-cp310-pt29-musa43-x64-3.3.5-server:202608262100 \ + sleep infinity +docker exec -it flagos /bin/bash +``` + +### Start the Server +```bash +bash /workspace/official-main-align/start_worker32081_bs16_graph_experimental.sh +``` + +## Service Invocation +### Invocation Script +```bash +curl -X POST http://127.0.0.1:30015/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "Qwen3.8-Flash-Next", + "messages": [ + {"role": "user", "content": "中国的首都是哪里?"} + ], + "temperature": 0.7, + "max_tokens": 500 + }' +``` + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response + +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of . This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. + +## FlagGems +FlagGems is a high-performance, generic operator library implemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutral kernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. + +## FlagTree +FlagTree is an open source, unified compiler for multiple AI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. For upstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. + +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to support the entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. + +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. + +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** +FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: +- **Multi-dimensional Evaluation**: Supports 800+ model evaluations across NLP, CV, Audio, and Multimodal fields, covering 20+ downstream tasks including language understanding and image-text generation. +- **Industry-Grade Use Cases**: Has completed horizontal evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support + +# License +The model weights are derived from Qwen/Qwen3.8-Flash-Next and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt + + diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Qwen3.8-Flash-Next-BF16-nvidia-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Qwen3.8-Flash-Next-BF16-nvidia-FlagOS.md new file mode 100644 index 000000000..d5305d6d2 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Qwen3.8-Flash-Next-BF16-nvidia-FlagOS.md @@ -0,0 +1,134 @@ +--- +language: +- zh +- en +license: apache-2.0 +--- + +# Introduction +On August 26, Alibaba open-sourced its multimodal MoE model Qwen3.8-Flash-Next. The FlagOS community completed Day-0 synchronized adaptation across multiple chips, having accomplished multi-chip adaptation, precision alignment, and deployment validation based on the FlagOS unified open-source technology stack on 8 AI chips — including T-Head (Pingtouge), NVIDIA, Moore Threads, Ascend, MetaX, Kunlunxin, Hygon, and Iluvatar CoreX. The first batch uniformly provides a BF16-precision version. The multi-chip versions have been open-sourced on the ModelScope and Hugging Face platforms, where developers can directly obtain out-of-the-box solutions for their respective chips. + + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Nvidia** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | Qwen3.8-Flash-Next-Nvidia-Origin | Qwen3.8-Flash-Next-Nvidia-FlagOS | +|--------------|--------------------------------|--------------------------------------| +| GPQA_Diamond | 92.9 | 91.3 | +| MuSR | 78.57 | Evaluating | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker version 24.0.0, build 98fdcd7 | +| Operating System | 22.04.4 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/qwen3.8-flash-next-nvidia003-gems5.3.3-tree0.5.0-cxnone-plugin0.3.0-vllm0.24.0-cp312-pt211-cu129-x64-580.126.20:202608262106 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/Qwen3.8-Flash-Next-BF16-nvidia-FlagOS --local_dir /data/Qwen3.8-Flash-Next +``` + +### Start the Container +```bash +set -euo pipefail +IMAGE_URI='harbor.baai.ac.cn/flagrelease-public/qwen3.8-flash-next-nvidia003-gems5.3.3-tree0.5.0-cxnone-plugin0.3.0-vllm0.24.0-cp312-pt211-cu129-x64-580.126.20:202608262106' +MODEL_DIR='/data/Qwen3.8-Flash-Next' +CONTAINER='qwen38_flash_next_flagos_release' +test -d "$MODEL_DIR" +docker run -d --name "$CONTAINER" --network host --shm-size 64g \ + --device /dev/nvidia0 --device /dev/nvidia1 --device /dev/nvidia2 --device /dev/nvidia3 \ + --device /dev/nvidia4 --device /dev/nvidia5 --device /dev/nvidia6 --device /dev/nvidia7 \ + --device /dev/nvidiactl --device /dev/nvidia-uvm --device /dev/nvidia-uvm-tools \ + --device /dev/nvidia-nvlink --device /dev/nvidia-nvswitch0 --device /dev/nvidia-nvswitch1 \ + --device /dev/nvidia-nvswitch2 --device /dev/nvidia-nvswitch3 --device /dev/nvidia-nvswitchctl \ + --device /dev/nvidia-caps/nvidia-cap0 --device /dev/nvidia-caps/nvidia-cap1 --device /dev/nvidia-caps/nvidia-cap2 \ + -v /usr/lib/x86_64-linux-gnu/libcuda.so.580.126.20:/driver/libcuda.so:ro \ + -v /usr/lib/x86_64-linux-gnu/libcuda.so.580.126.20:/driver/libcuda.so.1:ro \ + -v /usr/lib/x86_64-linux-gnu/libnvidia-ml.so.580.126.20:/driver/libnvidia-ml.so:ro \ + -v /usr/lib/x86_64-linux-gnu/libnvidia-ml.so.580.126.20:/driver/libnvidia-ml.so.1:ro \ + -v /usr/lib/x86_64-linux-gnu/libnvidia-ptxjitcompiler.so.580.126.20:/driver/libnvidia-ptxjitcompiler.so:ro \ + -v /usr/lib/x86_64-linux-gnu/libnvidia-ptxjitcompiler.so.580.126.20:/driver/libnvidia-ptxjitcompiler.so.1:ro \ + -v "$MODEL_DIR:/models/Qwen3.8-Flash-Next:ro" \ + -e NVIDIA_VISIBLE_DEVICES=all -e NVIDIA_DRIVER_CAPABILITIES=compute,utility \ + "$IMAGE_URI" -lc 'exec sleep infinity' +``` +### Start the Server +```bash +docker exec -d qwen38_flash_next_flagos_release bash -lc 'export FLAGGEMS_ATEN_PLAN_CACHE=0 VLLM_PLUGINS=fl USE_FLAGGEMS=1 VLLM_FL_PREFER=flagos VLLM_FL_PREFER_ENABLED=true QWEN4_QSA_FUSED_COMPRESS=1 VLLM_WORKER_MULTIPROC_METHOD=spawn VLLM_ENABLE_V1_MULTIPROCESSING=1 VLLM_MOE_USE_DEEP_GEMM=0 VLLM_USE_DEEP_GEMM=0 VLLM_USE_FLASHINFER_MOE_FP8=0 CUDA_DEVICE_ORDER=PCI_BUS_ID OMP_NUM_THREADS=1 OPENBLAS_NUM_THREADS=1 MKL_NUM_THREADS=1; export VLLM_FL_FLAGOS_BLACKLIST=index_put_,index_put,_index_put_impl_,nonzero,copy_,to_copy,index,index_select,conv1d,_conv_depthwise2d,conv2d,pad,constant_pad_nd,mul; export LD_LIBRARY_PATH=/driver:/usr/local/cuda/lib64:${LD_LIBRARY_PATH:-}; exec python3 -m vllm.entrypoints.cli.main serve /models/Qwen3.8-Flash-Next --served-model-name Qwen3.8-Flash-Next --host 0.0.0.0 --port 8001 --tensor-parallel-size 8 --distributed-executor-backend mp --max-model-len 102400 --max-num-seqs 64 --max-num-batched-tokens 32768 --limit-mm-per-prompt '\''{"image":16}'\'' --mm-encoder-tp-mode data --gpu-memory-utilization 0.90 --no-enable-prefix-caching --disable-custom-all-reduce --moe-backend triton --compilation-config '\''{"mode":"NONE","cudagraph_mode":"FULL","cudagraph_capture_sizes":[1,2,4,8,16],"max_cudagraph_capture_size":16,"pass_config":{"fuse_allreduce_rms":false}}'\'' >/tmp/qwen38-flagrelease-server.log 2>&1' +``` + +## Service Invocation +### Invocation Script +```bash +curl --fail --silent --show-error http://127.0.0.1:8001/health >/dev/null +curl --fail --silent --show-error http://127.0.0.1:8001/v1/chat/completions \ + --header 'Content-Type: application/json' \ + --data-raw '{"model":"Qwen3.8-Flash-Next","messages":[{"role":"user","content":"Reply with exactly: OK"}],"temperature":0,"max_tokens":8}' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., “Explain the basics of quantum computing”) +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a “develop once, run anywhere” workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from Qwen/Qwen3.8-Flash-Next and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt + diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Qwen3.8-Flash-Next-BF16-sunrise-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Qwen3.8-Flash-Next-BF16-sunrise-FlagOS.md new file mode 100644 index 000000000..6b1897ae0 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Qwen3.8-Flash-Next-BF16-sunrise-FlagOS.md @@ -0,0 +1,125 @@ +--- +license: apache-2.0 +language: +- zh +- en +--- + +# Introduction + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Sunrise** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + +# Evaluation Results +## Benchmark Result +| Metrics | Qwen3.8-Flash-Next-Nvidia-Origin | Qwen3.8-Flash-Next-Sunrise-FlagOS | +|--------------|--------------------------------|--------------------------------------| +| GPQA_Diamond | 92.9 | Evaluating | +| MuSR | 78.57 | Evaluating | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker version 20.10.12 | +| Operating System | 22.04 LTS | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/qwen3.8-flash-next-sunrise001-gems5.0.2-treenone-cxnone-plugin0.2.3-vllm0.20.2-sglangnone-sglangflnone-cp310-ptpu0.2.3-tang0.25.0-x64-0.25.0:202609030312 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/Qwen3.8-Flash-Next-BF16-sunrise-FlagOS --local_dir /data/Qwen3.8-Flash-Next +``` + +### Start the Container +```bash +docker run -d --privileged --network host --name "flagos" \ + -v /home/flagos/code/models/:/code/models \ + -v /usr/local/tangrt:/usr/local/tangrt \ + -v /usr/local/pccl:/usr/local/pccl \ + "harbor.baai.ac.cn/flagrelease-public/qwen3.8-flash-next-sunrise001-gems5.0.2-treenone-cxnone-plugin0.2.3-vllm0.20.2-sglangnone-sglangflnone-cp310-ptpu0.2.3-tang0.25.0-x64-0.25.0:202609030312" /bin/bash -c "tail -f /dev/null" + +``` +### Start the Server +```bash +export VLLM_FL_FLAGOS_BLACKLIST="linear,mm,mm_out,bmm_out,bmm,addmm,addmm_out,add,sub,copy_,to_copy,_to_copy,mul" + +vllm serve /data/Qwen3.8-Flash-Next --tensor-parallel-size 8 --max-num-batched-tokens 16384 --max-num-seqs 16 --gpu-memory-utilization 0.8 --port 7528 --skip-mm-profiling --compilation-config '{"mode": 0, "cudagraph_mode": "FULL_DECODE_ONLY"}' + + +``` + +## Service Invocation +### Invocation Script +```bash +curl -X POST http://localhost:7528/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "Qwen3.8-Flash-Next", + "messages": [ + {"role": "user", "content": "Hello, who are you?"} + ], + "max_tokens": 50, + "temperature": 0 + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of . This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from Qwen/Qwen3.8-Flash-Next and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Qwen3.8-Flash-Next-BF16-tsingmicro-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Qwen3.8-Flash-Next-BF16-tsingmicro-FlagOS.md new file mode 100644 index 000000000..e08bcbb11 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Qwen3.8-Flash-Next-BF16-tsingmicro-FlagOS.md @@ -0,0 +1,136 @@ +--- +base_model: +- "" +frameworks: +- "" +language: +- zh +- en +license: apache-2.0 +--- + +# Introduction +On August 26, Alibaba open-sourced its multimodal MoE model Qwen3.8-Flash-Next. The FlagOS community completed Day-0 synchronized adaptation across multiple chips, having accomplished multi-chip adaptation, precision alignment, and deployment validation based on the FlagOS unified open-source technology stack on 8 AI chips — including T-Head (Pingtouge), NVIDIA, Moore Threads, Ascend, MetaX, Kunlunxin, Hygon, and Iluvatar CoreX. The first batch uniformly provides a BF16-precision version. The multi-chip versions have been open-sourced on the ModelScope and Hugging Face platforms, where developers can directly obtain out-of-the-box solutions for their respective chips. + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Tsingmicro** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + +# Evaluation Results +## Benchmark Result +| Metrics | Qwen3.8-Flash-Next-Nvidia-Origin | Qwen3.8-Flash-Next-Tsingmicro-FlagOS | +|--------------|--------------------------------|--------------------------------------| +| GPQA_Diamond | 92.9 | Evaluating | +| MuSR | 78.57 | Evaluating | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker version 20.10.12 | +| Operating System | 22.04 LTS | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/qwen38-flash-next-test-tsingmicro001-gems4.2.1-treenone-cx0.1.0-plugin0.0.0-vllm0.20.2-cp310-pt211-raisa0.2927-x64-v0.29277.8:202608271915 + +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/Qwen3.8-Flash-Next-BF16-tsingmicro-FlagOS --local_dir /data/Qwen3.8-Flash-Next +``` + +### Start the Container +```bash +docker run --privileged -dit \ + --name flagos \ + --shm-size=128g \ + --network host \ + --ipc=host \ + -v /sys:/sys \ + -v /dev:/dev \ + -v /lib/modules:/lib/modules \ + -v /mnt/nvme_data:/mnt/nvme_data \ + -v /data:/data \ + harbor.baai.ac.cn/flagrelease-public/qwen38-flash-next-test-tsingmicro001-gems4.2.1-treenone-cx0.1.0-plugin0.0.0-vllm0.20.2-cp310-pt211-raisa0.2927-x64-v0.29277.8:202608271915 \ + /bin/bash +docker exec -it flagos /bin/bash +``` +### Start the Server +```bash +cd /home/secure/260629145101/qwen38-flash-next/ +source qwen_env.sh +bash qwen3.8_flash_next_server.sh +``` + +## Service Invocation +### Invocation Script +```bash +curl -X POST http://localhost:9005/v1/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "/mnt/nvme_data/models/Qwen3.8-Flash-Next/", + "prompt": "中国的首都是哪里?", + "max_tokens": 10, + "temperature": 0.0 +}' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of . This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from Qwen/Qwen3.8-Flash-Next and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt + + diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Qwen3.8-Flash-Next-BF16-zhenwu-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Qwen3.8-Flash-Next-BF16-zhenwu-FlagOS.md new file mode 100644 index 000000000..d29899ce4 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Qwen3.8-Flash-Next-BF16-zhenwu-FlagOS.md @@ -0,0 +1,138 @@ +--- +language: +- zh +- en +license: apache-2.0 +--- + +# Introduction +On August 26, Alibaba open-sourced its multimodal MoE model Qwen3.8-Flash-Next. The FlagOS community completed Day-0 synchronized adaptation across multiple chips, having accomplished multi-chip adaptation, precision alignment, and deployment validation based on the FlagOS unified open-source technology stack on 8 AI chips — including T-Head (Pingtouge), NVIDIA, Moore Threads, Ascend, MetaX, Kunlunxin, Hygon, and Iluvatar CoreX. The first batch uniformly provides a BF16-precision version. The multi-chip versions have been open-sourced on the ModelScope and Hugging Face platforms, where developers can directly obtain out-of-the-box solutions for their respective chips. + + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Zhenwu** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | Qwen3.8-Flash-Next-Nvidia-Origin | Qwen3.8-Flash-Next-Zhenwu-FlagOS | +|--------------|--------------------------------|--------------------------------------| +| GPQA_Diamond | 92.9 | 89.9 | +| MuSR | 78.57 | 78.04 | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker version 28.1.0, build 4d8c241 | +| Operating System | Ubuntu 24.04.2 LTS | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/qwen3.8-flash-ppu001-gems0.0-tree0.6.1-cxnone-plugin0.2.0-vllm0.24.0-cp312-pt210-x64-none:202608271317 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/Qwen3.8-Flash-Next-BF16-zhenwu-FlagOS --local_dir /data/Qwen3.8-Flash-Next +``` +### Start the Container +```bash +sudo docker run --privileged -dit \ + --network=host \ + --device=/dev/infiniband \ + --ipc=host \ + --device=/dev/alixpu_ctl \ + --device=/dev/alixpu \ + --ulimit memlock=-1 \ + --ulimit stack=67108864 \ + --init \ + -v /data:/data \ + -w /mnt/ \ + --name flagos \ + harbor.baai.ac.cn/flagrelease-public/qwen3.8-flash-ppu001-gems0.0-tree0.6.1-cxnone-plugin0.2.0-vllm0.24.0-cp312-pt210-x64-none:202608271317 +docker exec -it flagos /bin/bash +``` +### Start the Server +```bash +export VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=6000 +vllm serve /data/Qwen3.8-Flash-Next/ \ + --port 8000 \ + --trust-remote-code \ + --served-model-name Qwen3.8-flash \ + --tensor-parallel-size 8 \ + --max-model-len 100000 \ + --gpu-memory-utilization 0.85 \ + --compilation-config '{"cudagraph_mode": "FULL_AND_PIECEWISE", "cudagraph_capture_sizes": [1, 2, 4, 8, 16, 24, 32, 40, 48, 56, 64, 72, 80, 88, 96, 104, 112, 120, 128]}' + +``` + +## Service Invocation +### Invocation Script +```bash +curl http://127.0.0.1:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "messages": [{"role": "user", "content": "中国首都是?"}], + "max_tokens": 128 + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., “Explain the basics of quantum computing”) +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a “develop once, run anywhere” workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from Qwen/Qwen3.8-Flash-Next and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt + +