gspo
Here are 8 public repositories matching this topic...
Pure-Rust LLM runtime: one binary serves (OpenAI-compatible), runs local agents, and distills models on their own rollouts — on Apple Silicon and NVIDIA. No Python on the hot path.
-
Updated
Aug 20, 2026 - Rust
Reinforcement learning for text generation on MLX (Apple Silicon)
-
Updated
Feb 15, 2026 - Python
Sync vs fully-async agentic RL on verl: multi-turn GRPO, long-tail rollout profiling, staleness ablations — quantifying when async pays off.
-
Updated
Aug 20, 2026 - Python
OPSG-based test refinement for Java: Stable RL approach to generate maintainable, high-quality unit tests with 98.6% compilation success.
-
Updated
Jan 11, 2026 - Jupyter Notebook
Auditable online post-training lab for MiniMind: GRPO, CISPO, Dr.GRPO and GSPO with reproducible Math/Tool RLVR experiments.
-
Updated
Aug 17, 2026 - Python
RL agents across LangChain & LangGraph, 50+ MCP registry tools, PGVector stores, deployment w Dynamo and inference with vLLM & SGLang
-
Updated
Jul 29, 2026 - Python
Improve this page
Add a description, image, and links to the gspo topic page so that developers can more easily learn about it.
Add this topic to your repo
To associate your repository with the gspo topic, visit your repo's landing page and select "manage topics."