Fixing GRPO training collapse in long-horizon multi-tool agents. A lightweight PRM-Lite + LATA joint approach achieves +37% over vanilla GRPO on τ-bench airline (50-task, multi-turn).
-
Updated
Jun 27, 2026 - Python
Fixing GRPO training collapse in long-horizon multi-tool agents. A lightweight PRM-Lite + LATA joint approach achieves +37% over vanilla GRPO on τ-bench airline (50-task, multi-turn).
Enterprise-Grade Agent Orchestration Platform: 10 microservices, 5 architectural planes, 812 tests. Chaos-engineered, cost-optimized, security-audited. Built to survive real-world agentic AI production.
ToolWeave: Structured synthesis of complex multi-turn tool-calling dialogues: an end-to-end pipeline that builds tool graphs with real data dependencies, samples and verifies workflows, plans parameter provenance, and generates multi-step tool-calling dialogues for LLM fine-tuning.
To associate your repository with the multi-turn-agents topic, visit your repo's landing page and select "manage topics."