Specification for a reproducible, provenance-bound multi-lane LLM benchmark suite.
-
Updated
Aug 22, 2026 - Python
Specification for a reproducible, provenance-bound multi-lane LLM benchmark suite.
Public benchmark harnesses and reproducible evaluations from PixelSpaceAI
Pre-generation tool-call gating via linear probes on LLM hidden states. F1 ≈ 0.91–0.94 on BFCL v4, 14–22× faster than full generation. Cross-architecture transfer across Llama / Qwen / Phi / Mistral (3B–7B) with ≥96% retention.
What 200 steps of fully simulated multi-turn tool-use RL do to a 4B policy: every 10th checkpoint scored on BFCL v4, with the pipeline that produced the measurement.
A bf16 LoRA fine-tune of Qwen 3.5 4B for function calling on xLAM. v1.0 ships below the BFCL gate with full per-category failure analysis.
Canonical IR, schema validation, and deterministic delivery for function calls.
Open LLM leaderboard featuring Xiaomi MiMo v2.5 & MiMo 100T head-to-head with GPT-5, Claude, Gemini, DeepSeek, Llama 4. ARC-AGI · SWE-Bench · MMLU-Pro · GPQA · HumanEval · BFCL.
Independent audit of a fine-tuned LLM tool-calling PoC — BFCL regression decomposition, inference stack risk assessment, and production recommendation for a FinTech client. Qwen-2.5, LoRA, SGLang, H100.
Paired metamorphic evaluation of tool-calling robustness under realistic user phrasing, built on BFCL.
Task-conditioned tail reliability for tool-using agents under equivalent interfaces
OpenEuroLLM snapshot of the Berkeley Function Calling Leaderboard evaluation harness and OLMo evaluation orchestration.
A compact, model-first notation for LLM tool definitions. Re-encodes JSON Schema with a median ~30% input-token reduction and behavior-preserving fallback. Includes the specification, a reference converter, a deployment protocol, and the full evaluation.
QLoRA fine-tune of Qwen2.5-1.5B for tool calling on one 8 GiB GPU: +8.5% on BFCL call categories, -55 points on the ability to decline. Refusal data recovers ~49 of them, reproduced across 3 seeds.
Tool-call reliability fine-tuning lab with open-weight model training, benchmark evaluation, and serving notes.
To associate your repository with the bfcl topic, visit your repo's landing page and select "manage topics."