Official Repository: A Comprehensive Benchmark for Logical Reasoning in MLLMs
-
Updated
Jun 17, 2025 - Python
Official Repository: A Comprehensive Benchmark for Logical Reasoning in MLLMs
Live Deep Research Bench. A challenging, objective benchmark for deep research tasks.
First-Person Agent Memory Bench. 10 Categories including fact recall, multi-hop links, temporal reasoning, fact overwrites, speaker traps, refusal, credibility, and agentic tool usage. 540K token / 60 session corpus, all in first person. Dynamic output-answer-key portion. Comprehensive report with visuals and miss breakdown.
Benchmark for evaluating advanced reasoning, recursive dependency resolution, and robustness capabilities of large language models in dynamic, noisy, and structurally challenging environments.
Non-Western symbolic reasoning benchmark for LLMs. Multi-step rule-following inference within the Ba Zi (八字) formal system. Frozen lookup tables, Python reference implementation, mechanically verified gold CoT cases.
Local-first benchmark for measuring reasoning improvement in small language models through critique and debate.
Public preview of Society of Thought, a Qwen adapter that reasons through visible multi-persona debate, with benchmark evidence, raw traces, and demo.
CandyBench Notebook: visually select OpenAI-compatible relay models, run 10-way reasoning tests, and preserve every raw reply.
To associate your repository with the reasoning-benchmark topic, visit your repo's landing page and select "manage topics."