Skip to content
Gbolagade Ishola
Project

Local Agent Panel and Framework Benchmark

A three-agent judgement panel and a tool-calling benchmark, both running on a small open-weight model on a laptop.

The problem

Agent frameworks are easy to adopt and hard to compare. I wanted to know what AgentScope and Qwen-Agent add over a plain API call, and whether a small open-weight model on a laptop can run a multi-agent workflow at all.

What I built

A panel of three agents, a Researcher, a Critic, and a Chair, that review a piece of work. Each panellist records a structured verdict before seeing anyone else's, then they exchange views once, and the Chair gives a final verdict that names any dissent it overruled. Every message is written to a log. Alongside the panel is a benchmark that runs the same tool-calling prompts through a raw API call, AgentScope's ReAct agent, and Qwen-Agent, all on the same local Qwen model served by Ollama.

My approach

Verdicts are recorded before the exchange because a panellist who sees another view first tends to agree with it, and the log then holds one opinion written three times. Both frameworks also drive an MCP server, and any write through it waits for a person to confirm.

The result

The frameworks matched the raw API call on accuracy and added some latency to each call. The repository has the numbers, the failures, and my notes on what I would use each framework for.

Stack

PythonAgentScopeQwen-AgentOllamaMCP

Connect

Building something, or thinking about it? Book a 1:1 and let's dig in.

Gbolagade Ishola