Skip to content

Method

The carpenter's question

The question "which AI tool is best" has no answer until you finish the sentence, and the guide's own coverage data shows what the unfinished half costs.

The Agentarius Review Desk9 min read

A rack of twenty-six upright tool shapes on a dark ground. One at the left is filled solid brass and the other twenty-five are drawn as thin outlines of varying heights.
One generalist, twenty-five specialists

Nobody would ask a carpenter which tool is best for building a house, yet everybody asks IT which AI model is best for a company. The carpenter has a toolbox. "Which tool" only has an answer once you say "for what."

The question gets asked that way because it sounds like a procurement question, and procurement questions are supposed to have single answers. One vendor, one contract, one thing to roll out. But the sentence is unfinished, and finishing it changes the answer every time. Best for drafting a contract is not best for cutting a video. Best for a founder writing their first pitch is not best for an engineer refactoring a service that forty people depend on.

The failure is quiet

A hammer used as a screwdriver fails loudly. You can see it. The screw is mangled, the wood is split, and nobody has to be persuaded that something went wrong.

A general-purpose assistant does not fail that way. Point it at almost any knowledge task and it produces something reasonable: a draft that reads well, a summary that covers the material, code that runs. The output is acceptable. That is precisely the problem, because acceptable output ends the search. Nobody goes looking for a better tool for a job that appears to be done.

Its breadth can also hide the advantages of specialist tools.

The guide's own note on ChatGPT for people getting started

That sentence is not ours as an opinion. It sits in the catalogue as a boundary statement on the tool itself, written by the same review method that produced every other line about it. The risk it names is not weakness. It is plausibility.

What the catalogue actually shows

Every tool in the guide is assigned to the contexts it belongs in: the roles it serves and the categories it competes in. A tool with one assignment does one job. A tool with many is claimed to work across many.

In the Agentarius catalogue, ChatGPT is assigned to 26 distinct contexts. Across all 129 published tools the average is 4.1. Claude sits at 26 as well, exactly level, and no other tool comes closer than 16.

A bar chart of how many contexts each tool is assigned to. Most tools cluster between one and five, the tallest bar being forty-seven tools at three contexts. A dashed line marks the mean at 4.1. After sixteen there is a long empty gap, then a single brass bar of two tools at twenty-six.
Agentarius catalogue, 129 published tools, measured 4 August 2026

Be clear about what that number is and is not. It is a fact about this guide's own coverage judgement, not an external measurement of capability. It says where the review method decided a tool deserved a considered entry, and a different guide with different editors would draw different boundaries. It is checkable, though, which is the point: the 129 tool pages are public, and the assignments are on them.

Reporting Claude's 26 alongside ChatGPT's matters for the same reason. Leaving it out would let the piece read as a case against one product, and it is not. Two generalists sit at the same width, and both carry the same trap. The argument is about breadth itself.

The other half of that data is the part worth dwelling on. Each of those 535 context cells across the catalogue carries a documented boundary, a plain statement of when the tool is the wrong choice. Not one of them is empty. Here is what ChatGPT's own entries say, in the guide's words:

  • On writing and research: it "is not a substitute for verified sources, original reporting or expert review," and can "misread evidence, invent details or present uncertain claims too confidently."
  • On presentations and documents: "detailed layout, advanced charts, brand control and final visual polish still require specialist tools."
  • On acting as a general productivity assistant: "a specialised tool wins where a task is deep rather than broad."

These are not criticisms bolted on for balance. They are the boundaries the tool has, recorded next to the things it does well.

The toolbox

Here is what sits in the rest of the box. Each group leads with the reason a specialist wins there, because the reason is the transferable part. The tools are examples of it.

Search and knowledge retrieval: grounding

A generalist answers from whatever it can reach, and its confidence does not change when it cannot reach much. A retrieval tool inverts that. It answers from a defined body of material and shows you which part it used.

Perplexity is designed as an answer engine: a direct, current answer with numbered citations to the sources behind it. Gemini Notebook goes the other way and closes the world deliberately, answering strictly from documents you upload, cited to the passage. Glean applies the same idea inside a company, indexing the organisation's own applications and returning answers that respect each person's permissions. Elicit is built for academic literature, searching a very large corpus of papers and extracting findings into structured tables you define.

Writing and research: the last mile of language

Drafting is the easy part. What is hard is the final pass, where the standard is not "reads well" but "is correct in this language, in this register, for this reader."

DeepL is built around translation, working across whole documents rather than pasted fragments. Grammarly is designed to live where you already type, in the browser, the document and the mail client, catching errors in place rather than in a separate window you have to remember to open.

Coding: the repository is the context

Code is not judged as prose. It has to compile, hold together across files, and not break something written two years ago by somebody who has left. That means the codebase itself is the context, and tools built around it start from a different place.

Cursor is an editor rebuilt around models, where an indexed repository grounds the answers and agentic edits span multiple files. Claude Code works from the terminal, reading a repository, planning changes across it, and running commands and tests as it goes. GitHub Copilot sits across editors and wires into the issue, branch and pull request flow teams already use.

Presentations and documents: layout is a constraint system

A deck is not prose with pictures. It is a constraint problem, where hierarchy, spacing and brand rules interact, and where the answer has to survive being projected in a room.

Gamma generates decks, documents and simple pages from a prompt or outline, arriving already structured and styled. Beautiful.ai holds design rules as content changes, so slides adjust their own layout rather than drifting each time somebody adds a bullet.

Marketing content: the target is live

Writing for search is not writing into a blank page. The target is a set of pages currently ranking, and it moves.

Surfer SEO scores a draft in real time against the pages winning a given search. Clearscope grades content against the same competitive set, with the vocabulary a topic is expected to cover, which is also how a brief gets made consistent across several writers.

Meeting notes: the input is the room

The constraint here is capture, not composition. If the record is wrong, everything downstream is wrong, and no amount of language ability fixes a conversation nobody recorded.

Granola captures from the computer's own audio, so no bot appears in the call, and enhances the notes you were typing anyway. Otter.ai produces speaker-labelled transcripts live, in calls or in the room. Fathom is built around a free tier that stays useful for individuals rather than expiring into a demo.

Automation: it has to run unattended

The hard part of automation is not the happy path. It is the third-party API that times out at 2am, the malformed record, and the question of who finds out when a run fails.

Zapier is built for breadth of integration, connecting the long tail of applications that nothing else supports. Make treats branching, loops and error handling as first-class parts of a visual canvas. n8n is open source and self-hostable, which is the answer when data residency, not convenience, is the deciding constraint.

Creative work: the output has a downstream life

An image or a cut is rarely finished when it is generated. It goes into a layout, a brand system or an edit, and how it arrives determines how much work follows.

Midjourney is built for artistic and cinematic direction, with consistency of style and character across a set. Recraft generates native vector output that stays editable, which is what design work actually ships. Runway offers genuine camera direction, from dollies to rack focus, plus editing of existing footage. ElevenLabs is built for speech, including dubbing into other languages. Descript made the transcript the timeline, so cutting a sentence cuts the video.

When the generalist is the right answer

The argument above is not that breadth is a defect. It is that breadth is a specification, and it is the correct specification more often than a specialist-first reading of this piece would suggest.

Reach for the generalist when the shape of the problem is not yet known. Exploration is the clearest case: you do not know what you are looking for, the question changes as you learn, and switching tools every time it moves would cost more than the imprecision does. A specialist assumes the problem is already framed. When it is not, that assumption is wrong before you start.

First drafts are the second case. The value of a first draft is that it exists and can be reacted to, and a generalist gets you there across any subject without a decision about tooling first. The same holds for one-off work. A task you will do once, that nobody depends on, does not repay the cost of learning a dedicated tool, and treating every job as worth the optimal instrument is its own kind of waste.

Then there is the work that spans domains inside a single thought: reading a contract, drafting a reply, checking a calculation and summarising it for somebody else. Handing that between four specialists costs more in context than it gains in depth. The generalist holds the thread, and holding the thread is the job.

And there is a real cost on the other side, which the specialist case tends to skip. Every additional tool is another subscription, another login, another thing to learn and to govern, and another place where company data goes. A dozen specialists nobody has adopted is worse than one generalist everybody actually uses. Breadth has a genuine advantage in adoption, and adoption is what decides whether any of this produces value.

The round hole

The square peg goes into the round hole. That is the part people forget about the phrase. With enough force it fits, it holds, and from a distance the job looks done.

Two shapes side by side on a dark ground. On the left a brass square is forced inside a thin outlined circle, its corners crossing the outline. On the right a brass circle sits exactly within a circle of the same size.
The cost is not the failure. It is the comparison never made

The cost is not that it fails. The cost is that you never learn the round hole existed, and never find out what the work looks like when the tool is shaped for it. That is not a hypothetical failure mode. Parts of this guide were built by people who used the wrong tool for a long time, produced acceptable results with it, and only discovered what they had been missing when something forced a comparison. The output was never bad enough to prompt the question.

So the question is worth asking properly. Not "which AI tool is best," which cannot be answered, but the carpenter's version: what am I building, how often, and what does the tool made for this do that the one I already have does not?

The tools guide exists to answer the second half of that sentence, and the comparisons exist for the moment two tools both look reasonable. Every tool page states what the tool is for, who it suits, and where it stops, with the date its facts were last checked.

KEEP READING

The AI tools guide

Every tool in the guide carries what it is for, who it suits, what it costs and the date its facts were last checked.

Published 4 August 2026