Agents/For/Work

Independent, evidence-based comparisons of AI agents that do real work. Reviewed 16 September 2026.

Best starting point · reviewed 16 September 2026

ChatGPT (agent mode)

by OpenAI · Supervised autonomy · Web, desktop and mobile apps

ChatGPT agent mode is the easiest agent to start with and the most reliable at research-shaped tasks, but reviewers consistently describe it as slow and uneven on workflows that need many precise clicks.

Key facts

Best at
Research, data gathering and one-off browser chores
Works with
Sandboxed browser, uploaded files, connected apps and connectors
Cost model
Included with paid ChatGPT plans; a free tier exists but agent runs are limited
Handles
Deep research · Comparison and data collection · Form filling · Spreadsheet and slide building

What it actually does

Agent mode gives ChatGPT a sandboxed virtual computer with a browser. It can search, open pages, fill forms, read your uploaded files, use connected apps, and assemble the result into a spreadsheet or a set of slides — inside the app you already have open.

Where it is strongest

Gather, compare, summarise, tabulate. Anything that is fundamentally a research problem is where it performs best, and it shows each step, so a run that goes wrong is easy to diagnose and re-steer.

Where it struggles

Logins, paywalls and CAPTCHAs stop it. Unfamiliar interfaces that need a long chain of exact clicks produce inconsistent runs. Published hands-on reviews commonly land around six or seven out of ten for genuine end-to-end work, and slowness is the single most frequent complaint.

How to get good results from it

Use it for information work, not for operating systems on your behalf. Log in yourself before handing over a session, keep each task to one clear objective, and ask for the output format explicitly.

Strengths and weaknesses

Strengths

  • Zero setup — it is a mode inside a product most people already have open.
  • Very strong on research-shaped tasks: gather, compare, summarise, tabulate.
  • It shows its steps, so it is easy to see where a run went wrong.

Weaknesses

  • Reviewers consistently describe runs as slow, and long tasks can stall or time out.
  • It gets stuck on logins, paywalls and anything with a CAPTCHA.
  • Reliability drops on tasks that need many precise clicks in an unfamiliar interface.

What users report

User-reported

Published hands-on reviews and user threads land in a similar place: excellent for research, uneven for anything that has to click through a real workflow end to end. Several reviewers rate it around six or seven out of ten for real work.

Who it suits

  • Anyone trying agents for the first time
  • Research-heavy roles: strategy, marketing, procurement, journalism
  • People who want no new tool, no new subscription and no setup

Who should skip it

  • Anyone needing high-reliability, repeatable operations work
  • Teams who want the agent's output to land inside a shared system of record

Verdict

Pick it if you want to try agents today without changing tools, and most of your tasks are research and information gathering.

Sources used

Affiliate disclosure: the monday.com and n8n links on this site can earn us a commission if you sign up. Every other link goes straight to the vendor and earns us nothing. No vendor pays for placement, ranking or wording here.

Not sure ChatGPT (agent mode) is the right fit?

Send us the task you have in mind and we'll tell you whether this is the agent for it — or which of the other four is.

We reply within two working days. We never sell or share your details.