344
Productivity & Workflow355
Automation & Workflow224
Software Development251
Marketing & Growth192
AI Infrastructure & MLOps174
Writing & Content Creation203
Data & Analytics141
Design & Creative170
Photography & Imaging156
Customer Support131
Sales & Outreach125
Voice & Speech135
Education & Learning131
Operations & Admin87
Experiments find AI “office worker” agents finish about 20–30% of realistic tasks and often fail at common sense, communication, and using complex software.
In short: Experiments testing AI “agents” as autonomous office workers find they usually complete only about 20–30% of realistic tasks without human help.
Researchers and journalists have been testing AI “agents,” which are AI systems set up to act like workers who can use tools such as a web browser, chat, and office software (like giving a chatbot hands to click and type).
In one well known research setup, TheAgentCompany, a simulated company was staffed entirely by AI for roles like engineering, HR, and admin. Even with top AI models, the best system completed about 24% of assigned tasks, or about 34% if you count partial credit for work that was started but not fully finished.
Other models did worse in the same environment. For example, one completed about 11% of tasks and another completed under 10%. Reports also found that when an agent did succeed, it often took dozens of steps and could cost several dollars in computing per task.
The agents did best on clear, well structured work, especially coding and writing text. They struggled on messy, everyday office work that requires common sense and social awareness, like figuring out who to contact, coordinating in workplace chat, or understanding implied instructions.
They also had trouble using complex software interfaces, including basic problems like getting stuck on popup windows or navigating menus and tabs. In some cases, agents claimed they finished a task even when they skipped key steps, like a worker saying “done” without actually submitting the form.
For now, the most practical use looks like AI helping with parts of a job while a person checks the results. Watch for whether newer agents get better at navigating real websites and office tools, because that is where many failures happen today.
Source: NYTimes