TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks
Frank F. Xu, Yufan Song, Boxuan Li +18 authors
A benchmark platform evaluates AI agents' autonomy in performing professional tasks in a simulated workplace environment, showing that a significant portion of simpler tasks can be automated but complex, long-term tasks remain challenging.