FortunegeopoliticaBEARISHMEDIUM
Alibaba President: AI agents can talk, but can they actually do the work?
The strongest frontier model we tested successfully completed 61.7% of the tasks — high enough to be useful and low enough to be a warning.
The strongest frontier model we tested successfully completed 61.7% of the tasks — high enough to be useful and low enough to be a warning.