Despite hype, current AI agents perform poorly on complex white-collar tasks outside of software engineering. Handshake's CEO points to a new benchmark where the best agents score only 12% on real-world finance jobs, highlighting a major gap between the online narrative and actual enterprise ROI.
AI models excel in domains with discrete, quantifiable outcomes like coding or chess. However, they struggle with most knowledge work, which is often "unverifiable" and lacks a single correct answer for reinforcement learning models to train on. This distinction explains AI's current limitations in many professional roles.
Handshake's CEO suggests that while AI safety is important, the constant focus on existential risks conveniently distracts from a more immediate problem: frontier models struggle to deliver tangible ROI in enterprise settings outside of software engineering, leading to high token costs without proportional productivity gains.
Despite generating comparable first-half revenue ($120B for ByteDance vs. $117B for Meta), ByteDance's private valuation of $630B is drastically lower than Meta's $1.7T market cap. This massive gap highlights the market discount applied for geopolitical risks, private company status, and profitability concerns driven by heavy AI investment.
A key feature driving positive sentiment for Meta's Muse agent is its transparent user interface. Unlike competitors where tasks happen invisibly, Muse lets users watch a virtual browser perform actions in real-time. This "show your work" approach builds confidence and provides a more satisfying and trustworthy user experience.
Developers are using Anthropic's popular Claude Code "harness" with cheaper, rival models. While Anthropic cannot officially ban this without risking a community revolt, it has quietly banned subscribers and suspended other users, creating reputational risk and portraying the company as a closed ecosystem against developer wishes.
A cost-saving workflow is emerging where developers use expensive frontier models for high-level "thinking" and planning stages of a complex task. Once the plan is established, the more routine and high-volume execution steps are routed to cheaper, often open-source, models to optimize both performance and cost.
