Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

Instead of relying on predefined tools passed in a system prompt, models should be given a virtual environment to write and execute code directly. This approach offers more freedom and power, as models can use loops and conditionals, moving beyond simple chained tool calls to perform complex tasks.

Related Insights

The new paradigm for building powerful tools is to design them for AI models. Instead of complex GUIs, developers should create simple, well-documented command-line interfaces (CLIs). Agents can easily understand and chain these CLIs together, exponentially increasing their capabilities far more effectively than trying to navigate a human-centric UI.

A practical hack to improve AI agent reliability is to avoid built-in tool-calling functions. LLMs have more training data on writing code than on specific tool-use APIs. Prompting the agent to write and execute the code that calls a tool leverages its core strength and produces better outcomes.

Instead of placing agents inside a pre-set environment, a more powerful approach for reasoning models is to start with just the agent. Then, give it the tools and skills to boot its own development stack as needed, granting it more autonomy and control over its workspace.

The power of tools like Claude Code comes from giving the AI access to fundamental command-line tools (e.g., `bash`, `grep`). This allows the AI to compose novel solutions and lets product teams define new features using simple English prompts rather than hard-coded logic.

Poolside views model building as an industrialized, end-to-end process, optimizing for the speed from a researcher's idea to a trusted experimental result. This engineering-first approach uses thousands of components to streamline everything from data pipelines to training and reinforcement learning, treating models as artifacts of the process.

Instead of giving an LLM hundreds of specific tools, a more scalable "cyborg" approach is to provide one tool: a sandboxed code execution environment. The LLM writes code against a company's SDK, which is more context-efficient, faster, and more flexible than multiple API round-trips.

Once a universal code execution environment becomes the standard 'super tool' for AI agents, creating new capabilities will no longer require custom code. Instead, 'building a tool' will mean writing a detailed prompt that instructs the LLM on how to sequence actions using an already-exposed, comprehensive API SDK.

The next step for agentic AI is a 'cyborg' model. Instead of juggling numerous pre-defined tools, the LLM will have one primary tool: a code execution environment. It will write code against a company's SDK to perform tasks, which is more flexible, faster, and context-efficient than traditional tool calling.

"Code Mode" is not an alternative to MCP but a more efficient way to use it. Instead of multiple sequential tool calls, the model generates a single script that executes multiple actions in a sandbox. MCP still provides the core benefits of authentication, discoverability, and a standardized, LLM-friendly API.

The true capability of AI agents comes not just from the language model, but from having a full computing environment at their disposal. Vercel's internal data agent, D0, succeeds because it can write and run Python code, query Snowflake, and search the web within a sandbox environment.