Since AI can perform well-specified implementation tasks that were once typical for interns, Anthropic now assigns them novel, ambiguous problems. Interns work on challenges no one has solved before, such as developing new evaluation metrics for model performance.
There's a clear division of labor for writing at Anthropic. AI handles rote summarization and data readouts. But for documents requiring original thought or vision—like strategic essays—using AI is culturally discouraged, as the act of writing is inseparable from the act of thinking.
Contrary to past best practices, providing explicit examples within tool descriptions or system prompts can now degrade performance in advanced models. These models are imaginative enough to understand intention without being constrained by specific examples, which can negatively bias their output.
At Anthropic, AI models like Claude handle most technical onboarding questions. This shifts the role of human onboarding buddies to focus on social and cultural aspects, like team dynamics and getting buy-in, rather than code-level guidance.
Engineers at Anthropic are culturally encouraged to operate at a higher abstraction level. Their job is not just to build a product, but to create the automated systems and harnesses that allow an AI to build the product, even if it takes investment time.
AI agents performing tasks via a computer's UI are fragile because each action alters a state that is difficult to undo. Unlike code where state is fully controlled and reversible (e.g., git undo), a wrong click on a website creates a new, complex problem to solve.
As AI handles implementation, the new limiting factor is the user's "taste" and deep domain knowledge. To get great subjective output, like design, the user must provide high-quality references and have the expertise to judge the AI's work, knowing when to push for better results.
To get better results on complex tasks, tell the AI it has permission to use more resources, like sub-agents or workflows. Models are often tuned for speed for the average user. Explicitly stating "this is a hard problem, feel free to use workflows" overrides this default behavior.
Contrary to intuition, as AI models become more capable, the tooling (or "harness") around them must become more sophisticated to manage complex, long-running tasks. Features like "auto mode" are complex software built to leverage the model's advanced abilities.
Technical users gain a significant advantage by reframing knowledge work—like accounting or video editing—as a series of code-based tasks. They then use AI agents with tools like Python scripts or FFmpeg to automate these tasks, moving beyond traditional software like Excel.
Expert prompting is less about the final text command and more about architecting the entire context available to the AI. This includes building the right tooling ("harness"), defining custom "skills," and providing relevant data. Small prompts can seem magical only because of this extensive prep work.
With AI dramatically increasing code velocity, maintenance shifts from stylistic debates to robust verification. Anthropic's approach is to build ~100x more testing infrastructure than is typical, including fixtures from production data and recordings of the AI using the feature it just built.
