Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

Astra achieves long-term task persistence not by summarizing its context window, but through a novel mechanism. It maintains a 'long lived notes file' that it can update and has the ability to search its entire session history, effectively managing a much larger context than its token limit implies.

Related Insights

According to Harrison Chase, providing agents with file system access is critical for long-horizon tasks. It serves as a powerful context management tool, allowing the agent to save large tool outputs or conversation histories to files, then retrieve them as needed, effectively bypassing context window limitations.

Instead of relying on lossy LLM-based summarization, architect agent memory into three tiers: an ephemeral scratchpad for immediate tasks, a deterministic state machine for history (e.g., Redis), and a semantic anchor (e.g., vector store) for global knowledge lookup.

For complex coding tasks, Astra introduces an experimental method of keeping "running notes" across multiple context windows. Unlike summarizing, which can lose detail, this approach keeps earlier context searchable, preventing critical information (like why a previous fix failed) from being compressed away and lost.

To manage context effectively, an AI OS can run a nightly routine ('dreaming') that reviews daily memory files, compresses key information, and saves it into a long-term memory file. This process mimics human memory consolidation, preventing context loss over time.

To manage huge context sizes, Lindy uses "recursive context buckets" organized in a self-balancing tree. This data structure allows an AI agent to access information from a context of billions of tokens with just two LLM calls, effectively solving the context window limitation for complex tasks.

Instead of just expanding context windows, the next architectural shift is toward models that learn to manage their own context. Inspired by Recursive Language Models (RLMs), these agents will actively retrieve, transform, and store information in a persistent state, enabling more effective long-horizon reasoning.

Seemingly complex features like long-term memory and skill creation are fundamentally clever systems for managing an AI's limited context window. The "harness" efficiently loads and unloads relevant information (memories, skills) at the precise moment it's needed, rather than keeping it all in context constantly.

To enable long-horizon tasks, Cursor incorporates "self-summarization" directly into its RL loop. The model learns to compact its own history and restart its context window with the summary. This allows it to operate over millions of tokens despite a nominal 200k context limit.

Tasklet completely re-architected its agent, moving from feeding chat history into the LLM to treating the file system as the primary context. The agent now receives hints and pointers to relevant files, enabling it to handle infinitely long histories and larger contexts beyond the token window.

To make agents useful over long periods, Tasklet engineers an "illusion" of infinite memory. Instead of feeding a long chat history, they use advanced context engineering: LLM-based compaction, scoping context for sub-agents, and having the LLM manage its own state in a SQL database to recall relevant information efficiently.