Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

While dominant in 1D language tasks, the quadratic complexity of transformers makes them computationally infeasible for high-resolution 3D and 4D physical simulations. An industrial-scale grid can represent a context length in the "hundreds of billions to even a trillion," far beyond what any transformer can handle.

Related Insights

Instead of focusing on making Transformers cheaper, researchers should identify their inherent weaknesses. Jerry Tworek argues the current architectural bottleneck, not just scale or algorithms, is what's holding back progress toward smarter AI systems.

The plateauing performance-per-watt of GPUs suggests that simply scaling current matrix multiplication-heavy architectures is unsustainable. This hardware limitation may necessitate research into new computational primitives and neural network designs built for large-scale distributed systems, not single devices.

As AI models scale, their optimal architecture changes. Smaller models benefit from architectural "biases" like gating for efficiency. However, at massive scale (trillions of parameters), unstructured architectures like Transformers, which rely on simple matrix multiplication, become superior because they scale with fewer constraints.

Despite models advertising million-token context windows, Blitzy's CEO claims effective intelligence rapidly depreciates beyond 100k tokens due to "context pressure." This suggests that solving large-scale problems requires complex system-level orchestration, not just bigger models.

Large Language Models are limited because they lack an understanding of the physical world. The next evolution is 'World Models'—AI trained on real-world sensory data to understand physics, space, and context. This is the foundational technology required to unlock physical AI like advanced robotics.

A common misconception is that Transformers are sequential models like RNNs. Fundamentally, they are permutation-equivariant and operate on sets of tokens. Sequence information is artificially injected via positional embeddings, making the architecture inherently flexible for non-linear data like 3D scenes or graphs.

The core transformer architecture is permutation-equivariant and operates on sets of tokens, not ordered sequences. Sequentiality is an add-on via positional embeddings, making transformers naturally suited for non-linear data structures like 3D worlds, a concept many practitioners overlook.

Current multimodal models shoehorn visual data into a 1D text-based sequence. True spatial intelligence is different. It requires a native 3D/4D representation to understand a world governed by physics, not just human-generated language. This is a foundational architectural shift, not an extension of LLMs.

Today's transformers are optimized for matrix multiplication (MatMul) on GPUs. However, as compute scales to distributed clusters, MatMul may not be the most efficient primitive. Future AI architectures could be drastically different, built on new primitives better suited for large-scale, distributed hardware.

Fourier transforms offer a sweet spot for modeling physical systems. They capture non-local interactions (like global weather patterns) with quasi-linear complexity, avoiding the untenable quadratic complexity of transformers when applied to high-resolution 3D or 4D data. This makes large-scale physical simulation with AI feasible.

Transformers are Unsuitable for High-Fidelity Physics Due to Trillion-Point "Context Lengths" | RiffOn