Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

Moving data between warehouses is costly and often unnecessary. The industry is shifting to a "zero-copy" model where data is accessed and activated where it lives. This allows companies to tap into existing data investments without redoing integrations, focusing budgets on value creation instead of data movement.

Related Insights

The concept that data is too large and costly to move is an illusion created by legacy pipelines that repeatedly copy entire datasets. Fivetran's CEO asserts that modern change-data-capture techniques make data movement small-scale and inexpensive.

The significant barrier of messy, legacy data is being overcome by AI. Snowflake is developing "agent-driven migrations" that automate the process of moving data from old systems onto modern platforms. This drastically reduces project timelines from multiple years to just a few weeks.

The long-standing trend of centralizing all data into a single warehouse is incompatible with the speed of AI. Large-scale data migrations are too slow. The future architecture will involve AI models operating closer to data sources for faster, decentralized operation.

Denodo's logical approach is significantly faster because it fetches only the specific query results needed for an analysis, rather than physically moving entire datasets into a central repository. This is analogous to getting a single cup of water from a pitcher instead of carrying the entire heavy pitcher, explaining a 75% reduction in integration time.

AI agents make it dramatically easier to extract and migrate data from platforms, reducing vendor lock-in. In response, platforms like Snowflake are embracing open file formats (e.g., Iceberg), shifting the competitive basis from data gravity to superior performance, cost, and features.

The key "no-regret" move for pharma data teams is to abandon serving all use cases from one giant table. Instead, they should structure data into layered products: foundational (transactions), functional (KPIs), fit-for-use-case (decisions), and fit-for-AI (semantic context).

A key differentiator is that Katera's AI agents operate directly on a company's existing data infrastructure (Snowflake, Redshift). Enterprises prefer this model because it avoids the security risks and complexities of sending sensitive data to a third-party platform for processing.

Excel Data's CEO, Rohit Choudhary, contends that the long-held strategy of migrating all data to a central lake or warehouse is too slow for the AI era. The future is decentralized, requiring AI models to be brought to the data where it resides, rather than the other way around.

The idea of consolidating all enterprise data into one place is a fallacy. A more effective approach is to build an integration and semantic layer that creates a virtual, unified view of distributed data, enabling insight without costly and futile migration projects.

The traditional approach of building a central data lake fails because data is often stale by the time migration is complete. The modern solution is a 'zero copy' framework that connects to data where it lives. This eliminates data drift and provides real-time intelligence without endless, costly migrations.

Zero-Copy Architecture Makes Expensive Data Migrations Increasingly Obsolete | RiffOn