Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

Relying on a single metric to evaluate a complex system, like the Value at Risk (VaR) model in finance, invites manipulation and disaster. Instead, systems should be managed with a 'cockpit'—a dashboard of diverse metrics that provide a holistic view and prevent stakeholders from gaming one number.

Related Insights

A Palantir architect argues that standalone dashboards are a 'field of dreams' that nobody uses. Instead, metrics and KPIs should be byproducts embedded directly within operational applications, informing the user at the point of action and decision.

Standard qualitative risk assessments like 'heat maps' are flawed. A quantitative approach, calculating the dollar value of a risk materializing, allows for objective comparison against the cost of controls and potential business value, a method borrowed from military risk management.

A complex spreadsheet model is often brittle; a single questionable assumption can cause stakeholders to reject the entire analysis. To counter this, models should make key assumptions transparent and easily adjustable, like with a slider, to allow for sensitivity analysis rather than outright dismissal.

Beyond quantitative benchmarks, METR's assessment of AI's catastrophic risk relies heavily on qualitative evidence. This includes watching model transcripts for "derpy" mistakes, observing their inability to use resources well, and relying on the intuition that a new model is only incrementally more capable than the previously non-dangerous one.

Single-factor models (e.g., using only CPI data) are fragile because their inputs can break or become unreliable, as seen during government shutdowns. A robust systematic model must blend multiple data sources and have its internal components compete against each other to generate a reliable signal.

It's tempting to think you can intuit the few factors a decision hinges on. This is often wrong. Complex systems have non-obvious leverage points. The process of building an explicit model reveals which variables have the most impact—a discovery you can't reliably make with intuition alone.

Attempting to make all data from every source perfectly accurate is a recipe for failure. A more effective data strategy is to identify the 100-300 most critical business metrics and invest in making that subset a 'gold standard' single source of truth. This provides reliable intelligence without an impossible scope.

Executive dashboards often present a "watermelon" status: green on the surface due to vanity metrics like velocity, but red underneath when you examine actual business outcomes. This false sense of security hides deep-seated performance issues and punishes those who look deeper.

In emerging markets, where 'six sigma' events happen frequently, statistical risk models like Value at Risk are ineffective. A more robust approach is scenario analysis, stress-testing portfolios against specific historical crises like 1998 or 2008 to understand true vulnerabilities.

The typical reaction to metrics being gamed is to introduce more leading and lagging indicators. However, this is a trap that falls prey to Goodhart's Law. It doesn't solve the underlying issue of goal fixation and instead just creates more numbers for teams to manipulate, further obscuring business reality.