We scan new podcasts and send you the top 5 insights daily.
A player's seasonal statistics can be heavily skewed by outlier events, like scoring multiple goals against a team reduced to nine men. To create predictive models, analysts must identify and censor this "bad data" from irregular game states to get a true picture of a player's ability under normal conditions.
The challenge of modeling a fluid game like soccer is solved by "discretizing" continuous play. Analysts define a series of distinct micro-events (e.g., a 2.5-second pass-and-receive sequence), which turns an overwhelming stream of coordinate data into analyzable, aggregated metrics.
A significant part of a player's value, particularly on defense, comes from actions that prevent scoring opportunities. Analytics can now quantify this by measuring how effectively a player closes passing lanes, essentially calculating the value of a negative outcome that was successfully averted.
To moderate Fernandes' high-risk shots, manager Erik ten Hag presented him with a data board visualizing his success rate from different positions. This data-driven coaching method proved more effective than simple instruction, persuading Fernandes to focus on higher-percentage opportunities without stifling his creativity.
Complex AI models in soccer don't "speak English." Instead of feeding raw data to a coach, a specialized analyst interprets model outputs (e.g., moments with high goal probability) to find corresponding video clips. This translates complex analytics into a familiar medium coaches can act upon.
While talent identification is a goal, a primary function of analytics for clubs is defensive: avoiding catastrophic outcomes. This includes preventing relegation, which has huge financial consequences, and not wasting millions on underperforming players, making it a key risk management tool.
For decades, TV broadcasts have featured stats like possession percentage and yellow cards. However, deeper analysis reveals these common metrics are not predictive of a match's outcome. They serve primarily as entertainment but offer little actual insight for serious analysis.
Contrary to the "more data is better" mantra, scaling with bad data actively degrades model performance. Undeduplicated data makes models "forgetful" and less intelligent over time. You cannot overcome poor data quality simply by adding more compute; better, cleaner data is more effective.
Instead of continuous recording, Metal's software lets gamers save the last 30 seconds *after* an interesting event. This behavior, similar to Tesla's bug reporting, automatically filters the data, creating a massive dataset composed almost entirely of noteworthy, high-skill, or out-of-distribution moments, which is ideal for AI training.
When building revenue models, AI can quickly analyze infinite data slices to spot outliers that skew metrics, such as zero-day service renewals or old opportunities creating survivorship bias. This leads to a more accurate model, representing a performance gain, not just an efficiency one.
Both fields involve making high-stakes decisions based on imperfect, often non-predictive data. The inherent variance in soccer outcomes, from game results to the financial impacts of relegation, mirrors the distributional nature of financial volatility.