We scan new podcasts and send you the top 5 insights daily.
To set realistic expectations with business stakeholders, describe an AI model's performance by its error rate (e.g., "wrong 15% of the time") rather than its accuracy rate (e.g., "correct 85% of the time"). This reframing highlights the real-world impact of imperfections.
Technical metrics like "accuracy" are often the wrong measure for ML projects and can mismanage expectations. To achieve success, projects must be evaluated using business KPIs like profit, savings, or ROI. This aligns data science with business goals and reveals the true value of imperfect predictions.
Leadership's expectation of perfection from AI systems is a major red flag. Organizations ready for AI treat inevitable errors as data points for learning and tuning. If a leader would view a 95% accuracy rate as a failure without context, the company culture is not yet prepared for AI deployment.
Don't wait for AI to be perfect. The correct strategy is to apply current AI models—which are roughly 60-80% accurate—to business processes where that level of performance is sufficient for a human to then review and bring to 100%. Chasing perfection in-house is a waste of resources given the pace of model improvement.
The benchmark for AI performance shouldn't be perfection, but the existing human alternative. In many contexts, like medical reporting or driving, imperfect AI can still be vastly superior to error-prone humans. The choice is often between a flawed AI and an even more flawed human system, or no system at all.
To evaluate an AI model, first define the business risk. Use precision when a false positive is costly (e.g., approving a faulty part). Use recall when a false negative is costly (e.g., missing a cancer diagnosis). The technical metric must align with the specific cost of being wrong.
Don't aim for a 100% accurate evaluation system. A good system reveals a 'healthy percentage' of incorrect outputs. Getting excited when evals are wrong is key, as each failure is a clear, actionable opportunity to improve your AI agent.
Teams often fall into the trap of optimizing for model accuracy, a metric popularized by academic settings like Kaggle. In business, this is misleading. A highly accurate model might be too passive and miss opportunities. The focus must shift from pure accuracy to real-world business outcomes and ROI.
AI21 Labs' CMO Sharon Argov suggests openly discussing AI's potential for mistakes. This shifts the conversation from the technology's flaws to how an organization can manage the 'cost of error,' turning a negative into a strategic discussion about risk management and trustworthiness.
Leaders championing AI for efficiency often overlook the devastating brand and business impact of the small percentage of interactions where AI fails. The key is not to expect perfection, but to have a robust strategy for managing these inevitable failures.
Both humans and AI make mistakes. Instead of claiming AI is perfect, a more effective argument in regulated fields is that AI makes fewer mistakes and helps humans catch their own errors more quickly. This shifts the focus from perfection to improved safety and efficiency.