A less accurate LSTM model was deployed over a superior BERT model due to hardware limits (no GPUs), faster retraining needs, and the fact that a minor accuracy boost did not alter the essential human-in-the-loop workflow. Real-world constraints often outweigh marginal performance gains.
The team over-optimized model inference, which accounted for only 0.3% of the total processing time. The real bottleneck was the multi-minute human review step. Optimizing the user interface to save reviewers seconds would have been far more impactful than improving the model's speed.
Error analysis showed nearly half the model's mistakes came from a flawed data schema that forced a single label onto dual-intent text (e.g., a combined complaint and request). The root cause wasn't a modeling failure but a data modeling decision made before any training, highlighting that schema design is critical for accuracy.
The human review process, often seen as a temporary bottleneck, should be viewed as a valuable, continuously-running data labeling pipeline. By systematically capturing operator corrections, the team created a powerful, low-cost feedback loop that steadily improved the model's performance on live data.
