Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

In natural datasets, '1' is the leading digit about 30% of the time. Fabricated data often shows an unnaturally even distribution of leading digits, which violates this statistical law and can be a strong signal of fraud for auditors and data scientists.

Related Insights

Corporate financials require maker-checker systems, audit trails, and severe penalties for fraud. Scientific research data often lacks these controls, with no audit trails or meaningful penalties for errors. This disparity suggests we should apply at least as much skepticism to academic papers as to financial reports.

An LLM's core training objective—predicting the next token—makes it sensitive to the raw frequency of words and numbers online. This creates a subtle but profound flaw: it's more likely to output '30' than '29' in a counting task, not because of logic, but because '30' is statistically more common in its training data.

Deceivers hijack our trust in precision by attaching specific numbers (e.g., "13.5% of customers") to their claims. This gives a "patina of rigor and understanding," making us less likely to question the source or validity of the information itself, even if the number is arbitrary.

Stripe's AI model processes payments as a distinct data type, not just text. It analyzes transaction sequences across buyers, cards, devices, and merchants to uncover complex fraud patterns invisible to humans, boosting card testing detection from 59% to 97%.

A fraud operation can be brilliant at exploiting systemic weaknesses while being comically bad at faking basic evidence, like having one person forge dozens of signatures. This paradox is not surprising and reflects a division of labor similar to legitimate businesses, with different skill levels for strategy versus execution.

AI can be a powerful fraud detection tool by comparing a company's public statements against alternative data. For example, it can analyze satellite imagery of shipping traffic or factory activity and flag discrepancies with management's guidance.

The 1863 False Claims Act created a financial incentive to report fraud, but its impact was limited by the difficulty of detection. Modern AI solves this information processing bottleneck, finally allowing companies to act on the law's incentive at a massive scale.

Large-scale fraud operates like a business with a supply chain of specialized services like incorporation agents, mail services, and accountants. While some tools are generic (Excel), graphing the use of shared, specialized infrastructure can quickly unravel entire fraud networks.

A defender's key advantage is their massive dataset of legitimate activity. Machine learning excels by modeling the messy, typo-ridden chaos of real business data. Fraudsters, however sophisticated, cannot perfectly replicate this organic "noise," causing their cleaner, fabricated patterns to stand out as anomalies.

A core conceit of fraud is faking business growth. Consequently, fraudulent enterprises often report growth rates that dwarf even the most successful legitimate companies. For example, the fraudulent 'Feeding Our Future' program claimed a 578% CAGR, more than double Uber's peak growth rate. This makes sorting by growth an effective detection method.

Benford's Law Uncovers Fraud by Spotting Unnaturally Uniform Leading Digits in Data | RiffOn