We scan new podcasts and send you the top 5 insights daily.
Use Jev to perform pairwise comparisons on thousands of pull requests, asking "are these related?" to automatically form thematic clusters. A cheap LLM then labels these clusters (e.g., "tech debt," "new features"), providing a fast and accurate overview of engineering efforts for pennies.
Scanning millions of lines of code is infeasible. Mozilla uses a simple LLM to act as a 'judge,' scoring files on criteria like 'likelihood of a bug' and 'accessibility from the web.' This prioritizes where to focus the more expensive and time-consuming agentic analysis.
Don't ask an LLM to perform initial error analysis; it lacks the product context to spot subtle failures. Instead, have a human expert write detailed, freeform notes ("open codes"). Then, leverage an LLM's strength in synthesis to automatically categorize those hundreds of human-written notes into actionable failure themes ("axial codes").
Teams can measure the negative side-effects of AI adoption by tracking specific Git metrics. A drop in commit message length to 20-30 characters or a surge in single-commit PRs with 500+ lines are quantifiable signals that AI is amplifying poor practices and increasing technical debt.
A surprising side effect of using AI at OpenAI is improved code review quality. Engineers now use AI to write pull request summaries, which are consistently more thorough and better at explaining the 'what' and 'why' of a change. This improved context helps human reviewers get up to speed faster.
Jev can audit an entire website for internal linking opportunities by treating it as a large-scale classification problem. For every pair of pages, it answers a simple question: "Does this page have a real reason to link to that one?" This is far cheaper and faster than using a full LLM.
The scale of AI-driven development is staggering. GitHub saw 17 million agent-created pull requests in March alone and projects 14 billion total commits for the year, a 14x increase from the 1 billion in the previous year. This signals a shift to developers working with teams of AI agents.
LLMs can both generate code analysis tools (measuring metrics like cognitive complexity) and then act on those results. This creates a powerful, objective feedback loop where you can instruct an LLM to refactor code specifically to improve a quantifiable metric, then validate the improvement afterward.
Jev processes tasks up to 400x cheaper than LLMs, with costs as low as cents for thousands of complex queries. This economic shift makes it feasible to analyze entire archives (emails, ads) for deep insights, a task previously too expensive or time-consuming.
The primary obstacle to analyzing engineering output was the technical difficulty of synthesizing massive, unstructured data from disparate sources like code repositories, documents, and Slack. It wasn't a cultural issue or lack of tools; it was a data fragmentation problem that AI can now solve.
Use Jev, a fast and cheap decision model, for large-scale data classification and clustering. Then, apply more expensive, powerful LLMs like Astra to these refined datasets for deep analysis. This hybrid approach dramatically reduces costs and unlocks complex data products that were previously cost-prohibitive.