gg2
earendil.com

Measuring the sloppiness of code

Archived — this story has rotated out of today’s deck. It is kept here in full.

The gist

An Earendil engineer proposed metrics showing AI agent code is roughly twice as verbose and eroded.
The post argues correct code can still degrade a codebase, so human taste still.

Background

The post is framed around the question of what happens if coding is "solved" by AI agents, and how to tell whether the resulting code is actually good. The author, an Earendil engineer, tested several sloppiness metrics, including change in lines of code, verbosity (duplicated or unnecessary lines), and erosion (mass concentrated in a few large, complex functions). Comparing code generated in the SlopCodeBench evaluation against established repositories, the agent code scored roughly twice as high on both verbosity and erosion. The piece was shared on Hacker News, Lobsters, Reddit and X, where discussion focused on why correct code can still erode a codebase and why human intuition and taste still matter.

How it unfolded

  1. Sep 9, 2026The blog post "If coding is solved, what now?: Measuring the sloppiness of code" is published on earendil.com.
  2. Sep 11, 2026The post is shared on X by @pidotdev, describing it as a new blog post from Earendil engineer @SebastianBaye on why quantifying code sloppiness is difficult.
  3. Sep 11, 2026The post is submitted to Hacker News and Lobsters, and discussed on r/theprimeagen.

Who’s saying what

Author
Change in lines of code is a surprisingly effective sloppiness metric, with the ironic caveat that optimizing for it would make it meaningless.
Author
Verbosity and erosion separated legacy codebases from LLM-generated code well, with agent code roughly twice as verbose and eroded as human code.
Community
Discussion framed around why correct code can still erode a codebase and why human intuition and taste still matter.

Still unverified

The claim that verbosity and erosion "separated legacy code bases from LLM-slop quite well" rests on the author's own tests and the SlopCodeBench evaluation; the post is a single-source blog account, and the retrieved material gives no independent replication or peer review.

Sources

See today’s stories in the app gg2 — free on the App Store