Blog
Notes on building the benchmark, what the runs keep showing, and where agentic Apple development actually breaks.
-
Open the harness, keep the answers
Closing the whole benchmark defends against one threat and ignores the other. Here is the split, and why rotation matters more than secrecy.
-
Why AppleBench exists
Coding benchmarks measure whether a model can write a diff. Apple development is mostly the part after the diff, and nothing was measuring that.