Your Prompts Were Written for the Last Model: What Changed With Opus 5 and GPT-5.6

Anthropic recently deleted more than 80% of Claude Code’s system prompt for its Claude 5 generation of models. Their coding evaluations did not get worse. They did not get better either. Four fifths of the instructions that a frontier lab had built up around its own flagship product turned out to be doing nothing at all. The obvious question is how they were able to cut so much. The more useful question is why the rest of us have not, because the same kind of instruction is still sitting in your CLAUDE.md, in your agent’s system prompt, and in the sentences you type out of habit.

Read More

The Metric Is the Spec: Choosing and Designing Evaluation Metrics

Two models forecast daily demand for a bakery. Model A has the better RMSE. Model B makes the bakery more money, every single week. If that sentence sounds impossible, this post is for you: the evaluation metric you pick is not a scorecard bolted on after training, it is the specification your whole modelling effort ends up satisfying. Kaggle grandmasters internalise this to the point that studying the competition metric is their first act in any competition, and the habit transfers directly to real projects, where the difference between a metric that encodes what the business actually loses and one that merely sounds standard can make or break the use case.

Read More

Several Breakthroughs Away

The closest I came to the future in 2021 was a waiting list I never joined. I was new to the field then, and a colleague told me about a model you could reach through an API that would, he said, answer almost any question you put to it. There was a queue for access; he may have signed up, but I told myself it would take a while and never got around to it. I did skim the announcement post, and I remember the thought that crossed my mind before I filed it away: once the masses find out about this, it is going to be huge. Then I lost interest and went looking at other things. That model was GPT-3, and what stays with me now is not the model but the shrug: I saw a corner of what was coming, labelled it correctly, and still felt nothing move. It would be nearly two years before I understood that the ground had shifted anyway, and that the thing which shifted first was not the machine but what a small number of people had let themselves believe was possible.

Read More