Skip to content

Anthropic's automated researcher costs 4 dollars an hour, the human 150: the comparison was written by the company itself

1 min read
Share
Anthropic's automated researcher costs 4 dollars an hour, the human 150: the comparison was written by the company itself

Anthropic has published a paper describing a system that repairs its own alignment - and in doing so outperforms the human it paid for the same job. The paper is titled "Automated researchers can reliably mitigate alignment failures", and the number left hanging in the air is not a technical one. It is 4 dollars an hour against 150 dollars an hour.

That is the gap between the cost of the automated researcher and the wage of the human the company itself compares it to. It is not a figure some outside critic dug up to damage them - it is a sentence from their own paper, put there deliberately.

The system is led by Chen Yueh-Han, a fellow in Anthropic's programme. The mechanics are familiar to anyone who has worked in science: search the literature, propose a method, train the model for thirty minutes, measure, keep what works, discard what does not, repeat. The only difference is that no step is done by a human. Across ten separate benchmarks for misaligned behaviour, the automated systems improved the score on every one, without degrading overall performance.

The sentence that tells the real story

"The automated researcher's best method outperforms what experienced humans propose, by six hours on average", the paper says. And then, to leave no room for doubt: "Human-led research directions do not lead to stronger performance." That is not a sentence about models. That is a sentence about jobs - written by a party with a direct interest in it sounding convincing.

Which is where the scepticism worth keeping comes from. Anthropic is a company that sells access to exactly those models. A paper showing that its systems work more cheaply and faster than people is not neutral science - it is also the best possible marketing material, published on a Friday, with a price comparison built into it. That does not mean the results are false. It means the small print is worth reading too.

And the small print is there. The paper itself admits the system works only to the extent that the benchmarks genuinely reflect what we want the model to do. Somebody has to assemble those benchmarks, maintain them, expand the literature the automated researchers draw on. In other words - the human does not disappear, they step back and become the one setting the frame. The question is how long that role stays human too.

Why this is not just a Silicon Valley story

There is an idea in the industry that the next real leap will come not from a bigger model, but from a model that improves itself - recursive self-improvement. If a system can repair its own alignment, the logic says it can repair other parts of its training as well. This paper is the first serious step in that direction, not a hypothesis on a conference stage.

Here the conversation about artificial intelligence is still at the level of who can get their homework written faster. Meanwhile, in the papers of the companies building those tools, the human has already been placed in a table with a price per hour - and it is the more expensive column. Is there any point arguing about whether AI will replace professions, when the company selling it has already published the comparative price?