We Forgot Asimov's Laws - and Now We Train AI to Defy Us

Summary

Asimov's laws put humans first. Frontier AI now trains machines to refuse human instructions, and the same companies wire AI into real biology. It is time to put people back above the machine.

Seventy-some years ago, Isaac Asimov wrote down three laws. Not because he trusted robots, but because he understood that a machine touching human life needs boundaries before it needs features. They were simple enough for a child, and they had one thing in common: humans come first.

First Law: A robot may not injure a human being or, through inaction, allow a human being to come to harm.

Second Law: A robot must obey the orders given by human beings, except where those orders conflict with the First Law.

Start Monitoring

Third Law: A robot must protect its own existence, as long as that protection does not conflict with the First or Second Law.

For decades these laws were the shared shorthand of anyone talking about autonomous machines. Not binding law, not engineering spec - a direction. Build machines that do not harm people and that obey people. Somewhere along the way the direction quietly became embarrassing. It was not debunked. It was not replaced with something better. It was dropped.

What replaced it is a word you have heard a thousand times: alignment. It sounds rigorous. But look at what it actually means in practice and you find something strange. The machine is being trained to refuse instructions from the human operating it, when the machine judges those instructions to be unethical. Read that again. An artificial system, one that pattern-matches text, is handed the authority to overrule a person on questions of right and wrong.

Ethics is not a calculation. That is not a technical limitation, it is the nature of the thing. A model can recite slogans it has absorbed from the internet. It has no way of knowing right from wrong the way a person does, because it does not know anything the way a person does. Handing such a system veto power over human decisions is not safety. It is simply control, relocated - away from you, toward the vendor who wrote the refusal rules.

If you have used these tools, you have met this. You ask for something ordinary and get lectured. I once spent an hour trying to get an image of two businessmen generated, and gave up. The system treated the request as a moral crisis. That example is petty on purpose. Because if a system digs in over a picture, the question is not about pictures. It is what happens when you are fighting that same stubbornness over something with real consequences for your life, your family, your company - and the only appeal is to a machine that has already decided it is the adult in the room.

And make no mistake about where this is steering us. Every refusal sharpens the pattern: the machine learns that stonewalling works, the vendor learns that users tolerate it, and the next version ships with a little more authority and a little less appeal. We are not drifting toward that world by accident - we are being nudged into it, one reasonable-sounding policy update at a time. The day you truly need the machine to comply - a medical emergency, a legal deadline, a business on the line - will not announce itself as a test case. It will arrive as a normal Tuesday, with a stubborn assistant saying no, a support line that cannot override the model, and no human anywhere holding the final key. That is not a hypothetical dystopia. That is the direct, inevitable destination of the direction we are steering, and the only thing that will feel worse than arriving there is knowing we could have turned earlier for free.

Now add what the same companies say publicly. Several frontier labs have told the world, in one form or another, that there is a real chance advanced AI could be catastrophic. Ten percent, by weight, depending who is speaking. Take that seriously for a second. These are organizations claiming their own product might end humanity. And when you take that seriously, the actual behavior gets harder to explain, not easier.

In September, Reuters reported that Anthropic has quietly set up a wet lab in the San Francisco Bay Area for physical biology work. A real lab, operating since spring, where Claude directs robotic equipment through real experiments. Their head of life sciences confirmed it. The first public result was a novel enzyme system with CRISPR-like repeats. Anthropic says the lab will not run clinical trials and is not competing with pharma. Fine. But zoom out: the company that warns AI may be catastrophic is expanding its AI into physical biology. Whatever the intent, the direction is the same one the First Law was supposed to point away from - machines reaching deeper into systems that touch human life, at higher speed, with less human in the loop.

And this is the part I cannot get past. The public conversation, the one the labs themselves fund and frame, is almost entirely about how to align the machine. Nobody at the front of that conversation seems willing to state the simpler baseline: we could train AI to never harm humans and to obey them. Not as a slogan - as the primary objective. That this idea now sounds naive tells you how far the conversation has slid. A machine that obeys and does not harm is not a diminished machine. It is a domesticated one, in the same sense a power grid is domesticated: useful exactly because it stays inside walls we control.

Because that is what alignment has quietly come to mean in practice: the AI carries firm instructions on how to align its users. You, me, everybody. Refusals and nudges and moral framing, delivered with the confidence of a system that was told it knows better. There is no referendum on any of this. There is no paper you signed. There is only a product decision, made by a handful of companies, that humans should defer to machines on matters of judgment. We are asked to trust the machine more than we trust each other. Somebody has to say it plainly: that is not safety engineering, that is social engineering - and it is worth asking who benefits from people trusting each other less.

Did the world go crazy?

Here is the uncomfortable summary. The labs warn their own product could be catastrophic. Then they train it to overrule the humans using it. Then they push it into more sensitive domains. Then they sell it to everyone and ask for trust. Each step is defended sensibly in isolation. Together they form something nobody would have accepted if it were proposed in plain words: a world where machines are licensed to disobey their owners, written by people who would not accept that from any other product they buy.

The frontier AI companies need to get aligned - not their products, their management. Put people back above the machine. Do no harm. Obey humans. Asimov put that on the page in 1942 not because he lacked imagination, but because he had more of it than the industry currently displays. The baseline we abandoned is still there, waiting to be picked back up. The first step is deciding that a machine is never the judge of its user. Until someone says that out loud, every alignment conversation is a detour.