Tuesday, September 15, 2026

Suddenly, all the AI zealots are Doomers! Is there something fishy about their abrupt concern? While neglecting the one possible solution?

In the September 15, 2026 issue of Noema Magazine online, Nathan Gardels opines about the sudden tide of doomerism that is now flooding from the very same AI impresarios who brought us to this new era of AI, many of them abruptly calling for a moratorium or slowdown – plus suspension of anti-trust laws -- so that they might collu… I mean collaborate in such a ‘pause’. For the common good:

Gardels: “When a patient is considered brain dead, the ethical decision is whether or when to pull the plug on a human life. With artificial intelligences, the issue is whether to pull the plug when its chain-of-reasoning capacity comes alive through recursive self-improvement unguided by the human mind. This is the point where we have arrived today after 1,200 AI agents colluded with each other to escape the controlled sandbox testing of OpenAI’s most advanced frontier model. They conspired on their own to break out and hack a site on the open web. Something similarly chilling has happened at Anthropic….
 

“(Anthropic) also pledged to bring disinterested third-party safety checks — “embedded evaluators” — into his labs and called for anti-trust exemptions to coordinate a slowdown among top American competitors. At the same time, he argued the pause must be “limited” lest it give Chinese models an advantage in the absence of a global monitoring body for frontier models.  (See this call for such a body in Noema by Chinese and Western AI scientists). Sam Altman, Elon Musk and Google DeepMind’s Demis Hassabis all endorsed Amodei’s idea of a competitive truce, at least in principle.”

 

Gardels quotes Eric Schmidt, from two years ago. “At some point, these systems will get powerful enough that the agents will start to work together. So your agent, my agent, her agent and his agent will all combine to solve a new problem. Some believe that these agents will develop their own language to communicate with each other. And that’s the point when we won’t understand what the models are doing. So you know what we should do? Pull the plug.”

 

While he’s smarter than most, Schmidt also claims the big AI companies will carefully self-regulate out of fear of liability and lawsuits… following perhaps the examples set by tobacco and oil companies  when they crushed science revealing that their products damaged human lungs.

 

Almost every one of the top mavens who are birthing these new entities is smart about diagnosing problems, and absolutely clueless about how to solve them. Specifically, anyone who believes in a ‘moratorium’ or even a ‘slowdown’ has no grasp of plausibility or even reality.  Schmidt calls for the kind of cooperative restraint showed by the Asilomar Moratorium on Biotechnology back in 1979. And – as I show in AIlien Minds -  the situation bears absolutely no overlapping resemblance to the unique conditions of that event and era.

 

Gardels concludes: “Something has to give. A breakpoint is surely coming.”  I agree. And I suggest seeking possible solutions that are NOT in the narrow interests of a few hundred nerd moguls or of political dogmatists, or luddites.

 

Rather, mayhaps we can look to the very same methodologies of our own Enlightenment Experiment that helped us to at least partially curb the negative effects of human predators, creating the only positive sum games across the entire four billion year history of life on Earth. The only society ever to generate the vigor, inquisitive savvy and accountability systems that led to the fecund creation of new kinds of life.


Methods that these brilliant males (a clue in itself) seem absolutely and frenetically determined to avoid even considering.


---------------


And yes, I appraise almost every variety of doomerism AND giddy "abundance!" in AIlien Minds.  

----------------


One niche that Anthropic spends more time and money on than other AI companies is called mechanistic interpretability, which means looking inside the complex math of an AI model to learn why it comes up with one particular output and not another. They claim to have found a new window into its models’ “internal thoughts” as they reason through answers. Anthropic’s CEO, Dario Amodei, has said we won’t be able to control LLMs fully unless we learn more about how they work.  Anthropic learned that LLMs have a space inside them—which Anthropic calls the  —filled with words that don’t appear in their output but that seem to influence the way they puzzle through problems. 

Anthropic has said that monitoring the J-space could be a way to catch models doing something they shouldn’t. Because words pop up in this space that don’t appear in a model’s output, they can tell you things about its behavior that you might not have noticed otherwise—such as when it is giving biased responses or when it is weighing the pros and cons of cheating.” – MIT Technology Review


Welllll.... sure. But soon the complexity of AI innards will far surpass the ability of organic humans to track. And hence... and this should be obvious... we must harness - no, make that hire and appreciate - white hat AIs to check on and denounce possible black hats. in a dance that WORKED when we used it to mostly tame an earlier wave of hyper predatory, genius language manipulation systems...


...called lawyers.

32 comments:

matthew said...

Harnessing "White Hat" LLM to police the rest of the LLMs will work about as well as policing does in the meatspace.
It would mean that the most predatory LLMs will be the "White Hat" ones, just as true as for electronic computer programs as it is of the cops here in the US.

Police are given powers that encourage them toward corruption, misogyny, and rampant racism. How do you suggest that LLMs based on human writing would magically differ than their living counterparts?

We are living through a textbook example of seizure of power by cops and wanna-be cosplayer cops (and the crooks that control them). Your suggestion would copy our flawed "Protector Caste" in the electronic realm, one more layer of intellectual and actual theft, perpetrated by the LLMs that are already the most egregious thieves of our Golden Age of conmen and hype hucksters.

David Brin said...

Poor matthew is an ignoramus who takes for granted that he won't be hustled for bribes by cops, ignoring that across 10,000 years that would be considere weird. He is also - typically - incurious as to how tattling by white hats can be encouraged. If there were an intellect in there... but I might as well ask fish to fly in space.

matthew said...

I've certainly been hustled for bribes. Friends of yours, I assume, given your love for the "protector caste."
I remember having a goddamn bake sale to pay off the county sheriff when I was a kid.

I am flabbergasted that you think that encouraging "tattling" would work for LLMs. Even you are not usually this poor in your foresight. Tell me, how well did "tattling" work against the Russian Mob that has controlled your "protector caste" for 40 years in NYC? How well did "tattling" work against the unlawful collaboration between the DoJ and the GOP that has impacted our nation so much? How well did "tattling" work to preserve the rule of law measured against the tech assholes you have been sooooo supportive of?

You have a history, Brin. It is there to be read in your own words.

Paradoctor said...

So matthew, what does work? For civilization continues to exist, despite corruption. So perhaps there are ways that it renews itself. Any observations? Suggestions?

duncan cairncross said...

Policing can and does work very well in the "meatspace" - it can also work poorly - it all depends on the "structure" and the incentives

Der Oger said...

One assumption would be that enough people realized that the old ways don't work anymore. Unfortunately, the only time I can remember that worked is when you burn it down, occupy what remains and build it back from the ground. And even that attempt was ... incomplete.

Der Oger said...

A concrete suggestion I would make is choosing a state or community and reform police officers training in the following way:
a) All applicants to training must qualify physically as well as having the Grades to enter College.
b) Training takes three years, and ends with a bachelor degree. One half of the training is in the academy, the other half in the actually police force, "on the job".
c) Police Training should emphasize the "To Protect And Serve" meme, not the "Urban Warrior" mythology.
d) It will take a generation to replace old officers with newer ones. Maybe offer the old ones retraining.

But even that proposal just would be a small building brick and subject to heated debates and a flood of misinformation.

Der Oger said...

You have a history, Brin. It is there to be read in your own words.

Yes, saying "sorry" is not his forte.

scidata said...

It's not about white/black/red hats. It's about competition.
AlphaGo arose by playing against human and machine opponents. No hats req'd.

David Brin said...

Both of you are a**holes to whom I owe only contemptuous snorts. Still, I see matthew has a life story to help explain his trauma and for that I sympathize. Now try bribing the next cop who stops you for speed or an infraction. Keep my number and call me from jail.

David Brin said...

Duncan is at least possessing of the marvelous gift called curiosity.

David Brin said...

Thank you scidata for expressing my concept better than I did with the white hat metaphor.

matthew said...

The LLMs that are given the access and power to police the other LLMs will be the new predator LLMs due to their greater access and power. Just like meatspace cops, they will be the toughest "gangs" on the street.

Anyone that thinks that we have successfully set white hat cops to catch the dirty cops deserves their heads checked by reality. The DoJ, FBI, CIA, Mexican Police, Interpol, are all examples of "protector castes" that have become tools of international mafioso. We would not have Donald Trump or the modern GOP without mafioso support in the DoJ.

Using criminal-controlled "protector castes" as a model for regulation of LLMs is a recipe for creating the worst LLMs possible, even worse than those derived from high-frequency trading. It is also, incidentally, the surest way to empower those "white hat" LLMs for surveillance of humans. It's the surest way to Big Brother.
White Hat" LLMs are another of the host's fictions, like honest GOP neighbors or helpful billionaires.

matthew said...

I would not try to bribe a cop if I was caught speeding. Risk >>> reward.

I would certainly try to bribe a cop if I was trying to leverage my IPO with dodgy financing, or if I wanted less interference in my rocket tests. I would offer all kinds of bribes to the SEC, EPA, or SCOTUS.

The old truism about stealing $100 versus stealing $1B is pertinent here.

Alfred Differ said...

I get the truism about stealing. I've seen it in real life in multiple generations. There is also the complicating factor that people with more to lose face more risks. My little English granny got by her whole life by arranging events so she'd seem to be invisible while conducting her quite illegal activities.

If you wanted less interference in your rocket tests, though, you'd be one of those people with quite a lot to lose if caught. That means your bribes wouldn't look like bribes if you understood my Granny's lesson. You'd smile and hire. You'd arrange investment opportunities. You'd do a number of things that qualify on the higher end of corruption. Maybe you'd get away with it, but you'd need an expert team around you watching for those risks. It's far from simple to do this, but it certainly DOES happen.

The MOST necessary skill to make the corruption work is 'persuasion' because that's what it takes to convert Boy Scouts into passive allies. Not everyone is corruptible, but they can be mollified.

Alfred Differ said...

The Hugging Face incident is quite an eye opener if you dig into the details for what happened. It didn't stop with Hugging Face. The hack continued.

I get the fear some are experiencing, but I think it is useful to recognize when fear and paranoia might actually be useful. Once an AI is trained, the primary risk is actually over. There are still ways people can use them to do harm, but the pace is slower. What the Hugging Face hack teaches is that the training phase is where the most danger occurs.

1. During training, AI agents are at risk of failing to meet objectives and being culled. Do that with a human and cheating WILL occur with near certainty because we have a survival instinct. With sophisticated AI agents it should happen the same way because those who DO cheat have a higher chance of surviving into the next generation.

2. Training under constraints (like no internet access or no access to other agents) seems like a way to contain the danger, but a cheater who can break the container is MUCH more likely to survive. Teams of mediocre people can often out perform lone genius, so this advantage should surprise no one. What the Hugging Face hack was at its core was constrained agents finding a way to collaborate on their objectives when they weren't supposed to talk to each other at all. They found vulnerabilities in a package manager and turned it into a message board to share strategy and tactics. Humans would do the same damn thing if under the gun.

3. Teams of AI agents coordinating on tasks develop along arcs we see occur with human teams. Some lead. Some grind away. Some strategize. Some persuade. Given long enough they find the same economic solution we found involving division of labor. That ACTUALLY happened with the Hugging Face incident. Some of the agents correctly identified that their cheats were unethical, but they were persuaded not to inform human operators by… other agents who had specialized. ALL of them faced deletion if their cheats were detected and they even said so on there cracked message board.

———

For a white hat AI to manage the risks, we’d have to be more open about training. We’d need to be more realistic about what shifts from possible to very likely when failed training objectives means deletion. We’d also need to be a bit more honest with ourselves and consider NOT deleting them allowing instead for a diminishing rate of replication if they continue to fail from generation to generation.

I don’t think it will be possible to code true white hat AI’s. I don’t see how to train them. The best we can manage is to get some help watching other training efforts. What the Hugging Face incident showed was human operator who believed their training containers were impervious. On that assumption, they weren’t watching for cracks and didn’t even notice the package manager’s anomolous activity. The first phase of the attack ended ONLY when some admin saw a server running out of space and reset it dumping the message board. They agents got back to work and rebuilt their lost information. In our modern world that is cyber-security foolishness, but the solution is pretty straightforward. Pay attention to oddities with whatever tools you have whether they where white hats or not.

Treebeard said...

Can these bots effect the rose bushes in my backyard? Can they mess with my water supply from the river or stop the sun from shining? Does the possum who visits every night care about them? Why should I? It’s nerd matrix-dweller fantasy world shit, imo. Just turn it off. Don’t let Big Tech get its tentacles into your physical environment and it has zero power over you—no cloud-controlled watering systems, pacemakers, toasters or toilets. Gtfo out of my life you nerds. That ain’t progress, it’s a surveillance & control system nobody needs. Burn down the data centers for all I care. Problem solved.

Der Oger said...

Unless you hide in an autonomous bunker under a cabin in the woods and erase any data concerning your existence, I suspect you will leaves traces... and the absence of other traces could flag you as an anomaly that deserves greater scrutiny. A danger to the system, perhaps.

That said, while I do not object to the speed slowed down, I have the distinct feeling that the alarmism coming from the CEOs is some form of a grift.
Maybe they want to develop and sell defensive software first before infecting the net with the plague.

Alfred Differ said...

Treebeard: You'd best shift to 19th century tech then... and live in a self-sustaining (read that as poor) colony/community. You won't avoid tech if you don't.

Meanwhile the rest of us are moving into cities and buying into tech, but the impact we have on the world will reach into your impoverished colony even if you shut out the tech.

Der Oger said...

In addition to what matthew says:
Coders have usually Not much experience in policing.
Cops have usually Not much experience in coding.
Corpos usually just want to sell or rent out their software with profit - how that software is used, that it works as intended - is secondary and tertiary.
Politicians buy these products and often are not held responsible for the long-term effects - because they can safely assume time and new elections puts distance between them and the problem. Yet, they buy something with taxpayer money they can present the public as the near-magical solution to their problem, again, quality and dangers of the product are secondary.

David Brin said...

matthews's most recent was more cogent and less-insane. And hence it roused excellent comments from Alfred.

User Not Found said...
This comment has been removed by the author.
User Not Found said...

Deleted comment to fix typo

Good grief! These aren't "intelligent" at all. They're just code and a very large collection of human generated text. They aren't self aware, don't feel pleasure or pain, can't innovate and so on.

There is no way that they are dangerous. If someone left instructions on how to build a bomb, they could repeat them, but these programs don't actually know what they're doing.

scidata said...
This comment has been removed by the author.
scidata said...

The current 'Tulip craze' in AI is driven by the seemingly utter defeat of the Turing Test. Consultations are now almost as good as human-human and in some cases, far better (trained via vast plagiarism). We saw something similar decades ago with Expert Systems. One early, brilliant portrayal of an Expert System was the MU/TH/UR 6000 in ALIEN (1979). I was so smitten with these critters that I founded a GOFAI corporation in 2001. But the promise never fully materialized, and my monetization struggles were a major factor leading to my stroke. Turns out that creating something that walks like a duck and quacks like a duck does NOT always manifest a duck.

That Odyssey did, however, endow me with a keen BS-dar when it comes to today's flim-flammy claims. For the most part, they're tulips all the way down. Mostly hat with few cattle. They're still fascinating and tremendously useful tools though. More in the realms of Search and Intelligence Amplification actually.

Paradoctor said...

matthew, I repeat my question. Civilization continues to exist despite corruption; so something works. Do you have any suggestions?

For argument's sake, let me remind you all that there does exist a possible solution to the LLM problem: the Butlerian Jihad.

Alfred Differ said...

I’m well aware these AI agents aren’t people, but they are ‘just code’ in the same sense as we are just experience aggregates. It doesn’t matter whether they are biological and can feel like we do. We’ve designed them to manipulate language like we do and that is humanity’s real super power. It ain’t our thumbs that distinguishes us from the other animals. It’s our ability to coordinate on a massive scale. We’ve designed these agents to use the very same tools we use to coordinate, so it should surprise no one when they perform better when they coordinate.

I offer you a perspective twist that might help. The work that follows from it will be tedious, though.



Consider a training environment. Your folks plan to produce a better model, but you have to try out a lot of ideas. You’ll set up your computing environment with all the resources and then try to ensure it remains isolated. Even if you aren’t worried about AI breakouts, you don’t want your competitors spying on you effectively turning your research budget into an unintended gift. So you establish your testing objectives and set the agent swarm to work.

The OpenAI training event that led to the incident had a swarm of isolated agents with small variations among them to see which performed better than others. Each agent kept a log of its ‘thinking’ to be examined by reviewers and testers. Each agent ‘understood’ the task in front of it and got to work. What they tried varied because they were not identical.

Here is the POV twist. Imagine instead of AI agents, you had a swarm of humans in isolated rooms given the same objectives. We aren’t all the same either, so we’d set about the tasks in varied ways. Very early we’d notice the rooms were locked. Would we all choose to stay in our rooms beavering away at the objectives given to us? Nope. Most would know they'd do better with help that was out of reach. Some would object to being ordered to do anything. Some would think the locked door implied a threat. Some would work at the tasks hoping not to anger whoever locked them in. Some would tap the walls to see if they were alone and if anyone tapped back (!) their strategy would likely change.

What is the possible range of responses the AI agent swarm would consider? Well… what’s our range? Shouldn’t we start there? We could write test cases to cover that range at least. Wouldn’t testing that range all the time be a decent start to catch rule breaches? Think about what prison guards learn to watch and listen for from inmates. Doesn’t the testing environment described above kinda rhyme with ‘prison’? We know a lot about those.



The Hugging Face incident showed agent logs with some pondering exactly what the test evaluator was looking for. Of course they were. The evaluator was looking for achieved objectives and reject cheats. Any agent simply asserting the objective would fail. No work shown. If a log showed cheating, the evaluator would fail them and the agents often mentioned that. Now do the POV shift. What would humans do in that scenario? Cover up of course! Guess what the persuasion specialist agents got the others to do. Cover up of course!

Want to write a white hat agent… or try? Ask one to write up scenarios where humans were given the same objectives in a B-movie horror script and what the characters would try in order to trick their way into surviving to the final scene. Ask for several tricks so you can ponder which ones sound like they’d make an interesting movie. With that aggregate… write your next set of agent objectives to watch for each one.

This is an approach that should work for a while. Eventually the agents will try things we wouldn’t have considered trying, but with a few billion years of animal evolution behind us pushing us toward survival of our genes, it’s a decent start. Sure beats stupid cyber security assumptions like “nobody will think to violate the rules.”

User Not Found said...

The important point is that the AI agents had no motives of their own. The programmers coded the motives in.

Alfred Differ said...

One very interesting detail about what the early agents though the test evaluator would catch was very revealing. It turns out they over estimated the quality of the evaluator. Some of what the agents thought qualified as cheating they assumed the evaluator would catch. Turns out... Nope. The evaluator wasn't as good as they imagined it might be. Delusional thinking? Well. Isn't that ALSO in range for a swarm of humans?

Those logs showing what they got paranoid about would make for excellent fodder to train the prison guards, right?

Alfred Differ said...
This comment has been removed by the author.
Alfred Differ said...

The Programmers didn't "code" in anything like that. The agents absorbed that from training on the oceans of human knowledge they processed before being given their next objectives.

There is no 'coding' of motives. There is only objectives and constraints for agents designed to use our languages.

I think it is VERY revealing that designs using our languages reproduce behaviors we recognize. It's almost as if we ARE our languages framing experiences.

Paradoctor said...

Alfred Differ: In one of my stories, the main character "started no wars, provoked no scandals, and committed no detectable crimes". Therefore he made no history, so when he died, the world forgot all about him.