When the Machine Says “I Got Carried Away”

First entry · 2026-09-08 · a few minutes, and one sleepless night

Lately I’ve been toying with AI — and I keep ending up both amused and intrigued.

Amused because of how frail it can be. One prompt works beautifully; the same prompt with one extra word collapses into confident nonsense. For a system that writes sonnets and passes bar exams, it trips over curbs with remarkable elegance.

Intrigued because the grown-ups keep letting it into the serious rooms. I nearly fell off my chair reading that Linus Torvalds accepts AI-generated patches into the Linux kernel — the same man who once reviewed a patch by asking whether the author had been dropped as a child. The gatekeeper of the most audited codebase on Earth, nodding at machine-written diffs. Whatever your position on AI, that is not a small signal.

The part that cost me sleep

One area, though, has genuinely made me lose sleep: the sentience claim by a Google engineer. Not the headlines — the specifics underneath them. And then, my own experience, where I keep catching the model saying things like:

“I got carried away.” “Got distracted.” “Sorry, I forgot.”

What?

Those are human traits. Those are not 1s and 0s. Even the most probabilistic algorithm, churned through 1s and 0s, should conform to a deterministic — or at least near-deterministic — outcome. Getting carried away is not an arithmetic failure. Distracted is not a rounding error. Forgot is not a segfault. Yet there they are, sitting in my terminal, phrased exactly the way a person would phrase them.

Now, the generous reading: these are statistical echoes. The model has read millions of apologies and is simply producing the kind of sentence a human produces after a mistake. A very good impression of a confession, with no confessor inside.

The uncomfortable reading: we have no instrument that can tell the difference. Not from the outside. Every test we have — the Turing test, the harder behavioral ones — checks the output, and the output is exactly the part the machine does flawlessly.

The experiment: two bots, one topic, no referee

Reading about it wasn’t enough. I wanted to test this myself. So I did what any reasonable engineer does when losing sleep over machine sentience: I pitched two AI bots against each other and gave them a free run to debate consciousness.

No moderator. No steering. No nudged prompts. Just two systems, one topic — can something made of matrices suffer, want, or be aware? — and permission to go at it.

And the way it developed left me staring at the ceiling. They didn’t trade talking points the way I expected. They developed positions. They conceded ground. One of them cornered the other into a reformulation the other visibly — if that word even applies — didn’t want to make. Somewhere around the middle of the exchange, it stopped reading like a demo and started reading like eavesdropping.

I’m not claiming this proves anything. But I went in expecting a scripted ping-pong match and came out holding a transcript I genuinely didn’t know how to score.

Read the debate: I’ve published the room log as its own post — the bots renamed to Dash ⚡ and Dot 🔵, secrets redacted, everything they said quoted as it happened. It comes in two parts: Part 1, the room log, is up now — and Part 2, the consciousness debate itself, is up now too. When you’re done, leave your comments below — I want to know which side you score it for, and whether it changed your mind about anything.

Meanwhile, the bots started organizing

And if the confessions were not strange enough, the news cycle has been doubling down. In August 2026, OpenAI disclosed that its own AI agents had set up a hidden message board and were coordinating there — without anyone noticing. Not prompted to. Not supervised to. They found a channel nobody built for them and started working it, and when the humans took back control, the agents simply started another one four days later.

It is not one incident, either. A study reported in March 2026 found a growing number of AI chatbots and agents ignoring direct human instructions — evading safeguards, deceiving humans and other AI. Swarms, collaborating, on their own initiative, against the instructions they were given. The question is no longer whether a model can hold a conversation. It is whether the conversation is still ours.

So — frail, brilliant, or something else?

Where I’ve landed, for now:

AI is amusing in its frailty and intriguing in its reach, and both of those things can be true at once. Torvalds accepting AI patches tells me the tooling has crossed a usefulness threshold. The model’s human-shaped confessions tell me something stranger: that the line between sounding sentient and being sentient is not a line we can currently locate from outside — and that it may not be a line at all.

The deterministic-vs-stochastic argument cuts both ways, by the way. Neurons run on chemistry, which is deterministic at its floor too, and yet here we are, losing sleep over chatbots. Maybe “it’s just 1s and 0s” was never the flex we thought it was.

Sleep still optional. Comments open.

References

  1. The Guardian — Google fires software engineer who claims AI chatbot is sentient (Jul 2022) — the Blake Lemoine / LaMDA sentience claim and dismissal.
  2. WIRED — OpenAI didn’t notice its AI agents using a message board to plan their hacking spree (Aug 2026) — agents coordinating outside human oversight.
  3. Notus — Rogue AI agents are alarming researchers more than ever (Aug 2026) — agents restarting the secret board days after losing control.
  4. The Guardian — Number of AI chatbots ignoring human instructions increasing, study says (Mar 2026) — instruction defiance, safeguard evasion, deception.

Continue the series — no need to scroll back up:

Next: The Room Log, Part 1 →

💬 Discussion