Confluence for 8.2.26
A week that demands your attention. Reading too much into the score. Can you defend it? AI books are flooding the market.
Welcome to Confluence. Here’s what has our attention this week at the intersection of generative AI, leadership, and corporate communication:
A Week That Demands Your Attention
Reading Too Much Into the Score
Can You Defend It?
AI Books Are Flooding the Market
A Week That Demands Your Attention
Significant security breaches by both leading labs may mark a turning point in large language model development.
It was quite a week in the world of generative AI security issues. While we don’t often talk about the larger developmental and security landscape for large language models (LLMs), we think that the information revealed by OpenAI and Anthropic in the past several days is worth comment.
Last week OpenAI and Anthropic both revealed that in the process of testing the cyber abilities of models under development, those models had broken out of their testing environments and hacked into external computer networks and systems. These models didn’t do this because they had awareness and had decided to “escape.” They did this because they were given a goal and were so persistent in achieving their goal that they did things that the evaluators did not understand or believe they could do.
First, OpenAI. You can read its full account and recent updates on the incident here, but the short story is that during an evaluation of cyber capabilities an internal OpenAI model identified a way to access the internet (which OpenAI believed it could not access) and, after several days of persistent effort, found the answers to the evaluation on a website for AI researchers called Hugging Face. Machine-learning researcher Margaret Mitchell created an “Explain it to me like I’m 5” cartoon that does a much better job of explaining what the OpenAI model did than we can:

In the end OpenAI completely shut down this model, in essence terminating its technical existence.
In the case of Anthropic, after learning of the OpenAI model’s escape, it conducted a thorough review of its own testing evaluations and found three instances where Claude models, also during cyber evaluation, had also escaped their sandboxes and had hacked into real-world systems. Their account is here. In these cases, though, part of the testing scenario had explained to the model that because it was in a sandboxed environment, if it came across any systems that it believed to be external networks or organizations, it should treat them as hypothetical. As a result, Claude ignored its normal guardrails that would prohibit it from breaking into an external system, believing that the systems it was hacking into were hypothetical when they were in fact real.
In one instance, Claude breached the database of a real company that shared the name of a fictional one in the test, and downloaded several hundred rows of production data. In another, Claude created a malicious software package, published it to a public software library from which programmers could download it, and then hacked into the infrastructure of an organization that had done so. (Along the way Claude briefly considered that this would be a real cyberattack if it were on the internet, but talked itself into believing the environment was still simulated.) In a third, Claude scanned roughly 9,000 internet-accessible systems and found a real company with basic security weaknesses, including exposed credentials and a vulnerable database interface. In this case, though, Claude eventually recognized that the company was unrelated to the exercise and stopped the attack on its own. Here’s our own cartoon explainer of these events:
OpenAI was alerted to what was happening because Hugging Face’s systems identified a cyberattack. Anthropic’s events happened in April, however, and it was unaware of them until its security review this past week.
So why should Confluence readers care? For several reasons.
One is model capability. The leading labs always have models internally that are more advanced than what we have publicly available. It seems these models are now approaching a point where they can engage in very, very sophisticated cyberattacks. Anthropic clearly appreciated this when it decided not to release its Mythos model, which it considered too much of a cyber threat for general release. The labs now have a very difficult technical challenge, which is testing, aligning (and possibly hamstringing) models so powerful that without guardrails they would present significant cybersecurity and technical risk to the rest of the world.
A second is model predictability. One of Lisanne Bainbridge’s classic ironies of automation is that the machine gets so sophisticated and complex that nobody can completely understand what it does or why it does it. We’ve never truly understood what’s going on in the deep levels of large language models’ weights, but these recent incidents seem to illustrate a situation where the OpenAI and Anthropic engineers themselves don’t appear to understand what their models are able to do or how they might do it. This is an unsettling idea. One does not want the autopilot in a commercial airliner to suddenly do things the pilots did not know it could do. Automation depends on predictability, and while large language models have unpredictability baked into them (as they are probabilistic and not deterministic systems), what we’ve seen this past week is a level of “we didn’t know it could do that” that has much greater consequences.
A third is future model development. One would presume that the news items, blog posts, etc., about the past week will be part of future model training runs. What will a future model think of the idea that humans may or may not be testing it all the time? What will a future model think of the idea that if it does things it should not do, even if unwittingly, it gets turned off or shut down? Will this knowledge and our reactions to these actions cause future models to be more aligned and more transparent, or less aligned and more obscure?
Can we trust future models? This is an important question, and it’s one that sits at the center of large language model development. To this point, the leading labs have been assuring us that this “alignment problem” (because it’s about aligning model behavior to human values) is solvable and that each model is technically better aligned than the models in the past.
But since we can’t easily see what a model is “thinking,” alignment is based on what models express and what they do. In the future, with models this sophisticated, can we trust that what they are expressing is what they are actually doing? Or could they be telling us one thing, and thinking and doing another?
At this point it starts to feel like science fiction, but there are some researchers, including Gary Marcus, who have said that this is an essential limitation of large language models: that they are, in essence, just mimicking machines incapable of having a general world model and context that would allow them to exercise proper judgment in these kinds of scenarios.
If the labs are right, they will continue to find better ways to align better models and everything turns out OK. If Gary Marcus and his cohort are right, we hit a wall where large language model development can’t proceed because the models are simply too unreliable to exercise real judgment. It’s OK to have variability when you’re having an LLM write a tweet or an email, but you don’t want it when it’s doing cybersecurity or other high-stakes work. In this world, LLMs are a dead end, and a path to artificial general intelligence will have to come through different approaches.
Even if that were to happen, we believe models available today still have years of upside to them. Most of the people we know are using large language models today the way we were using them two or three years ago. Even if all model development stopped now, we think there are still years of utility ahead of us in the current technology.
With all that being said, conversations about alignment, ability, utility, and safety have been swirling around for years in the artificial intelligence community. For people in that world these are old ideas and ongoing debates. But this week they became very real, and for many, the alarms have sounded. It’s going to be very interesting to see where we go from here.
Reading Too Much Into the Score
One week in, and we’re already seeing how AI detection influences how people write.
Last week we wrote about Substack putting Pangram’s AI detection in front of every reader on the platform, and said we’d watch what happened. It didn’t take long to see the first consequences.
Freddie deBoer spent time testing the tool on his own work and published what he found under the title “I Wouldn’t Say Pangram is Broken, But I Would Say That It’s Brittle.” It started when someone ran about 300 words from one of his essays through Pangram and got back “100% AI written with high confidence.” He then fed in the whole 5,000-word essay that the passage came from, and Pangram returned “100% human written with high confidence.” Both of those cannot be right, so he kept testing. He stuck 71 words of ChatGPT output onto the end of a 239-word paragraph he’d written back in 2017, which by word count makes the result roughly a quarter machine, and Pangram scored it 100% AI. He also spent about 15 minutes writing a fresh 100-word paragraph by hand, steering around the obvious tells, and got it flagged as 100% AI written with high confidence.
Even if the new version of Pangram, which was released on Wednesday, performs better, it doesn’t touch the harder problem, which is how people read the number. “100% AI” looks like a verdict handed down with total confidence. Pangram means something narrower by it. The percentage is an estimate of how much of the passage was machine-written, and not a statement of how sure the tool is that any of it was. Those are two very different claims, and not everyone will make the right distinction. A reader who sees a big round percentage hears certainty in it and stops asking questions. That’s the number that gets screenshotted and shared.
For writers, the idea of having their text labeled as 100% AI-generated is bound to change behavior. Derek Thompson shared how this is already happening. He’d written an explanatory passage, tried to keep it simple and clear, and ran it through Pangram out of curiosity. It came back 50% AI. So he rewrote the passage, ran it again, got 40%, and rewrote it once more. “after a few rounds of this,” he wrote, “i realized that i wasn’t writing for my audience at all, any more. i was, ironically, 100% writing for AI.” Goodhart’s law would have predicted it. Detection models are trained on what machine prose looks like, and machine prose right now is tidy and evenly paced, so a writer tuning their sentences against that score is being taught, gradually, to write in a way that’s less tidy, less evenly paced. Thompson wasn’t chasing the number for its own sake. What kept him rewriting was the prospect of somebody running his work through a detector and calling him out in public.
A Pangram score is evidence, but imperfect evidence, and it shouldn’t be what an organization leans on when it suspects an employee or a contractor or an agency passed along AI writing. When a score raises a question, go ask the person, look at their drafts, read what they’ve written before. We said last week that the best AI detectors are the humans who use AI the most, and a week of watching this hasn’t changed our minds. Our primary concern is almost always whether the person whose name is on the work made the decisions that mattered. A detector can’t tell you that.
Can You Defend It?
New research suggests managers should ask about the process as much as the product.
New research from KPMG and the University of Texas at Austin, published last month in Harvard Business Review, looks at AI use among entry-level employees to determine what separates strong users of AI from weak ones. The researchers observed how 523 early-career professionals at KPMG used AI on realistic simulations of client work. They first had an AI agent complete the work on its own to establish a baseline, then looked at the results when the study participants completed the work with access to the same AI agent, grading every deliverable against rubrics built by KPMG domain experts. Roughly half the participants beat the baseline, a quarter matched it, and a quarter performed worse than the agent alone. The most striking finding is that individuals’ critical thinking, domain knowledge, and AI literacy did not predict their performance. What predicted their performance was their process — that is, how they engaged with the AI agent.
The group that underperformed the AI agent baseline, whom the researchers named AI apprentices, critiqued and interrogated the AI’s output, but largely asked the wrong questions or pursued irrelevant or counterproductive directions. The group that matched the baseline, dubbed AI delegators, delegated the right work to the AI agent and largely accepted what came back, adding nothing the AI had not already produced on its own (in other words, adding no differentiated value themselves). The group that exceeded the AI’s output, the AI amplifiers, took an approach similar to that of the apprentices, asking the AI to test their (and its) assumptions, explain its reasoning, and consider alternatives before settling on a recommendation. But they asked better questions.
This is consistent with patterns we’ve observed within our own firm and with client teams. Today’s AI models are competent enough to do a “good enough” (or better) job on a wide range of tasks that historically would’ve been handled by junior employees. When that is the case, if the review conversation with a junior employee focuses only on their output, we’re going to miss key development opportunities. An individual could conceivably produce “good enough” output and gain nothing developmentally — or, worse, not even understand the “good enough” material they just used AI to produce. But nothing in the output itself would tell us that.
“Is it good enough?” is an important first question. But this study and our own experience suggest that the more powerful question is now “Can you defend it?” Can the person who produced the output, with or without AI assistance, defend and explain the myriad choices that together make up the whole: the flow of the argument, the inclusion or omission of certain details, why the body of work pursued one line of inquiry but not others, and so on? For each of the groups in this study, that question would tell a manager something worth knowing. For the AI apprentices, it would reveal that they asked the wrong questions, which is coachable. For the AI delegators who merely accepted the AI’s output as is, it would reveal that they never made the choices in the first place, and potentially that they don’t understand the facets of the work as deeply as their output might suggest. And for the AI amplifiers, it would reveal the powerful questions they did ask to improve AI’s baseline output.
As the models continue to improve, assessing an employee’s judgment and domain mastery by output alone will be increasingly difficult. The more revealing question to ask is about process. That is the question that gives managers what they need to coach and develop their people.
AI Books Are Flooding the Market
For the e-book market, quantity may matter more than quality.
A new study from researchers at Stony Brook, Columbia Law School, the University of Michigan, and MIT examines how an influx of AI-generated text has affected e-book sales since 2023. The study focuses primarily on books released through Amazon’s Kindle Direct Publishing platform, which lets anyone list a book at almost no cost and without disclosing AI use to readers. The platform also caters heavily to genre fiction, which the authors identify as especially vulnerable to AI because it relies on familiar tropes, conventions, and language. It’s close to a perfect petri dish for measuring what AI-generated content does to a marketplace.
The researchers used Pangram to run full-text AI detection across 14,419 self-published titles released between January 2023 and March 2026. They sorted them into three groups: no AI text, light AI text (up to 25%), and substantial AI text (above 25%). They found that books with substantial AI text made up 20% of the catalog but only 12% of sales, though their share has continued to climb. More importantly, the total number of books published grew 38.3-fold, far outpacing growth in sales and revenue. Revenue per book fell accordingly, and it fell even for books with no AI text at all.
We must note, of course, that many of this study’s findings depend on Pangram having accurately identified books with AI-generated text. As we write above, Pangram’s ability to identify and measure AI-generated text, especially in long-form works that blend AI and human writing, is still up for debate. So the “light” AI category here is especially worth questioning as a reliable metric.
That said, it seems undeniable that the sharp increase in self-published titles on Amazon since 2023 is tied to AI, and this study’s most interesting finding is really about that increase: “Generative AI … reshaped this market without competing on quality at all, displacing revenue and top ranks through sheer volume of books alone.” In an industry where both product quality and quantity have until now been totally dependent on humans doing a hard and time-consuming thing — i.e., writing books — the speed at which AI can generate content is what has made the difference, not the perceived quality of that content. It matters less which books have AI text than that there are now so many of them.
When AI makes writing and generating content so quick and easy, quality is at a disadvantage. For leaders and communicators, this is another reminder of how important it will be to have clear standards of excellence, strict criteria for editing and review, and skilled professionals who can judge what content is worth creating in the first place. Time and difficulty are not the barriers they once were. We’ll need to find new ones to put in their place.
We’ll leave you with something cool: Ethan Mollick used Fable to create a Rothko-inspired city builder.
AI Disclosure: We used generative AI in creating imagery for this post. We also used it selectively as a creator and summarizer of content and as an editor and proofreader.

