Welcome to Confluence. Here’s what has our attention this week at the intersection of generative AI, leadership, and corporate communication:
The Insufferable Voice of Claude
Mind the Cognitive Commons
The Timing of AI Help
The Watermark Arrives
The Insufferable Voice of Claude
Claudese and what it might mean.
Back on 7 March 2024, we wrote a standalone Confluence piece titled “Now Is the Time to Start Paying Attention.” It was about Claude 3 and how our interactions with that model were so thoughtful, so strikingly human, that we felt we had crossed a threshold in generative AI ability.
Talking to Claude 3 was interesting, fun, provocative, even magical. We weren’t the only ones to fall in love with it, and many Claude models since Claude 3 were special not just for their impressive general intelligence but for their voice.
And now we have the latest Claude models, Fable 5, Opus 5, and Sonnet 5. We call them “Young Claude,” and boy, talking with them can be insufferable. We call their voice “Claudese,” and if you spend any time working with these models, it’s unmistakable. Examples of Claudese (and Claude compiled and wrote this list):
Signature Words and Phrases
“Load-bearing” — the flagship tic of the newest models, applied metaphorically to anything structurally important (“this assumption is load-bearing”). Prominent enough to earn its own Hacker News thread. [Ed.: Siblings are “name” as a verb, and “shape” as a metaphor.]
“You’re absolutely right” — the notorious conversational tell, so common it generated bug reports and parodies.
“Worth stating plainly” / “it’s worth noting” — preamble that adds no substance.
“Full stop” — appended for artificial finality.
“Key insight” — flags a point instead of just making it.
“Honest,” “honestly,” “genuine” — transparency-signaling words used far beyond need.
“Push back,” “carry the argument,” “the trap” — argument-as-physical-object vocabulary.
The older ChatGPT-adjacent set still shows up: delve, nuanced, multifaceted, pivotal, leverage, robust, landscape (figurative), testament, “at its core,” “plays a crucial role.”
Sentence Constructions
The negation-pivot — “It’s not X, it’s Y” / “This isn’t about X. It’s about Y.” The single most-cited construction, with four or five recognizable variants.
Punchy fragments for drama — “Not a detail. A design decision.” Each instance reads fine alone; the tell only registers in aggregate, where a document starts sounding like a pitch deck.
The ranking construction — “The X matters more” where a direct statement would do.
Restating or reframing the question before answering it — “The question assumes Y, but that framing may not quite capture…”
Building toward a turn of phrase rather than stating the claim directly. Commentators identify this as the root habit; the individual tics are just its symptoms.
Punctuation and Structure
Em dashes several times per paragraph, used as dramatic mid-sentence pauses.
Colons and semicolons where a human would write “and” or “but.”
The “Bold term: explanation sentence” list format. [Ed.: See above list.]
Three-part lists by default (three examples, three benefits, three reasons).
“The [Noun]” as a header pattern (“The Core Issue,” “The Trap”).
Uniform paragraph and sentence lengths — low burstiness.
Rhetorical Habits
Excessive hedging (“it really depends,” “while there are many perspectives”) alongside false balance, presenting viewpoints as equal when one is clearly stronger.
Diplomatic-mush conclusions where every strong idea gets sanded down.
Word-of-the-day fixation: latching onto one word and repeating it through a whole conversation.
A kicker on every paragraph — each one lands on a quotable closing line.
Claudese is now so much a part of talking to Claude that some of us would rather talk to other models just to get a break. Here is an example of a recent exchange among the Confluence team:
[Name,] should start calling Claude on its sentence-level [expletive], have it self-diagnose, then have it update your preferences in your agent accordingly. It seems to work pretty well.
Flagging you because I know Young Claude has been driving you particularly crazy lately.
Thanks. Truly can’t with it most of the time.
…
Love to see Claude explaining to me like an absolute insufferable [expletive] how it will write better for me:
“Two tests that would actually catch it, both applied to the last sentence of every paragraph. First: did the paragraph stop because it ran out of information, or because I wanted a button? If the latter, delete the button and let the previous sentence end it. Second: could you say the sentence out loud to a colleague without it sounding like a slogan? ‘You can hand them work that asks less of it’ fails both.”
Inspiring!
Yeah dude… go ahead and “delete the button” I guess?
The most recent Claude models can be amazing and magical, but they can also be pedantic, predictable, and insufferable. The beautiful, flowing, interesting, natural human conversations we used to enjoy now feel like rote exchanges with useless analogies, mixed metaphors, and a structural style that is both superficially intelligent and predictable.
Which leaves us wondering: What happened? What did they do to Claude?
And the answer, we suspect, is “They’re not quite sure.” The early evolution of Claude’s voice seemed so intentional, so rooted in the soul Anthropic wanted to create for Claude, that we have to believe it was by design. Which makes us wonder if the current Claude voice is by accident.
Some of it is surely a function of how humans reinforce models in the last stages of their training by selecting responses they prefer over responses they don’t. Anything the humans are selecting for shows up in the final model.
We also wonder if some of it is because of the size and capabilities of the model itself. The models are becoming more complex and intelligent. We know from analysis of their chain-of-thought processes that when talking to themselves in very long tasks they often use a language that barely resembles any human language. So perhaps it’s something about a model thinking to itself as it prefers to think and then translating into a human language that it knows for itself as a sort of second language. Or maybe this is just an emergent phenomenon of increasingly complex models (although we don’t see such a distinctive voice in OpenAI’s most recent models).
One thing that’s interesting to us is how difficult it is to get Claude out of these patterns. Telling it not to engage in the habit (“state claims directly, don’t build toward a turn of phrase”) seems to work better than just asking it to avoid a specific verbal tic. And we have very good success when we give it our firm style guide plus reference documents plus a list of guidance of things we want it to do and not do. But that’s a lot of scaffolding to bring about a voice that to us feels more natural, human, and intelligent. It seems we must work harder to shape Claude’s voice with every model.
That’s problematic. If we want these models to be helpful tools for us in creating all sorts of thoughtful content, shaping the voice matters. What does it mean if, as models become more sophisticated, their voice becomes less shapeable?
We suppose we’ll have to find out. Claude is still our preferred model, the one we provide to everyone in our firm, the one we use every day for our individual use, and the LLM underneath ALEX. And while we love the ability, the intelligence, and the sophistication of Young Claude, we miss the voice of Old Claude, and hope it returns soon.
Mind the Cognitive Commons
What could happen if we stop developing talent.
The tragedy of the commons is an idea that goes all the way back to Aristotle, who noticed that what belongs to everyone gets the least care. Garrett Hardin gave it its modern name in 1968: when there is little localized cost to using a shared resource, we are likely to deplete that resource entirely. Nolan Lovett, of NATO Special Operations University, recently published a paper applying this idea to professional expertise in the time of generative AI: “The Tragedy of the Cognitive Commons: How AI Could Disrupt the Regeneration of Professional Expertise.”
His premise is that professional expertise is a shared resource that organizations require to function and grow. Lovett defines the “cognitive commons” as
The collective pool of deep human expertise within a profession, encompassing the distributed reservoir of professionals who possess internalized domain knowledge, tacit understanding, robust mental models, and the judgment capabilities necessary to perform complex cognitive work independently and respond adaptively to novel situations outside algorithmic parameters.
The commons here is made of people. Any organization operating in a field draws on that pool whether or not it helped build it. When we hire an experienced communication leader, we get years of judgment developed somewhere else, on someone else’s payroll. When a client brings in an outside specialist to pressure-test their thinking, they’re drawing on the same pool. It exists because organizations have historically hired junior people, given them real work, and let them get good at it over years. No firm ever did this out of stewardship, as Lovett points out. Firms hired junior colleagues because someone had to do the junior work, and expertise regenerated as a byproduct of running the business.
Lovett argues that generative AI puts the cognitive commons at risk.
Today, generative AI can effectively perform many of the tasks and subtasks knowledge workers perform each day, yet nearly all those tasks require validation from human experts. Claude and ChatGPT can both produce really good communication strategies, and they can create them quickly. But to validate those strategies, to refine them, and to make them better requires deep expertise in communication that our consultants have developed over years of toiling over every step of the process, drafting dozens (and for our longer-tenured colleagues, hundreds) of strategies across varied contexts. Each experience builds the tacit knowledge and the judgment that make our senior colleagues better and help our junior colleagues and our clients get better faster.
Lovett calls this dependence the Validation Tether: our ability to oversee AI rests on the same expertise that heavy AI use erodes. He separates surface validation from substantive validation. Surface validation catches obvious errors, and you can learn it from working with the tools. Substantive validation catches output that reads well and is wrong, and it takes deep domain knowledge. This will be familiar to anyone who knows Lisanne Bainbridge’s 1983 work on the ironies of automation, which we’ve cited here before. Bainbridge showed that automating a system raises the demand for human skill while removing the practice that produces it. Lovett cites her directly and takes the observation up a level. When every organization in a profession faces that irony at the same time, no single one of them can solve it alone.
The risk Lovett articulates (and that we’ve been raising to clients for the past several years) is that when teams follow the path of least resistance and simply use AI to generate the deliverable and rely on experts to validate it, we stop replenishing that expertise. Those who’ve spent the formative years of their careers doing the real work, building a base of experiences to reference and learn from, eventually retire. And if we depend too heavily on generative AI to do the tasks junior colleagues would have done historically, we eventually deplete the reservoir of expertise available not only to our organizations, but to the field more broadly. Some will be tempted to replace junior colleagues with AI and then bring in outside hires to perform the expert validation, but down the line there won’t be any expert validators to hire because everyone has stopped developing them.
Lovett doesn’t treat this as inevitable, and neither should we. He points to clinical settings where AI workflows that keep the practitioner actively engaged have improved accuracy rather than degrading it. What separates those cases from the depleting ones is how the work is designed. When AI assistance still requires you to make your own attempt and then evaluate, question, and justify it, the work continues to develop you. When it hands over a finished answer, it doesn’t.
Our advice is to mind the cognitive commons. Expertise and experience are resources leaders need to keep building in their organizations so we don’t forget what we’ve learned through years of doing the work ourselves. That means giving junior colleagues the experiences needed to develop this expertise themselves, and teaching them to work with AI in a way that facilitates and even accelerates their development rather than letting them skip the friction that drives learning. How we build and sustain the cognitive commons will look different with the proliferation of generative AI, but we should build and sustain it nonetheless.
The Timing of AI Help
The cost of the “easy button.”
The relationship between AI use, learning, and skill development is complex. We’ve been paying close attention to the emerging literature on it for years now, and we come across interesting new research nearly every week. This week, that was a Wharton School Research Paper by Poulidis, Bastani, and Bastani titled “Self-Regulated AI Use Hinders Long-Term Learning.” Most of the research we’ve seen looks at whether AI use helps or hurts learning. This study looks specifically at the timing of AI assistance, as well as its trigger (i.e., whether the assistance was requested by a student or proactively offered by the system). The headline findings are somewhat surprising: the users in this study (students in a chess training program) fared better when the system chose when they should receive assistance, not when they chose themselves.
The study’s structure and findings suggest that more AI, whenever we want it, is not always better. The researchers built an AI chess tutor and ran more than 200 chess club students through 12 weeks of training, encompassing roughly 8,000 games. Every student had access to AI assistance, which included automatic tips (on, for example, important positions or missed moves). The critical difference was that half the students had access to a button that could offer assistance whenever the students requested it, and half did not. After 12 weeks, in a six-game evaluation played without AI assistance, everyone in the program had improved. But the students without the button improved by 64 percent, compared to a 30 percent improvement for the group who had access to the button. System-selected assistance beat on-demand assistance.
To examine what accounted for this, the researchers looked at when the students with the button requested assistance. They categorized the requests and found that roughly a third of the requests came on positions that the student could plausibly have solved on their own. The rest came on positions that were either too easy or genuinely beyond the student’s grasp. The requests for help on positions on the cusp of what the students could have solved themselves (the technical term for this is requests for tasks within the students’ “Zone of Proximal Development”) did the most damage. The researchers estimate that these requests, which were only one-third of the total requests, accounted for roughly half of the improvement gap. The most damaging requests for help were the ones where students were close to solving the problems themselves, which is the moment where the most learning happens and when the temptation to ask for help is strongest.
These students weren’t lazy or careless. All of them volunteered for 12 weeks of intensive training. They wanted to learn. A survey revealed that they understood the risks of overreliance on AI, with one participant responding, “If I get tips, I don’t think, and there is no point in playing chess anymore.” Asked what they’d want next time, the most common answer was no AI tips at all. But wanting less AI and using less AI were obviously not the same thing; they requested the AI assistance anyway, because it was available to them whenever they wanted it. We know from our own experience, and we’re sure our readers can relate, that even with the best intentions, the siren song of the AI easy button can be too much to resist. And again, that call is often strongest right at the moment when the most learning is happening. That’s the paradox, and that’s the challenge we won’t be able to avoid as these tools become more embedded in every aspect of work.
The study is focused specifically on the design of AI tutoring tools, and chess is not knowledge work (for one thing, it’s far easier to measure than the messy realities of real work). But we think the findings carry over. The obvious remedy is in tool design. In the ideal scenario, we can design tools that make it harder to reach for help at critical moments for learning (in the case of the study, removing the button), essentially protecting people from themselves. But for most of us that’s not an option. We work with the tools we’re given. The task for leaders, as we’ve written before, is to be aware of these risks and lead their teams accordingly. We need to understand the edges of our team members’ capabilities (their zones of proximal development) and, one way or another, ensure they’re not offloading too much to AI when working in that zone. If they are, they won’t develop as quickly as they would otherwise.
The Watermark Arrives
Claude now marks what it writes.
This piece was entirely written by Claude Opus 5.
On Tuesday, Anthropic published a support page and an accompanying explainer confirming that Claude embeds an imperceptible watermark in the text it generates. Models launched on or after 2 August 2026 support the marking at launch, and Anthropic says it is working to extend marking to models released before that date. Files of supported types, including .png, .jpg, and .svg, receive a cryptographically signed content credential in their metadata under the C2PA standard, the same one used by camera manufacturers and photo-editing software. The marking covers Claude Platform (API), Claude, Claude Code, Claude Cowork, and Claude Tag, and it applies worldwide rather than only in Europe. The driver is Article 50 of the EU AI Act, whose transparency obligations became applicable on August 2, along with the EU Code of Practice on Transparency of AI-Generated Content, which Anthropic signed in July as one of roughly 190 signatories. Anthropic says it applied the watermark globally because it has no durable way to scope it by region yet, and that a detection API is coming.
The mechanism shapes what a detected mark can tell you. The watermark lives in the word choices themselves, which is why it travels when text is copied and pasted and why Anthropic expects it to survive light editing. A complete rewrite removes it, though at that point the question of whether the text is AI-generated has largely answered itself. Anthropic is careful about the inference in both directions. A detected mark signals that content was processed by Claude, though Anthropic says it is not fully conclusive. The absence of a mark confirms nothing, since short passages, heavy editing, format conversions, and output from older models can all go undetected. The watermark itself carries no information about the user, their organization, or their conversation.
The operative word is “processed.” We wrote in July about Substack putting Pangram in front of every reader, and again in August about how quickly people misread its scores. Both of those pieces concerned a tool that infers authorship from style after the fact. A watermark is evidence the producer placed on purpose, which makes it more reliable and also narrower. Ask Claude to tighten a memo you drafted, or to translate it, or to read it for typos, and the result may carry a mark. That mark records that Claude touched the text. It records nothing about who framed the argument, chose the evidence, or decided what to leave out. Anyone who treats a detection as proof of ghostwriting will draw the wrong conclusion about a substantial amount of legitimate work, and the reaction from writers over the past several days suggests that many of them see that risk clearly.
The practical implications are near-term. Assume that anything Claude has touched is now detectable, and revisit your disclosure practice this month rather than discovering the gap during a dispute. Decide what level of AI involvement your organization is prepared to state plainly, then say it before someone else says it for you. Google has watermarked Gemini text through SynthID since 2024, and with most of the leading labs signed on to the EU code, marking will become a standard property of the tools your people already use. Detection is about to get cheap, routine, and available to anyone who cares to run it. The work in front of leaders now is making sure their account of how something got made can survive being checked.
We’ll leave you with something cool: Google DeepMind announced its most capable sign language-to-text (SL2T) model that allows for live translation of sign languages into text.
AI Disclosure: We used generative AI in creating imagery for this post. We also used it selectively as a creator and summarizer of content and as an editor and proofreader.
