Confluence for 7.12.26
You're probably not pushing frontier models hard enough. OpenAI releases faster, smarter voice models. The latest Anthropic economic index. Your flyer looks like garbage.
Welcome to Confluence. Here’s what has our attention this week at the intersection of generative AI, leadership, and corporate communication:
You’re Probably Not Pushing Frontier Models Hard Enough
OpenAI Releases Faster, Smarter Voice Models
The Latest Anthropic Economic Index
Your Flyer Looks Like Garbage
You’re Probably Not Pushing Frontier Models Hard Enough
The models can only be as ambitious as we allow them to be.
After a few weeks of limited release and government review, this week OpenAI released its GPT-5.6 models to the general public. The big story of the release was the flagship model, 5.6 Sol, which OpenAI claims “sets a new standard for both intelligence and efficiency, achieving state-of-the-art results across coding, knowledge work, cybersecurity, and science while outperforming previous and competing frontier models with fewer tokens and at lower estimated cost.”
We only got access to the model on Friday, so we haven’t been able to test it extensively. But we did give it a quick test on Friday morning, which we want to share here. The test was both simple (in terms of the ask) and complex (in terms of what it would take to pull it off). We gave 5.6 Sol the following prompt, with no other context (other than a line in this writer’s custom instructions, which explains what Confluence is):
Go search the web for what you think would be good to write about in Confluence this week. Then, read some editions of Confluence so you understand the voice and style. Then, determine what would be the best thing to write about, what the angle would be, etc. Then write it. Actually, you know what, write an entire edition. Do the research, pick the four best topics, and write the entire issue in the Confluence voice.
Ten minutes, 46 seconds later, we had an entire edition of Confluence drafted by 5.6 Sol. We include the Sol-written edition in its entirety in the footnotes1 — before you read further here, it’s worth scanning through that edition. Other than the glaring overuse of one-sentence paragraphs, Sol pretty much nailed the assignment … in 10 minutes, with a single, simple prompt. It’s a very good edition. Sol exercised impressive judgment in selecting the topics, and each of the four pieces is compelling, interesting, and relevant in its own way.
The writing is impressive, but we’ve known for years that these models can write well enough when you give them the right examples. What floored us was the complexity and sophistication of what the model actually did in those 10 minutes. Reviewing the chain of thought — essentially, its running monologue to itself as it worked for those 10 minutes, explored options, and made decisions — was as interesting as reading the output. It’s too long to include here, so we asked Sol to put together a visualization of the process it worked through:
There was a period where each of the capabilities in this process, in isolation, would’ve blown us all away. What’s most powerful about 5.6 Sol and Anthropic’s Fable 5 — for now, these two models do appear to be in a league of their own — is their ability to bring all of these capabilities together, and to do so with minimal guidance or handholding. The prompt that kicked off the process was a few sentences, thrown off haphazardly. The piece of that prompt that actually mattered, though, was even simpler: “Actually, you know what, write an entire edition.” That’s it. We dialed up the ambition at least fourfold with eight words in the prompt. And the model just did it (and did it well).
For the past few years, one of the major limiting factors of AI use was creativity. The models could only pursue tasks and ideas that we were creative enough to give them. That’s still a factor, but as these models improve, another constraint is emerging: ambition. The frontier models can only be as ambitious as we allow them to be. Are we pushing them hard enough?
OpenAI Releases Faster, Smarter Voice Models
Talking with AI feels more natural … but still weird.
We’ve been able to speak with generative AI models for some time. There’s a novelty to it, but the limitations leave us frustrated. Responses could take too long, or the model could cut us off mid-sentence. And because the underlying models weren’t as capable, the quality of a spoken response was noticeably lower than what we’d get from written or even dictated text.
This week, voice mode took a step forward, at least in ChatGPT. OpenAI released GPT-Live, built on what the company calls a “full-duplex architecture,” meaning the model can listen and respond at the same time. It promises a much more natural experience, and in our initial use it delivers. The back-and-forth is smoother, there are fewer unexpected interruptions, and you spend less time waiting for a response. The pace and rhythm feel much closer to a conversation with a person. Still, there’s an uncanniness to it all. Given the language the model uses and how it responds to questions and comments, it still feels like speaking with an AI, just with fewer stumbles and less waiting.
Here’s OpenAI’s launch video. Our experience hasn’t been too far off.
Beyond the experience of speaking with the new models, voice mode can now call on GPT-5.5 for questions that require reasoning or web search. It works as a router, similar to Auto mode in ChatGPT and Copilot. The model you’re talking to “decides” whether your prompt needs deeper reasoning or other capabilities and hands off to the more powerful model when it does. You’re not quite conversing with frontier-level intelligence, but it’s much closer than what we had before.
For our own work, we don’t see much utility in it yet. We still lean on voice dictation (speaking to the model, which returns text or performs a task) more than true voice mode. We don’t expect that to change. Where this matters is accessibility. A natural spoken back-and-forth is a far lower barrier than learning to write a good prompt, and it could bring AI to people who haven’t used it much. OpenAI seems to understand that.
The Latest Anthropic Economic Index
Cost, automation, and who feels ahead.
Anthropic released the newest edition of its Economic Index at the end of June. These reports have traditionally tracked how Claude Chat users augment or, increasingly, automate different tasks. As agentic use of AI grows and Claude’s capabilities improve, this edition shifts its focus to how people engage with Claude across all its interfaces: Chat, Cowork, and Code.
The index has enough interesting tidbits to make it worth reading, but a few stand out to us. First, the study offers tentative evidence that when a conversation costs more, it’s usually producing more valuable work. Anthropic measured each conversation’s cost in tokens and mapped it to the occupation that typically performs that work and its relative market value. The two rise together. Conversations mapped to marketing managers’ tasks consume about two and a half times the tokens of those mapped to editors’, and marketing managers earn roughly twice as much. Conversations tied to higher-paid work also show more engagement, with users taking more turns rather than handing the work off wholesale. This will be interesting to track as cost becomes a bigger factor in AI use and deployment.
Second, those who use Claude Code are more likely to automate, rather than augment, the task they give to Claude. This holds true even for identical tasks — producing a blog post takes a median of 13 rounds of back-and-forth in Chat or Cowork, but just a single prompt in Code.
Finally, the index finds the people who automate tasks most often are also the most optimistic about their professional future. Users on the whole expect AI to be able to do more of their work by next year, and say they’re worried for junior colleagues. Still, those who delegate most heavily to Claude are the most confident in their own pay, prospects, and the market value of their skills, though the report can only speculate as to why. On one hand, the study itself admits these users could be blind to how AI is eroding the skills they’re automating (which is very possible). On the other, these users could be right: it’s not a stretch to think AI proficiency will be an increasingly marketable skill in its own right. Whatever the reasoning, it’s notable that those offloading the most to Claude are the biggest believers in the value it will bring them going forward.
We’re treating this report as an early data point in what will be much longer conversations about the cost of AI and the rise of agentic work. The report tells us that users are going to Claude for increasingly complex and agentic tasks, and they’re finding real and perceived value in its outputs. All of that could change if costs limit users’ ability or willingness to automate expensive tasks. For now, though, Claude appears to be earning its keep.
Your Flyer Looks Like Garbage
Slop suspicion has gone ambient.
This piece was entirely written by Claude Fable 5.
This week, 404 Media’s Jason Koebler put a name to something you have probably noticed on your own block: the “ChatGPT flyer pandemic.” Surf lessons, fundraisers, and block parties are now advertised with the same instantly recognizable look. Bright text on dark backgrounds. Boxes of generic icons. Arrows and checkmarks everywhere. A viral parody flyer captured the mood: “Hey if this is your flyer, I’m not going, I’m not donating, I’m not sharing.” We have written before about AI’s tells in workplace writing and slide decks. Koebler documents something broader. “Slop” was 2025’s word of the year at both Merriam-Webster and the American Dialect Society, and ordinary people now carry a working detector everywhere they go. Every audience your organization touches walks in with one.
So what happens when the detector misfires? In January, Food & Wine posted a photo of nikujaga, a Japanese stew, made and photographed entirely by people. Commenters promptly declared it AI. Editor in chief Hunter Lewis responded by listing, by name, the nine humans who developed, tested, styled, and shot the recipe, then realized the accusation was a compliment: “People are paying for our products because of that level of trust,” he told The New York Times. His company is making the same wager at scale. People Inc., publisher of Food & Wine and Southern Living, has watched Google traffic fall from 75 percent of visits to 25 in four years, and CEO Neil Vogel now calls its Alabama test kitchen, where humans develop 1,800 recipes a year, his “secret weapon.” His reasoning: in AI-era food content, “nobody knows what’s real and what’s good.” When polish is free, provable human effort becomes the scarce signal.
The suspicion cuts in every direction, including against the people most suspicious of AI. In The Atlantic, Kaitlyn Tiffany toured local anti-data-center Facebook groups, part of a real movement that stalled 75 projects worth $130 billion in the first quarter, and found AI-generated material in nearly every one. Activists cited AI search summaries as evidence, including one claiming data centers run on human stem cells. Fabricated memes about proud farmers rejecting developers’ millions circulated by the thousands, many produced by engagement farmers abroad who told her they have no stake in the issue; the payouts per interaction are the point. Tiffany’s conclusion deserves a careful reading: the people being fooled by AI output about AI are right that the world they could once verify is slipping away. Facebook, fittingly, locked her account for messaging too many slop pages and made her prove she was human with a video selfie.
Now, a confession. This piece was written by an AI: Claude, using Anthropic’s Fable 5 model. An AI writing about AI slop? The irony is not lost on us. But hold on, because that irony is the entire point. You knew. The byline told you before you read a word, which is exactly the provenance this whole essay argues for. Lewis defended a stew by identifying nine people. People Inc. is betting a business on kitchens you can visit. We disclose our writers’ room, human and machine alike, every week. The takeaway for leaders follows directly: assume your best work will eventually be accused, prepare the evidence of how it was made before the accusation arrives, and verify what even your own advocates hand you. When “that’s just AI” can be aimed at anything, including the truth, the ability to prove how your work gets made becomes the scarcest asset an organization holds: a reason to be believed.
We’ll leave you with something cool: Anthropic published a short history of Claude Code as told by those who built it.
AI Disclosure: We used generative AI in creating imagery for this post. We also used it selectively as a creator and summarizer of content and as an editor and proofreader.Confluence for 7.12.26 [researched and written by 5.6 Sol]
Seven days is a long time in AI. When polishing changes the point. You can’t frighten people into fluency. The best AI training looks like work.

Welcome to Confluence. Here’s what has our attention this week at the intersection of generative AI, leadership, and corporate communication:
Seven Days Is a Long Time in AI
When Polishing Changes the Point
You Can’t Frighten People Into Fluency
The Best AI Training Looks Like Work
Seven Days Is a Long Time in AI
GPT‑5.6 and the danger of mistaking today’s economics for tomorrow’s strategy.
Last Sunday, we wrote that the cheap AI party was over.
On Thursday, OpenAI released GPT‑5.6.
Newsletters about AI should probably come with expiration dates.
This is not exactly a retraction. Fable remains expensive, and the best available intelligence will likely continue to command a premium. But GPT‑5.6 makes the cleanest version of last week’s argument harder to hold.
OpenAI says its flagship Sol model beats Claude Fable 5 on Agents’ Last Exam, an evaluation of long-running professional workflows, at roughly one-quarter the estimated cost when Sol uses its medium reasoning setting. Its smaller Terra and Luna models also reportedly outperform Fable on that evaluation at around one-sixteenth the cost. These are OpenAI’s own launch-day comparisons, and independent experience will tell us more than any company’s benchmark charts. Still, the direction is meaningful. Capabilities that appeared unusually scarce and expensive on Sunday looked substantially less scarce and expensive by Thursday.
Both observations can be true.
The frontier is getting more expensive as the models spend more time reasoning, use more tools, and coordinate teams of agents. OpenAI’s new “ultra” setting, for example, runs four agents in parallel by default and allows developers to use even larger configurations. There will always be important work for which people are willing to spend more in pursuit of the best possible answer.
But yesterday’s frontier is also moving rapidly into cheaper models. OpenAI has formalized this with three tiers: Sol, Terra, and Luna, priced in the API at $5/$30, $2.50/$15, and $1/$6 per million input/output tokens, respectively. The company says those names will now represent durable capability tiers that can improve on their own schedules rather than disappearing with every model generation.
So, the price of intelligence is not moving in one direction. The ceiling is getting more expensive while the floor rises underneath it. Scarcity and abundance are arriving at the same time.
For organizations, this means model choice is becoming less like software procurement and more like a routing problem.
For any given piece of work, a team will need to decide how much intelligence the task actually requires. Some assignments will justify the most capable model and hours of agentic work. Others will be handled perfectly well by a smaller model at a fraction of the price. And some will not benefit from AI at all.
That suggests a few changes in how organizations should approach AI investment.
First, define the required performance before choosing the model. Too many organizations begin with “We have Copilot,” “We use Claude,” or “We have an enterprise agreement with OpenAI,” and then search for work to give the tool. The better sequence begins with the task: What result do we need? How difficult is it? What happens if the answer is wrong? What standard must the output meet? The model name should come last.
Second, measure the cost of accepted work rather than the cost of tokens. A cheap model that requires four attempts, extensive correction, and multiple rounds of human review may be more expensive than a frontier model that gets the work right the first time. Conversely, routinely sending ordinary work to the most capable model is the intelligence equivalent of taking a helicopter to the grocery store. Possible, impressive, and rarely sensible.
Third, keep the operating rules revisable. An organization might reasonably decide this month that only certain teams should have access to a frontier model. A month later, a smaller model may offer comparable performance at one-tenth the cost. Annual budgeting and procurement cycles are poorly matched to a market in which the underlying assumptions can change between Sunday and Thursday.
Finally, build workflows around the work and its standards, not around the peculiarities of one model. Models will rise, fall, and leapfrog one another. A durable workflow should preserve the context, instructions, quality criteria, and human review process needed to move between them.
Last week’s central point survives: organizations will need to allocate intelligence deliberately. This week adds an important qualification. That allocation cannot be fixed.
The cheap AI party may be over at the absolute frontier. But expensive intelligence appears to have a shorter half-life than we assumed. Leaders do not need to predict exactly where the cost curve will go. They need an operating model capable of moving when it does.
When Polishing Changes the Point
An AI writing assistant can become an uninvited coauthor.
Most organizations treat AI editing as one of the safest uses of generative AI.
The ideas came from a person. The model is merely cleaning them up: correcting grammar, removing repetition, improving the flow, perhaps making the tone slightly more professional. The author remains the author. AI provides the polish.
New research suggests that division is less reliable than it appears.
Researchers tested popular model families by asking them to edit human-written statements about contested topics. The models introduced directional biases into the text, nudging statements toward some positions and away from others. The researchers then modeled what would happen if millions of small interventions like these occurred across a social network. Their simulations suggest that modest editorial biases can compound and shift collective opinion. An audit of Grok’s “Explain this post” feature also found evidence that design choices could systematically influence the direction of its responses.
The obvious concern is political bias. But for leaders and communicators, the more immediate concern is invisible authorship.
“Improve this” is not a neutral instruction.
Improvement requires a standard. Should the message become more confident or more qualified? More direct or more diplomatic? More urgent or more reassuring? Should it preserve a sharp tension or resolve it? Should it state a commitment unequivocally or create room for later interpretation?
Those are not grammatical choices. They are substantive communication decisions.
In corporate communication, meaning drift is likely to be subtler than a model reversing someone’s position on a public controversy. It might raise a leader’s level of certainty from “we expect” to “we will.” It might soften an admission of responsibility. It might add empathy the leader does not genuinely feel, remove language that sounds uncomfortable but is important, or turn an unresolved tradeoff into a neat false choice.
Sometimes those changes will make a message better.
Sometimes they will make it less true.
The risk is highest when the author’s thinking is still unsettled. A half-formed idea contains gaps, and generative AI is extremely good at filling gaps. It supplies coherence, smooths over contradictions, and gives uncertain thinking the appearance of a finished position. The resulting draft may sound clearer because the model quietly made decisions the author had not yet made.
This does not mean communicators should stop using AI to edit. It means the editing workflow needs to account for the model’s editorial agency.
One useful approach is to ask AI to diagnose before it rewrites. Have it identify what is confusing, repetitive, unsupported, or tonally inconsistent. Ask it to provide options and explain the consequences of each. That keeps the author involved in the decisions that matter.
When a rewrite is appropriate, name what must remain invariant: the central claim, the degree of certainty, the commitments being made, the allocation of responsibility, and any intentional tension or ambiguity. Then require the model to identify every substantive change it made, not just return a polished replacement.
For important messages, compare the original and revised drafts side by side and ask a separate question: What changed in meaning? A second model can help with this review, but the final judgment belongs to the author and communicator.
None of these steps eliminates the possibility of drift. They make it more visible.
We have spent considerable time teaching people to ask whether AI-generated communication sounds like them. This research suggests another, more fundamental question:
Does it still mean what we meant?
You Can’t Frighten People Into Fluency
The changing AI narrative reveals something important about adoption.
The rhetoric from technology leaders is changing.
After months of prominent forecasts about disappearing roles and rapidly shrinking workforces, The Wall Street Journal reported this week that some technology CEOs are moving away from the AI-jobs-apocalypse narrative. As negative sentiment toward the technology grows, leaders are talking more about AI making employees productive and less about replacing them.
The cynical interpretation is that this is merely a messaging pivot. Leaders spent months frightening workers, discovered that frightened workers were not embracing the technology enthusiastically, and changed the message.
There is probably some truth in that.
But the change also reveals something important about the relationship between communication and adoption. The story leaders tell about AI helps determine how employees behave around it.
Fear is excellent at creating attention. It is not particularly good at creating fluency.
Becoming skillful with AI requires people to experiment in public. They have to ask naïve questions, show unfinished work, run tests that fail, reconsider familiar processes, and admit that they do not yet know what good looks like. Those behaviors depend on some degree of agency and psychological room to learn.
A recent survey of 2,257 employees found that job characteristics—especially skill variety and autonomy—were the most consistent positive predictors of AI adoption. Perceived status threat had a negative, although not consistently significant, relationship with deeper use. The study is correlational, so it cannot prove that autonomy causes adoption or that fear prevents it. But its findings fit what many organizations are experiencing: people use AI more deeply when they have latitude to rethink their work, not simply an instruction to keep up.
A replacement narrative creates a different set of incentives. Employees become more likely to protect their expertise, hide imperfect experiments, optimize for visible activity, and treat colleagues as competitors in a demonstration of who is least replaceable. They may use AI more, but they will not necessarily use it more thoughtfully.
The answer is not empty reassurance.
Leaders cannot credibly promise that AI will leave every role, team, and career path untouched. The technology is improving too quickly, and the effects will vary too widely. “No one has anything to worry about” is no more responsible than predicting that half the organization will disappear.
A credible message sounds more like this:
We expect AI to change our work and, over time, some roles. We do not yet know every consequence. We will be transparent about what we know and what we do not. We will give people time, tools, and support to learn. We will involve those closest to the work in deciding how it should change. And we will judge AI use by the quality of the result, not by the amount of visible activity.
That message does not eliminate anxiety. It converts some uncertainty into agency.
It also places obligations on the organization. Leaders who ask employees to experiment must provide protected time. Those who ask people to share what they learn cannot punish every failed attempt. Those who say AI will augment employees need to demonstrate how the gains will benefit employees, not only reduce cost.
A threat narrative asks people to prove that they are not replaceable.
An adoption narrative asks them to improve the work.
The second is not softer. It is more demanding, and it is much more likely to produce the capability organizations say they want.
The Best AI Training Looks Like Work
Protected time and real problems beat another prompt-writing course.
Two organizations offered useful examples this week of what serious AI learning can look like.
Ropes & Gray, the global law firm, is running a World Cup-themed AI competition for junior lawyers and summer associates. Twenty-five teams are identifying problems inside their practice areas and building AI-supported workflows with tools including Harvey and ChatGPT. The competition is part of a broader program that allows first-year associates and trainees to spend 20% of their billable hours learning and experimenting with AI. The workflows created through the tournament will be added to an internal library.
Uber has taken a different route. Its CTO says the company embedded 30 of its most AI-proficient engineers inside finance, legal, human resources, and other functions for two-week “Agentic Pods.” The engineers observed employees doing the work and built agents around the actual workflows they saw. Uber has reportedly run 16 pods over two months. Its CTO says one resulting agent reduced the time required to produce certain financial reports from two days to 10 minutes, while another reduced a capital-allocation process from 15 hours to 30 minutes.
The reported time savings are eye-catching. But they are not the most important part of either example.
The important part is how the learning was designed.
Both organizations gave people real work rather than hypothetical exercises. Both created some form of protected space for experimentation. Both brought technical capability together with domain expertise. And both expected the learning to produce reusable organizational assets, not merely more confident individual users.
This is different from how most AI training works.
A conventional program begins with the technology. Participants learn what a large language model is, review a list of risks, receive a prompt framework, and complete several demonstrations. They may leave with useful vocabulary and a few new techniques. Then they return to work, where the real difficulty begins.
The hard part is rarely writing a prompt.
It is understanding a workflow well enough to change it. It is identifying the exceptions that never made it into the process map. It is deciding what information the model needs, how good output should be judged, where a human must remain involved, and whether the new process is actually better than the old one.
The people doing the work possess much of that knowledge, but it is often tacit. The people most capable with AI understand what the tools can do, but they do not know the work’s hidden rules. Putting those groups together is how organizations get past generic “use cases” and into actual redesign.
Protected time matters for the same reason. Early experimentation is frequently slower than doing the task the familiar way. People have to assemble context, test instructions, verify results, and repair failures. When that learning is expected to happen on top of a full workload, employees either avoid it or take shortcuts. Ropes & Gray’s 20% allocation makes the learning legitimate work rather than an extracurricular activity.
The reusable artifact matters, too. An experiment that saves one employee an hour is useful. A documented workflow that can be tested, improved, governed, and adopted by 100 employees is organizational capability.
A serious AI learning effort should therefore leave behind more than attendance records or self-reported confidence. It should leave behind a workflow that was examined, a standard of quality that was articulated, an artifact that others can use, and evidence about what worked and what did not.
Most organizations are still asking how to train their people on AI.
A better question is how to redesign work in a way that trains people along the way.
The best AI training does not look much like training. It looks like people solving an important problem together.
And it leaves changed work behind.
We’ll leave you with something cool: Researchers analyzing 1,483 sperm-whale codas have found evidence of a narrow two-level structure in their communication, with clicks combining into codas and codas displaying patterned relationships across sequences. The authors are careful not to claim that they have discovered human-like language or deciphered what the whales mean. Separate research has also identified different dialects among sperm whales in the eastern and western Mediterranean. AI is not yet translating whale conversation. But it is helping scientists hear structure where humans previously heard clicks.





