MacOS 26.7 Tahoe Release Candidate Contains a Video Demonstrating Camera-Equipped AirPods in Action
Oops.
Some follow-up to this weekend’s stemwinder “Anthropic’s ‘Watermark’ Text Adulteration in Claude Is a Perversion of Writing”:
Contra a bunch of idiots at Hacker News and elsewhere, I understand that popular LLMs do not just pick the “best” token (word) at each decision point. Counterintuitively, always selecting the highest-probability option produces undesirable results. So the models apply some randomization, and “temperature” is the term for the weighting that’s applied so that the “better” (higher-ranked by the model) choices have a higher chance of being chosen.
With a temperature of 1, models use their built-in probability distribution. With a temperature greater than 1, this distribution gets flatter — less-likely alternatives get a higher probability of being selected, and more-likely alternatives lower. With a temperature lower than 1, the probability distribution leans more toward the higher-ranked options. And with a temperature of 0, the highest-ranked option is always chosen. A temperature of 0 generally produces undesirable results — too predictable, too likely to get stuck. Like over-smoothing an image from a camera sensor, eliminating all noise makes the overall result worse, even if each single bit of “noise”, evaluated in isolation, is in some sense wrong.
The temperature-based randomness — which is what makes LLM output non-deterministic — is in place to help make the output better. The prose is clearly better with a temperature of 1 (with weighted randomness) than at temperature 0 (with no randomness). The watermarking schemes, on the other hand, are applying predictable-with-the-secret-key randomness for an entirely different purpose than improving the quality of the output, and thus, I believe, inherently make the output at least slightly worse.
Advocates of LLM watermarking schemes for text argue that the schemes don’t necessarily lower the quality of the generated prose, because they don’t change the temperatures — they only change the source of the randomness. Daniel Jalkut wrote a good piece today about this. I hope that’s true. I believe it’s possible that it is true. I think it’s highly unlikely that it is true. I do not see how a detectable signal can be added encoded in the choice of words without affecting the meaning of the prose. If it were true I think they’d show examples proving that it’s true. Also, Anthropic itself admits that it can’t properly watermark text that is programming language code:
For the same reason, code — which in very many cases has to be exact — has generally less watermarking than some other forms of text.
Having said that, in areas where there is an arbitrary choice between particular words or terms within the code, the watermark can be used, such as comments within code. But by definition, it will have a negligible effect on the actual code produced.
I hold that good prose is much more like programming code. Exactness in word choice, phrasing, tone, and even punctuation is always better than imprecision. The difference is that sloppy programming code doesn’t run, or doesn’t run correctly. The human brain, on the other hand, is adept at parsing and making sense out of inexact, even sloppy, prose.
I do not believe these schemes can work without degrading prose quality, if only slightly. Again, though, I am open to being proven wrong. But even if we concede for the moment that such watermarking schemes do not necessarily degrade the quality of generated prose — not one iota — I still object to their use when they are being applied secretly, behind users’ backs. A useful watermark would be one that anyone can check. These SynthID “watermarks” are entirely dependent upon secrets held by the LLM providers (so far, Anthropic/Claude and Google/Gemini). I find that unacceptable, for reasons I hopefully made clear in my essay.
The people in favor of this watermarking for text have been sold a pipe dream, a fantasy. I’ve encountered dozens of comments from angry AI haters (many of them on Bluesky in particular, but also Threads and Hacker News) who are convinced that the only people who could be against the watermarking of AI-generated text are those who are duplicitously passing off AI-generated text as their own writing — and thus that I must be upset only because the jig will soon be up for me too. This of course is not true. I don’t even use AI to write text messages or emails for me, let alone a single sentence of my work.
But I find it funny that so many people who claim to believe that LLMs only produce “slop” and never anything useful also seem 100 percent convinced that the same LLMs are capable of watermarking their output in reliable ways. These people so desperately want to be able to point a finger at AI-generated text that they’ve fallen hook, line, and sinker for the argument from Google and Anthropic that, thanks to them, they’ll be able to.
I don’t want to spend too much time thinking about this because it’s a waste of time, but how exactly do these people think the existence of these mandatory watermarks and detection tools will change anything for the better? Let’s say you work at an office and you suspect that numerous of your colleagues are using AI to write emails and other work-related messages. Their messages are too long, too prolific, and lack lucidity. What are you going to do now? Copy and paste each of their messages into the watermark detectors from Anthropic, Google, and OpenAI? There cannot exist a single detector for all LLMs. And even if you find out that it says it’s a match, that an email or blog post or Slack message was very likely generated by, say, Claude, what are you going to do? March into your colleague’s office and tell them you caught them?
Anyone in a situation where “getting caught” would matter — students, say — is going to use non-watermarking LLMs or run their watermarked text through paraphrasing tools like Declaude.
No practical good is going to come of this, even if these watermarking schemes work as promised (and to be clear, I don’t believe any of it is going to work as promised).1
My advice is not to care whether anything was written by an AI or a human. The only thing worth evaluating is what we human readers are naturally good at determining: whether it is good or bad. If it’s good, read it. If it’s not, don’t. If you’ve got a job where you’re surrounded by colleagues filling your inbox with AI-generated messages that you can’t abide, get a new job or learn to live with it. Hidden secret watermarking signals — even if they work — aren’t going to make things go back to the way they used to be. If you read something and enjoy it, and subsequently find out it was generated by an LLM, don’t feel bad. You read something good that you enjoyed.
I read something earlier today that claimed most of the posts on LinkedIn are generated by AI. That the whole platform is just inundated with AI slop. Maybe it is, but I wouldn’t know, because I never look at LinkedIn because it’s always been filled with crap. If it smells like crap it’s crap, whether the turds came out of a human anus or a turd-generating robot.
Dan Moren, writing at Six Colors today, “LLMs Aren’t Writing”:
LLMs do not care about the words that they pick because they cannot care about anything.
Speaking of two things that are not the same, John rightly points out the difference between the phrases “he leaped at the chance” and “he jumped at the opportunity”. Those are indeed distinct — if semantically similar — phrases, each of which might be more apt in a particular situation; or, to put it in another fashion: the use of each of those phrases tells us something different, whether about the person being described or the writer.
But the LLM doesn’t know which of those phrases is the right phrase to use. It has a guess, based on its models and weights and inputs. But the ultimate choice of those phrases tells us nothing about the writer because there is no writer.
Moren’s is a fine retort to my post, but I fundamentally disagree — albeit at a philosophical level. If you’re reading a written work only to gain insight into the mind that produced it, there is no mind on the other end of AI-generated text. But the work itself exists. My disagreement with Moren starts and effectively ends with his (wonderfully summative) headline. I say if you can read something, it was necessarily written.
Again, this is philosophical. Was a photorealistic image generated by AI photographed? No, I would say it was not. Photography, I would say, is the act of focusing light through a lens onto a capturing sensor, capturing, to some extent, reality. I think Moren is arguing that writing is like that. If photography captures a physical scene from reality, writing captures thoughts from an actual mind. That something you can read that was produced by an LLM was merely generated in a way that doesn’t qualify as writing. Semantics. I just care about the article of text. Moren argues that LLMs are not writing; I say they are. But we’re disagreeing only over what the word writing means, not what is being produced.
As for “caring” about the difference between semantically similar but tonally different phrases, like “he leaped at the chance” versus “he jumped at the opportunity”, no, of course the LLM doesn’t “care”. But I, the reader, care very much. I wrote a column back in November on ChatGPT changing (and renaming) the “personalities” it allows users to choose from. These personalities generate text with strikingly different styles and tones. Because I use ChatGPT, I care very much about the tone and style of its responses to my queries. Not because I’m ever going to pass them off as my own writing, but because I’m the one who is reading them.
Moren, near the end of his column:
In the end, I can’t summarize it any better than to ask: if you care so much about word choice, why are you using AI to generate text?
If this does truly make AI-generated text worse, well… good. A lot of people are already willing to accept what an LLM churns out as “good enough” and, if I’m being realistic, I don’t think this will change anything. But if it does lead to more people being dissatisfied with the pablum they’re being fed and turning instead to writing and editing their own text, then that would actually be a positive outcome. Maybe it’d even mean fewer human writers being put out of jobs.
I sympathize, but I must disagree that it can possibly be seen as a net good for LLMs to produce worse prose. I read the output of LLMs every day. I use AI to generate text because I ask it questions (in text). I want the answers that I read to be cogent, lucid, accurate, blessedly terse — and ideally to strike a consistent tone that is pleasant to my reading ear. The genie is not going back in the bottle.
Lastly, here’s an interesting point to ponder. English is the most expressive language in the world. Don’t take my word for it — it’s the only language I speak (despite four years of Spanish in high school). Take the word of famed 20th century author Jorge Luis Borges, an Argentine polyglot whose first language was Spanish. In 1977 he was the guest on William F. Buckley’s “Firing Line”. You can (and should) watch the interview on YouTube, but here’s a transcript of the relevant portion from Jordan M. Poss:
Borges: I have done most of my reading in English. I find English a far finer language than Spanish.
Buckley: Why?
Borges: Well, many reasons. Firstly, English is both a Germanic and a Latin language. Those two registers — for any idea you take, you have two words. Those words will not mean exactly the same. For example if I say “regal” that is not exactly the same thing as saying “kingly.” Or if I say “fraternal” that is not the same as saying “brotherly.” Or “dark” and “obscure.” Those words are different. It would make all the difference — speaking for example — the Holy Spirit, it would make all the difference in the world in a poem if I wrote about the Holy Spirit or I wrote the Holy Ghost, since “ghost” is a fine, dark Saxon word, but “spirit” is a light Latin word. Then there is another reason. The reason is that I think that, of all languages, English is the most physical of all languages.
Buckley: The most what?
Borges: Physical. You can, for example, say “He loomed over.” You can’t very well say that in Spanish.
Buckley: “Asomó?”
Borges: Well, no, no, they’re not exactly the same. And then you have, in English, you can do almost anything with verbs and prepositions. For example, to “laugh off,” to “dream away.” Those things can’t be said in Spanish. To “live down” something, to “live up to” something — you can’t say those things in Spanish. They can’t be said. Or really in any Romance language.
I’ve seen this interview before, but watched it again today after an email exchange with Kirk McElhearn. Quoting (with permission) from McElhearn’s email to me:
For many years, I worked as a French → English translator, and there is one key difference between the two languages. France is a Romance language, and English is a language with both Germanic and Romance (mainly French) influence. This means that English often has synonyms where other languages may not.
Using your example, “He leaped at the chance” and “He jumped at the opportunity”, both would be translated in French as “Il a sauté sur l’occasion.” Meaning that someone writing in French wouldn’t have the same range of words to choose from. It’s maybe not the best example, because both are clichés, but there are many examples of French words where English has both a Romance equivalent and a Germanic equivalent: pig and pork, sheep and mutton, beef and cow. Food words are just one example, but English also has many more verb choices than French, since it has a larger vocabulary coming from both influences.
English gleefully borrows from any and all other languages. McElhearn wonders whether English is thus more fingerprintable than other languages, because of its richer vocabulary of roughly equivalent synonyms, and its multitude of idioms.
However, this vein of pro-watermarking support from people opposed to AI in general has opened my eyes to the notion that Anthropic is throwing its support behind this in order to get people who despise AI off their backs. ↩︎
The Savant is a political thriller series starring Jessica Chastain that was supposed to debut a year ago. Apple “postponed” it, apparently out of fear of upsetting extremist right-wing nut jobs because the show is about an undercover investigator (Chastain) hunting down extremist right-wing nut jobs. Chastain was not happy about the show being delayed.
Last we heard about the show was back in April, when Marc Malkin reported this for Variety:
“Before it was like, ‘I don’t know if we’re going to see it,’ but now I can say, ‘We’re going to see it,’” Chastain told me exclusively on Saturday at the Breakthrough Prize ceremony in Santa Monica.
As for when, sources tell me that Apple is planning for a July release.
Given that it’s now the middle of August, I think that a July release is looking less and less likely by the day.
The Financial Times, back on July 1, with the transcontinental byline “Michael Acton in San Francisco and Barbara Moens in Brussels” (non-paywalled summaries from 9to5Mac and MacRumors):
Apple chief executive Tim Cook and EU tech chief Henna Virkkunen held “constructive” talks on Tuesday as the two sides aim to lower the temperature in a bitter dispute over the iPhone maker’s new “Siri AI.”
An EU spokesperson said the virtual meeting had involved a “constructive exchange on topics of common interest, on which the work continues”. The meeting included a discussion of how Apple can launch its reinvented Siri in Europe while avoiding millions of dollars in fines for violating the bloc’s flagship competition rules, according to two people familiar with the talks.
[9 paragraphs of explanatory backstory on the Siri AI/DMA standoff elided ...]
The dispute triggered a fierce public backlash against the commission, with European officials reporting hundreds of emails from consumers accusing Brussels of depriving Europeans of a new technology. One EU official said that a commission spokesperson had received a stream of abusive messages, including several death threats.
If there are actual kooks who made credible death threats, of course they should be investigated, identified, and arrested. But mentioning this in the context of the “fierce public backlash” feels like fishing for sympathy over a deeply unpopular policy of zero benefit to Europeans. That there are hundreds or thousands of complaints is no surprise. This policy sucks and is indefensible on practical grounds.
In November, Apple first proposed a technical fix to the EU it later dubbed a “Trusted System Agent” — a layer of software between a user’s device data and a third-party AI model. It would allow rival AI assistants to draw on personal information from the device without giving them full access to the data. However, Apple has yet to build the agent, and said it was looking for assurances from the EU before it starts.
A commission official said its contact with Apple on the idea was limited, and that it lacked a concrete proposal or details on how such an agent would work beyond the general concept. They said Apple “focused on obtaining a green light to delay the compliance”.
That’s the nut of it. Last winter Apple sent an entire team, including engineers, to Brussels to present a proposal for the TSA (maybe that’s another acronym that needs rethinking on the basis of prior art) to ask, basically, “If we build this, would you deem it DMA-compliant?” and the Commission’s response was basically, “Build it first and then we’ll tell you, after we get feedback from your competitors on what they think of it.” And Apple doesn’t want to spend up to two years building a complex system, exclusively for the EU, only to find out then whether it was all for naught.
As for how detailed Apple’s Trusted System Agent proposal was, we have a he-said/she-said dispute. Apple said at a press briefing at WWDC that it was quite detailed. An anonymous “commission official” here told the Financial Times “that it lacked a concrete proposal or details on how such an agent would work beyond the general concept”. Apple has more credibility here, if only because their statements claiming the proposal was detailed weren’t from anonymous sources. They were on the record.
In the meantime, we’ve seen what the European Commission is demanding of Google with Android regarding third-party LLMs. I can’t see Apple ever agreeing to such a system for iOS. We don’t know the details of Apple’s TSA proposal, but I’d be rather flabbergasted if it enabled the things the EC is demanding for third-party AI providers on Android, like unrestricted background execution.
When I wrote this week about Anthropic’s announcement that all Claude models, worldwide, would soon begin “watermarking” everything they generate, including text, to comply with this EU regulation, we were left to speculate how this was going to work, because Anthropic offered not even a vague description of how it would work — despite the fact that the title of the announcement was, absurdly and insultingly, “How Claude Marks AI-Generated Content”.
My initial speculation was that maybe they’d hide invisible non-printing Unicode characters in the text. Just spitballing. Turns out that’s not what they’re going to do. What they’re going to do is apply a form of steganography, where the choice of words (or other token output) at inference time will leave fingerprints that can later, maybe, be detected probabilistically.
I initially guessed “invisible characters” not because I didn’t think of the semantic word-choice technique, but because I was a fool who took Anthropic at its word in their description of what they would do. Their original support document claims:
When a supported Claude model generates text, it weaves an imperceptible watermark directly into the text itself. You won’t see it, and it doesn’t change the meaning, quality, or readability of Claude’s response.
They say “imperceptible” and “doesn’t change the meaning, quality, or readability”. Their words. Not almost imperceptible. Not slightly changes the meaning, quality, or readability. That made sense to me, because that’s absolutely what I want — nay, demand — from any tools I use personally. It’s unacceptable for a tool to sacrifice an iota of clarity, coherence, meaning, quality, etc. for the purpose of embedding hidden clues within the text to suggest its provenance. That’s what I would and will demand. And Anthropic’s (original) support document unambiguously claims that’s what their system will enable. So if that were true, I couldn’t see what was left other than hiding invisible characters within the text.
My error was believing Anthropic that their system wouldn’t adulterate and corrupt the semantics of the text their models generate. That is in fact exactly what they plan to do. I should have my head examined for believing a single word of a document titled “How Claude Marks AI-Generated Content” that doesn’t explain, at all, how Claude marks (or will mark) AI-generated content.
Yesterday, on an entirely different website than the original “How Claude marks AI-generated content” article (the one that didn’t explain anything at all about how it works), Anthropic published “How Claude’s Text Watermark Works”, which does actually explain in layman-accessible terms how it’s going to work. I will return to Anthropic’s new highly euphemistic and slightly misleading description below.
There’s a bunch of research on this topic, some of which I have also linked to below. But the very best description of the general idea behind the technique is an interactive essay by James Padolsey, “How AI Text Watermarking Works”. It’s a wonderfully cogent read, and the interactive elements splendidly illustrate the main concepts. A+ work. If you have any interest in this at all, I dare say you must read — and play with — Padolsey’s piece.
But here’s my stab at a layman’s high-level summary. If you toss a coin N times and note the results, you can determine with a degree of certainty whether the coin is fair or biased. LLMs are, in their popular incarnations, non-deterministic. Ask the same question of the same model and you often get at least slightly different answers. Maybe the same meaning, but different phrasing. At each decision point for generating the next token, the model makes a choice. With these semantic watermarking techniques, they make different choices for some tokens based on word lists that could be called “green” and “red”. At each decision point, they’re a little more likely to pick a word from the green list than the red list. That doesn’t mean they never choose words from the red list. Just that they’re less likely to than they would if the adulterated marking technique weren’t in place. (Same way that a crooked 51-49 coin will still land “wrong” side up 49 times out of 100 on average.)
Words or word phrases are sorted into the green and red lists deterministically on the fly, at each “next token” generation point. So sometimes a specific word will be on the green list, and other times it will be on the red list. Someone with the secret key can determine which list a word will be on at each token generation point (which is how the watermarking is detected); those without the secret key cannot. This means there will never be a list of words that Claude prefers or eschews.
With coin flipping, the higher N is — the more times you flip — the more confident you can be that the coin is fair or biased. So too with this semantic watermarking. The more words in the text, the more accurate the analysis will be that the text was generated by a specific AI model or not. With too few coin flips, you can’t achieve any confidence at all regarding a coin’s fairness. With too few words (or tokens), there’s no way to achieve any confidence whether a string of text was AI-generated or not.
Given a string of text to examine for signs of a specific watermarking system, if there are more words tagged as green and fewer tagged as red than would otherwise be expected, the text can be flagged — with some degree of confidence — as having been generated, or merely modified, by the AI system that applies the specific secret-key watermarking system. The amount of confidence in the determination will obviously vary, significantly, based on the size of the text string and randomized weights given to words on the green and red lists. But only Anthropic will be able to determine if text was seemingly generated by Claude, and Anthropic will only be able to detect the watermarks that are applied by Claude. Claude can’t detect the hidden watermark signals generated by, say, Gemini, and Gemini can’t detect the hidden watermark signals created by Claude, because each implementation is predicated on secret keys held only by the LLM provider.
One of my fundamental problems with this is that no two synonyms carry the exact same meaning. “He leaped at the chance” and “He jumped at the opportunity” are very similar sentences expressing the same general sentiment, but they are not the same. The exact words we choose when writing matter. I want any LLM I use to choose the very best, most precise words at every single decision point. An obvious constraint that I accept is time and computation. Within the constraint of executing inference quickly, and at a certain cost per token, I want the best words. This constraint matches human writing. I could surely write a better column by taking longer to write it. I write with a sense of how much care I should put into every word and punctuation choice I make. I take more time with certain paragraphs, sentences, or even individual word choices when my gut feeling says I should.
In other words, these are necessary trade-offs. These factors are all in my interest: speed, cost, quality. Ideally I would like perfect writing, at instantaneous generation speed, at zero cost. None of those things are possible. Computation is not free of charge (and cloud-based LLM inference with leading models is actually expensive). Inference is not instantaneous. And great writing, whether natural or artificial, can only approach perfection.
The idea that anything other than my needs should factor into the generation of text for me is patently offensive.
This isn’t just about text one might generate with the intention of passing it off as their own natural work. This isn’t even about LLM proofreading of work written by hand. Anthropic is saying that all new Claude models are going to adulterate every single bit of text longer than 200 tokens (~150 words) they generate, including everything it presents to its users to read. So even in a private conversation between a user and Claude, which will never be read by anyone other than the user, Claude will begin making word choices in the name of marking its output in statistically predictable ways rather than maximizing clarity and precision.
Even today’s so-called frontier models are already decidedly lacking in lucidity. Claude, ChatGPT, Grok, et al. are “better writers” than most humans and produce better prose than the median human. But: no shit. Most people are terrible writers. The “average person” is pretty stupid and half of all people are stupider than that. And there are many smart, interesting people who are miserable writers. So as impressive as LLMs are, the bar is low. The best writing I see come out of these models is worse than anything I would choose to read for pleasure. And now Anthropic is saying they’re going to make it worse, on purpose, for purposes that do not benefit me in any way? Even if only slightly worse?
Get fucked.
Speaking of objections, the relevant EU regulation motivating all of this, “Code of Practice on Transparency of AI-Generated Content”, is red-tape nanny-state pipe-dream nonsense. Here’s Ben Thompson’s summary from a paywalled Stratechery update this week:
- The regulation applies to text longer than 200 tokens.
- The provider must mandate in their terms-of-service that users not remove the watermarking.
- The solution should be robust in terms of evading “typical processing solutions” like screen shots, scanning and OCR, copy-and-pasting, translations, etc.
Taken literally, compliant LLM terms of service must forbid users from rephrasing the output from models that comply with this regulation, because the word choices are the marks. But it’s not the European Union that is trying to impose their absurd, impractical, witch-hunt-fueling regulation on the entire world. That falls on Anthropic.
Complying with this, particularly with regard to text, is only going to create problems for honest users. Dishonest users attempting to pass off AI-generated text as their own writing (students, employees, whoever) will simply circumvent detection through non-compliant AI paraphrasing tools.
James Padolsey — whose interactive visual explanation of how these schemes work I linked to above — explains this in a post titled “Anthropic’s Weak Watermarks Appease a Weak Law” (which, if it rings a bell, I linked to in a standalone post earlier today):
The same thought that led to this law could have applied to calculators at the time of their inception, had their outputs revealed themselves through artefacts. Thankfully, a sum borne of the brain is treated no differently from one produced by a calculator. Likewise with spellcheckers. To make assistance suspect only once the tool becomes capable enough to compose a whole sentence is not a principled boundary. It is a moral premium placed on difficulty itself.
Anthropic has nevertheless chosen a blanket, model-level implementation that appears broader than the law’s minimum requirement. That may be convenient compliance engineering, but it discards distinctions the law expressly attempted to preserve. The result is a signal broad enough to implicate harmless and assistive use, yet fragile enough to be removed by a motivated person through substantial recomposition. It risks concentrating suspicion on ordinary and assistive users while remaining weakest against deliberate deception.
Padolsey is the creator of Declaude, a delightfully simple web app that allows you to “Paste in AI-flavored text and get the same content back as plain prose”. Declaude’s original purpose is cleaning the saccharine Claude personality stink from text (whether it was created by Claude or any other LLM), but, if Anthropic persists in its stated plan to begin adulterating all text Claude generates, Declaude will also serve as a copy-paste single-extra-step way to eliminates those marks. Declaude is interesting and useful already, but it exemplifies how ill-considered and futile this EU regulation is when it comes to prose.
Google has a watermarking system in place that they call SynthID, which they apply to AI-generated images, video, audio, and text. I’m concerned in this article only with text. With multimedia, embedded watermarks can be metadata within files, and truly not affect the experiential quality of the work when viewed or listened to. With text, we are talking about the actual words that are chosen. From the “AI-generated text” section of Google DeepMind’s own description of SynthID:
We’ve expanded SynthID to watermarking and identifying text generated by the Gemini app and web experience. Large language models generate text one word (token) at a time. Each word is assigned a probability score, based on how likely it is to be generated next. So for a sentence like “My favorite tropical fruits are mango and…”, the word “bananas” would have a higher probability score than the word “airplanes”. SynthID adjusts these probability scores to generate a watermark. It’s not noticeable to the human eye, and doesn’t affect the quality of the output.
In a group chat, a friend of mine quoted the above, and I responded that if a chatbot wrote “My favorite tropical fruits are mango and airplanes”, I’m pretty sure I’d fucking notice. Another friend then responded with this:
Days later, that still cracks me up.
But Google’s absurd description puts the lie to their own claim that it isn’t noticeable, and it serves to show just how little regard the people behind these generated-text fingerprinting schemes have for the actual craft of writing. Of course bananas has a higher probability score than airplanes, because airplanes aren’t fruit. But what about pineapple? Should the sentence complete to “mango and bananas” or “mango and pineapple”? That’s a good question, and the only acceptable answer for why an LLM should choose bananas instead of pineapple (or coconut, or guava, or papaya...) is that it has determined that it’s the best fit for the intended meaning, tone, and sentiment of the text. Not because bananas is on the watermarking “green” list and pineapple is on the “red” list, even though pineapple might be the better fit. Google’s own supposedly jocular description of how SynthID works in fact captures how the scheme perverts the text it generates.
They’re saying you won’t notice because if it only chooses bananas over pineapple for these fingerprinting purposes, well, they’re both tropical fruits and who cares. But it’s utter nonsense that the difference is “not noticeable to the human eye”. The semantic difference between banana and pineapple is just as noticeable to the human eye as the taste of the two are to the human tongue.
If it did produce “My favorite tropical fruits are mango and airplanes”, it’d be incredibly stupid, but it wouldn’t be offensive because we’d all recognize that something completely off-key happened. What’s offensive is that with a system like SynthId in place, where the fingerprinting decisions are motivated by a secret key, we have no idea whether it completed to “mango and bananas” because bananas was determined to be the best next token, or because bananas is in the “green” bucket of words. It calls every single word choice into question.
Here’s a paper published in Nature where Google’s team behind SynthID published their work, after putting it into production with Gemini (née Bard):
We analysed approximately 20 million watermarked and unwatermarked responses and computed the thumbs-up and thumbs-down rates (both as a fraction of the total number of thumbs-up and thumbs-down feedback received). We found that the thumbs-up rate for the two models differed by 0.01% (with the watermarked model being higher); and the thumbs-down rate differed by 0.02% (with the watermarked model being lower). We found both of these differences to be statistically insignificant, and well within the 95% confidence intervals.
From this experiment, we conclude that over a wide variety of real chatbot interactions, the difference in response quality and utility, as judged by humans, is negligible. Subsequently, non-distortionary SynthID-Text has been productionized and is currently watermarking responses in Gemini and Gemini Advanced. To the best of our knowledge, this evaluation represents the first systematic watermarking investigation of its kind within a large-scale production system.
To this I say:
Gemini/Bard’s thumbs-up/thumbs-down buttons are not a good experiment for evaluating the effect on quality. If a chatbot tells me “My favorite tropical fruits are mango and bananas” instead of “mango and pineapple”, I’m not going to give the response a thumbs down because of the fruit it chose. I’d give it a thumbs down if it said “airplanes”, yes, but that’s a strawman. (The paper in Nature even uses “My favourite tropical fruit is ...” as an illustration, but in the paper, the only four next tokens considered are, in order of probability distribution, mango, lychee, papaya, and durian. No airplanes. And, conveniently, in the paper’s example, the “winner” of the watermarking “tournament” just happens to be mango, the one that would have been selected as the best if the watermarking weren’t in place.)
A “difference in response quality and utility, as judged by humans” that is “negligible” does not mean imperceptible. What they really mean is that it’s only slightly worse and that everyone is either too stupid to notice or too indifferent to care.
It’s widely considered that Gemini is behind ChatGPT and Claude in quality. Perhaps the fact that they’ve put SynthID-text into production is one of many reasons why. I personally agree that Gemini’s prose is inferior. Maybe the use of SynthID has nothing to do with the fact that I, along with the general public consensus, consider Gemini to be a second-rate chatbot — but in that case, maybe it’s the fact that Gemini is a second-rate chatbot that makes the difference “negligible” when Google started mixing in SynthID-motivated tokens in its results. It’s a lot more likely that your restaurant customers won’t notice that you replaced your regular coffee with Folgers Crystals if your regular coffee is second-rate to start with.
Now, finally, back to Anthropic’s new “How Claude’s Text Watermark Works”, published yesterday. I have some comments.
To summarize:
We use a method of watermarking that does not have any practical impact on the quality or content of Claude’s outputs;
The difference between watermarked and un-watermarked text will not be distinguishable to readers;
Translation: Specific words do not matter and we don’t think anyone reads anything closely.
- Nothing is added to the text and there are no hidden characters;
This would have been worth clarifying at the outset.
- Watermarking won’t be specific to Claude. As of August 2, the EU requires AI providers serving its market to mark AI-generated content. Other major model developers have signed the same Code of Practice and will be implementing their own watermarks.
No other AI provider has stated that they will apply such marking, adulterating all generated text, outside the EU.
Take the sentence “The weather today was cold and…”. The next word is very unlikely to be “sugary.” But it is quite likely to be “overcast” or “grey.” Under most circumstances, it doesn’t matter much to the reader which of these latter two words the model ultimately chooses — the meaning of the sentence is largely the same either way. In cases like this, the choice is settled by a random number.
Arguing that grey vs. overcast “doesn’t matter much to the reader” is the crux of my argument that this entire endeavor is a perverse adulteration of what it means to write — or to read. That it’s subtle in some ways makes it more perverse, because it’s sneaky.
In internal testing, we’ve seen no impact of watermarking on the content, level of creativity, or readability of Claude’s text. In the SynthID-Text paper, which introduced the technique we use, Google DeepMind tested this impact by serving a model that used watermarking to a portion of their Gemini traffic and comparing thumbs-up and thumbs-down ratings. They found no statistically significant differences from the unwatermarked model. And in a controlled study, human raters comparing watermarked and unwatermarked answers side-by-side saw no difference in quality.
See above for my argument that this thumbs-up/thumbs-down data is absolutely worthless in evaluating whether the SynthID-style word-bias watermarking makes text worse. By definition it must make text worse, unless the underlying LLM model’s scoring is wrong, because the nature of the watermarking algorithm requires it to sometimes increase the probability of selecting a worse word choice and decrease the probability of selecting the model’s best choice. It’s only a question of how much worse. What Google’s thumb-counting data shows is only that it isn’t so much worse as to make Gemini users click the thumbs-down button.
Watermarking doesn’t change the meaning or experience for the person reading it, but if you wanted to check after the fact whether the text was likely generated by Claude, the watermark allows you to do so.
No, it does not. Because the entire scheme is tied to secret keys held only by the AI provider, it only allows Anthropic, not “you”, to check anything.
When Claude proofreads text written by a person, what it gives back has generally only been lightly edited; because nearly all the words are the person’s, there’s very little (if anything) for the watermark to attach to. Depending on the length of the text and how heavily Claude has edited it, those changes might not be enough to make Claude’s involvement detectable. The more Claude writes, the more decisions it has to make, and the more space there is for a watermark.
Translation: No one can ever again use Claude for proofreading their own prose unless they’re willing to risk that the whole thing might be flagged as having been generated by Claude.
For example, once the model has written “2 + 2 =”, there is a very clear best choice for the next token (if the model is completing the sum, there isn’t an answer that’s equally as good as “4”; if it’s talking about George Orwell’s Nineteen Eighty-Four, there isn’t an answer that’s equally as good as “5”). The “nudge” of the watermark wouldn’t be applied here. For the same reason, code — which in very many cases has to be exact — has generally less watermarking than some other forms of text.
Having said that, in areas where there is an arbitrary choice between particular words or terms within the code, the watermark can be used, such as comments within code. But by definition, it will have a negligible effect on the actual code produced.
Translation: We value precision in programming code; we do not in prose.
And it is exceedingly rich to cite George Orwell’s Nineteen Eighty-Four, approvingly, in the context of justifying a text adulteration scheme premised on the notion that specific words do not matter. I mean what the actual fuck? Orwell!
Lastly, as to why they’re doing this:
We’re implementing watermarking to comply with the EU AI Act. Anthropic, along with several other major AI model providers and around 190 total signatories, signed the EU Code of Practice on Transparency of AI-Generated Content in July 2026. This requires AI system providers to use methods of “marking” AI-generated text. We’re applying watermarking globally at launch because we don’t yet have a durable way to scope it by region.
This, from a company that the Financial Times just reported is weeks away from an IPO with an intended valuation of $2 trillion, which would make it one of the 10 highest-valued companies in the world — as of today, placing it at #7, between TSMC ($2.2T) and Broadcom ($1.9T).
This leaves us to believe that one of the following must be true:
It’s perfectly reasonable that a technology company valued on par with Amazon and TSMC is technically incapable of complying with an EU regional law only within the EU itself.1 Not a cause for concern at all.
Anthropic is in over their heads, wields shockingly little control over their own tech stack, and their imminent IPO is likely to be remembered only as a new high-water mark in the manic global AI bubble.
Also, what happens if another major global market makes it unlawful for AI to secretly watermark generated text?
From an OpenAI support document titled “Provenance Signals (Content Credentials, SynthID) in OpenAI-Generated Content”:
Consistent with our commitments under the European Commission’s Code of Practice on Transparency of AI-generated content, our goal is to expand provenance signals to all modalities including text, so customers and developers have clear ways to meet their own transparency obligations as standards and tooling continue to mature.
There’s a lot of wiggle room in this brief statement, and it could just as well mean that OpenAI models will only adulterate text with fingerprint markers when users or developers ask for it. Or that it will only be mandatory for users in the EU. If I were at OpenAI I’d go hard on this and publicly say that ChatGPT will never watermark text it generates unless you ask it to, and that if you want tools that secretly work behind your back without telling you how they work to flag your words in ways you can’t see, go ahead and use Claude.
Three papers on ArXiv:
I will admit that while I’m profoundly offended by the idea of personally using tools that attempt to leave such watermarks in text they produce or touch, the mathematics behind it are fascinating.
Michael Lopp, at Rands in Repose, “RIP Claude”:
As a human who has had to wrangle with EU regulations in the past, I am abundantly clear what’s involved in the laborious bureaucratic process. I can guess what threats Anthropic is facing. However, this is a tone-deaf, clumsy, and alarming opening salvo in their watermark strategy. [...]
My writing is my work, and Anthropic’s current strategy is aggressively writer-hostile.
Jeff Gamet, “Anthropic’s Claude Watermark Is Akin to an AI Poison Pill”:
To be clear, the watermarking is embedded in pretty much any text Claude touches. Along with text Claude generates, it also applies to text it processes, such as proofreading and summarizing. I expect we’ll see too many inaccurate accusations of using Claude to write documents where the content was human-written, but AI-proofread.
The watermarking sticks with documents through copy-and-paste, too. Imagine copying text from a blog post or email only to have what you wrote tagged as potentially AI-generated. In fact, that could very well happen with this post. I personally write all of my content without AI tools, but I copied the quote at the top of this piece directly from Anthropic’s website. Does that mean what I wrote here will show as AI-generated? If they used their own models to generate or edit what I quoted, then the answer is very likely “yes.”
One of the papers published at ArXiv I cited above claims that such watermarking even persists when an article of text originally generated in English is translated into German.
Secrets are the poison here. When only Anthropic holds the secret keys that both produce the watermarking and perform the probabilistic detection of those marks, we’re all left to wonder. To wonder if what we’re reading is secretly watermarked, what we’re quoting is secretly watermarked, and whether what we ourselves are writing will be unjustly accused of being AI-generated based on secrets we don’t know and can’t see. Poisonous is exactly the right word.
Or should I say toxic? Or airplanes?
This is the side that noted savant Jim Cramer is on. ↩︎
Cory Weinberg, Jemima McEvoy, Jessica E. Lessin, and Stephanie Palazzolo, writing for the paywalled-without-gift-links The Information:
As Amodei has hopscotched the globe to preach about the potential and risks of AI — from New Delhi to Davos to Sun Valley — Clark has almost always been near his side. Several people who know the couple describe Clark as Amodei’s emotional ballast, someone he has sought counsel from during the turbulence of Anthropic’s growth and clashes with Washington over the future of AI.
Clark may well be the most consequential person within Anthropic’s orbit who doesn’t have a formal role at the startup. You might even call her the first lady of Anthropic. [...]
And before Female Algorithm Technologies, Clark tried her hand at building a pornographic film company, Eddice, that planned to make movies with high-production values and female protagonists. She also hoped it would have online shopping capabilities. In 2011, she tried to raise money for Eddice from Jeffrey Epstein, who’d pleaded guilty to charges of soliciting a minor for prostitution three years earlier, according to emails made public as part of the Epstein files released by the Department of Justice earlier this year. She emailed Epstein the script for the first of a series of films they hoped to make (title: “American Girl in Paris”).
“We thought you and the ladies might enjoy,” Clark wrote. She then appended a winking warning: “A little nsfw,” an acronym for “not safe for work.” Epstein didn’t invest.
While enormous attention has fallen on AI and the people behind the companies developing it, Clark has largely flown under the radar.
James Padolsey, on the Claude-text-watermarking-to-comply-with-an-EU-regulation imbroglio:
The same thought that led to this law could have applied to calculators at the time of their inception, had their outputs revealed themselves through artefacts. Thankfully, a sum borne of the brain is treated no differently from one produced by a calculator. Likewise with spellcheckers. To make assistance suspect only once the tool becomes capable enough to compose a whole sentence is not a principled boundary. It is a moral premium placed on difficulty itself.
Anthropic has nevertheless chosen a blanket, model-level implementation that appears broader than the law’s minimum requirement. That may be convenient compliance engineering, but it discards distinctions the law expressly attempted to preserve. The result is a signal broad enough to implicate harmless and assistive use, yet fragile enough to be removed by a motivated person through substantial recomposition. It risks concentrating suspicion on ordinary and assistive users while remaining weakest against deliberate deception.
Padolsey is the creator of Declaude, a delightfully simple web app that allows you to “Paste in AI-flavored text and get the same content back as plain prose”. Declaude’s original purpose is cleaning the cutesy Claude personality stink from text (whether it was created by Claude or any other LLM), but, if Anthropic persists in its stated plan to begin adulterating all text Claude generates, Declaude will also serve as a copy-paste single-extra-step way to eliminates those marks. Declaude is interesting and useful already, but goes to show how ill-considered this EU regulation is.
Padolsey also wrote and programmed a splendid interactive essay that explains and illustrated how the “watermarking” scheme Anthropic is adopting (and Google Gemini has already adopted) works.
The Wall Street Journal (gift link):
To help alleviate the supply crunch, Apple is looking to Chinese manufacturers.
“The Trump administration is not in favor of that,” Lutnick said in an interview after touring an Apple manufacturing facility in Houston. There have to be “other solutions to the memory issue, but it’s not great American companies using Chinese memory.”
Asked if he has relayed that message to Apple, Lutnick said “plainly.”
But:
U.S. government rules require American companies to secure a license before sharing product information with CXMT and YMTC. The Chinese chip makers would need such information if they were to make customized chips for Apple. But Apple is free to buy off-the-shelf parts from the companies, and to negotiate prices with them.
Apple’s chief operating officer, Sabih Khan, declined to confirm that Apple was testing Chinese memory chips in an interview Thursday. But he said that given the magnitude of the supply shortage, “we have to look at all options,” including working with existing suppliers to increase U.S. production.
XCancel:
XCancel is an instance of Nitter.
Nitter is a free and open source alternative Twitter front-end focused on privacy and performance. The source is available on GitHub at https://github.com/zedeus/nitter [...]
Using an instance of Nitter (hosted on a VPS for example), you can browse Twitter without JavaScript while retaining your privacy. In addition to respecting your privacy, Nitter is on average around 15 times lighter than Twitter, and in most cases serves pages faster (eg. timelines load 2-4× faster).
I personally don’t care for Elon Musk’s X-rebranded Twitter. I also realize that some of you outright despise it, and, worse, that under Musk, they make it difficult to view content if you’re not signed into an account or if you have JavaScript disabled. So, sometimes, when I link to posts or threads on Twitter/X, I include a link to the same content at XCancel.
XCancel isn’t perfect, alas. Just this week I linked to a tweet from Google Design that contained an animation of an Android app that they were inexplicably proud to show off. The XCancel cache of that tweet only showed it as animated to some people. Others just saw a static image. I don’t know why. That’s XCancel’s problem, not mine.
Also, it’s tiresome and repetitive to include an extra link to XCancel every time I link to a tweet on Twitter/X. If you would prefer to view x.com links on xcancel.com, you should automate the redirection. If you use Safari, XCancel Redirect is a free extension that claims to do just what you think it does. (I haven’t tried it.) Or you can use a general-purpose redirect extension. RedirectWeb is a good one that I can recommend. Also, Jeff Johnson’s excellent StopTheMadness — an extension I rely on for multiple purposes and frequently recommend. With StopTheMadness, you can redirect all x.com URLs to xcancel.com with the following rules (you need the leading and trailing slashes in the pattern to tell STM that it’s a regex pattern, not a plain text string):
Pattern: /https?://(www[.])?x[.]com/
Replacement: https://xcancel.com
Enabled on platforms: All
I understand not wanting to visit Twitter/X. But if you feel that way, I do sympathize, but that’s your decision, and you should use tools to redirect Twitter/X links automatically. And if you’re wondering why I don’t just refuse to link to anything on Twitter/X, that just isn’t practical. I wish it weren’t so, but many people, including senior Apple executives, and many organizations post interesting content exclusively to Twitter/X. If something posted to Twitter/X is also available elsewhere, I link to the elsewhere version. But if it’s exclusively available on Twitter/X and interesting, I link to the original.
My linking to original content on Twitter/X does not perpetuate Twitter/X’s relative popularity. Original content appearing exclusively on Twitter/X does.
My thanks to Drata for sponsoring last week at DF. Their message is short and sweet: Leverage autonomous AI agents to automate compliance, manage internal and third-party risk, and continuously prove your security posture.
Following up on my post yesterday about modern iPhones and scratch resistance, and my personal habits of (a) almost never using an iPhone case, and (b) never setting the iPhone face down except on soft (cloth) surfaces.
I should have anticipated this, because I’ve had this conversation with real-world normies repeatedly in recent years, but I got a bunch of questions from people who say they’re wary of ever putting their iPhone down on its back because they’re worried about scratching the camera lenses. That’s not a silly thing to worry about. It looks like those lenses are glass; glass scratches; and it sure seems like a scratched camera lens might forever ruin all the photos and videos you take with the camera.
There are a couple misconceptions here though. First, the glassy flat circles exposed on the back of your phone aren’t the camera lenses. Those are covers over the lenses, which are smaller, spherical (convex, not flat), and recessed. You can look through the covers and see the actual spherical lenses inside. The exposed lens covers are made of sapphire, not glass, and are thus incredibly scratch-resistant. They’re by far the most scratch-resistant parts of the phone. I’ve never once found even a tiny scratch on any of my iPhone camera lens covers, and I’ve never taken any particular care to avoid scratching them. I mean, I don’t drag the lenses face-down on surfaces. I’m not trying to scratch them. But I set my phones down on hard surfaces lenses-down all the time and they never seem to pick up even fine scratches.
Zack “JerryRigEverything” Nelson is the guy on YouTube who makes videos where he scratches the hell out of products to see how durable they are. (And he bends them, burns them, and abuses them in other gruesomely creative ways.) In his iPhone 17 Pro video, starting around the 4:10 mark, he tries scratching the sapphire lens covers with a sharp razor blade. No effect. Sapphire is far more likely to shatter than scratch, because it’s so hard. (You may recall that a decade ago Apple pursued using sapphire for the displays of iPhones but it didn’t work out.) Don’t try scratching your lenses with a diamond hardness pick and you’ll be fine.
Also, believe it or not, a fine scratch on the lens cover will almost certainly not affect image quality at all. Try sticking a strand of hair (which is probably much thicker than a typical scratch) to the surface of your iPhone 1× main camera lens. Take a picture. Clean the hair off the lens. Retake the same picture. You almost certainly won’t see any difference at all. Here’s a great video from DPReview back in 2020 showing just how much dust or scratching you need to put on the outside of a lens to degrade image quality. It’s counterintuitive but the surface of the lens is not where light is focused — the sensor, inside the camera, is.
This page from Google Design on their “Material 3” UI language came to my attention after my snarky post about the ungainly new to-do app they bizarrely bragged about on Twitter/X this week. I don’t think this “Material 3” page is new — I think it’s a few years old — but I’d never seen it before.
First, it’s crazy that (in desktop browsers with a mouse cursor) they change the I-beam cursor for text selection to ... a circle. I guess it looks kind of fun but the whole point of the I-beam cursor is to enable precise character selection. They changed that to something that makes precise selection as difficult as possible. It’d be better to just use an arrow pointer than to replace the I-beam with a goddamn circle. This, from the UI design team.
Second, this paragraph made me honestly wonder if I was reading a spoof, a years-old unfunny April Fool’s “prank”. But as far as I can tell this is totally straight, not a prank:
These factors can be quantified in users’ responses to new M3 Expressive designs. We found a 32% increase in subculture perception, which indicates that expressive design makes a brand feel more relevant and “in-the-know.” We also saw a 34% boost in modernity, making a brand feel fresh and forward-thinking. On top of that, there was a 30% jump in rebelliousness, suggesting that expressive design positions a brand as bold, innovative, and willing to break from convention.
Nothing says rebellious and “in-the-know” like assigning precise percentages to “rebelliousness”, “modernity”, and “subculture perception”.
Update: Believe it or not there is video of the Google Design team putting this research together. Worth a watch.
Philip Michaels, writing last September for Tom’s Guide:
iPhone 17 torture test videos done by JerryRigEverything indicate that Ceramic Shield 2 certainly resists scratching, with scratch testing leaving only light scratches at level 7 on the Mohs scale of hardness. Scratches typically show up on glass at levels 5 or 6 on that scale.
“Ceramic Shield 2 is indeed the best we’ve ever seen,” JerryRigEverything remarks in the iPhone 17 Pro testing video.
That backs up Apple’s own Ceramic Shield 2 video, in which a mineral tip can be seen rubbing against a Ceramic Shield 2-coated display. There’s residue left on the screen, but it’s material from that tip, as it wipes away fairly easily.
I mentioned yesterday that I try never to place my iPhones face down, to avoid scratches, and that my nearly year-old iPhone 17 Pro seemingly has not one visible scratch on the front glass. Not even a single micro abrasion I can see, even when tilting it with the display off to hunt for them. Surely Apple’s Ceramic Shield 2, which is intended to provide best-ever scratch resistance, is a factor too — and likely the biggest factor. Apple at its best.
A few years ago I’d have looked at this post and maybe leaned toward the idea that a precocious 8th grader somewhere hacked into the @GoogleDesign Twitter account and tried to pass off their little to-do app as having come from Google’s design team. But this is apparently real. I almost hope it’s AI slop and that there aren’t any human designers there who think anything in this app has appropriate proportions or is aesthetically pleasing.
(Here’s an XCancel link for the x.com averse. You can just change any x.com/* URL to xcancel.com/*. But, alas, the XCancel mirror of the original tweet loses the animation showing the UI in motion — for some, but not all, users.)
Joanna Stern, writing at The New Things (gift link):
I used to love the blinking notification light on my BlackBerry, and later my Droid 2. It was a simple way to know I had a message without actually looking at my messages. Then BlackBerry let you customize the color, and it was a rainbow dream.
Google’s HiLight takes it a step further by letting you assign different colors to VIP contacts. So when your phone is face down, you can tell who’s trying to reach you without picking it up. It looks cool. The big bummer? At launch, it only works for phone calls.
Google told me it’s “continuing to invest in this technology” and that messaging notifications are coming. Still, it’s odd they aren’t there at launch, especially since Google found plenty of colors for Gemini. The ring changes hue depending on whether the AI is listening, thinking or responding.
I have no interest in this because a few years ago I stopped ever putting my iPhone on any surface face down. I don’t use a case, so setting the phone face down is a good way to pick up scratches. I stopped doing that, ever, and now I’ve got a nearly year-old iPhone 17 Pro that seemingly doesn’t have even a micro abrasion on the display glass.
And it seems downright goofy that it’s launching with support only for phone calls. Google can’t even launch a blinking light right on these Pixel phones.
David Imel, The Verge:
But in our current moment, the photos people are drawn to are not flawless — and that’s created a real problem for the people making smartphone cameras. “The gap between what two random people want from their camera is growing dramatically,” says Isaac Reynolds, who leads the Pixel camera team at Google. Some people want a perfectly optimized photo, Reynolds says. “But there’s a growing number of people who want something that they feel is more authentic or traditional, by their definition.”
That’s what Reynolds and the Pixel team are setting out to solve with Camera Looks, a total rethink of how the camera captures and styles an image on the Pixel 11 series. The phones include a new suite of processing tools aimed at giving people much more control over the camera’s output. At first glance, they look a lot like filters, with names like Black Tie, Classic, and most strikingly, Digi, in reference to the growing trend of buying cheap digital cameras from the early aughts. But unlike a filter, which generally applies editing tweaks after the processing happens, Camera Looks is changing how image data is processed at the sensor level.
This sounds a lot like Apple’s “Photographic Styles” in the iPhone 13 and later, but with the addition of some control over grain and sharpness. I love grain, and the thing I like least about Apple’s built-in photo processing is the way it attempts to eliminate all noise. I like digital cameras that embrace noise and try to make it pleasing, rather than over-processing to try to pretend these tiny sensors aren’t inherently noisy. So I say kudos to Google for loosening up on “everyone gets the over-sharpened, over-smoothed look”.
But, that said, what’s made me happy over the last year is shooting almost all my photos with third-party apps — in particular, Halide, Not Boring Camera, and Analogue — that give me low-level control over processing through LUTs (or in Halide’s parlance, “looks”). These apps give me the looks I want just through point-and-shoot photography. No need to process images after shooting. But because they (optionally) shoot RAW, I can reprocess in post without losing any original sensor data. I can shoot in black-and-white and subsequently change my mind and reprocess with a color look later. These apps enable honest film-look photography without looking at all gimmicky, and without baking any sort of filtering into the image data in my camera roll. It’s the best of both worlds. It’s also way, way fussier than the built-in Camera app should be.
Anyway, all this Pixel announcement news today has me thinking that these days, I never hear much about Pixel cameras or image processing. In the early years of the Pixel lineup, there was a common view that Pixel phones were the best camera phones in the world. I never believed that, but I did think they were an interesting alternative, taking different but equally valid philosophical choices than Apple. Apple has always insisted on instant processing. You snap the shutter and the iPhone saves your image instantly. Pixel cameras often undertook seconds-long processing. That enabled features the iPhone couldn’t offer.
I think it’s pretty easy to pinpoint when things changed for Pixel — when Marc Levoy left Google for Adobe in 2020 (where he now leads the team making the very intriguing Indigo camera app, that is exclusively available for iPhone and which offers a slew of innovative computational photography features).
When Levoy was leading Google’s camera photography team, there was a credible argument that Pixel phones were the best still camera phones in the world. (They were never even close on video.) Ever since, I never hear about them except in the middle of August each year, when Google announces new Pixel hardware. See you in 12 months.
Ivan Mehta, TechCrunch:
With this year’s Pixel 11 series launch, Google is thinking more agentic AI to complete your tasks. With Gemini, U.S.-based users will be able to order groceries, book rides, or get coffee, for instance. Plus, Gemini can call businesses on users’ behalf for table reservations or appointments. Google said that users can take over or stop tasks at any point in time and also review transcripts for AI-operated calls.
Over the next few weeks, a new slate of connected apps will be able to work with Gemini to get things done, including Granola, Otter.ai, Wix, Fever, Get Your Guide, Localiza, OpenTable (U.K.), Ticketmaster, iHeartRadio, Pandora, Angi, Thumbtack, and Zocdoc.
I’ve never heard of most of those apps, or, if I have heard of them, I thought they went defunct years ago. The one app amongst those I do use, OpenTable, apparently only supports this Gemini integration in the U.K.?
Another knockout year for the Pixel phones and Android.
Sam Rutherford, writing for Engadget:
Granted, the P11 Pro Fold still includes IP68 dust and water resistance, which is better than what you get on the Z Fold 8 line (IP48). But after making an absolute tank of a foldable phone with last year’s model, I said Google really needed to cut weight and thickness this generation, and it just hasn’t. Instead, it seems like Google has leaned even more into the Pro Fold’s sturdiness with a new glass fibre material for its rear panel that the company says is nearly impossible to crack. This is great in theory, but obviously I wasn’t in a position to smash one of Google’s demo units in order to test that claim.
I have never seen anyone, anywhere, with a folding Pixel phone. Pretty sure I won’t this year either.
I played with a Pixel 10 Fold or whatever its name was last year at the Google Store on Newbury street in Boston, which store felt like a small library or museum. The Pixel Fold display was so plasticky that I had a moment where I wondered if the display units were fake props, like at Ikea. But no, it was the real thing. And it cost $1,800 to start.
Victoria Song, The Verge:
The $399 Google Pixel Watch 5 isn’t about the hardware. Sure, there’s a new satin pyrite case finish, a few new strap colors, and a Steph Curry Special Edition. Under the hood, there’s a slightly faster Qualcomm processor and an itty-bitty battery bump. There’s a $50 price hike from last year, too, because the Pixel Watch 5 isn’t immune to RAMageddon — none of us are. Otherwise, no one would blame you for looking at this watch and thinking absolutely nothing’s changed. That’s because the big updates this year are all software-based.
I have never seen anyone in real life wearing a Pixel Watch. (I’ve only seen a handful of people using Pixel phones, and every single one of them was at a tech event of some sort.)
You can also use AI to generate watchfaces. I created a watchface featuring “fat cats and an elegant font.” It didn’t quite do that. The fat cats weren’t characters, so much as pleasantly plump, cat-shaped clock numbers. It’s a gimmicky way to shoehorn the Nano Banana model onto the watch… but if that’s your thing, it works.
Well I’m sure now I’ll see plenty of people wearing Pixel Watches.
Mia Sato, The Verge:
Earlier this summer, Amazon customers began noticing that emails related to their online orders looked sparse: Order confirmation emails didn’t name specific items anymore, and instead listed only item categories.
“Your Beauty item is confirmed!” an email about my retainer cleaning tablets read. Shoppers have posted other iterations of the redacted emails as well: “Ordered: 1 Hardware item,” “Your Drugstore, Shoes, and other items are here!” and “1 Nutrition & Wellness, 1 Wireless Accessories,” for example. The emails have clip art-style illustrations of general product categories, and a shopper has to exit their email and go to Amazon to see what item it’s actually referring to.
Complaints about the vague emails appear to have started in July; orders placed as recently as June contain thumbnails and names of the exact item purchased. [...]
“As customers shop with us more frequently, including on their phones, and use the ‘Your Orders’ page in the Amazon app to get real-time, consolidated order details and delivery status, we’ve simplified several order-related emails to direct customers to our app and website for the latest information on their orders,” spokesperson Maxine Tagay said in an email. “This also reduces customer information shared outside the Amazon app and website to further improve customer privacy.”
That statement from Amazon is almost honest, if you squint at it and ignore the bit about it having anything to do with “simplification”. If you order a Torx screwdriver, what’s simple is getting a confirmation email with the exact name, brand, and photo of the item you ordered, not “a hardware item”. The honest part is “reduces customer information shared outside the Amazon app”.
The fact is, a lot of people use apps like Gmail for email. In fact, a lot of people use Gmail in particular. There’s a good chance you do. When you use Gmail, Google learns everything in your email. Amazon sees this as a competitive problem for them, not a privacy problem for users. Everyone who uses Gmail basically knows that Google “knows” what’s in their email, and they’re voluntarily agreeing to that. Amazon just doesn’t like that any of its competitors — perhaps especially Google — can glean so much information about what people are buying from Amazon. So Amazon, in an act of competitive spite that corrupts its relationship with its own customers, is limiting all useful purchase information to channels that it controls — its website and its own apps.
If customers relied solely on the “Your Orders” page at Amazon, they wouldn’t need to send any emails at all. The obvious fact is, people go to Amazon when they want to buy something. They don’t come back until they want to buy something else. They enjoy getting confirmation and shipping notifications in email because they check their email regularly.