With the Navier–Stokes Millennium Prize Problem apparently solved, and potential solutions to one or two of the remaining problems on the horizon, it’s an understatement to say AI is good at math. AI has gotten so good, it has even called into doubt if math a profession can survive, with a collective of Field Medalists discouraging its use in a recent open letter.
This leaves writing as the final frontier. Interestingly, when Chat GPT launched, and as recently as early 2025, its math skills were poor compared to writing, but now this has reversed. AI has never quite gotten to the level of producing prose that doesn’t, well, sound like AI, which is the ultimate test.
A successful implementation would be something like this: you input a writing sample (longer is better), a target wordcount, and a style to emulate, from which it does the background research and fills in the gaps by inferring what you meant, and adhering to the wordcount and style. It can do this with math: I have tested this firsthand by having it fill in the gaps on a math paper I am working on, with great success. The downside is the paper appears obviously AI-generated, but this is acceptable, as correctness matters more than style with math. With writing, however, the use of AI invites complaints about “slop”.
So the ultimate goal is to produce an output which passes as human. This is hard, because human writing tends to have certain tics or idiosyncrasies. For example, in my post, “The sad truth about metabolism: why nothing works,” I mention “shit genetics” as a reason for some people remaining obese. No AI would ever say that unless otherwise specifically told to. Its training discourages both profanity and subjective judgments about genetics as “good” or “bad.”
The use of metaphor also makes human writing unique. In another post, I liken Twitter’s paid blue checkmark to a participation trophy. It’s not obvious how to make the connection between the two, since participation trophies are often used in the context of sports. But they both represent a low barrier of entry, in which participation or a inexpensive monthly fee create a false sense of exclusivity.
Or the following passage from another metabolism-related post, “Sure, ‘too much NEAT’ or ‘too much waste-heat’ will not earn one a Heisman trophy, but not being ridiculed on social media for having a fat profile pic is nice too. You cannot see waste heat energy like replaying an NFL highlight reel, but it’s every bit as genetic.”
No AI is ever going to come up with something like that. It connects seemingly unrelated ideas: metabolism, early athletic achievement, and the indignity of being mocked for your weight on social media, in which your achievements offer no protection. Those connections were obvious to me when I wrote it, but I doubt an AI would spontaneously make the same associations. Or contrasting invisible “waste heat energy” with a visual football highlight reel.
Overall, the inability to make non-obvious connections is why AI-generated writing all begins to sound the same once you read enough of it. It is competent in terms of the grammar and content–much as AI is competent at math–but it fails to make the non-obvious associations that makes human writing unique.