"Sign of the future: GPT-5.5" is an essay by Ethan Mollick (Wharton; author of Co-Intelligence; writer of the One Useful Thing newsletter), published April 23, 2026 on oneusefulthing.org. Drawing on early API access to GPT-5.5, it argues that AI capability is best understood across three simultaneously advancing layers — models, apps, and harnesses — rather than reduced to "the model," and uses hands-on testing of GPT-5.5 to illustrate both rapid gains and persistent irregularity in what current systems can do. Mollick states he takes no money from any AI lab and did not share the post with OpenAI in advance.
Models, apps, and harnesses
The essay's central argument is that three things are advancing at the same time and should be tracked separately to understand AI capability:
- Models — the reasoning engine; Mollick's examples include Opus 4.7, Gemini 3.1, and GPT-5.5.
- Apps — the products built around the model, such as chatgpt.com and claude.ai, and increasingly desktop apps including Claude Code, Claude Cowork, and OpenAI Codex.
- Harnesses — the tools a model is given access to and how, including image generation, code execution, file access, web browsing, and computer use.
Mollick writes: "The real magic happens when you combine harnesses, apps, and models on a real problem." The essay presents this three-layer framework in a condensed form, arguing that subsequent reporting on capability gains needs to disentangle which layer is doing the work.
What GPT-5.5 demonstrates
Mollick describes GPT-5.5 Pro as "plain good" and reports several hands-on tests:
- Code generation. Asked to "build me a procedurally generated 3D simulation showing the evolution of a harbor town from 3000 BCE to 3000 AD," GPT-5.5 Pro was the only model in his test to actually model evolution; earlier models generated new building replacements over time rather than evolving the town. GPT-5.5 Pro completed the task in 20 minutes, against 33 minutes for GPT-5.4 Pro.
- Image generation. A new "gpt-imagegen-2" model rendered high-quality text in images and arbitrary scenes. Mollick's "Otter Test" — an otter on a plane using wifi — was generated reliably; the model also produced a fake academic paper on the Otter Test and a gallery of otter images styled after Klimt, Rothko, Matisse, Monet, Picasso, Titian, Rembrandt, and O'Keeffe.
- Multi-step research workflow. Mollick fed Codex (GPT-5.5-backed) hundreds of anonymized crowdfunding survey and data files (STATA, CSV, XLS, Word) with four prompts: "Help me sort it out and generate a new hypothesis…test it in sophisticated ways and write an academic paper." He called the result "very impressive, especially after I asked GPT-5.5 Pro to comment on the paper and fed those results back into Codex," adding that he "would have been very happy if this paper was the outcome of a 2nd year PhD project."
- Tabletop game generation. Codex one-prompted a complete D&D-style fantasy game, including a 101-page illustrated PDF. Mollick: "The setting is interesting and novel, and the rules appear to make sense."
The jagged frontier
Alongside these gains, Mollick describes capabilities that remain weak, illustrating the irregular "jagged frontier" (Jagged Frontier):
- Long-form fiction has the same problems as prior generations: "a love of the uncanny; overly complex ideas that do not fully pay off; weird metaphors ('weather and architecture are the same argument at different speeds'); too many ornate sentences…dialogue where every character speaks in the same clipped tone; and the name 'Mara.'"
- In the academic-paper workflow, the literature review and statistics were real, but the hypothesis was "not that interesting and there are some standard concerns about causation."
Context
GPT-5.5 was OpenAI's entry in a model release cycle that included Anthropic's Claude Opus 4.5 (late 2025), Opus 4.6, and Opus 4.7 (1M context, 2026). The essay is Mollick's first hands-on assessment of GPT-5.5 (GPT-5.5 ('Spud')). Mollick's earlier framing of the shift from "chatbots that say things" to "agents that do things" is cited as inflection-point shorthand in So, About That AI Bubble — Rogé Karma (The Atlantic, May 2026).
Mollick writes on enterprise and educational AI adoption, and his essays are read by Fortune 500 CIOs, university faculty, journalists, and policy researchers; his track record on capability inflection points is documented on the Ethan Mollick page.
Provenance and caveats
Mollick disclosed his early API access and that he received no compensation; the essay's hands-on impressions represent one researcher's assessment rather than benchmark results. The "near-PhD quality" characterization is Mollick's own, and academic-paper quality is hard to benchmark: the essay supports the statement that GPT-5.5 plus Codex produced material Mollick assessed as PhD-quality, not that it produces PhD-quality papers as an established fact. The 3D-simulation evolution task is a single test, and the 20-versus-33-minute speedup figures are not statistical. The finding that long-form AI fiction remains poor has persisted across model generations.
Relationships
- supports: GPT-5.5 ('Spud') — primary hands-on capability assessment
- supports: Jagged Frontier — Mollick's coined term; reinforced and extended here
- supports: AI Coding Agents — Codex/GPT-5.5 workflow demonstration
- supports: Ethan Mollick — adds to the body of Mollick essays
- related: So, About That AI Bubble — Rogé Karma (The Atlantic, May 2026) — Karma quotes Mollick's "say things vs. do things" framing as the inflection-point shorthand
- related: Agentic AI — the three-layer (model/app/harness) framework is conceptual scaffolding for understanding agentic capability gains