>_ ANALYSIS
GPT-5.2’s benchmark jump matters because delegation now has a better chance of paying off
GPT-5.2 appears to change the economics of AI-assisted expert work. Ethan Mollick says GPT-5.2 Thinking and Pro tied or beat human experts on GDPval 72% of the time, and that matters most for managers and specialists deciding when AI is worth the review overhead.
GPT-5.2 appears to change the economics of AI-assisted expert work. Ethan Mollick says GPT-5.2 Thinking and Pro tied or beat human experts on GDPval 72% of the time, and that matters most for managers and specialists deciding when AI is worth the review overhead.
The mechanism is simple: if a task is slow for a human but fast for the model, the deciding factor is no longer raw generation speed. It is the probability that the first draft clears your bar after review. That is why a stronger win-or-tie rate can make delegation rational on work that still looks messy, because the savings come from reducing baseline labor, not eliminating supervision.
But the evidence here is narrow. Mollick is citing GDPval in commentary, not presenting the benchmark paper itself, so the 72% figure should be treated as an attributed claim rather than a fully audited result. The underlying task mix, judging rules, and failure modes are not described in the passage, which matters because benchmark gains often hide large variation across tasks.
The practical takeaway is not “AI now replaces experts.” It is more specific: for work with a long human baseline and a tolerable review burden, GPT-5.2 may be good enough to shift the default from doing everything by hand to trying, checking, and retrying. For short, high-stakes, or tightly specified tasks, the overhead can still erase the advantage. The right test now is not whether the model can help at all, but whether the review cost stays below the time it saves.
Source: https://www.oneusefulthing.org/p/management-as-ai-superpower
