OpenAI has simply been launched GPT-5.2is probably the most superior Frontier mannequin for skilled work and long-term company, deployed throughout ChatGPT and APIs.
GPT-5.2 is a household of three variants. In ChatGPT, customers see ChatGPT-5.2 Prompt, Pondering, and Professional. Within the API, the corresponding fashions are: gpt-5.2-chat-latest, gpt-5.2and gpt-5.2-pro. Prompt targets on a regular basis help and studying, Pondering targets complicated multi-step duties and brokers, and Professional allocates extra computing to tough technical and analytical duties.
Benchmark profiles from GDPval to SWE bench
GPT-5.2 Pondering is positioned because the mainstay of real-world data work. GDPval, an evaluation of well-specified data duties throughout 44 occupations in 9 massive industries, outperforms or ties industry-leading consultants in 70.9% of comparisons, producing outcomes greater than 11 occasions quicker and at lower than 1% of the knowledgeable’s estimated price. For engineering groups, which means fashions can observe structured directions and reliably produce deliverables reminiscent of shows, spreadsheets, schedules, and diagrams.
For the Junior Funding Banking Spreadsheet Modeling job inner benchmark, common scores elevated from 59.1 p.c for GPT-5.1 to 68.4 p.c for GPT-5.2 Pondering and 71.7 p.c for GPT-5.2 Professional. These duties embody a three-statement mannequin and a leveraged buyout mannequin with formatting and quotation constraints, that are consultant of many structured enterprise workflows.
In Software program Engineering, GPT-5.2 Pondering reaches 55.6 p.c in SWE-Bench Professional and 80.0 p.c in SWE-bench Verified. SWE-Bench Professional evaluates repository-level patch technology throughout a number of languages, whereas SWE-bench Verified focuses on Python.
Lengthy contexts and agent workflows
Lengthy context is a core design purpose. GPT-5.2 Pondering establishes a brand new cutting-edge expertise for OpenAI MRCRv2. This can be a benchmark that inserts a number of an identical “needle” queries right into a “haystack” of lengthy interactions and measures whether or not the mannequin can reproduce the right reply. That is the primary mannequin reported to achieve almost 100% accuracy on a four-hand MRCR variant as much as 256,000 tokens.
For workloads past that context, GPT-5.2 Pondering integrates with Responses. /compact Endpoint. Carry out context compaction to increase the lifetime of tool-intensive, long-running jobs. That is related if you’re constructing an agent that calls instruments repeatedly over many steps and desires to take care of state past the uncooked token restrict.
By way of instrument utilization, GPT-5.2 Pondering reached 98.7% on Tau2 Bench Telecom, a multi-turn buyer assist benchmark the place the mannequin should coordinate instrument calls throughout a sensible workflow. Official examples since OpenAI’s launch present eventualities reminiscent of delayed flights, missed connections, misplaced luggage, and vacationers with medical seat necessities, the place GPT-5.2 manages reservation adjustments, particular seats, and compensation in a constant order, whereas GPT-5.1 leaves steps unfinished.
imaginative and prescient, science, arithmetic
It additionally improves the standard of your imaginative and prescient. GPT-5.2 pondering roughly halves the error charge of chart reasoning and person interface understanding benchmarks reminiscent of CharXiv Reasoning and ScreenSpot Professional when Python instruments are enabled. This mannequin reveals improved spatial understanding of photos. For instance, when labeling motherboard parts with approximate bounding containers, GPT-5.2 identifies extra areas with a tighter association than GPT-5.1.
For scientific workloads, GPT-5.2 Professional scored 93.2 p.c and GPT-5.2 Pondering 92.4 p.c on GPQA Diamond, and GPT-5.2 Pondering solved 40.3 p.c of FrontierMath Tier 1 to Tier 3 issues with Python instruments enabled. These benchmarks cowl graduate-level physics, chemistry, biology, {and professional} arithmetic, and OpenAI highlights early makes use of of GPT-5.2 Professional to assist show statistical studying theories underneath human validation.
Comparability desk
| mannequin | Major positioning | Context window / most output | data blockage | Noteworthy benchmarks (Pondering / Professional vs GPT-5.1 Pondering) |
|---|---|---|---|---|
| GPT-5.1 | The flagship mannequin for coding and agent duties with configurable reasoning duties | 400,000 token contexts, as much as 128,000 outputs | 2024-09-30 | SWE-Bench Professional 50.8 p.c, SWE Bench Validated 76.3 p.c, ARC-AGI-1 72.8 p.c, ARC-AGI-2 17.6 p.c |
| GPT-5.2 (thought) | New flagship mannequin for industry-wide coding and agent duties and long-running brokers | 400,000 token contexts, as much as 128,000 outputs | 2025-08-31 | GDPval wins or attracts 70.9 p.c in opposition to {industry} consultants, SWE-Bench Professional 55.6 p.c, SWE-bench Verified 80.0 p.c, ARC-AGI-1 86.2 p.c, ARC-AGI-2 52.9 p.c |
| GPT-5.2 Professional | The superior computing model of GPT-5.2 for probably the most difficult inference and scientific workloads produces smarter, extra correct responses. | 400,000 token contexts, as much as 128,000 outputs | 2025-08-31 | GPQA Diamond 93.2 p.c vs. 92.4 p.c for GPT-5.2 pondering, 88.1 p.c for GPT-5.1 pondering, 90.5 p.c for ARC-AGI-1, and 54.2 p.c for ARC-AGI-2 |
Necessary factors
- GPT-5.2 Pondering is the brand new default workhorse mannequin: Replaces GPT-5.1 Pondering as the primary mannequin for coding, data work, and brokers, delivering considerably larger benchmark efficiency throughout GDPval, SWE-Bench, ARC-AGI, and Scientific QA whereas sustaining the identical 400k context and 128k most output.
- Considerably higher accuracy than GPT-5.1 at related scales: In the primary benchmarks, GPT-5.2 Pondering will increase from 50.8 p.c to 55.6 p.c on SWE-Bench Professional, from 76.3 p.c to 80.0 p.c on SWE-bench Verified, from 72.8 p.c to 86.2 p.c on ARC-AGI-1, and from 17.6 p.c to 52.9 p.c on ARC-AGI-2 whereas retaining token limits related. rose to %.
- GPT-5.2 Professional targets high-end reasoning and science: GPT-5.2 Professional is a complicated computing variant that primarily improves tough reasoning and scientific duties, for instance, it scores larger within the ARC-AGI tier, reaching 92.4 p.c in GPT-5.2 Pondering and 88.1 p.c in GPT-5.1 Pondering, in comparison with 93.2 p.c in GPQA Diamond.
Asif Razzaq is the CEO of Marktechpost Media Inc. As a visionary entrepreneur and engineer, Asif is dedicated to harnessing the potential of synthetic intelligence for social good. His newest endeavor is the launch of Marktechpost, a synthetic intelligence media platform. It stands out for its thorough protection of machine studying and deep studying information, which is technically sound and simply understood by a large viewers. The platform boasts over 2 million views per 30 days, demonstrating its recognition amongst viewers.


