AI Policy Wiki
Dashboard

GPT-4 Technical Report (OpenAI, arXiv 2303.08774, March 2023)

high confidence · updated 2026-07-31

OpenAI's technical report for GPT-4 — a multimodal Transformer accepting image and text inputs, reported at human-level performance on professional and academic benchmarks including a simulated bar exam around the top 10% of test takers, and notable for the predictable-scaling claim that some GPT-4 performance was forecast from models trained with under 1/1,000th the compute.

This is OpenAI's technical report accompanying GPT-4, submitted to arXiv on 15 March 2023 (arXiv:2303.08774, cs.CL) and last revised 4 March 2024 as version 6. It is authored institutionally as "OpenAI" followed by a roster of roughly 283 named contributors, and runs to 100 pages; the version-6 comment records the update as "updated authors list; fixed author names and added citation."

Reported claims

The abstract states that GPT-4 is "a large-scale, multimodal model which can accept image and text inputs and produce text outputs," and that while "less capable than humans in many real-world scenarios, GPT-4 exhibits human-level performance on various professional and academic benchmarks, including passing a simulated bar exam with a score around the top 10% of test takers."

Architecturally it is described only as "a Transformer-based model pre-trained to predict the next token in a document," with a post-training alignment process that "results in improved performance on measures of factuality and adherence to desired behavior." No parameter count, training-data composition, or compute figure is disclosed — an omission that became the reference case in the subsequent argument over what a frontier "technical report" is obliged to contain.

The predictable-scaling claim

The abstract identifies as "a core component of this project" the development of "infrastructure and optimization methods that behave predictably across a wide range of scales," and states that this "allowed us to accurately predict some aspects of GPT-4's performance based on models trained with no more than 1/1,000th the compute of GPT-4." This is the report's most consequential methodological claim for governance: if frontier capability can be forecast from small-scale runs, then capability thresholds are in principle knowable before a training run completes, which is the premise underlying compute-threshold rules in EU AI Act (Regulation 2024/1689) and the capability-threshold structure of frontier safety frameworks. See Scaling Laws.

Position in the record

GPT-4 anchors the 2023 capability wave. It arrived roughly three and a half months after the ChatGPT launch, and its combination of headline professional-exam results with withheld architectural detail set the template that later system cards both followed and were measured against. The model page is at GPT-4 Family (OpenAI).

Provenance

Retrieved July 31, 2026 from the arXiv abstract page using the vault's bin/fetch-source.py; Firecrawl was rate-limited during this cycle. Title, arXiv identifier, submission and revision dates, subject classification, author roster, page count and abstract were captured from that page. This summary rests on the abstract and citation metadata; the 100-page body, including the benchmark tables and the system-card appendix, was not retrieved in this pass.

Relationships