GPT-4: Multimodal LLMs Reach Human-Level Performance

Michael BrenndoerferAugust 18, 20258 min read

Part of History of Language AI

Covers GPT-4, including multimodal capabilities, improved reasoning abilities, enhanced safety and alignment, human-level performance on standardized tests.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

2023: GPT-4

OpenAI released GPT-4 in March 2023 as a model that accepted text and image inputs and produced text. Its technical report emphasized benchmark gains over GPT-3.5, including reported scores around the 90th percentile on the Uniform Bar Exam and 88th percentile on the LSAT. These results measured performance on specific test sets; they did not establish general reasoning ability or professional competence.

GPT-3 and GPT-3.5 had shown that one pretrained model could perform many tasks from instructions or examples in the prompt. They also produced inconsistent answers, hallucinated facts, and accepted text rather than images. GPT-4 was presented as an attempt to improve benchmark performance, instruction following, and safety behavior while adding image input.

OpenAI reported using supervised fine-tuning, reinforcement learning from human feedback (RLHF), adversarial testing, and additional safety mitigations. These methods improved measured behavior but did not eliminate hallucinations, harmful outputs, or performance variation across tasks.

The release drew attention to multimodal input, stronger benchmark results, and systematic safety evaluation in a general-purpose language model. It was not evidence that artificial general intelligence had been achieved.

The Problem

GPT-3 and GPT-3.5 could answer one problem correctly and fail on a nearby variant. This sensitivity to wording and sampling made their outputs difficult to rely on without verification, especially in professional or educational settings.

Multi-step mathematics and logical deduction were common failure modes. Producing a plausible chain of text did not guarantee that every step was valid or that the final answer followed from the premises.

Earlier GPT interfaces accepted text but not images. That excluded tasks in which the prompt depended on a chart, diagram, photograph, or other visual source.

Safety and alignment presented ongoing challenges. While GPT-3.5 incorporated reinforcement learning from human feedback (RLHF) to align behavior with human preferences, the models still occasionally produced harmful outputs, refused reasonable requests, or provided incorrect information confidently. The balance between being helpful and avoiding harm was difficult to achieve, and models sometimes erred on both sides: refusing legitimate requests while occasionally producing problematic content.

GPT-3.5 also scored below GPT-4 on several standardized tests reported in the GPT-4 technical report. Exam scores offered a reproducible comparison, though they covered only selected tasks and prompting conditions.

The training process itself posed challenges. Scaling language models required massive computational resources, careful data curation, and sophisticated training techniques. Finding the right balance of data quality, model scale, and training objectives to produce both capable and safe models was an ongoing research challenge. The field needed better methods for training models that were simultaneously more capable, more reliable, and better aligned with human values.

The Solution

OpenAI disclosed few architectural or dataset details for GPT-4. The public technical report focused on capabilities and predictable scaling, with separate sections on evaluation and safety mitigations.

GPT-4 accepted prompts containing text and images, then generated text responses. Demonstrations included questions about charts, diagrams, and photographs. OpenAI did not publish enough architecture detail to specify how visual representations were connected to the language model.

GPT-4 outperformed GPT-3.5 on many reported exams and benchmarks, including tasks involving mathematics and logical questions. It still made reasoning errors, and the public report did not attribute the gains to a specific architectural change.

The reported post-training process included supervised fine-tuning and RLHF, followed by safety evaluation and mitigation work. These steps changed response behavior but could not guarantee reliability or alignment with every user's values.

Training and Safety Process

The development process targeted behavior often summarized as helpful, harmless, and honest. OpenAI used red-team testing, safety evaluations, post-training, and deployment mitigations to reduce harmful responses. The technical report also documented remaining risks and known failure modes.

OpenAI reported GPT-4 around the 90th percentile on the Uniform Bar Exam and 88th percentile on the LSAT, along with gains on other exams. Percentile comparisons depend on the test version, prompting method, and scoring procedure; they do not measure all skills needed in professional practice.

Image input let GPT-4 answer questions about visual material or combine it with text. The feature broadened the kinds of prompts the model could accept, while retaining familiar failure modes such as incorrect descriptions and unsupported inferences.

These capabilities supported experiments in tutoring, writing assistance, software development, document analysis, and other applications. Each use still required task-specific evaluation and safeguards rather than assuming benchmark gains would transfer directly.

Applications and Impact

Educational applications used GPT-4 in tutoring interfaces to explain material and generate practice. They also used it for feedback. Its fluent answers were useful for interaction, but factual and reasoning errors meant that students and teachers still needed verification.

Professional organizations also tested the model for drafting, research support, and training. Exam performance motivated these trials but did not establish accuracy on confidential, domain-specific work.

Writers and other content teams used GPT-4 to generate ideas, draft passages, and revise text. These workflows treated model output as material for review rather than finished work.

Image input supported experiments with scientific figures and document screenshots, including charts. The model could describe or extract visible information, but its answers required checking against the source image.

Organizations could adapt one hosted model to several workflows through prompts, retrieval, tools, or fine-tuning rather than train a model from scratch for each task. This reduced some development costs while adding evaluation and data-handling work. Latency and vendor dependence created separate operational concerns.

GPT-4's reported exam and benchmark scores became common comparison points for later model releases. Its architecture and data were not disclosed in detail, nor was its training compute, so other projects could compare outputs more readily than reproduce the system.

The release also placed red-team findings, safety evaluations, and deployment mitigations alongside capability results. Publishing these limitations did not make the model safe in every context, but it made risk evaluation a visible part of the release process.

Limitations

GPT-4 was costly to train and, for high-volume applications, costly to serve through an API. OpenAI did not disclose the full training compute, but access to comparable development resources was limited to well-funded organizations.

Biases in web and corpus data could appear in GPT-4's outputs and affect downstream applications. Mitigation required targeted evaluation, changes to data or post-training, and monitoring in the intended deployment context.

GPT-4 still failed on logical deduction, symbolic manipulation, and multi-step problems. A detailed explanation could contain an invalid intermediate step even when its prose sounded confident.

Image input also had failure modes. Small details, spatial relations, unusual diagrams, and ambiguous images could produce incorrect descriptions or inferences.

The context window bounded how much text or conversation history a request could include. Applications handling longer material had to split, retrieve, or summarize content, each of which could lose relevant information.

Safety mitigations reduced some harmful behavior but did not prevent every unsafe, biased, or factually incorrect output. They could also cause false refusals on benign requests.

Performance varied by domain and prompt; results also depended on the evaluation method. Strong exam scores therefore could not substitute for testing the model on the actual distribution and failure costs of an application.

The model also offered little traceability from a response back to training examples or internal computations. This complicated debugging and auditing. It also constrained use in settings that required an accountable explanation.

Legacy

GPT-4 combined image input, stronger benchmark results, and a detailed safety report in a widely used commercial model. Later releases were frequently compared with its reported capabilities and limitations.

Its exam scores were easy to communicate, but they also illustrated the limits of benchmark-centered evaluation. A percentile says little about calibration, hallucination rates, sensitivity to prompt changes, or performance in a deployed workflow.

Image input helped make multimodal evaluation a common part of general-purpose model comparisons. Later systems expanded the set of accepted and generated media, but their designs were not necessarily derived from GPT-4's undisclosed architecture.

The technical report documented red-team findings and model-assisted safety evaluation alongside benchmark results. It also made clear that post-training and deployment controls reduce risk rather than certify a model as safe for general use.

GPT-4's release intensified debate about workplace automation and classroom use. Questions about authorship and AI governance also depended on how organizations deployed the model rather than on capability demonstrations alone.

Subsequent model releases competed on cost, latency, context length, benchmark performance, multimodal input, and safety evaluations. GPT-4 served as a comparison point, not a public blueprint: its detailed architecture and training data remained undisclosed.

General-Purpose AI Assistants

GPT-4 could be adapted through prompting and retrieval, sometimes with external tools. That flexibility reduced the need to train a separate model for every workflow. Production systems still required task-specific data and evaluation, plus controls appropriate to the deployment.

GPT-4's exam results provided reproducible evidence on selected tasks. Its uneven performance elsewhere showed why evaluations should test sensitivity to perturbations and measure calibration. Safety and the intended deployment domain required separate evaluation as well.

GPT-4 became a prominent reference point for multimodal general-purpose models in 2023. Its benchmark gains and broad interface attracted adoption. Hallucinations and opacity remained concerns, while cost and safety failures reinforced the need for application-specific verification.

Quiz

The following questions review GPT-4's image input, reported benchmarks, post-training, and limitations.

GPT-4 Quiz

Question 1 of 50 of 5 completed
What was one of the key innovations that distinguished GPT-4 from previous GPT models?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2025gpt4, author = {Michael Brenndoerfer}, title = {GPT-4: Multimodal LLMs Reach Human-Level Performance}, year = {2025}, url = {https://mbrenndoerfer.com/writing/gpt4-multimodal-language-models-reach-human-level-performance}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-27} }
APAAcademic
Michael Brenndoerfer (2025). GPT-4: Multimodal LLMs Reach Human-Level Performance. Retrieved from https://mbrenndoerfer.com/writing/gpt4-multimodal-language-models-reach-human-level-performance
MLAAcademic
Michael Brenndoerfer. "GPT-4: Multimodal LLMs Reach Human-Level Performance." 2026. Web. September 27, 2026. <https://mbrenndoerfer.com/writing/gpt4-multimodal-language-models-reach-human-level-performance>.
CHICAGOAcademic
Michael Brenndoerfer. "GPT-4: Multimodal LLMs Reach Human-Level Performance." Accessed September 27, 2026. https://mbrenndoerfer.com/writing/gpt4-multimodal-language-models-reach-human-level-performance.
HARVARDAcademic
Michael Brenndoerfer (2025) 'GPT-4: Multimodal LLMs Reach Human-Level Performance'. Available at: https://mbrenndoerfer.com/writing/gpt4-multimodal-language-models-reach-human-level-performance (Accessed: September 27, 2026).
SimpleBasic
Michael Brenndoerfer (2025). GPT-4: Multimodal LLMs Reach Human-Level Performance. https://mbrenndoerfer.com/writing/gpt4-multimodal-language-models-reach-human-level-performance

About the author

Continue with the full handbook

This chapter is part of History of Language AI. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore History of Language AI
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.