
OpenAI announced on the 6th that it has unveiled its latest large language model (LLM), GPT-5.4.
GPT-5.4 is a frontier model deployed across OpenAI's major products including ChatGPT, API, and Codex. The model integrates reasoning capabilities, coding performance, and agent-based tasks into a single unified system.
GPT-5.4 incorporates the coding capabilities of GPT-5.3 Codex while improving how it utilizes various tools and software in work environments such as spreadsheets, presentations, and documents. This enables more accurate and efficient execution of complex real-world tasks while reducing the iterative work required to achieve user-desired outcomes.
In terms of performance, GPT-5.4 showed meaningful improvements across major benchmarks. On GDPval, a benchmark that evaluates AI agents' ability to perform actual knowledge-based work, GPT-5.4 achieved results equal to or better than industry experts in 83% of all task comparisons. This represents a significant improvement over GPT-5.2's 71.0%. GDPval assesses models' real-world work performance based on tasks from 44 job categories representing major industries in the U.S. GDP.
During GPT-5.4's development, OpenAI particularly strengthened capabilities in spreadsheet, presentation, and document creation and editing. On an internal benchmark evaluating spreadsheet modeling tasks at the level of junior analysts at investment banks, GPT-5.4 scored an average of 87.5%, significantly exceeding GPT-5.2's 68.4%. Presentation creation also saw improvements in design completeness, visual diversity, image generation utilization, and factual accuracy.
Additionally, GPT-5.4 is the first general-purpose model among OpenAI's releases to feature built-in computer use capabilities. In Codex and API environments, AI agents can manipulate software in real computer environments, navigate across multiple applications, and execute complex workflows. GPT-5.4 supports up to 1 million tokens of context, making it suitable for building agent systems that plan, execute, and verify long-duration tasks. These capabilities have been validated with strong performance across various benchmarks including web browsing, desktop environment manipulation, and multimodal understanding.






