
"AI Tops University of Tokyo Entrance Exam"
OpenAI's latest model, ChatGPT, has posted scores surpassing the highest-scoring human test-takers on Japan's top university entrance exams. Just two years ago, AI had fallen short of passing scores on the same tests.
According to Kyodo News and the Nikkei on Thursday, GPT-5.2 scored 452 points in the humanities and 503 points in the sciences on this year's University of Tokyo entrance exam, which has a maximum score of 550. Both figures exceed the top scores announced by the university for actual test-takers, which were 434 for humanities and 453 for sciences. The model received a perfect score in mathematics.
The results were similar for the Kyoto University exam. GPT-5.2 scored 771 points in the law faculty exam and 1,176 points in the medical faculty exam, surpassing the top scores of 734 and 1,098, respectively. The scores effectively correspond to "top-ranked admission."
The assessment was conducted by Japanese AI startup LifePrompt and the Kawaijuku cram school, among others, based on actual questions from the 2026 academic year. The AI wrote the answers, which were then graded by a team of professional instructors. Teachers involved in the grading said the reasoning process had become "much more sophisticated and close to model answers." The AI also completed most answers within 30 minutes, far faster than human test-takers.

Even more striking is the pace of progress. According to LifePrompt, ChatGPT-4 failed to meet the minimum passing score on the University of Tokyo exam in 2024. In 2025, GPT-o1 barely cleared the passing threshold. In 2026, GPT-5.2 surged past the highest scores. In just two years, the model jumped from "failing" to "top-ranking."
Other AI models were not far behind. Anthropic's Claude Opus 4.5 and Google's Gemini 3 Pro Preview also recorded strong overall scores on the University of Tokyo and Kyoto University exams.
Each model showed distinct strengths, however. OpenAI excelled in STEM subjects such as math and chemistry, while Google scored relatively higher in humanities areas such as Japanese language and history. Some models posted top scores in specific subjects.
Limitations remain. Evaluators noted weaknesses in logical development on essay-style questions such as world history, and formal errors such as exceeding the allotted answer length were also found. The models still struggle with questions requiring contextual interpretation, such as those involving metaphor or satire.
Still, experts note that AI's learning capabilities have already surpassed those of top-ranked test-takers. The analysis is that competition is shifting from "how well AI can solve problems" to how humans utilize, design, and supervise these systems.






