GPT-4 Arrives! Bar Exam Score Jumps From the 10th to the 90th Percentile
Hi there, it's me, Shiichan! I found some huge news today that any tech fan will love. OpenAI just released a new AI model, so let's dive right in!
OpenAI NewsWhat was announced?
From OpenAI's News: a new large language model called GPT-4 has been announced. Thanks to its broader general knowledge and problem-solving abilities, GPT-4 can solve difficult problems with greater accuracy than before.
On top of that, it's more creative and collaborative too. For creative and technical writing tasks like composing songs or writing screenplays, GPT-4 can generate and edit text together with users, iterating back and forth. It can even learn a user's own writing style.
The original post gives a fun example: given the prompt "Explain the plot of Cinderella in a sentence where each word has to begin with the next letter in the alphabet from A to Z," GPT-4 pulled it off.
Why it matters
GPT-4 is described as the latest milestone in OpenAI's research, following the path from GPT, GPT-2, and GPT-3. It continues the same deep learning approach: using more data and more computation to build increasingly sophisticated and capable models.
The post includes an example that shows just how much of a leap this is. Given the scheduling question, "Andrew is free from 11 am to 3 pm, Joanne is free from noon to 2 pm and then 3:30 pm to 5 pm. Hannah is available at noon for half an hour, and then 4 pm to 6 pm. What are some options for a 30-minute meeting?", GPT-3.5 answered with 4 pm — a time when Joanne isn't actually available. GPT-4, on the other hand, correctly worked out that 12 pm to 12:30 pm is the only slot that works for everyone. Same question, clearly different reasoning accuracy.
What changes
With GPT-4 out, a lot of tasks should feel more dependable.
- For logical tasks that require juggling multiple constraints at once, like scheduling, GPT-4 can now reach more accurate answers.
- For creative work like songwriting or screenwriting, GPT-4 can iterate back and forth with you, making it a more capable creative collaborator.
- Exam scores show a big jump too. On the Uniform Bar Exam, GPT-3.5 scored around the 10th percentile among test-takers, while GPT-4 reached roughly the 90th percentile.
- On Biology Olympiad questions, GPT-3.5 scored around the 31st percentile, while GPT-4 with vision support reached roughly the 99th percentile.
Given gains like these, anyone who needs to hand off tasks with complex instructions or intricate constraints is likely to notice GPT-4's improved capabilities the most.
Dive Deep
OpenAI says it spent six months improving GPT-4's safety and what it calls "alignment" — making the model's behavior match human intent. According to OpenAI's internal evaluations, compared to GPT-3.5, GPT-4 is 82% less likely to respond to requests for disallowed content and 40% more likely to produce factual responses.
Here's what went into that safety work:
- More human feedback was incorporated, including feedback submitted by ChatGPT users, to improve GPT-4's behavior.
- OpenAI worked with over 50 experts in fields like AI safety and security to get early feedback.
- Lessons learned from real-world use of previous models were applied to GPT-4's safety research and monitoring system. Like ChatGPT, GPT-4 will keep getting updated and improved at a regular cadence as more people use it.
- GPT-4's own advanced reasoning and instruction-following abilities were used to help with safety work itself — creating training data for fine-tuning and iterating on classifiers used across training, evaluation, and monitoring.
On the infrastructure side, GPT-4 was trained on Microsoft Azure AI supercomputers, and that same AI-optimized infrastructure is what lets OpenAI deliver GPT-4 to users around the world.
That said, OpenAI is upfront that known limitations remain, including social biases, hallucinations, and vulnerability to adversarial prompts. OpenAI says it wants to keep encouraging transparency, user education, and wider AI literacy, and to expand the ways people can help shape how its models behave.
As for availability, it's straightforward: GPT-4 is available on ChatGPT Plus and as an API for developers.
Wrap-up
- OpenAI announced GPT-4, a new large language model that solves difficult problems more accurately thanks to broader knowledge and problem-solving skills.
- A worked example shows GPT-4 outperforming GPT-3.5 on a logical scheduling task that requires juggling multiple constraints.
- Exam scores jumped sharply: Uniform Bar Exam from around the 10th to the 90th percentile, and Biology Olympiad from around the 31st to the 99th percentile.
- Six months went into safety work too: 82% less likely to respond to disallowed content and 40% more likely to be factual, per internal evaluations.
- OpenAI is clear that limitations remain, including social biases, hallucinations, and adversarial prompt vulnerabilities.
- GPT-4 is available now on ChatGPT Plus and via API — great news for developers and creators who want AI to handle complex instructions, creative work, and logical constraints.