OpenAI GPT-6 Astra identifies IKEA assembly errors with 80% accuracy
2 sources · Google News · The Decoder- Boom: GPT-6 Astra reaches 80% accuracy identifying IKEA furniture assembly errors from photos
- Boom: Best model scored only 28% on the same task in November 2025
- Doom: Epoch AI says processing speed is still too slow for real-time assembly guidance
- Boom: Epoch AI reports the speed gap to real-time performance is closing
The story in full
OpenAI's GPT-6 Astra model can examine a photo of assembled IKEA furniture and identify assembly mistakes, reaching 80 percent accuracy on the task, according to a report published September 26, 2026. Epoch AI evaluated the model's performance on this benchmark.
The 80 percent figure compares to 28 percent achieved by the best available model in November 2025, marking a substantial gain over roughly ten months. Epoch AI noted that the model's processing speed is not yet sufficient for real-time step-by-step assembly guidance, though the gap between current performance and that threshold is narrowing.
Analysis
356 wordsOn September 26, 2026, Epoch AI published an evaluation showing that OpenAI's GPT-6 Astra model can identify assembly mistakes in IKEA furniture from photographs with 80 percent accuracy. That figure is measured against a benchmark where the best available model scored only 28 percent in November 2025, meaning accuracy more than tripled in roughly ten months. Epoch AI also noted that the model's processing speed remains below what would be needed for real-time, step-by-step assembly guidance, while adding that the gap between current speed and that threshold is narrowing.
The jump from 28 to 80 percent is significant because spatial and mechanical reasoning from images has historically been one of the harder tasks for vision-language models. IKEA assembly errors involve understanding three-dimensional relationships, part orientation, and sequential logic from a flat photograph, which makes the benchmark a reasonable proxy for broader physical-world visual reasoning. The speed limitation matters practically: a model that can review a finished photo is useful, but one that can guide a user in real time during assembly would be far more valuable as a consumer or industrial tool. Whether the speed gap closes fast enough to matter competitively is a genuine open question.
No reactions from the Pro-AI, Anti-AI, or Middle Ground camps have been published on this story yet. Typically, the Pro-AI camp would treat an accuracy gain of this size over ten months as evidence that capability curves remain steep and that practical consumer applications are approaching faster than expected. The Anti-AI camp would likely emphasize the speed shortfall as a reminder that benchmark performance and real-world deployment readiness are different things, and might question how well the 80 percent figure holds across furniture types or image conditions not covered by the evaluation. A middle-ground position would probably acknowledge the progress as genuine while cautioning that closing the last gap to real-time performance often proves harder than the early gains suggest.
The clearest thing to watch is whether Epoch AI or OpenAI publishes updated speed benchmarks showing the model reaching real-time processing thresholds, and whether independent evaluations replicate the 80 percent accuracy figure across a wider range of assembly scenarios.
Where do you stand?
Add your take
0 reader votesSign in with Google to pick a side and post. Your vote moves the story's Doom / Boom score.


