OpenAI vs Claude Prediction Battle: Week 2
All About AIgo watch the original →
the gist
The author compares the predictive performance of GPT-4o and Claude 3 Opus on Polymarket by feeding them event criteria without market price data to see which model better forecasts real-world outcomes.
Methodology for AI Forecasting
The author conducts a weekly comparison between OpenAI's GPT-4o and Anthropic's Claude 3 Opus by tasking them with predicting the outcomes of three specific events: Ariana Grande's album debut sales, the July unemployment rate, and the peak temperature in London. To prevent market contamination, the author uses a scaffolding technique that provides the models with event rules and outcome ranges while explicitly withholding current Polymarket odds and price data. Both models are set to high reasoning modes, and the author clears the context window between runs to ensure independent predictions.
Prediction Comparison
For the second week of the challenge, the models largely converged on similar forecasts. Both models predicted a peak London temperature of 23 degrees Celsius and an unemployment rate of 4.2 percent. The primary divergence occurred in the Ariana Grande album sales prediction, where the models selected different sales brackets. The author places bets on the Polymarket platform based on these outputs to track performance, noting that Claude 3 Opus led the competition 1-0 following the first week of testing. The author intends to follow up with a tutorial on building custom machine learning models for Polymarket prediction to provide a more systematic edge.