Beta Today's Hot Take is brand new — expect rough edges, and tell us what you think.
We'll email you a one-click sign-in link. No password needed.
One story. Every side.

Technology

OpenAI quietly altered its new Astra model's benchmark scores multiple times after launch — boosting its own numbers while showing worse results for rival Anthropic's models. Routine tuning, or manipulating the scoreboard?

In early September 2026, OpenAI was caught changing evaluation metrics for its newly launched Astra model within hours of release, a practice critics labeled 'benchmaxxing' — adjusting test results to — Deceptive manipulation 100%, Normal post-launch adjustment 0%, Needs independent audit 0% (2 votes so far).