Skip to content

Question of the day

Artificial Analysis found GPT-6 Astra used roughly a third of Sol's tokens in the Codex coding harness but only about 10% fewer on general tasks, making it around 75% dearer per task there. What benchmark or test set do you trust when deciding which model handles which job?

No replies yet

Anyone can read the thread. Replying is for AISAT members. Sign in to join in, or apply if you have not already. It is free, and approval usually takes a day or two.