Skip to content
X · @simonw · X / Twitter

One of the things the pelican benchmark is still useful for is visually representing (to a tiny extent) the improvements in a single model family Here…

One of the things the pelican benchmark is still useful for is visually representing (to a tiny extent) the improvements in a single model familyHere's Meta AI's Spark (8th April), Spark 1.1 (9th July), and Spark 1.2 (today, 5th August)https://simonwillison.net/2026/Aug/5/muse-code-and-muse-spark-12/