7
Why does everyone keep ignoring how dated benchmark data is for AI models?
I jumped on the AI bandwagon back in 2022 and I keep seeing posts where folks compare models using numbers from six months ago. Last week I saw a thread talking about a specific LLM's coding score from January and acting like it still matters. The thing is, these models get updated every few weeks, sometimes with huge jumps in performance. I run my own tests on a simple Python script I wrote for summarizing local news and the difference between April and June was night and day. Companies patch weaknesses fast so relying on old benchmarks is like judging a car by its 2023 safety rating when they added new brakes in 2024. How often do you all actually retest the tools you use versus just quoting those charts?
3 comments
Log in to join the discussion
Log In3 Comments
dakota4157d ago
Yeah six month old benchmarks are basically useless. I run my own tests every couple weeks on a few things I actually use like summarizing articles and writing emails and the difference between versions is huge. Companies are shipping updates so fast that those old numbers just don't tell you anything about what the model can do right now.
9
evan_green527d ago
Right there with you. I've noticed the same thing with coding tasks, the stuff that was terrible three months ago is almost usable now.
4
markh855d ago
Keep a running log of your own test prompts. It's the ONLY way to really track progress now since official benchmarks change every week.
7